Children language narrative ability evaluation method and tool based on multi-modal analysis

By integrating speech, text, and visual behavioral features through multimodal data analysis, this approach addresses the problems of time-consuming, inefficient, and feature-ignoring traditional assessment methods, enabling a more comprehensive and accurate assessment of children's language narrative abilities and personalized intervention guidance.

CN121237120APending Publication Date: 2025-12-30TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511319652.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing methods for assessing children's language narrative rely on manual annotation and sentence-by-sentence transcription, which are time-consuming and inefficient, making it difficult to support large-scale screening or longitudinal tracking. Furthermore, they neglect non-verbal features such as phonological rhythm, facial expressions, gestures, and gaze, and cannot fully reflect children's true narrative abilities in natural communicative situations.

Method used

By collecting multimodal data such as speech, text, and visual behavior, a multimodal analysis method is constructed. Combining automatic speech recognition and computer vision analysis, speech prosody, text language, and visual behavior features are extracted to generate multidimensional quantitative scores. These scores are then compared with a norm database to output personalized intervention suggestions.

Benefits of technology

It enables a more comprehensive and accurate reflection of children's overall language abilities, improves the ecological validity, accuracy and efficiency of the assessment, supports large-scale screening and longitudinal tracking, provides personalized intervention guidance, and enhances the reliability and clinical applicability of the assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237120A_ABST
    Figure CN121237120A_ABST
Patent Text Reader

Abstract

The invention discloses a children's language narrative ability assessment method and tool based on multi-modal analysis, and relates to the technical field of natural language processing, and the technical scheme is characterized in that audio data and video data of a testee during a narrative process are collected, and voice rhythm features, text language features and visual behavior features are extracted for fusion; the method comprises the steps of inputting a multi-modal scoring model to generate multi-dimensional quantitative scores of a macrostructure, a microstructure, a language organization and a pragmatic function of the language narrative ability of a testee child, generating a structured evaluation report for the testee child according to the scores, and giving personalized intervention suggestions. According to the method, multi-modal data such as voice, language texts and visual behaviors are collected, so that the narrative ability of the children can be comprehensively analyzed from multiple dimensions such as macrostructures, microstructures, language organizations and pragmatic functions, and the method is closer to the real natural expression state of the children; therefore, the comprehensive language ability and communication performance of children can be reflected more comprehensively and truly.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of natural language processing, more particularly, it relates to a children's language narrative ability evaluation method and tool based on multi-modal analysis. BACKGROUND

[0002] Children's language narrative ability is a high-level stage and core component of children's language development, which goes beyond the isolated word and sentence level, referring to the comprehensive ability of children to organize and logically tell events, express intentions, and construct storylines in a specific context. This ability not only reflects the degree of children's mastery of language structure, but also closely related to their cognitive development, social communication ability, academic performance (such as reading comprehension and writing), and is an important indicator for identifying developmental language disorders (DLD), autism spectrum disorders (ASD), attention deficit hyperactivity disorder (ADHD), and other neurodevelopmental disorders, and provides key evidence for the development of subsequent intervention paths.

[0003] Existing children's language narrative evaluation generally uses structured story retelling or free narration tasks, and combines artificial scoring systems for analysis. This method usually relies on standardized picture stimulus materials (such as "Frog Story") and pre-set scoring scales (such as Story Grammar Scoring), and the evaluation focuses on the macrostructure of the narrative, i.e., whether it contains background, starting event, internal reaction, plan, attempt, result, and ending, etc. For the microstructure of the narrative, the research field has established a relatively mature index system, such as mean length of utterance (MLU), the proportion of use of subordinate and complex sentences, lexical diversity (such as MTLD, HD-D), and the use of cohesion means (pronouns, time markers, and logical connectives), which can effectively reflect the development level of children in syntactic complexity, lexical use, and discourse coherence.

[0004] However, in practical application, these analyses often rely on manual annotation and sentence-by-sentence transcription, which is time-consuming and inefficient, making it difficult to support large-scale screening or longitudinal tracking. Although in recent years, tools such as automatic speech recognition (ASR) have been introduced and can achieve automatic transcription to some extent, the recognition accuracy and seamless integration with syntactic and semantic analysis are still insufficient due to the particularity of children's speech (such as pronunciation variation, unstable speech speed, and mixed non-verbal sounds). At the same time, existing evaluations generally ignore phonetic prosody, facial expressions, gestures, and gaze, making it difficult to fully present children's real narrative ability in natural communication situations.

[0005] Therefore, the present application aims to comprehensively and objectively capture the narrative characteristics of children in a natural context by using multi-source data (voice, text, image, behavior), and to solve the above problems by constructing an evaluation method and tool with development tracking and intervention guidance capabilities. SUMMARY

[0006] The present application aims to provide a child language narrative ability evaluation method and tool based on multi-modal analysis. By collecting multi-modal data such as voice, language text and visual behavior, the present application can comprehensively analyze the narrative ability of children from multiple dimensions such as macro structure, micro structure, language organization and pragmatic function, and more closely reflect the real and natural expression state of children, so as to more comprehensively and truly reflect the comprehensive language ability and communication performance of children.

[0007] The above technical purpose of the present application is achieved by the following technical scheme: a child language narrative ability evaluation method and tool based on multi-modal analysis, comprising the following steps:

[0008] S1, presenting standardized narrative stimulus materials to the subject children, and synchronously collecting audio data and video data of the subject children during the narrative process;

[0009] S2, processing the collected audio data, extracting its speech rhythm features such as speech rate, pause and fundamental frequency, and converting the audio data into text data through automatic speech recognition technology;

[0010] S3, performing natural language processing analysis on the converted text data, and extracting text language features such as average sentence length, number of different word types, syntactic complexity and narrative structure elements;

[0011] S4, performing computer vision analysis on the collected video data, and extracting visual behavior features such as facial expression changes, gaze directions and gesture actions;

[0012] S5, fusing the extracted speech rhythm features, text language features and visual behavior features, and inputting them into a trained multi-modal scoring model to generate multi-dimensional quantitative scores of the macro structure, micro structure, language organization and pragmatic function of the language narrative ability of the subject children;

[0013] S6, comparing the generated multi-dimensional quantitative scores with a pre-stored normal database grouped by age, calculating the standard scores of each dimension, and generating a structured evaluation report;

[0014] S7, according to the ability short board dimension in the structured evaluation report, matching and outputting individualized intervention suggestions from an intervention strategy knowledge base.

[0015] The application is further provided: the standardized narrative stimulus material in step S1 includes picture story retelling, free narration, personal event narration and procedural narration task, wherein the synchronous acquisition is realized through a high-fidelity microphone and a high-definition camera, a synchronous time stamp is generated at the start of acquisition to ensure the time sequence alignment of the multi-modal data, and the duration of a single evaluation task is controlled within 5-10 minutes.

[0016] The application is further provided: the extraction of the prosodic features in step S2 specifically includes noise reduction and endpoint detection preprocessing of the audio data, extraction of the average fundamental frequency F0 and its standard deviation F0-Std by calculating the speech rate SR and the articulation rate AR, and identification and statistics of the pause frequency and average duration longer than 200 ms, wherein the calculation formula of the speech rate SR is:

[0017] SR=N words / T total

[0018] wherein N words is the total number of words, T total is the total speech duration, and the unit is word / minute.

[0019] The calculation formula of the articulation rate AR is:

[0020] AR=N phonemes / T phonation

[0021] wherein N phonemes is the total number of phonemes, T phonation is the total phonation duration, and the unit is phoneme / second.

[0022] The application is further provided: the syntactic complexity in step S3 is quantified by calculating the average dependency distance MDD, and the calculation formula is:

[0023]

[0024] wherein N is the total number of dependency relations, DependencyDistance i is the linear distance between words in the i-th dependency relation.

[0025] The application is further provided: the feature fusion in step S5 is realized by using a model based on an attention mechanism, the model assigns dynamic weights to each modal feature and then performs weighted summation to obtain the fused feature representation H fused , and the calculation formula is:

[0026] H fused =α v ×H v +α t ×Ht +α p ×H p

[0027] Among them, H v H t H p These are the encoded feature vectors of speech, text, and visual modalities, respectively, α. v α t α p These are the weight coefficients corresponding to each mode, calculated using an attention network.

[0028] The present invention is further configured such that the formula for calculating the standard Z-score in step S6 is:

[0029] Z = (X - μ) age ) / σ age

[0030] Where X is the raw score of the child in the study, μ age and σ age These are the mean and standard deviation of the normal sample of the same age, respectively.

[0031] This invention also provides a multimodal analysis-based assessment tool for children's language narrative ability, including a data acquisition module, a data processing module, a speech processing submodule, a text analysis submodule, a visual analysis submodule, a multimodal fusion scoring module, a norm comparison and report generation module, and an intervention suggestion module;

[0032] The data acquisition module is used to present standardized narrative tasks and simultaneously collect children's audio and video data;

[0033] The data processing module is communicatively connected to the data acquisition module, and is used to receive audio and video data, and perform the functions of the following sub-modules:

[0034] The speech processing submodule is used to extract speech prosody features and perform speech-to-text conversion;

[0035] The text analysis submodule is used to extract text language features;

[0036] The visual analysis submodule is used to extract visual behavioral features;

[0037] The multimodal fusion scoring module is communicatively connected to the data processing module and is used to receive speech prosody features, text language features, and visual behavior features, and output a multi-dimensional quantitative score after fusion analysis.

[0038] The norm comparison and report generation module is communicatively connected to the multimodal fusion scoring module and has a built-in norm database for generating a structured evaluation report containing standard scores and capability profiles.

[0039] The intervention suggestion module is connected to the norm comparison and report generation module and has a built-in intervention strategy knowledge base, which is used to output personalized intervention suggestions based on the assessment report.

[0040] The present invention is further configured such that: the data acquisition module includes a high-fidelity microphone with a sampling rate of not less than 44.1kHz and a bit depth of 16-bit; a high-definition camera with a resolution of at least 1080p and a frame rate of not less than 30fps; and a synchronization controller for controlling the synchronous triggering of the microphone and the camera, with a synchronization timestamp accuracy of ±10ms.

[0041] The present invention is further configured such that: the visual analysis submodule specifically includes a facial expression recognition unit, a gaze estimation unit, and a gesture recognition unit;

[0042] The facial expression recognition unit is configured to use a convolutional neural network (CNN) model to output the probability distribution of seven basic emotions and calculate the frequency of expression changes.

[0043] The gaze estimation unit is configured to estimate the child's gaze point and calculate the percentage duration during which the gaze deviates from the area of ​​the narrative stimulus material.

[0044] The gesture recognition unit is configured to recognize predefined narrative gestures based on a hand key point detection model.

[0045] The present invention also provides a children's language narrative ability assessment device based on multimodal analysis, comprising at least one processor; and a memory communicatively connected to at least one of the processors; wherein the memory stores instructions executable by the processor, the instructions being executed by the processor to implement a children's language narrative ability assessment method based on multimodal analysis.

[0046] In summary, the present invention has the following beneficial effects:

[0047] 1. This invention integrates and synchronously collects multimodal data such as speech, language text, and visual behavior, breaking through the limitations of traditional methods that rely solely on text or a single modality. It can comprehensively analyze children's narrative abilities from multiple dimensions such as macro-structure, micro-structure, language organization, and pragmatic function, making it closer to children's true and natural expression state, greatly improving the ecological validity of the assessment, and thus more comprehensively and realistically reflecting children's comprehensive language ability and communicative performance.

[0048] 2. This invention replaces the traditional manual scoring method that relies on the subjective experience of assessors by constructing a multimodal fusion analysis algorithm and an automated scoring engine based on machine learning models. This effectively reduces the errors and biases caused by human subjective factors, ensures the objectivity, consistency and repeatability of the scoring results, significantly improves the reliability and accuracy of the assessment, and provides comparability for assessment results between different children and at different time points.

[0049] 3. This invention achieves a high degree of automation in the entire process from multimodal data acquisition, feature extraction, fusion analysis to report generation, avoiding the disadvantages of traditional methods that are cumbersome, time-consuming and labor-intensive, greatly improving the evaluation efficiency, enabling rapid screening of large-scale populations, school surveys and frequent clinical evaluations, and effectively reducing human and time costs.

[0050] 4. This invention is designed with the language and cognitive characteristics of children aged 3-6 years in mind. It features fun and standardized narrative tasks and uses speech recognition and visual analysis models optimized for children to ensure applicability and assessment accuracy in this age group. At the same time, the system tools have built-in developmental norms and longitudinal tracking functions, which can visualize the child's ability development curve, providing a continuous and dynamic basis for monitoring the intervention effect and adjusting rehabilitation strategies.

[0051] 5. This invention can not only assess children's language narrative ability, but also identify the shortcomings of children's language narrative ability based on the automated assessment results through the built-in intervention strategy knowledge base. Then, it can automatically match and generate personalized and actionable intervention suggestions. By implementing the assessment results and intervention guidance at the same time, a scientific process of "assessment-intervention-reassessment" is formed, which greatly improves the clinical applicability and translational value of this invention. It can directly provide clear guidance for the rehabilitation education and family training of speech therapists, teachers and parents.

[0052] 6. The system tool of this invention has a high degree of integration and great potential for productization and promotion. By integrating functions such as multimodal data acquisition, intelligent algorithm analysis, norm comparison, and report generation, it provides a complete system tool solution. It is also equipped with a user-friendly interface and standardized operating procedures, making the system tool easy to deploy and apply in various scenarios such as hospital rehabilitation departments, special education schools, ordinary kindergartens, and families. It has good productization prospects and broad social promotion potential, which helps to promote the early identification and accurate intervention of children's language disorders. Attached Figure Description

[0053] Figure 1 This is a flowchart illustrating a method for assessing children's language narrative ability based on multimodal analysis in Embodiment 1 of the present invention;

[0054] Figure 2This is a schematic diagram of the module structure of a children's language narrative ability assessment tool based on multimodal analysis in Embodiment 2 of the present invention;

[0055] Figure 3 This is a schematic diagram of the data processing module in Embodiment 2 of the present invention;

[0056] Figure 4 This is a schematic diagram of the framework of a children's language narrative ability assessment device based on multimodal analysis in Embodiment 3 of the present invention. Detailed Implementation

[0057] The following is in conjunction with the appendix Figures 1-4 The present invention will be described in further detail below.

[0058] Example 1: A method for assessing children's language narrative ability based on multimodal analysis. The hardware configuration used in this example is as follows:

[0059] Main computer: Equipped with an Intel Core i7-12700K processor, 32GB DDR4 memory, 1TB NVMe solid-state drive, NVIDIA GeForce RTX 3080 graphics card (12GB VRAM), running Windows 11 operating system. This configuration provides computing power for real-time processing of multimodal data and deep learning model inference.

[0060] Audio acquisition equipment: Shure MV7 professional USB microphone, sampling rate set to 48kHz, sampling depth to 24-bit, frequency response range to 20Hz-20kHz, microphone gain set to 70%, placed at a distance of 30±5cm from a child's mouth, with a pop filter to reduce airflow impact noise.

[0061] Video capture equipment: A Logitech Brio 4K Ultra HD webcam was used, with a resolution of 1080p (1920×1080 pixels) and a frame rate of 30fps. The camera was placed directly above the monitor, with the center of the lens at eye level and a distance of approximately 60±10cm. This ensured that the captured image included the child's head, shoulders, and upper body movements. The lighting environment was controlled at 300-500 lux to avoid overexposure or backlighting.

[0062] Display device: 27-inch IPS LCD display with a resolution of 2560×1440, used to present narrative stimulating materials.

[0063] S1. Present standardized narrative stimulus materials to the children and simultaneously collect audio and video data of the children during the narrative process.

[0064] After the system tool is started, the assessor selects the "Puppy Finds the Bone" picture story task suitable for children aged 4-5. The system tool presents a series of 5 pictures on the screen in sequence, with each picture lasting 15 seconds. Then, the child is prompted to start telling the story. The system tool starts audio and video recording simultaneously through a software trigger and writes a uniform timestamp (accuracy ±10ms). The total recording time is strictly controlled within 7 minutes. If the child's storytelling exceeds the time limit, the system tool automatically stops recording and saves the collected data.

[0065] The parameter controls include: image presentation time: 15 ± 0.5 seconds / image, controlled by the system's internal timer; total acquisition duration: 420 seconds (7 minutes), automatically terminating upon timeout; audio sampling parameters: 48kHz, 24-bit, mono, with accuracy ensured by the microphone firmware and driver; video acquisition parameters: 1920×1080@30fps, YUY2 color encoding, controlled by the camera driver; synchronization accuracy: ±10ms, achieved through the system's high-precision clock (QueryPerformanceCounteronWindows).

[0066] S2. Process the collected audio data, extract its speech prosodic features such as speech rate, pauses, and fundamental frequency, and convert the audio data into text data through automatic speech recognition technology.

[0067] The system calls the LibROSA library to preprocess the original audio. First, pre-emphasis (coefficient 0.97) is performed to compensate for high frequencies. Then, noise reduction based on spectral subtraction is used (the first 50 frames are noise, noise figure = 0.2). Speech segments are segmented using a dual-threshold endpoint detection algorithm based on short-time energy and zero-crossing rate (energy threshold: -50dB, zero-crossing rate threshold: 30). The silence segment length threshold is set to 200ms. Then, feature extraction is performed, including the calculation formula for speech rate (SR).

[0068] SR=N words / T total

[0069] Where, N words To determine the total number of words identified, T total Total speech duration, unit: words per minute;

[0070] The formula for calculating the rate of sound (AR) is:

[0071] AR = N phonemes / T phonation

[0072] Where, N phonemes T is the total number of phonemes. phonationTotal speech duration, unit: phonemes / second. In this embodiment, 85 words were specifically identified, with a pure speech duration of 38.2 seconds, SR = 133.5 words / minute; Average fundamental frequency (F0): extracted using the PYIN algorithm, frame length 25ms, frame shift 10ms. The mean F0 was calculated to be 328.5Hz (F0-Std = 42.1Hz); Pause characteristics: the number of pauses (12 times) and the average pause duration (0.68 seconds) were counted.

[0073] S3. Perform natural language processing analysis on the converted text data to extract textual language features such as average sentence length, number of different word types, syntactic complexity, and narrative structure elements.

[0074] The denoised audio was input into the Wav2Vec 2.0 ASR model, which is optimized for children's speech (fine-tuned on 500 hours of Chinese children's speech corpus), for transcription, and the output text was: "Then the puppy searched and searched in the garden (pause) and found its bone (laughter)."

[0075] Average Sentence Length (MLU): Based on morphemes, this example is divided into 4 sentences with a total of 19 morphemes, MLU = 4.75; Lexical Diversity: TTR = NDW / NTW = 15 / 85 = 0.176; Syntactic Complexity: Stanford Parser is used for dependency parsing to calculate the average dependency distance (MDD), MDD = 2.31 in this example; Narrative Structure: Based on a predefined StoryGrammar rule set, the CRF sequence labeling model is used to identify "character" (puppy), "scene" (garden), "attempt" (find), and "result" (find), but the "cause" is missing.

[0076] S4. Perform computer vision analysis on the collected video data to extract visual behavioral features such as facial expression changes, gaze direction, and hand gestures.

[0077] The video stream is processed frame by frame. First, the Dlib library is used for face detection and localization of 68 key points, and then feature calculation is performed.

[0078] Gaze direction: The gaze point was estimated using the appearance-based GazeNet model, and the percentage of time the gaze deviated from the picture area was calculated to be 22.4%.

[0079] Facial expressions: A Mini-Xception network (trained on FER2013) was used to output the probabilities of 7 expression categories, and expression entropy was extracted as a change indicator: H express =-∑(p i ·log2p i ), where p iGiven the probability of each type of expression, the average expression entropy in this embodiment is 1.82.

[0080] Gesture recognition: The MediaPipe Hands model was used to detect 21 key hand points. The "pointing" gesture was defined as the index finger being straight and the other fingers being bent. The gesture frequency was 5.2 times per minute.

[0081] S5. The extracted phonological prosody features, textual language features, and visual behavioral features are fused and input into a trained multimodal scoring model to generate a multidimensional quantitative score of the test children's language narrative ability, including macro-structure, micro-structure, language organization, and pragmatic function.

[0082] The 25-dimensional features extracted from the three modalities (6-dimensional speech, 8-dimensional text, and 11-dimensional vision) are Z-score normalized and then input into a pre-trained multimodal fusion model.

[0083] The features are first modally specific encoded through three independent fully connected layers (64 dimensions), that is, three independent fully connected networks (FCNs) encode the speech feature vector F. v Text feature vector F t Visual feature vector F p Mapping to the same latent space yields the encoded feature vectors H of speech, text, and visual modalities. v H t H p .

[0084] A multimodal fusion network based on additive attention mechanism is used to calculate the attention weights of each modality feature:

[0085] α v =softmax(W a ×tanh(W v ×H v +b v ))

[0086] α t =softmax(W a ×tanh(W t ×H t +b t ))

[0087] α p =softmax(W a ×tanh(w p ×H p +b p ))

[0088] Among them W a Wv W t H p Let b be the learnable weight matrix, and b be the bias term. Then, the fused feature representation H is calculated. fused The calculation formula is as follows:

[0089] H fused =α v ×H v +α t ×H t +α p ×H p

[0090] The fused features H_fused (64 dimensions) are input into a 3-layer MLP (structure: 64->32->16->4) for regression prediction, and the output is the raw scores of 4 dimensions.

[0091] Scoring output:

[0092] Macro structure: Raw score = 68.2 / 100

[0093] Microstructure: Original score = 72.5 / 100

[0094] Language organization: Raw score = 58.4 / 100

[0095] Pragmatic function: Raw score = 76.8 / 100

[0096] S6. Compare the generated multi-dimensional quantitative scores with the pre-stored norm database grouped by age, calculate the standard scores for each dimension, and generate a structured evaluation report.

[0097] The system tools query the built-in norm database (containing data on 1250 children aged 4-5) to obtain the norm parameters (μ) for each dimension of this age group. macro =75.3, σ macro =12.1; μ fluency =71.6, σf luency =14.2;...).

[0098] Z-score calculation:

[0099] Z macro =(68.2-75.3) / 12.1=-0.59

[0100] Z fluency =(58.4-71.6) / 14.2=-0.93

[0101] Similarly, for other dimensions, if it is the child's second assessment, the system extracts the historical scores from the child's personal file, plots the ability development curve, and calculates the slope to assess the rate of progress.

[0102] S7. Based on the capability gap dimensions in the structured assessment report, match and output personalized intervention recommendations from the intervention strategy knowledge base.

[0103] The system tool populates child information, scores, Z-values, radar charts, and other content into a preset HTML report template, and then converts it into a PDF document.

[0104] Intervention strategy matching: The system identified Z fluency The lowest value (-0.93) is used, so strategies matching this tag are retrieved from the intervention knowledge base. The knowledge base is a graph structure, where nodes represent intervention activities and edges represent conditional logic. The matched suggestions are:

[0105] "Slow-speed imitation and repetition training" (match rate: 92%)

[0106] "Rhythmic tapping to aid fluency training" (match rate: 87%)

[0107] "Breathing support exercises" (match rate: 78%)

[0108] Final output: The generated structured report includes four parts: assessment results, comparison with norms, development curve, and intervention recommendations, and is presented directly to the assessors through the system interface.

[0109] Example 2: A multimodal analysis-based assessment tool for children's language narrative ability, including a data acquisition module, a data processing module, a speech processing submodule, a text analysis submodule, a visual analysis submodule, a multimodal fusion scoring module, a norm comparison and report generation module, and an intervention suggestion module.

[0110] In this embodiment, the data acquisition module uses the main control computer, audio acquisition device, video acquisition device, and display device as described in Embodiment 1 above. It is used to present standardized narrative tasks and simultaneously collect children's audio and video data. The module has built-in various narrative task stimulus materials (including the "Frog Story" series of pictures, scenario prompt cards, etc.), which are presented on the screen in a graphic and textual form to guide children to retell stories, tell stories freely, etc. The software controls the microphone and camera to trigger synchronously. A synchronization timestamp (accuracy ±10ms) is generated when the acquisition begins to ensure that all modal data streams are aligned on the timeline. It is recommended that the duration of a single assessment task be controlled within 5-10 minutes to avoid children's fatigue.

[0111] The data processing module communicates with the data acquisition module to receive audio and video data, and specifically includes the functions of the following sub-modules:

[0112] The speech processing submodule is used to extract speech prosodic features and perform speech-to-text conversion. This submodule first performs noise reduction, endpoint detection (VAD), and automatic speech recognition (ASR) preprocessing operations on the input raw WAV format audio data. In the noise reduction process, spectral subtraction or a deep learning-based noise reduction model (including Demucs) is used to improve the signal-to-noise ratio (SNR) by at least 10 dB. Endpoint detection uses a dual-threshold method based on energy and zero-crossing rate to accurately segment speech segments and non-speech segments and eliminate silent segments. In the automatic speech recognition, an ASR engine optimized for children's speech (including fine-tuning on a children's speech corpus using Wav2Vec 2.0 or Conformer models) is used to convert speech to text. The transcription accuracy (Word Error Rate, WER) is required to be less than 15% for typical developing children's speech and less than 25% for children with language disorders.

[0113] Then, prosodic acoustic features were extracted from the cleaned speech signal, with a sampling frame length of 25ms and a frame shift of 10ms. These features included: Speaking Rate (SR), in words per minute (normal range: approximately 80-180 words per minute for children aged 3-6); Articulation Rate (AR), in phonemes per second; Mean Fundamental Frequency (Mean F0): the fundamental frequency was extracted using the Praat script or PYIN algorithm, and the mean was calculated in Hz (typically 250-400Hz for children); Fundamental Frequency Standard Deviation (F0Std): characterizing the richness of intonation variation; Pause Frequency and Duration: silent segments longer than 200ms were defined as pauses, and the number of pauses per minute and the average pause duration were calculated; Acoustic Energy: the root mean square (RMS) energy was calculated to characterize changes in sound loudness.

[0114] The text analysis submodule is used to extract linguistic features from the text. It takes the text output by ASR, performs spell correction (based on a children's common word dictionary), word segmentation, and part-of-speech tagging (POS tagging), and then extracts linguistic features, including: Mean Length of Sentence (MLU): the ratio of morphemes to sentences; Number of Different Word Forms (NDW): the total number of different word forms used in the narrative; Word Class Ratio (TTR): representing lexical diversity; Syntactic Complexity: using SyntaxNet or StanforCoreNLP tools for dependency parsing, calculating the Mean Dependency Distance (MDD), a larger MDD usually indicates a more complex sentence structure; Connective Use: statistically analyzing the frequency of coordinating conjunctions (including "and," "then") and subordinating conjunctions (including "because," "therefore"); Narrative Structure Elements: based on a predefined narrative grammar model (including StoryGrammar), using rule matching or sequence labeling models (including CRF / BiLSTM) to identify whether the text contains "Setting," "Initiating Event," and "Internal Response." Elements such as "Response", "Attempt", and "Consequence" are included.

[0115] The visual analysis submodule is used to extract visual behavioral features, specifically including a facial expression recognition unit, a gaze estimation unit, and a gesture recognition unit. In the facial expression recognition unit, the input high-definition video stream is preprocessed, and face detection and alignment are performed using OpenCV or Dlib libraries. The face regions in each frame are normalized to 256x256 pixels, and then visual feature extraction is performed, including facial expression recognition: a convolutional neural network (including VGG16 or ResNet) pre-trained on the FER2013 or AffectNet dataset is used to output seven basic emotions (neutral, happy, surprised, sad, angry, etc.). The probability distribution of disgust and fear is analyzed, and the frequency of facial expression changes (the number of dominant facial expression changes per unit time) is extracted as a feature. In the gaze estimation unit, appearance-based methods (including the iTracker.x model) or geometry-based methods are used to estimate the child's gaze point (x, y coordinates on the screen) and calculate the percentage duration of the gaze deviating from the center of the screen (stimulus area). Head posture estimation: the pitch, yaw, and roll angles of the head are calculated by solving the PnP problem, and the head movement amplitude (standard deviation of angle change) is extracted. In the gesture recognition unit, the MediaPipe Hands model is used to detect 21 key points of the hand, and rules (including five fingers open, fist clenched, pointing) or a simple classifier are defined to recognize common narrative gestures.

[0116] The multimodal fusion scoring module communicates with the data processing module to receive speech prosody features, text language features, and visual behavior features. It takes a standardized feature vector (Z-score normalized) extracted from the three sub-modules as input and uses an attention-based multimodal fusion network to compute the fused feature representation H. fused Then, using the intelligent scoring engine, H... fused The input is fed into a multilayer perceptron (MLP) regressor, which outputs standardized scores (Z-scores) for four core dimensions, including:

[0117] Macro-Structure Score: Assesses the completeness and logic of the story structure;

[0118] Micro-Structure score: assesses the complexity of vocabulary and grammar;

[0119] Fluency score: assesses the fluency of speech flow;

[0120] Pragmatics score: assesses the appropriateness of nonverbal communicative behaviors.

[0121] The norm comparison and report generation module communicates with the multimodal fusion scoring module and has a built-in norm database, including a database based on a large sample (N>1000, stratified by age, gender, and region). The database stores the mean (μ) and standard deviation (σ) of the scores of children in each age group (3-3.5 years, 3.5-4 years, etc.) in four dimensions. Then, the Z-score is calculated based on the child's score X, which directly reflects the child's ability level relative to his / her peers. At the same time, the system tools create an independent profile for each child and have developmental tracking functions. After each assessment, the score will be recorded and plotted on a development curve graph, with the horizontal axis representing the assessment time and the vertical axis representing the Z-score, intuitively showing the trend of the child's ability over time.

[0122] The intervention suggestion module communicates with the norm comparison and report generation module, automatically populating the child's basic information, scores for each dimension (raw score, Z-score), ability profile radar chart, and comparative analysis with developmental norms to generate a structured PDF report. At the same time, the system has a built-in intervention strategy knowledge base, which automatically identifies the 1-2 dimensions with the lowest Z-score as target intervention areas based on the barrel principle, and matches preset intervention activity suggestions from the knowledge base (for example, if the "microstructure" score is the lowest, "sentence expansion game" and "connective word fill-in-the-blank exercise" are recommended).

[0123] Example 3: A device for assessing children's language narrative ability based on multimodal analysis, comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the processor, the instructions being executed by the processor to implement a method for assessing children's language narrative ability based on multimodal analysis, the method comprising: presenting standardized narrative stimulus materials to the test children and simultaneously collecting audio and video data of the test children during the narrative process; processing the collected audio data, extracting its speech rate, pauses, and fundamental frequency prosodic features, and converting the audio data into text data through automatic speech recognition technology; performing natural language processing analysis on the converted text data, extracting its average sentence length, different... The study analyzes textual language features, including word count, syntactic complexity, and narrative structure elements. It performs computer vision analysis on collected video data to extract visual behavioral features such as facial expression changes, gaze direction, and gestures. The extracted phonological features, textual language features, and visual behavioral features are then fused and input into a trained multimodal scoring model to generate multidimensional quantitative scores for the macro-structure, micro-structure, language organization, and pragmatic function of the children's language narrative abilities. These multidimensional quantitative scores are compared with a pre-stored, age-grouped norm database to calculate standard scores for each dimension and generate a structured assessment report. Based on the ability gap dimensions in the structured assessment report, personalized intervention recommendations are matched and output from an intervention strategy knowledge base.

[0124] This specific embodiment is merely an explanation of the present invention and is not intended to limit the invention. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they are within the scope of the claims of the present invention.

Claims

1. A method for evaluating children's language narrative ability based on multi-modal analysis, characterized in that: The method comprises the following steps: S1, presenting standardized narrative stimulus materials to the test children, and synchronously collecting audio data and video data of the test children during the narration process; S2, processing the collected audio data, extracting its speech rhythm features such as speech rate, pause and fundamental frequency, and converting the audio data into text data through automatic speech recognition technology; S3, performing natural language processing analysis on the converted text data, and extracting its text language features such as average sentence length, number of different word types, syntactic complexity and narrative structure elements; S4, performing computer vision analysis on the collected video data, and extracting visual behavior features such as facial expression changes, gaze directions and gesture actions; S5, fusing the extracted speech rhythm features, text language features and visual behavior features, inputting them into a trained multi-modal scoring model, and generating multi-dimensional quantitative scores of the macro structure, micro structure, language organization and pragmatic function of the language narrative ability of the test children; S6, comparing the generated multi-dimensional quantitative scores with a pre-stored norm database grouped by age, calculating the standard scores of each dimension, and generating a structured assessment report; S7, according to the short board dimension in the structured assessment report, matching and outputting individualized intervention suggestions from an intervention strategy knowledge base.

2. The method of claim 1, wherein the method is based on multi-modal analysis of children's language narrative ability. The standardized narrative stimulus materials in step S1 include picture story retelling, free narration, personal event narration and procedural narration tasks, wherein the synchronous collection is realized through a high-fidelity microphone and a high-definition camera, a synchronous time stamp is generated at the start of collection to ensure the time sequence alignment of multi-modal data, and the duration of a single evaluation task is controlled within 5-10 minutes.

3. The method of claim 1, wherein the method is based on multi-modal analysis. The extraction of speech rhythm features in step S2 specifically includes noise reduction and endpoint detection preprocessing of audio data, calculation of speech rate SR and pronunciation rate AR, extraction of average fundamental frequency F0 and its standard deviation F0-Std, identification and statistics of pause frequency and average duration with a duration longer than 200ms, the calculation formula of speech rate SR is: SR = N words / T total wherein N words is the total number of words identified, T total is the total speech duration, in words per minute; The calculation formula of pronunciation rate AR is: AR = N phonemes / T phonation where N phonemes is the total number of phonemes, T phonation is the total duration of utterance, unit: phonemes / second.

4. The method of claim 1, wherein the method is based on multi-modal analysis. The syntactic complexity in step S3 is quantified by calculating the average dependency distance MDD, and the calculation formula is: where N is the total number of dependency relations, DependencyDistance i is the linear distance between words in the ith dependency relation.

5. The method for evaluating children's language narrative ability based on multi-modal analysis according to claim 1, characterized in that: The step S5 is implemented by using a model based on an attention mechanism. The model assigns dynamic weights to each modality feature and performs weighted summation to obtain a fused feature representation H fused The calculation formula is: H fused = a v x H v + a t x H t + a p x H p wherein H v , H t , H p are the feature vectors of the speech, text, and visual modalities after encoding, respectively, and a v , a t , a p are the weight coefficients of each modality calculated by the attention network.

6. The method for evaluating children's language narrative ability based on multi-modal analysis according to claim 1, characterized in that: The calculation formula of standard score Z-score in step S6 is: Z = (X - μ age ) / σ age where X is the raw score of the child under test, μ age and σ age are the mean and standard deviation of the normative sample of the same age, respectively.

7. A tool for evaluating children's language narrative ability based on multi-modal analysis, applied to the method for evaluating children's language narrative ability based on multi-modal analysis in any one of claims 1-6, characterized in that: The system comprises a data collection module, a data processing module, a speech processing submodule, a text analysis submodule, a visual analysis submodule, a multi-modal fusion scoring module, a norm comparison and report generation module and an intervention suggestion module; The data collection module is used for presenting standardized narrative tasks and synchronously collecting audio and video data of children; The data processing module is in communication connection with the data collection module, used for receiving audio and video data, and performing the functions of the following submodules: The speech processing submodule is used for extracting speech rhythm features and performing speech-to-text conversion; The text analysis submodule is used for extracting text language features; The visual analysis submodule is used for extracting visual behavior features; The multi-modal fusion scoring module is in communication connection with the data processing module, used for receiving speech rhythm features, text language features and visual behavior features, and outputting multi-dimensional quantitative scores after fusion analysis; The norm comparison and report generation module is connected with the multi-modal fusion scoring module in communication, has a built-in norm database, and is configured to generate a structured assessment report containing a standard score and an ability profile. The intervention suggestion module is connected with the norm comparison and report generation module in communication, has a built-in intervention strategy knowledge base, and is configured to output individualized intervention suggestions according to the assessment report.

8. The tool for assessing children's language narrative ability based on multi-modal analysis according to claim 7, characterized in that: The data acquisition module includes a high-fidelity microphone with a sampling rate of no less than 44.1 kHz and a bit depth of 16-bit, a high-definition camera with a resolution of at least 1080p and a frame rate of no less than 30 fps, and a synchronous controller configured to control the microphone and the camera to be triggered synchronously, with a synchronization timestamp accuracy of ±10 ms.

9. The tool for assessing children's language narrative ability based on multi-modal analysis according to claim 7, characterized in that: The visual analysis submodule specifically includes a facial expression recognition unit, a gaze estimation unit, and a gesture recognition unit. The facial expression recognition unit is configured to output a probability distribution of seven basic emotions using a convolutional neural network (CNN) model and to calculate an expression change frequency. The gaze estimation unit is configured to estimate a child's gaze landing point and to calculate a percentage of time when the gaze deviates from a narrative stimulus material region. The gesture recognition unit is configured to recognize predefined narrative gestures based on a hand key point detection model.

10. A device for assessing children's language narrative ability based on multimodal analysis, characterized in that: The apparatus includes at least one processor and a memory connected with the at least one processor in communication, wherein the memory stores instructions executable by the processor, and the instructions are configured to be executed by the processor to implement the method of any one of claims 1-6.

Citation Information

Cited By

  • Data processing method and system based on language function detection system

    CN122090820A