A social training intervention system for children with autism based on a large model and VB-MAPP

Through a social training system for children with autism based on a large model and VB-MAPP, using Chat-GPT to generate personalized conversation content and combining it with multimodal evaluation, the problems of lack of professionalism and personalization in existing technologies are solved, and efficient intervention effects for children with autism are achieved.

CN119560107BActive Publication Date: 2025-10-03ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411598862.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-10-03
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

Existing intervention methods for children with autism lack professionalism and personalization, and the existing system lacks clinical evaluation basis, resulting in poor intervention effects.

Method used

A social training intervention system for children with autism based on a large model and VB-MAPP is adopted. Chat-GPT is used to generate personalized conversation content, and multimodal evaluation is performed by combining near-infrared and audio data. The system includes a perception module, a performance module, and a control module. The conversation process is designed according to the VB-MAPP paradigm, and the form of questions is restricted to improve children's enthusiasm for interaction.

Benefits of technology

The conversation quality and intervention effect of autistic children were improved, the effectiveness of the system was verified through multimodal data analysis, and the problem of scarce resources was partially compensated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119560107B_ABST
    Figure CN119560107B_ABST
Patent Text Reader

Abstract

The present invention discloses a kind of child autism social training intervention system based on large model and VB-MAPP, including perception module, performance module and control module; Perception module is used to collect information of testee, including content and audio of speech; Performance module is used to generate corresponding feedback content according to the information of testee and its own configuration; Control module is used to control the workflow of whole system, complete synchronization and record; Performance module builds system prompt words based on VB-MAPP paradigm, and uses large model to carry out dialogue generation according to system prompt words; System prompt words include three parts, the first part is the personal information of testee; The second part is that the current large model needs to play virtual image, including role description and ability limitation; The third part is the description of current dialogue topic. The present invention can generate personalized dialogue content for different autistic children, make dialogue richer, more easily attract children's attention, and enhance intervention effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of clinical intervention for autistic children, and in particular to a social training intervention system for autistic children based on a large model and VB-MAPP. Background Art

[0002] Autism spectrum disorder is a widespread neurodevelopmental disorder characterized by social impairments and repetitive, stereotyped behaviors. Other symptoms may also include difficulties with emotional regulation, language, and non-verbal communication. Individuals with autism struggle to care for themselves and integrate into social life, placing a significant burden on their families and society. Reports indicate that the prevalence of autism among school-age children worldwide is as high as 1 in 59, and the trend is increasing. Early diagnosis and effective intervention can, to a certain extent, improve the language, cognitive abilities, and behavioral habits of children with autism. Therefore, effective diagnosis and intervention for autism have become key areas of focus.

[0003] Due to insufficient and unbalanced medical resources, intervention resources are extremely limited in some areas. Therefore, researchers hope to use computer technology to assist professional interventionists in intervention or develop new systems and paradigms to complete independent intervention, so that more children with autism can receive intervention support. Current mainstream research often uses AR and robots as carriers, and uses manually designed intervention paradigms to guide autistic children to interact, and then collects indicators to evaluate the effectiveness of the intervention.

[0004] Such as the document "Social Intervention Strategy of Augmented Reality Combinedwith Theater-Based Games to Improve the Performance of Autistic Children inSymbolic Play and Social Skills." and the document "Effectiveness of a Robot-AssistedPsychological Intervention for Children with Autism Spectrum Disorder."

[0005] However, the current methods lack a certain degree of professionalism. The design of the paradigm method is not based on clinical assessment and intervention methods, and the intervention form lacks personalized content for subjects with different themes. Summary of the Invention

[0006] The present invention provides a social training intervention system for children with autism based on a large model and VB-MAPP, which can generate personalized dialogue content for different children with autism, making the dialogue richer, easier to attract children's attention, and improving the intervention effect.

[0007] A social training intervention system for children with autism based on a large model and VB-MAPP, comprising a perception module, a performance module, and a control module; the perception module is used to collect information from the subject, including speech content and audio; the performance module is used to generate corresponding feedback content based on the subject's information and its own configuration; and the control module is used to control the workflow of the entire system, completing synchronization and recording.

[0008] The performance module constructs system prompts based on the VB-MAPP paradigm and uses the large model to generate dialogue based on the system prompts. The system prompts include three parts: the first part is the subject's personal information; the second part is a description of the virtual image that the large model needs to play, including role description and ability limitations; and the third part is a description of the current dialogue topic.

[0009] The system's intervention process is divided into the preparation phase, intervention phase, and collection phase, as follows:

[0010] In the preparation stage, the personal information of the test child is input, and the intervention stage is officially entered after the virtual image is selected; in the intervention stage, the performance module combines the child's personal information, the virtual image of the large model and the current topic to construct the system prompt words, and inputs them into the large model, so that the large model can start limited multi-round dialogue intervention with the autistic child under the limited topic; the collection stage runs through the entire intervention stage. After the intervention on each topic is completed, the system uses the control module to save the near-infrared, text, and audio data collected by the perception module for evaluating the child's performance and the effectiveness of the system.

[0011] The Verbal Behavior Milestones Assessment and Placement Program (VB-MAPP) is a widely used assessment tool that divides items and assessment scores based on the subject's developmental level and provides guidance for intervention. Turn-taking is an effective intervention recommended by the VB-MAPP and a task format that excels at large-scale models. Therefore, based on the guidance provided by the VB-MAPP and the advice of practicing therapists, this paper designed a communication paradigm with predetermined themes. This paper customized five topics commonly used in daily communication: food, animals, toys, family, and colors.

[0012] During the conversation, the form of questions is limited to three types of questions: "who", "what" and "where", while questions of the form of "why", "how" and "when" are excluded. The present invention takes into account that the first three types of questions are related to specific memory images, which are relatively easy for subjects with autistic language disorders to answer; while "why, how and when" involve logical thinking, abstract generalization and time perception ability, and the questions are more complicated. The thinking and answer willingness of children with autism will decrease, and the intervention effect will not be achieved.

[0013] In the present invention, the large model can adopt Chat-GPT or other large language models.

[0014] Furthermore, in the first part of the system prompt words, the subject's personal information includes the subject's age, gender, nickname, preferences, interests and recent experiences. The content of the dialogue intervention topic is personalized based on this personal information, thereby increasing children's enthusiasm for interaction.

[0015] During the intervention phase, the large model began to conduct limited multi-round dialogue interventions with autistic children under limited topics, specifically:

[0016] The performance module asks questions after greeting the autistic child, and then provides different feedback based on the child's response. When the child does not respond, additional prompt words are input to allow the performance module to encourage the child to speak. When the child speaks, praise is given and a new round of conversation begins until the preset time is up.

[0017] During the intervention phase, a fixed prompt word was added before each round of dialogue to strengthen the limited dialogue generation designed based on the VB-MAPP paradigm.

[0018] During the intervention phase, after the conversation time is up, an ending prompt word will be input to allow the performance module to say goodbye to the autistic child.

[0019] During the intervention phase, in order to address the issue of autistic children not responding, the system uses sound detection to determine whether they are speaking. If there is silence for more than ten seconds, a pre-set encouraging voice will be triggered to encourage the child to speak.

[0020] During the collection phase, the text and audio of the conversation intervention were automatically saved by the system, and the near-infrared signal was collected by an 8-channel NirSmart system with a time resolution of 0.1 seconds, and the electrodes were located in the prefrontal cortex.

[0021] Compared with the prior art, the present invention has the following beneficial effects:

[0022] This paper builds on VB-MAPP, a current clinical autism intervention assessment method, to design the entire system paradigm and prompts. It also utilizes a large model as the system's conversation generation hub, thereby improving the quality of conversations during the intervention phase. The proposed system boasts strong conversational capabilities and supports personalized conversation content customization. Multimodal clinical data analysis demonstrates the effectiveness of the system's conversational intervention, potentially alleviating the scarcity of resources for autism intervention. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a schematic diagram of system prompt words based on VB-MAPP in an embodiment of the present invention;

[0024] Figure 2 is an intervention flow chart of the system of the present invention;

[0025] Figure 3 This is a diagram showing the analysis results of the conversation text during the intervention process;

[0026] Figure 4 Figure 2 is the analysis result of physiological signals during the intervention process. DETAILED DESCRIPTION

[0027] The present invention will be described in further detail below with reference to the accompanying drawings and examples. It should be noted that the following examples are intended to facilitate understanding of the present invention and do not have any limiting effect on the present invention.

[0028] To validate the effectiveness of the proposed paradigm and system, the present invention designed an experiment and collected data based on the following hypothesis: if the system can provide effective social intervention, then certain metrics of children's conversations with the system should be similar to those of professional autism interventionists. Therefore, the present invention conducted an experiment with 12 children diagnosed with autism. Eligibility criteria required that their therapists confirmed the ability to engage in multi-turn conversations. The present invention had the subjects engage in conversations with the proposed system, focusing on five topics designed based on the VB-MAPP, with a time limit of two minutes for each topic. Prior to the conversations, the present invention collected the subjects' preferences and personalized data, such as favorite foods or recent activities, from their parents. This data was then added to the system prompts as personalized content, making the conversations more targeted and personalized, thereby better engaging the children and enhancing the effectiveness of the intervention. Furthermore, the present invention had professional interventionists conduct conversations with these subjects, focusing on the same topics and constraints, with a total duration of less than 15 minutes. The effectiveness of the system was demonstrated through analysis of near-infrared, text, and audio data during the intervention period.

[0029] A social training intervention system for children with autism based on a large model and VB-MAPP. The specific implementation process is as follows:

[0030] (1) Implementation based on VB-MAPP paradigm: The paradigm is specifically applied to the construction of system prompts, such as Figure 1 The prompt is divided into three parts. The first part is the personal information of the subject, which mainly introduces the age, gender, and nickname of the subject. info It represents the child's personalized information, such as "His favorite food is dried tofu, and he hires an aunt to cook for him; he doesn't like small animals very much and has never raised one; he has a large family, including his parents, grandparents, and younger brother, whose nickname is Jiabao, and everyone plays with him; his favorite toys are cars, and he has many, including sedans, fire trucks, and so on; his favorite color is red, and he has a red water cup and a red car toy." The second part is the role description and ability restrictions, which are used to standardize the content and restrictions of the Chat-GPT dialogue generation; the third part is a description of the conversation topic. topic Describe specific topic information, such as food: "The topic of the conversation is food. Please develop your ideas around the theme. Here are some suggestions: staple food, snacks, drinks, desserts and snacks."

[0031] (2) The intervention process can be divided into three parts, such as Figure 2 As shown in the figure, during the preparation phase, an interface is used to input information and select system preferences. Based on information provided by parents and professional interventionists, the child's basic information and personal preferences are entered, and an avatar is selected to officially enter the intervention phase. During the intervention phase, the system inputs system prompts incorporating personalized information about the subject into Chat-GPT, allowing it to act as an interventionist through multiple rounds of conversation. Specifically, the system greets the autistic child and asks questions, then provides different feedback based on the child's response. If the child does not respond, a pre-recorded recording is played to encourage the child to speak. If the child speaks, praise is expressed, and a new round of conversation begins, continuing until the preset time for each topic (2 minutes) has expired. The collection phase continues throughout the intervention phase. After the intervention on all five topics is completed, the system saves the collected near-infrared, text, and audio data for use in evaluating the child's performance and the system's effectiveness. The text and audio recordings of the conversation intervention are automatically saved by the system. The near-infrared signals are collected by an 8-channel NirSmart system with a temporal resolution of 0.1 seconds, using electrodes located in the vmPFC.

[0032] (3) Implementation details of the social training intervention system for children with autism of the present invention: When the system conducts each round of dialogue (one question and one answer) with the autistic child, in order to enhance the context, an idle prompt will be added before the start of each round of dialogue to strengthen the dialogue generation based on the VB-MAPP paradigm design restrictions: "Continue to communicate around the topic and the content he mentioned. Only ask one question at a time, such as what, who, where, do not involve why, how, or when. Do not explain. If his speech is not clear, you can ask him to repeat it." After the dialogue time is up, there will be an endprompt input for ending, allowing the system to say goodbye to the autistic child: "The communication time is over, respond to my words, summarize our communication, and say goodbye to me." At the same time, in order to deal with the problem of autistic children not responding, the system uses sound detection to determine whether they have spoken. If there is silence for more than ten seconds, it will automatically generate "(I didn't speak)" as a response, and will trigger a pre-set encouraging voice to encourage the child to speak.

[0033] The system is written in python. Chat-GPT (gpt3.5-turbo-0125) is used as the conversation generation hub and is called through an API. Since Chat-GPT does not have the functions of voice-to-text and text-to-speech, the present invention uses the real-time speech recognition function of the iFLYTEK open platform to convert speech into text as a user prompt, and combines the system prompt input to Chat-GPT to complete the generation of conversation text, and then converts the text into speech through Microsoft's Azure service for output, and the voice role is selected as "zh-CN-YunxiNeural". When speech recognition is unsuccessful, "(unrecognizable speech, duration of about ${second} seconds.)" will be used as the filler for the user prompt, so that the system can reasonably respond to the speech content that cannot be correctly recognized.

[0034] (4) System effectiveness evaluation and result analysis: The experiment collected 12 autistic children aged 4-11 (11 boys and 1 girl). The intervention conversations between professional interventionists and autistic children were recorded by a voice recorder (XXJ1, HP Inc.), and text transcription was performed using paraformer and speaker recognition annotation was performed using cam++. To ensure the accuracy of the transcribed text, manual calibration and correction were performed based on the recognition results of the model. During the preprocessing process, the speech segments were separated from the audio and outliers were discarded. The raw intensity data of the near-infrared signal was first converted into optical density (OD) changes, and then motion artifacts were detected and sample interpolation was performed. An FIR filter (0.01–0.2 Hz) was used to extract low-frequency fluctuations, and the improved Beer-Lambert law was used to convert the OD data into hemoglobin concentration changes.

[0035] Analysis of texts such as Figure 3 As shown. Engagement reflects the level of children's participation in the intervention process. The higher the engagement, the more focused the child is. It indirectly reflects that ASD-Chat is able to attract children's attention and achieve the expected intervention effect. Engagement can be measured by text length and audio duration. We collected data on the average number of words (excluding punctuation) and average duration of each child's sentence in the five subjects. Figure 4 As can be seen, when interacting with ASD-Chat, children spoke more words and for longer periods of time than when interacting with a professional interventionist. On average, the number of words spoken when interacting with ASD-Chat was 13.11% higher and the duration of the interaction was 43.03% longer than when interacting with a professional interventionist. One possible explanation is that children find it easier to communicate with virtual cartoon characters than with adults. This is because they are accustomed to having similar conversations with toys in their daily lives. Communicating with virtual characters may put them in a more relaxed state, thereby increasing their enthusiasm and participation in the conversation. Response quality, measured by semantic similarity, reflects whether children can provide appropriate responses, rather than illogical nonsense. Results indicate that professional interventionists performed slightly better in response quality, with the system scoring 11.31% lower. The interventionists' extensive clinical experience enables them to effectively guide children to provide appropriate responses through language and gestures. However, the current system lacks this capability and requires improvement in this area.

[0036] The analysis of the audio is shown in Table 1 below.

[0037] Table 1

[0038]

[0039] At the audio level, analysis focuses on the subject's speech. Voiced segments are classified based on the frequency pattern of formants. Fundamental frequency (F0) and zero-crossing rate (ZCR) can, to a certain extent, represent the speech characteristics of a specific individual and their arousal level.

[0040] Table 1 lists the average F0 and ZCR for speech and voiced sounds, as well as the frequencies of the first three formants of voiced sounds. F1 is related to the degree of vowel opening, F2 to the front-back position of the tongue, and F3 to the position of the tongue tip and lip shape. Statistically, due to the different frequencies of the various voiced sounds in the comparison experiment, the mean values ​​of F0, ZCR, F1, F2, and F3 differed. Similar results were observed for speech. It can be seen that on these metrics, the same subjects performed similarly on the interventionist and proposed system. Specifically, in the ASD-Chat system experiment, speech F0 was 2% higher than the interventionist, speech ZCR was 48%, and voiced ZCR was 40%. Regarding voiced sound metrics, F0 was 12% higher, F1 was 11% lower, F2 was 5% higher, and F3 was 2% higher. Overall, the performance was very similar, validating the effectiveness of the system.

[0041] For near-infrared data, the first second of oxyhemoglobin (HbO) data was discarded for steady-state control and normalized using the z-score method to eliminate the influence of data units and facilitate comparisons across different conditions. Finally, the average HbO change for each channel was calculated for each subject across five conversation topics in both task scenarios to verify the effectiveness of the system.

[0042] This paper uses HbO amplitude induced by the same subject to illustrate the similarities between the ASD-Chat system and real clinical intervention scenarios. Blood oxygen amplitude is particularly related to social and communication characteristics, representing the core symptoms of autism spectrum disorder (ASD). Figure 4 As shown, the present invention calculated the HbO amplitude difference for each channel of all autistic children in five topic conversations in two scenarios. Compared with the clinician intervention scenario, the HbO amplitude changes of half of the channels in the ASD-Chat scenario were greater, while the brain activation intensity of four channels was weaker. It is worth noting that although the conversations in the intervention scenario led to relatively stronger fNIRS signal changes on some channels, the difference range between the two was only about 0.1. The difference between the two is almost negligible, which means that there is almost no difference in the HbO amplitude changes caused by the two scenarios. This proves that the ASD-Chat intervention system proposed in the present invention has similar effects to traditional intervention methods. It is expected to replace traditional intervention programs and save resources.

[0043] The embodiments described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A social training intervention system for children with autism based on a large model and VB-MAPP, characterized by: It includes a perception module, a performance module and a control module; the perception module is used to collect information about autistic children, including near-infrared, text, and audio data; the performance module is used to generate corresponding feedback content based on the information about autistic children and its own configuration; the control module is used to control the workflow of the entire system and complete synchronization and recording; The performance module constructs system prompts based on the VB-MAPP paradigm and uses the large model to generate dialogue based on the system prompts. The system prompts include three parts: the first part is the personal information of the autistic child; the second part is the virtual image that the large model needs to play, including role description and ability limitations; and the third part is a description of the current dialogue topic. The first part of the system prompt contains personal information about the child with autism, including their age, gender, nickname, preferences, interests, and recent experiences. This personal information is used to personalize the topic of the conversation intervention, thereby increasing the child's enthusiasm for interaction. The system's intervention process is divided into the preparation phase, intervention phase, and collection phase, as follows: In the preparation phase, the autistic child's personal information is input and an avatar is selected before the intervention phase officially begins. In the intervention phase, the performance module combines the autistic child's personal information, the large model's avatar, and the current topic to construct system prompts, which are then input into the large model, allowing the large model to begin limited, multi-round dialogue intervention with the autistic child under a limited topic. The collection phase runs through the entire intervention phase. After the intervention for each subject is completed, the system uses the control module to save the near-infrared, text, and audio data collected by the perception module for evaluating the performance of autistic children and the effectiveness of the system.

2. The social training intervention system for children with autism based on a large model and VB-MAPP according to claim 1 is characterized in that: In the third part of the system prompt words, themes include food, animals, toys, family and colors.

3. The social training intervention system for children with autism based on a large model and VB-MAPP according to claim 1 is characterized in that: During the intervention phase, the large model began to conduct limited multi-round dialogue interventions with autistic children under limited topics, specifically: The performance module asks questions after greeting the autistic child, and then gives different feedback based on the autistic child's response. When the autistic child does not respond, additional prompt words are input to allow the performance module to encourage the autistic child to speak. When the autistic child speaks, the module praises him or her and then starts a new round of conversation until the preset time is up.

4. The social training intervention system for children with autism based on a large model and VB-MAPP according to claim 1, characterized in that: During the intervention phase, a fixed prompt word was added before each round of dialogue to strengthen the limited dialogue generation designed based on the VB-MAPP paradigm.

5. The social training intervention system for children with autism based on a large model and VB-MAPP according to claim 4 is characterized in that: During the intervention phase, after the conversation time is up, an ending prompt word will be input to allow the performance module to say goodbye to the autistic child.

6. The social training intervention system for children with autism based on a large model and VB-MAPP according to claim 1 is characterized in that: During the intervention phase, in order to address the issue of autistic children not responding, the system uses sound detection to determine whether they are speaking. If there is silence for more than ten seconds, a pre-set encouraging voice will be triggered to encourage the autistic child to speak.

7. The social training intervention system for children with autism based on a large model and VB-MAPP according to claim 1 is characterized in that: During the collection phase, text and audio are automatically saved by the system, and near-infrared signals are collected by an 8-channel NirSmart system with a time resolution of 0.1 seconds. The electrodes are located in the prefrontal cortex.

Citation Information

Patent Citations

  • Dialogue machine self-service psychological intervention system and method

    CN115482912A

  • Self-adaptive intervention method, system and device based on social story training and medium

    CN115588485A