Intelligent AR glasses life auxiliary system for autism spectrum teenagers

By integrating environmental perception and real-time dialogue communication modules into smart AR glasses, and utilizing advanced algorithms and personalized data adjustments, the safety and communication needs of adolescents on the autism spectrum in complex environments are addressed, achieving highly integrated life assistance support.

CN121807167APending Publication Date: 2026-04-07NINGBO INST OF NORTHWESTERN POLYTECHNICAL UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-12
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies are insufficient in terms of real-time performance, accuracy, adaptability, and functional synergy when addressing the travel safety and social communication needs of adolescents on the autism spectrum, and cannot provide timely and appropriate integrated support in complex environments.

Method used

Design an intelligent AR glasses system that integrates an environmental perception and guidance module and a real-time dialogue communication module. Utilize convolutional neural networks and the YOLO object detection algorithm for environmental recognition, and generate cognitively adapted voice feedback through a large language model. Combine reinforcement learning and personalized data adjustment to achieve high system integration and collaborative operation.

Benefits of technology

It improves the accuracy and real-time performance of environmental perception, generates language feedback that matches the cognitive characteristics of adolescents on the autism spectrum, realizes intelligent linkage between safe travel and effective communication, and enhances user experience and system robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121807167A_ABST
    Figure CN121807167A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent AR glasses life auxiliary system for autism spectrum teenagers, and the system comprises an intelligent AR glasses terminal, the intelligent AR glasses terminal comprises an image collection unit, an audio collection unit, a display unit, an audio output unit and a processing unit, and the processing unit comprises an environment sensing and guiding module, an image processing module and an image processing module. The guidance module is configured to acquire an environment image and recognize traffic lights, vehicles and pedestrian elements in the environment image to generate augmented reality visual guidance information and synchronous voice guidance information for guidance; and the real-time dialogue communication module is configured to obtain voice input information of the user, input the voice input information into the large language model to generate voice feedback information conforming to autism spectrum teenager cognitive characteristics, and adjust the feedback speed and the feedback tone of the synchronous voice guide information according to the voice feedback information. The intelligent AR glasses life auxiliary system has the beneficial effects that the intelligent AR glasses life auxiliary system can realize high integration, can perform accurate environment perception in real time, and can output cognitive adaptability languages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of user-guided assistance systems, and more specifically, to a smart AR glasses life assistance system for adolescents on the autism spectrum. Background Technology

[0002] With the development of intervention and support technologies for autism spectrum disorder (ASD), a variety of products and solutions have emerged on the market aimed at improving the independent living abilities of ASD adolescents. However, these existing technologies still have significant technical bottlenecks and practical shortcomings when addressing the complex needs of ASD adolescents in the real world, which are intertwined with "travel safety" and "social communication".

[0003] In terms of environmental perception and safety guidance, existing solutions mostly rely on a single technological approach. For example, some wearable devices use GPS technology for geofence positioning or combine it with simple proximity sensors, but their perception granularity is coarse, unable to identify specific dynamic hazards such as vehicles running red lights or rapidly approaching pedestrians in real time, resulting in high false alarm and false negative rates and insufficient reliability in real traffic scenarios. Other studies have attempted to apply computer vision algorithms to mobile devices, but limited by the computing power of mobile devices, they often use simplified models, making it difficult to meet the stringent requirements of safety guidance in terms of recognition accuracy and real-time performance under complex lighting and occlusion conditions. Furthermore, their functions are limited and they do not integrate with communication assistance.

[0004] In terms of communication assistance, existing technologies mainly fall into two categories: one is the use of Preset Voice or Picture Exchange Communication Systems (PECS), whose content is pre-defined and limited, lacking flexibility and unable to handle open social conversations; the other is devices equipped with general-purpose intelligent voice assistants, such as those based on general-purpose language models trained on large-scale general corpora. However, sentences generated by general-purpose language models often contain complex clauses, abstract vocabulary, or implicit social contexts, which are seriously inconsistent with the cognitive characteristics of ASD teenagers who prefer literal, concrete, and structured information. Using such tools not only has limited assistance effects but may also exacerbate users' confusion and frustration due to information overload or misunderstanding.

[0005] A more prominent problem is that these functions are often fragmented in existing technologies. Users may need to switch between different devices or applications, or receive inappropriate calm voice prompts in dangerous situations. This fragmented and non-intelligent collaborative approach disrupts the continuity of the user experience and fails to provide timely, appropriate, and integrated support in real-world situations such as being approached by a stranger while crossing the street, where users need to simultaneously manage environmental threats and social information.

[0006] Therefore, there is an urgent need for a highly integrated life assistance system for adolescents on the autism spectrum that can accurately perceive the environment in real time and output cognitively adaptive language, in order to solve the shortcomings of existing technologies in terms of real-time performance, accuracy, adaptability and functional synergy, and truly meet the core assistance needs of adolescents on the autism spectrum in complex real-life situations. Summary of the Invention

[0007] The technical problem to be solved by this invention is how to realize a highly integrated intelligent AR glasses life assistance system that can perform real-time accurate environmental perception and output cognitively adaptive language. In order to overcome the defects of the above-mentioned prior art (or related technology), this invention provides an intelligent AR glasses life assistance system for adolescents on the autism spectrum.

[0008] This invention provides a smart AR glasses-based daily living assistance system for adolescents on the autism spectrum, comprising: A smart AR glasses terminal, comprising an image acquisition unit, an audio acquisition unit, a display unit, an audio output unit, and a processing unit, wherein the processing unit is connected to the image acquisition unit, the audio acquisition unit, the display unit, and the audio output unit, and includes: The environmental perception and guidance module is configured to acquire environmental images in real time through the image acquisition unit, identify traffic lights, vehicles, and pedestrians in the environmental images using a convolutional neural network and YOLO object detection algorithm, obtain recognition results, generate corresponding augmented reality visual guidance information based on the recognition results, and transmit it to the display unit for visual guidance; and / or Based on the recognition results, corresponding synchronous voice guidance information is generated and transmitted to the audio output unit for auditory guidance. The real-time dialogue communication module, connected to the environment perception and guidance module, is configured to acquire the user's voice input information through the audio acquisition unit and input the voice input information into a pre-trained large language model to generate voice feedback information that conforms to the cognitive characteristics of adolescents on the autism spectrum, and to adjust the feedback speed and tone of the synchronous voice guidance information according to the voice feedback information.

[0009] Compared with existing technologies, the intelligent AR glasses daily living assistance system for adolescents on the autism spectrum of this invention has the following advantages: This invention integrates environmental perception and communication assistance functions into a single wearable smart AR glasses terminal. It highly integrates image acquisition, audio acquisition, display, audio output, and processing units, providing non-invasive, real-time online life assistance support for adolescents on the autism spectrum. Utilizing convolutional neural networks and the YOLO object detection algorithm for environmental image recognition, it significantly improves the accuracy and real-time performance of perceiving key hazard elements such as traffic lights, vehicles, and pedestrians in complex dynamic environments, providing a technological foundation for safe travel. Furthermore, it integrates a large language model into the real-time dialogue communication module, enabling the system not only to respond to simple commands but also to understand and generate adaptive language that matches the specific cognitive patterns of adolescents on the autism spectrum, achieving a leap from machine broadcasting to adaptive communication. The environmental perception and guidance module and the real-time dialogue communication module work collaboratively to achieve intelligent linkage between contextual awareness and communication strategies, making life assistance more humane and effective.

[0010] In one possible implementation, the environment perception and guidance module is configured to execute a three-step guidance method of scene-element-hint, which includes: Step S1: Analyze the environmental image using a continuous image frame analysis algorithm and SLAM technology to determine whether the user's current scenario is a preset typical travel scenario. If so, proceed to step S2; If not, then exit; Step S2: Call the YOLO object detection algorithm to identify traffic lights, vehicles and pedestrians in the image frame corresponding to the user's scene and obtain the corresponding identification results; Step S3: Based on the recognition result, visual guidance is provided to the user by overlaying a warning box, directional arrow, or symbol on the display unit; and / or The audio output unit broadcasts concise, instructive voice commands for auditory guidance.

[0011] Compared with existing technologies, the above technical solution can simulate the thought process of engineers making safety judgments through a three-step guidance method of scenario-element-prompt, making the system's decision-making process clear and controllable. This reduces the risk of system misjudgment and ensures that the generation of augmented reality visual guidance information and synchronous voice guidance information is based on an accurate understanding of the current user's scenario, thereby improving the accuracy of visual and auditory guidance and user trust.

[0012] In one possible implementation, the training process of the large language model in the real-time dialogue communication module includes: Step A1: Construct a fine-tuned dataset consisting of real dialogue records of adolescents on the autism spectrum, simplified sentence patterns annotated by therapists, and emotional tone samples. Step A2: Train the large language model based on the fine-tuned dataset and inject predefined rules, which include at least splitting complex sentences into simple sentences, prioritizing the use of concrete nouns and action verbs, and prohibiting the use of metaphors and ironic words.

[0013] Compared with existing technologies, the above-mentioned technical solution can ensure the high adaptability and security of the system's output language feedback information. By fine-tuning through the use of real dialogue records and simplified sentence patterns annotated by therapists, the large language model can learn the real language patterns and interaction difficulties of adolescents on the autism spectrum. The injected rules of breaking down complex compound sentences into simple sentences, prioritizing the use of specific vocabulary, and prohibiting the use of metaphors and ironic words directly target the common obstacles in language comprehension of adolescents on the autism spectrum, thereby greatly improving the comprehensibility and effectiveness of the generated voice feedback information and avoiding secondary confusion or anxiety caused by poor communication.

[0014] In one possible implementation, the real-time dialogue communication module is further provided with a local lightweight voice interaction engine, which is configured to provide reassuring and emergency prompts through the audio output unit when the network connection is poor, based on preset local keywords and fixed phrases.

[0015] Compared with existing technologies, the above technical solution can significantly improve the robustness and reliability of the system. Poor network connectivity is a common challenge for mobile wearable devices. By setting up a local lightweight language interaction engine, the system can still provide the most critical synchronized voice guidance information, including reassurance and emergency prompts, when the network is interrupted, thus preventing users from panicking in dangerous scenarios.

[0016] In one possible implementation, the processing unit further includes a personalization adaptation and feedback module, connected to the environment perception and guidance module, configured to collect personalized data of the user through a collaborative management platform that interacts with the therapist or the user's parents, and adjust the generation strategy of the environment perception and guidance module for the augmented reality visual guidance information and / or the synchronous voice guidance information by combining the user's interaction data during use with a reinforcement learning algorithm.

[0017] Compared with existing technologies, the above technical solution can collect users' personalized data and use reinforcement learning algorithms to adjust the generation strategy of the environment perception and guidance module, so that the system can continuously track and adapt to each user's unique reaction patterns, learning progress and emotional fluctuations.

[0018] In one possible implementation, the operation flow of the personalized adaptation and feedback module includes: Step B1: Input the user's sensory sensitivity parameters, interest topics, known fear sources, and behavioral baseline data into the collaborative management platform as the personalized data; Step B2: Record the user's behavioral responses and emotional indicators to each visual guidance, auditory guidance, or voice feedback, and combine them with the personalized data to generate state-feedback-result data pairs; Step B3: The strategy optimization module based on the Proximal Policy Optimization algorithm adjusts the generation strategy of the environment perception and guidance module according to the state-feedback-result data.

[0019] Compared with existing technologies, the above technical solution can form state-feedback-outcome data pairs by recording behavioral responses and emotional indicators, providing high-quality and quantifiable training samples for reinforcement learning algorithms. This makes the optimization of the generation strategy of the environment perception and guidance module no longer a black box, but an iteration based on specific interaction evidence, which greatly improves the accuracy and efficiency of personalized adaptation.

[0020] In one possible implementation, the collaborative management platform is configured to allow therapists or parents to view system logs and key video clips, manually annotate or correct the personalized data, and adjust the weight of the sensory sensitivity parameters in the personalized data.

[0021] Compared with existing technologies, the above-mentioned technical solution can seamlessly integrate the professional knowledge of rehabilitation therapists and the daily observations of users' parents into the system. This not only enhances the authority and accuracy of life assistance functions, but also reduces the long-term burden on caregivers.

[0022] In one possible implementation, the smart AR glasses life assistance system further includes a cloud server connected to the processing unit, configured to store the environmental image, the augmented reality visual guidance information, the synchronized voice guidance information, the voice input information, and the voice feedback information.

[0023] Compared with existing technologies, the above technical solution can separate computing, storage and analysis, ensuring the system's powerful performance and scalability. By storing massive amounts of environmental images and interactive data on cloud servers, it relieves the computing pressure on smart AR glasses terminals, making them lighter and with longer battery life.

[0024] In one possible implementation, the image acquisition unit is a camera, the audio acquisition unit is a microphone, the display unit is an embedded display screen, the audio output unit is a bone conduction headset worn on the user's ear, and the processing unit is an embedded processor.

[0025] Compared with existing technologies, the above-mentioned technical solution can provide clear auditory guidance while keeping the user's ears open to external ambient sounds by using bone conduction headphones, thus avoiding the isolation and safety hazards that may be caused by traditional in-ear headphones. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the structure of the present invention; Figure 2 This is a flowchart illustrating the three-step guidance method of the present invention, which includes scenarios, elements, and prompts. Figure 3 This is a flowchart illustrating the steps of the large language model training process of the present invention; Figure 4 This is a flowchart illustrating the operation process of the personalized adaptation and feedback module of the present invention. Detailed Implementation

[0027] First, those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention. Those skilled in the art can make adjustments as needed to adapt to specific application scenarios.

[0028] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0029] See Figure 1This invention discloses a smart AR glasses life assistance system for adolescents with autism spectrum disorder (ASD). It comprehensively utilizes augmented reality (AR), artificial intelligence (AI), edge computing, and cloud computing technologies to provide a real-time, personalized, and multimodal life assistance solution for ASD adolescents in two core life scenarios: travel and communication, thereby improving their independent living abilities and social integration. The system physically comprises three main parts: a smart AR glasses terminal, a cloud server, and a collaborative management platform for therapists and parents. The smart AR glasses terminal, as the user's directly worn interactive device, is the system's front-end perception and execution unit. Its core hardware includes an image acquisition unit, a processing unit, a display unit, an audio acquisition unit, and an audio output unit. The image acquisition unit uses a 12-megapixel front-facing camera with high dynamic range (HDR) and a frame rate of at least 20fps to capture environmental images within the user's field of view in real time, providing high-quality visual data for 3D environment modeling and target recognition. The processing unit uses an embedded high-performance processor, such as an NVIDIA Jetson series module, integrating a GPU and a dedicated AI acceleration core (such as Tensor). The Core component handles real-time image processing, lightweight AI inference, and task scheduling for each module via the cloud (Alibaba Cloud). The display unit combines high-transmittance waveguide optical lenses with a micro-display, allowing the overlay of virtual augmented reality visual guidance information, such as warning icons and directional arrows, into the user's natural field of vision, achieving a visual enhancement effect. The audio acquisition unit includes a beamforming microphone array for clearly picking up the user's voice input. The audio output unit uses bone conduction headphones to output voice guidance and feedback, ensuring that while receiving synchronized voice guidance information, the user's ears can still maintain an open perception of external environmental sounds, which is crucial for travel safety.

[0030] In this embodiment of the invention, three core software modules run on the processing unit: an environmental perception and guidance module, a real-time dialogue communication module, and a personalized adaptation and feedback module. These three modules work closely together via a shared bus and data interface. The environmental perception and guidance module is configured to acquire environmental images in real time through an image acquisition unit, and use convolutional neural networks and YOLO object detection algorithms to identify traffic lights, vehicles, and pedestrians in the environmental images to obtain recognition results. Based on the recognition results, it generates corresponding augmented reality visual guidance information and transmits it to the display unit for visual guidance; and / or generates corresponding synchronous voice guidance information based on the recognition results and transmits it to the audio output unit. The unit provides auditory guidance; the real-time dialogue communication module is configured to acquire the user's voice input information through the audio acquisition unit and input the voice input information into a pre-trained large language model to generate voice feedback information that conforms to the cognitive characteristics of adolescents on the autism spectrum, and adjust the feedback speed and tone of the synchronous voice guidance information based on the voice feedback information; the personalized adaptation and feedback module is configured to collect the user's personalized data through a platform that interacts with therapists or parents, and combine the user's interaction data during use to adjust the generation strategy of the environmental perception and guidance module for augmented reality visual guidance information and / or synchronous voice guidance information through reinforcement learning algorithms.

[0031] In this embodiment of the invention, the core problem faced by the environmental perception and guidance module is the "risk identification - rule understanding - action execution" problem under sudden situations. Autistic spectrum adolescents often exhibit dysfunctional behaviors such as "getting stuck, stagnating, repeating behaviors, and escaping" when encountering sudden changes during real-world travel (such as the end of a traffic light countdown, a vehicle suddenly approaching or honking, crowded platforms, pushing and shoving, temporary transfers, inability to understand broadcast information, or interference from strangers). The core reason is the inability to extract risk clues, match situational rules, and generate executable action sequences in a short period of time, leading to a significant increase in safety risks. The outstanding technical innovation is the adoption of a lightweight "risk-rule-action" reasoning and generation mechanism for sudden situations, constructing a situational map / rule template (crossing the street, waiting for a bus, entering a station, transferring, avoiding, etc.) and risk triggering conditions for travel tasks. Based on lightweight VLA (Vision-Language-Action), it realizes a strategy of extracting key events from real-time visual scenes → judging risk levels → matching rules → generating step-by-step action instructions (e.g., stop - look left - wait for the green light - then go).

[0032] In this embodiment of the invention, the core problem faced by the real-time dialogue communication module is the "intention expression—semantic understanding—dialogue progression" problem in communication scenarios. Adolescents on the autism spectrum often exhibit problems such as incomplete language expression, difficulty in pragmatic comprehension, and unstable dialogue rhythm in daily communication. Especially in real social situations (such as communicating with drivers / ticket sellers, asking and answering questions in the classroom, and expressing needs and emotions at home), they are prone to dysfunctional behaviors such as "unable to speak clearly, unable to understand, giving irrelevant answers, remaining silent, repeating questions, or escalating emotions." The core reason is that common conversations in life often output complex sentence structures, implicit semantics, and contextual jumps, which are difficult to match with the communication patterns of adolescents on the autism spectrum who require "simple, clear, predictable, and repeatable confirmation," leading to low interaction efficiency, strong frustration, and increased social avoidance. The outstanding technological innovation is the "large model modulation + structured dialogue" generation method tailored to the language characteristics of adolescents on the autism spectrum. The system employs a mechanism based on a general-purpose large language model combined with efficient parameter fine-tuning methods such as LoRA and Prefix-Tuning to construct a dedicated conversational ability adapted to the language habits of adolescents on the autism spectrum. A dedicated corpus is built around high-frequency conversational data from scenarios such as transportation, home, and school. Three key features are explicitly constrained in the training objectives: firstly, grammatical simplification (controlling sentence length, avoiding complex clauses and metaphorical expressions, and outputting short sentences and fixed sentence structures); secondly, semantic explicitness (highlighting imperative and step-by-step expressions, providing direct verbal and action suggestions such as "You can say this / do this"); and thirdly, controllable pacing (supporting 1-2 second delayed responses, key sentence repetition, rhetorical confirmation, and multi-round paraphrasing). This forms a closed-loop interactive process of "input understanding—intent completion—template generation—repeated confirmation—dialogue termination," enabling the system to provide stable, predictable, and executable real-time communication support even under social pressure.

[0033] See Figure 2 In this embodiment of the invention, the environment perception and guidance module is configured to execute a three-step guidance method of scene-element-prompt, which includes: Step S1: Analyze the environmental image using a continuous image frame analysis algorithm and SLAM technology to determine whether the user's current scenario is a preset typical travel scenario. If so, proceed to step S2; If not, then exit; Step S2: Call the YOLO object detection algorithm to identify traffic lights, vehicles and pedestrians in the image frame corresponding to the user's scene and obtain the corresponding recognition results; Step S3: Based on the recognition results, visual guidance is provided to the user by overlaying warning boxes, directional arrows, or symbol icons on the display unit; and / or Auditory guidance is provided by broadcasting concise, instructive voice commands through the audio output unit.

[0034] In this embodiment of the invention, after the system starts up in step S1, the image acquisition unit continuously acquires video streams, and the processing unit runs two tasks in parallel: one is to use vision-based SLAM technology to estimate the spatial pose of the smart AR glasses terminal in real time and construct a sparse environment map; the other is to use a lightweight convolutional neural network to analyze the video stream in real time and identify whether the current user's environment belongs to a preset typical travel scenario such as an intersection, bus stop, or subway station entrance. The spatial information provided by SLAM technology and the semantic information provided by the lightweight convolutional neural network are fused together as the decision basis for activating the next step of guidance.

[0035] In this embodiment of the invention, in step S2, when the user is in a high-attention scene, the system calls the YOLO target detection algorithm deployed on the edge processor. Due to its advantages in balancing speed and accuracy, the YOLO target detection algorithm is particularly suitable for real-time detection of dynamic targets such as traffic lights, vehicles, and pedestrians. At the same time, combined with a simple tracking algorithm, short-term trajectory prediction is performed on the detected dynamic targets to determine their relative motion relationship with the user, such as whether they are approaching and being fused into the recognition result.

[0036] In this embodiment of the invention, in step S3, based on the recognition result of step S2, the system generates and outputs augmented reality visual guidance information and synchronous voice guidance information. The visual guidance is implemented through a display unit. For example, when a red light is detected, a soft, pulsating red halo is superimposed on the actual traffic light location; when a rapidly approaching vehicle is detected from the side, a flashing orange warning box is superimposed on the outline of the vehicle. The visual guidance elements are designed to be quickly understood by each adolescent on the autism spectrum, avoiding overstimulation. The voice guidance is played synchronously through bone conduction headphones. The voice commands are designed as concise, clear, and instructive short sentences, such as "Red light, stop" or "Car on the left, wait." The urgency of the synchronous voice guidance information can be reflected by fine-tuning the speech speed and volume.

[0037] See Figure 3 In this embodiment of the invention, the training process of the large language model in the real-time dialogue communication module includes: Step A1: Construct a fine-tuned dataset consisting of real dialogue records of adolescents on the autism spectrum, simplified sentence patterns annotated by therapists, and emotional tone samples. Step A2: Train the large language model based on the fine-tuned dataset and inject predefined rules. The rules include at least splitting complex sentences into simple sentences, prioritizing the use of concrete nouns and action verbs, and prohibiting the use of metaphors and ironic words.

[0038] In this embodiment of the invention, the core of the real-time dialogue communication module lies in using a large language model to provide cognitively adapted language understanding and generation support for ASD adolescents. The Deepseek large language model is used as the basis. In step A1, real records of ASD adolescents after desensitization are collected and labeled by rehabilitation therapists to form a fine-tuned dataset consisting of simplified sentence patterns, clear semantics, and emotional tone samples. The large language model is then fine-tuned in a supervised manner using this fine-tuned dataset. At the same time, preset rules are injected through prompting engineering. The core rules include: grammatical simplification (breaking down complex sentences), semantic clarification (using specific nouns and action verbs), and disabling metaphors and irony. Finally, an ASD-adaptive large language model is obtained.

[0039] In this embodiment of the invention, after the user's voice input information is converted into text, it is sent to the aforementioned ASD-adaptive large language model. The large language model combines the current brief environmental context, such as "in front of the supermarket shelf," to generate a text response that conforms to the cognitive characteristics of ASD adolescents. Subsequently, through the speech synthesis module integrated into the smart AR glasses terminal or cloud server, the text response is converted into voice feedback information with a gentle tone and moderate speaking speed, which is output through bone conduction headphones. For example, when the user whispers "milk...", the system may reply, "You want to find milk. It's in the refrigerated section. Please look in the direction of the green arrow."

[0040] See Figure 4 In this embodiment of the invention, the operation flow of the personalized adaptation and feedback module includes: Step B1: Input the user's sensory sensitivity parameters, interest topics, known fear sources, and behavioral baseline data into the collaborative management platform as personalized data; Step B2: Record the user's behavioral responses and emotional indicators to each visual, auditory, or voice feedback, and combine them with personalized data to generate state-feedback-outcome data pairs. Step B3: Based on the state-feedback-outcome data, adjust the generation strategy of the environment perception and guidance module using a reinforcement learning algorithm.

[0041] In this embodiment of the invention, the system collects personalized user data from therapists and parents through a collaborative management platform. This data includes sensory sensitivity parameters such as reactions to bright light or specific sounds, interests, and known fears, forming an initial user profile. After each interaction with the user, the system records state (environment and user input) - feedback (system output) - result (user behavior response) data pairs. This data is encrypted and uploaded to a cloud server. The cloud server uses a reinforcement learning algorithm to automatically optimize the decision-making strategies of the environmental perception and guidance module and the real-time dialogue communication module, with the goal of improving user task completion rate, safety, and emotional stability. For example, the system may learn that visual flashing cues are more likely to attract a user's attention than voice cues, and thus prioritize visual guidance in similar scenarios in the future.

[0042] In this embodiment of the invention, rehabilitation therapists and parents can view system logs and key video clips of key events after desensitization through the collaborative management platform. They can also manually annotate the validity of personalized data from a particular system feedback. This manual feedback serves as a high-value reward signal for reinforcement learning, directly participating in model iteration and forming a collaborative optimization closed loop of "AI assistance + expert guidance".

[0043] In this embodiment of the invention, the smart AR glasses terminal can also capture the visual scene seen by adolescents on the autism spectrum in real time through its built-in camera, and synchronize the video to the APP software on the mobile phones of parents and special education teachers. Parents and special education teachers can observe the scene seen by the child in real time on the APP software. This function is to reduce the anxiety of parents / special education teachers about uncontrollable risks and increase the possibility of letting children live and travel independently with peace of mind. In addition, the APP software can play back the video, allowing parents and special education teachers to review and analyze changes in the child's behavior to correct the child's behavior. The smart AR glasses terminal also has an emergency assistance remote access function. When the child encounters a situation that cannot be handled, he / she can call the parents or special education teachers for remote guidance with one click. The call can be made by using a specific wake word, such as "Little steps, little steps, call Mom". The smart AR glasses terminal also has a positioning function, which allows parents and special education teachers to see the location of the child wearing the smart AR glasses terminal on the APP software.

[0044] In the description of this invention, the references to "one embodiment," "some embodiments," "in this embodiment," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0045] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A smart AR glasses-based daily living assistance system for adolescents on the autism spectrum, characterized in that, include: A smart AR glasses terminal, comprising an image acquisition unit, an audio acquisition unit, a display unit, an audio output unit, and a processing unit, wherein the processing unit is connected to the image acquisition unit, the audio acquisition unit, the display unit, and the audio output unit, and includes: The environmental perception and guidance module is configured to acquire environmental images in real time through the image acquisition unit, identify traffic lights, vehicles, and pedestrians in the environmental images to obtain recognition results, and generate corresponding augmented reality visual guidance information based on the recognition results and transmit it to the display unit for visual guidance; and / or Based on the recognition results, corresponding synchronous voice guidance information is generated and transmitted to the audio output unit for auditory guidance. The real-time dialogue communication module, connected to the environment perception and guidance module, is configured to acquire the user's voice input information through the audio acquisition unit and input the voice input information into a pre-trained large language model to generate voice feedback information that conforms to the cognitive characteristics of adolescents on the autism spectrum, and to adjust the feedback speed and tone of the synchronous voice guidance information according to the voice feedback information.

2. The intelligent AR glasses life assistance system according to claim 1, characterized in that, The environment perception and guidance module is configured to execute a three-step guidance method of scene-element-prompt, which includes: Step S1: Analyze the environmental image using a continuous image frame analysis algorithm and SLAM technology to determine whether the user's current scenario is a preset typical travel scenario. If so, proceed to step S2; If not, then exit; Step S2: Call the YOLO object detection algorithm to identify traffic lights, vehicles and pedestrians in the image frame corresponding to the user's scene and obtain the corresponding recognition results; Step S3: Based on the recognition result, visual guidance is provided to the user by overlaying a warning box, directional arrow, or symbol on the display unit; and / or The audio output unit broadcasts concise, instructive voice commands for auditory guidance.

3. The intelligent AR glasses life assistance system according to claim 1, characterized in that, The training process of the large language model in the real-time dialogue communication module includes: Step A1: Construct a fine-tuned dataset consisting of real dialogue records of adolescents on the autism spectrum, simplified sentence patterns annotated by therapists, and emotional tone samples. Step A2: Train the large language model based on the fine-tuned dataset and inject predefined rules, which include at least splitting complex sentences into simple sentences, prioritizing the use of concrete nouns and action verbs, and prohibiting the use of metaphors and ironic words.

4. The intelligent AR glasses life assistance system according to claim 1, characterized in that, The real-time dialogue communication module also includes a local lightweight voice interaction engine, which is configured to provide reassuring and emergency prompts through synchronized voice guidance information based on preset local keywords and fixed phrases when the network connection is poor, and play it through the audio output unit.

5. The intelligent AR glasses life assistance system according to claim 1, characterized in that, The processing unit also includes a personalized adaptation and feedback module, which is connected to the environment perception and guidance module. It is configured to collect personalized data of users through a collaborative management platform that interacts with therapists or parents, and adjust the generation strategy of the environment perception and guidance module for the augmented reality visual guidance information and / or the synchronous voice guidance information by combining the user's interaction data during use with reinforcement learning algorithms.

6. The intelligent AR glasses life assistance system according to claim 5, characterized in that, The operation process of the personalized adaptation and feedback module includes: Step B1: Input the user's sensory sensitivity parameters, interest topics, known fear sources, and behavioral baseline data into the collaborative management platform as the personalized data; Step B2: Record the user's behavioral responses and emotional indicators to each visual guidance, auditory guidance, or voice feedback, and combine them with the personalized data to generate state-feedback-result data pairs; Step B3: The strategy optimization module based on the Proximal Policy Optimization algorithm adjusts the generation strategy of the environment perception and guidance module according to the state-feedback-result data.

7. The intelligent AR glasses life assistance system according to claim 6, characterized in that, The collaborative management platform is configured to allow therapists or parents to view system logs and key video clips, manually annotate or correct the personalized data, and adjust the weight of the sensory sensitivity parameters in the personalized data.

8. The intelligent AR glasses life assistance system according to claim 1, characterized in that, The intelligent AR glasses life assistance system also includes a cloud server connected to the processing unit, configured to store the environmental image, the augmented reality visual guidance information, the synchronous voice guidance information, the voice input information, and the voice feedback information.

9. The intelligent AR glasses life assistance system according to claim 1, characterized in that, The image acquisition unit uses a camera, the audio acquisition unit uses a microphone, the display unit uses an embedded display screen, the audio output unit uses bone conduction headphones worn on the user's ear, and the processing unit uses an embedded processor.

Citation Information

Patent Citations

  • AI interaction system and method for multi-modal fusion holographic image

    CN120610629A

  • Method and system for training dialogue intervention large model for children with autism

    CN121301930A

  • Children autism adaptive brain-computer fusion intervention system based on large model

    CN121331383A

  • Information processing system

    CN121600721A