ASD early screening system and method
The ASD screening system, which combines multimodal data acquisition and edge computing, solves the problems of insufficient screening resources, low accuracy, and inadequate privacy protection in grassroots and remote areas. It achieves an efficient, accurate, and secure screening solution that meets the needs of grassroots and remote areas.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies for ASD screening in grassroots and remote areas suffer from problems such as high resource dependence, low accuracy, insufficient privacy protection, equipment incompatibility, and complex operation, making it difficult to meet the actual needs of grassroots and remote areas such as agricultural and pastoral areas in Tibet.
The study employs multimodal data acquisition (facial images and voice data) combined with multimodal models (image, emotion, and audio models) for screening, utilizes edge computing for offline inference, generates preliminary conclusions through a dynamic weighted voting algorithm, provides intervention suggestions in conjunction with the DB-GPT framework, and uses federated learning for model iteration and privacy protection.
It improves the accuracy and reliability of screening, reduces equipment costs and resource dependence, adapts to environments without or with weak networks, simplifies operating procedures, protects the privacy data of children, and provides reliable intervention recommendations.
Smart Images

Figure CN121662352A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of ASD screening technology, and more specifically, to an ASD early screening system and method. Background Technology
[0002] Current autism screening for children mainly relies on two types of technical solutions, but both have significant limitations and are difficult to meet the actual needs of grassroots and remote areas (such as rural and pastoral areas of Tibet): 1. Traditional manual screening methods Technical principle: Screening is conducted by professional doctors using scales, manual observation of children's behavior and language expression, etc., and relies on the doctor's professional experience and judgment.
[0003] Core flaw: High dependence on resources: There is a shortage of child psychiatrists in grassroots areas (such as Tibet) and a lack of professional testing equipment such as MRI and brain imaging. Most grassroots medical staff have no experience in ASD diagnosis and are unable to carry out standardized screening. Highly subjective and low in accuracy: The scales are mostly derived from Western countries, which presents cultural compatibility issues. Furthermore, manual observation is easily influenced by emotions and experience, resulting in high rates of missed diagnoses and misdiagnoses. The scales are also inefficient and require parents to provide long-term behavioral descriptions. Parents in rural areas generally have low cooperation rates due to factors such as education level and time costs.
[0004] 2. Existing intelligent screening technology solutions Existing technologies include intelligent screening systems based on single-modal or dual-modal data, but they suffer from the following key shortcomings: Incomplete modality coverage and high cloud dependency: Existing systems rely heavily on cloud computing for model inference. Limited network bandwidth and frequent network outages in remote areas at the grassroots level can lead to screening interruptions or delays, making real-time feedback impossible. Risk of model "illusion": Some systems rely on general large language models to generate diagnostic suggestions, but ASD clinical data is not open source, and model training data is mostly sourced from the Internet, which can easily generate fictitious and inconsistent intervention suggestions, failing to provide reliable clinical guidance. Insufficient privacy protection: Existing systems often upload raw data directly to the cloud without using local processing or federated learning technology, posing a risk of leakage of children's privacy data and failing to comply with medical data security standards.
[0005] 3. Poor adaptability to grassroots scenarios Existing smart devices are mostly fixed (such as hospital-specific testing terminals), which are large, expensive, and require external power supplies, making them unsuitable for mobile screening needs in remote areas such as Tibet. At the same time, the penetration rate of information technology terminals in primary healthcare institutions is low, and existing systems are complex to operate, making it difficult for primary healthcare workers to quickly get started. Summary of the Invention
[0006] The embodiments of this application provide an early screening system and method for ASD to solve the technical problems existing in the prior art.
[0007] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0008] According to a first aspect of the embodiments of this application, an early screening system for ASD is provided, comprising: The data acquisition module collects multimodal data, including: children's facial image data and children's voice data; The inference module is deployed, a multimodal model is deployed, features of the multimodal data are extracted as input to the multimodal model, a dynamic weighted voting algorithm is used for inference of each model, and preliminary screening conclusions are generated based on the output confidence of each model. The multimodal model includes an image model, an emotion model and an audio model. The model iteration module acquires multimodal data in real time, classifies and stores it, and updates and iterates the multimodal model based on the real-time data.
[0009] In some embodiments of this application, based on the foregoing scheme, the acquisition module includes: The image acquisition submodule is used to acquire facial image data of children using a facial tracking algorithm. The audio acquisition submodule is used to collect children's voice data.
[0010] In some embodiments of this application, based on the foregoing scheme, the deployment inference module includes: The image model inference submodule is used for inference based on children's facial image data; The emotion model reasoning submodule is used for reasoning based on children's facial image data; The audio model inference submodule is used for inference based on children's voice data; Edge computing devices are used to deploy image model inference submodules, emotion model inference submodules, and audio model inference submodules, and to display the inference results of each model in real time.
[0011] In some embodiments of this application, based on the foregoing scheme, the model iteration module includes: The IoT management platform is used to receive anonymized multimodal data collected in real time and classify and store the multimodal data according to image features and audio indicators. The iterative update module is used to update the multimodal model in real time based on the classified and stored multimodal data.
[0012] In some embodiments of this application, based on the foregoing scheme, the following further methods are also included: The association module is used to generate targeted intervention suggestions by associating authoritative medical literature and historical case data based on the DB-GPT framework.
[0013] In some embodiments of this application, based on the foregoing scheme, the following further methods are also included: The management module is used to manage the start / stop of multimodal data collection, view screening results, and query / create cases; The display module provides access to regional screening data statistics, equipment status monitoring, RAG regional data, and autism management recommendations.
[0014] According to a second aspect of the embodiments of this application, an early screening method for ASD is provided, comprising: Collect multimodal data, including: children's facial image data and children's voice data; Deploy multimodal models, including: image models, emotion models, and audio models; Features of multimodal data are extracted as input to a multimodal model. A dynamic weighted voting algorithm is used to generate preliminary screening conclusions based on the output confidence of each model. Multimodal data is acquired in real time, classified and stored, and the multimodal model is updated and iterated based on the stored multimodal data.
[0015] The technical solution of this application improves the screening accuracy by using multimodal data fusion to overcome the problem of single-dimensionality in existing systems; it solves the screening interruption problem caused by poor network at the grassroots level by using edge computing and cloud collaboration, and realizes offline inference and real-time feedback; it avoids the risk of large model "illusion" by using RAG technology, while protecting the privacy data of children and providing reliable intervention suggestions.
[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings: Figure 1 A structural block diagram of an early screening system for ASD according to one embodiment of this application is shown. Detailed Implementation
[0018] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.
[0019] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0020] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0021] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0022] It should be noted that "multiple" in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such uses of these terms can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described.
[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0025] The following detailed description of some embodiments of this application will be provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0026] To address the technical problems in the prior art, embodiments of this application provide an early screening system for ASD, such as... Figure 1 As shown, it includes: The data acquisition module collects multimodal data, including: children's facial image data and children's voice data; The inference module is deployed, a multimodal model is deployed, features of the multimodal data are extracted as input to the multimodal model, a dynamic weighted voting algorithm is used for inference of each model, and preliminary screening conclusions are generated based on the output confidence of each model. The multimodal model includes an image model, an emotion model and an audio model. The model iteration module acquires multimodal data in real time, classifies and stores it, and updates and iterates the multimodal model based on the real-time data.
[0027] In some feasible embodiments, based on the foregoing scheme, the acquisition module includes: The image acquisition submodule is used to acquire facial image data of children using a facial tracking algorithm. The audio acquisition submodule is used to collect children's voice data.
[0028] For example, in this example, the acquisition module is located in the perception layer of the system, wherein: The image acquisition submodule uses a high-definition camera to capture facial expression image data.
[0029] The audio acquisition submodule uses an omnidirectional microphone to collect children's speech data (such as pronunciation frequency, vowel-consonant ratio, and intonation changes).
[0030] The HD camera and omnidirectional microphone are both connected to the OrangePie Kunpeng Pro edge computing device via USB.
[0031] In some feasible embodiments, based on the foregoing scheme, the deployment inference module includes: The image model inference submodule is used for inference based on children's facial image data; The emotion model reasoning submodule is used for reasoning based on children's facial image data; The audio model inference submodule is used for inference based on children's voice data; Edge computing devices are used to deploy image model inference submodules, emotion model inference submodules, and audio model inference submodules, and to display the inference results of each model in real time.
[0032] For example, in this example, the edge computing device is the Orange Pi Kunpeng Pro edge computing device; on it, image models, emotion models, and audio models are deployed.
[0033] In this example, the image model, sentiment model, and audio model are used in a multimodal fusion inference approach, namely, a "dynamic weighted voting algorithm". Based on the confidence scores of each model, a preliminary screening conclusion (high risk / medium risk / low risk) is generated.
[0034] Among them, the edge-end interactive interface developed based on PyQt5 can display screening results in real time and support offline viewing of historical cases.
[0035] In some feasible embodiments, based on the foregoing scheme, the model iteration module includes: The IoT management platform is used to receive anonymized multimodal data collected in real time and classify and store the multimodal data according to image features and audio indicators. The iterative update module is used to update the multimodal model in real time based on the classified and stored multimodal data.
[0036] It should be noted that in this embodiment, the multimodal data is classified and stored in an ECS service + MySQL spatiotemporal database. This database stores all the screened data and supports regional data statistics.
[0037] In some feasible embodiments, based on the foregoing scheme, the following further applies: The association module is used to generate targeted intervention suggestions by associating authoritative medical literature and historical case data based on the DB-GPT framework.
[0038] Understandably, the associated modules can be set up to bypass the model's visual interface.
[0039] In some feasible embodiments, based on the foregoing scheme, the following further applies: The management module is used to manage the start / stop of multimodal data collection, view screening results, and query / create cases; The display module provides access to regional screening data statistics, equipment status monitoring, RAG regional data, and autism management recommendations.
[0040] It is understandable that the management module simplifies the operation process and is suitable for grassroots medical staff.
[0041] For example, in this example, the management module is developed based on PyQt5; the presentation module is developed based on the Flask framework.
[0042] The following is an example of how this system runs.
[0043] 1. Offline screening process Primary healthcare workers input the child's basic information through the edge terminal interface and click "Start Collection"; The perception layer device starts up: the camera and microphone simultaneously acquire image and audio data; Data preprocessing: Edge devices perform noise reduction, keyframe extraction, and initial feature extraction on the collected data; Offline inference: Three specialized models are called for parallel processing, and screening results are generated through a dynamic weighted voting algorithm; Results feedback: The edge interface displays the screening results and annotations of abnormal features in each modality, and supports printing reports.
[0044] 2. Cloud synchronization and analysis process (in network scenarios) Edge devices automatically connect to the network and upload the de-identified multimodal data and screening results to the Huawei Cloud IoT platform; The cloud server receives data and writes it to the MySQL database. The RAG system then uses the knowledge base to generate regional data and autism management recommendations. Grassroots organizations can view regional data statistics, equipment status, and intervention suggestions through the web interface; Federated learning iteration: Multiple edge devices upload model parameters, the server aggregates and updates the global model, and distributes it to the edge devices to complete the upgrade.
[0045] Based on the same inventive concept, embodiments of this application also provide an early screening method for ASD, including: Collect multimodal data, including: children's facial image data and children's voice data; Deploy multimodal models, including: image models, emotion models, and audio models; Features of multimodal data are extracted as input to a multimodal model. A dynamic weighted voting algorithm is used to generate preliminary screening conclusions based on the output confidence of each model. Multimodal data is acquired in real time, classified and stored, and the multimodal model is updated and iterated based on the stored multimodal data.
[0046] In summary, this technical solution has the following advantages: 1. Reduce reliance on grassroots resources and improve accessibility for screening. No need for a professional ASD doctor: The system replaces manual judgment with multimodal objective data and model reasoning, and primary care medical staff can operate it after 1 hour of training; Adapted for scenarios with no or weak network: The edge layer offline inference function solves the problem of poor network in regions such as Tibet, and improves the screening rate; Low cost and portable: The entire set of equipment costs 4363 yuan, weighs ≤5kg, supports lithium battery power, and can be used for on-site screening.
[0047] 2. Improve the accuracy and reliability of screening. Multimodal data coverage: Compared to emotion, image, and audio multimodal diagnosis, it improves the accuracy of comprehensive screening; Avoiding model illusion: RAG technology links clinical literature with real cases, with no fictional content; Dynamic model iteration: Federated learning enables the model to adapt to the characteristics of children in different regions, improving the accuracy of regional screening.
[0048] 3. Protecting privacy and improving management efficiency Data security: Edge devices first perform local inference and data anonymization, then upload only model parameters instead of raw data to the cloud, and finally use a federated learning mechanism to allow models from different locations to optimize together without leaking local patient information.
[0049] Regionalized management: The web platform allows grassroots organizations to view regional screening trends, providing data support for public health decision-making; Operational efficiency: The time for a single screening has been reduced from 30 minutes in the traditional manual process to 5 minutes.
[0050] 4. Social Value In response to the national strategic need to "improve the mental health service system" and to address the "last mile" problem of providing mental health services for children at the grassroots level, especially to provide feasible early screening programs for ASD in remote areas such as Tibet, to help with early intervention and reduce the burden on families and society.
[0051] Other embodiments of this application will readily conceive of by those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. It should be understood that this application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. An early screening system for ASD, characterized in that, include: The data acquisition module collects multimodal data, including: children's facial image data and children's voice data; The inference module is deployed, a multimodal model is deployed, features of the multimodal data are extracted as input to the multimodal model, a dynamic weighted voting algorithm is used for inference of each model, and preliminary screening conclusions are generated based on the output confidence of each model. The multimodal model includes an image model, an emotion model and an audio model. The model iteration module acquires multimodal data in real time, classifies and stores it, and updates and iterates the multimodal model based on the real-time data.
2. The system according to claim 1, characterized in that, The acquisition module includes: The image acquisition submodule is used to acquire facial image data of children using a facial tracking algorithm. The audio acquisition submodule is used to collect children's voice data.
3. The system according to claim 1, characterized in that, The deployment inference module includes: The image model inference submodule is used for inference based on children's facial image data; The emotion model reasoning submodule is used for reasoning based on children's facial image data; The audio model inference submodule is used for inference based on children's voice data; Edge computing devices are used to deploy image model inference submodules, emotion model inference submodules, and audio model inference submodules, and to display the inference results of each model in real time.
4. The system according to claim 1, characterized in that, The model iteration module includes: The IoT management platform is used to receive anonymized multimodal data collected in real time and classify and store the multimodal data according to image features and audio indicators. The iterative update module is used to update the multimodal model in real time based on the classified and stored multimodal data.
5. The system according to claim 1, characterized in that, Also includes: The association module is used to generate targeted intervention suggestions by associating authoritative medical literature and historical case data based on the DB-GPT framework.
6. The system according to claim 1, characterized in that, Also includes: The management module is used to manage the start / stop of multimodal data collection, view screening results, and query / create cases; The display module provides access to regional screening data statistics, equipment status monitoring, RAG regional data, and autism management recommendations.
7. An early screening method for ASD, characterized in that, include: Collect multimodal data, including: children's facial image data and children's voice data; Deploy multimodal models, including: image models, emotion models, and audio models; Features of multimodal data are extracted as input to a multimodal model. A dynamic weighted voting algorithm is used to generate preliminary screening conclusions based on the output confidence of each model. Multimodal data is acquired in real time, classified and stored, and the multimodal model is updated and iterated based on the stored multimodal data.