A multi-modal data fusion vision detection and management method, device and medium

By using multimodal data fusion and artificial intelligence technology, the problem of insufficient accuracy in vision testing has been solved, enabling accurate risk prediction and personalized intervention plans for vision testing, and improving the accuracy of testing and the synergy of management processes.

CN122494094APending Publication Date: 2026-07-31QINGDAO PENGFENGCHENG MEDICAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO PENGFENGCHENG MEDICAL TECHNOLOGY CO LTD
Filing Date
2026-03-17
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Current vision testing technologies rely on eye charts and single optical tests, lacking physiological data on the eyes such as eye movement trajectories and pupillary responses. This results in insufficient accuracy and makes it impossible to make accurate risk predictions and personalized interventions.

Method used

A multimodal data fusion method is used to acquire vision test data, ocular physiological data, and detection environment data. Feature extraction and fusion are performed through an attention mechanism deep learning model. The risk of refractive abnormalities is predicted by combining LSTM neural network and the risk factors are quantified by SHAP value analysis to generate personalized intervention plans.

Benefits of technology

It has improved the accuracy of vision testing, increased the accuracy of refractive error risk prediction, enhanced the personalization of intervention plans, facilitated cross-institutional management process collaboration, and broken down information silos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122494094A_ABST
    Figure CN122494094A_ABST
Patent Text Reader

Abstract

This application discloses a method, device, and medium for vision detection and management based on multimodal data fusion. The method includes: acquiring multimodal vision data of the examinee; preprocessing the data to generate standardized multimodal data, judging based on thresholds and triggering alarms; extracting and fusing features from the standardized data using an attention mechanism deep learning model to output vision parameters; inputting historical multimodal data into an LSTM model to determine the risk level of refractive errors; quantifying the feature contribution through SHAP value analysis to output core risk factors; generating an intervention plan based on vision parameters, risk level, and core risk factors using a knowledge graph engine; and pushing the plan to the doctor's and examinee's terminals through a multi-terminal interaction module. This application achieves accurate collection, intelligent analysis, trend prediction, and collaborative management of vision data, and can be applied to ophthalmological clinical practice, public health screening, and family health monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of healthcare informatics technology, and in particular to a method, device and medium for vision detection and management using multimodal data fusion. Background Technology

[0002] Vision testing and management is an important topic in the fields of healthcare and public health. Currently, clinical and civilian vision testing mostly relies on vision charts or single optical testing devices, and the technical system has significant shortcomings.

[0003] The data collection is limited to basic information such as visual acuity chart values, lacking supporting physiological data on eye movement trajectories and pupillary responses. Furthermore, it fails to incorporate environmental parameters for calibration, resulting in insufficient evidence for visual acuity testing and diagnosis. Simultaneously, data processing relies primarily on manual recording and experience-based judgment, making it impossible to accurately predict the risks and development trends of vision problems such as refractive errors.

[0004] While existing vision screening technologies have achieved some digital improvements, such as computer vision-based children's vision screening systems focusing on eye movement analysis and AR testing, and ophthalmology screening and treatment systems achieving basic data uploading and evaluation, none of them have completed the deep fusion of multimodal data for vision testing. This results in insufficient accuracy in vision testing and an inability to generate dynamic intervention plans tailored to individuals. Summary of the Invention

[0005] To address the technical problems existing in the background art, embodiments of this application provide a method, device, and medium for vision detection and management based on multimodal data fusion. The method includes: acquiring multimodal vision data of a subject; the multimodal vision data includes vision test data, ocular physiological data, and detection environment data; preprocessing the multimodal vision data to generate standardized multimodal data, and judging the standardized multimodal data based on a preset threshold; if data does not meet the preset threshold, triggering a terminal alarm; and using a preset attention mechanism deep learning model to extract and fuse features from the standardized multimodal data to generate a fused feature vector. The system outputs visual acuity parameters based on the fused feature vector and a preset activation function. It then inputs the subject's historical multimodal data as a time series into a pre-trained LSTM neural network model to determine the subject's refractive error risk level within a preset future timeframe. Finally, it quantifies the contribution of standardized multimodal data features to the refractive error risk level using SHAP value analysis, outputting core risk factors. An intervention plan is generated based on the visual acuity parameters, the refractive error risk level, and the core risk factors using a preset knowledge graph engine. Finally, the intervention plan is visualized and pushed to both the doctor's and the subject's terminals via a preset multi-terminal interaction module.

[0006] In one example, a pre-defined attention mechanism deep learning model is used to extract and fuse features from the standardized multimodal data, generating a fused feature vector. Based on the fused feature vector and a pre-defined activation function, visual acuity parameters are output. Specifically, this includes: extracting visual function features from the visual acuity test data in the standardized multimodal data; these visual function features include at least optotype recognition accuracy and average pupillary reaction time; and extracting ocular physiological features from the ocular physiological data in the standardized multimodal data; these ocular physiological features include at least eye movement stability, pupillary constriction rate, pupillary dilation rate, blink frequency variation coefficient, and other parameters. The axis length is normalized; combined with the light intensity and temperature and humidity parameters in the standardized multimodal data, environmental interference correction is performed on the extracted visual functional features and ocular physiological features; the corrected visual functional features and ocular physiological features are input into a preset attention mechanism deep learning model, and the weight coefficients corresponding to each feature parameter are assigned through the attention mechanism to perform feature weighting; the weighted features are mapped to a unified feature space to generate a fused feature vector; the fused feature vector is input into a fully connected layer, and the visual acuity parameters are output through a preset Softmax activation function; the visual acuity parameters include refractive power, astigmatism degree, and astigmatic axis.

[0007] In one example, the subject's historical multimodal data is used as a time series and input into a pre-trained LSTM neural network model to determine the subject's refractive error risk level within a preset future time period. Specifically, this includes: acquiring the subject's historical test records sorted by test timestamp within a preset historical time period; the historical test records include standardized multimodal data; inputting the historical test records as time series data into the pre-trained LSTM neural network model; capturing the temporal dependency of visual parameters changing over time in the time series data through the gating unit of the LSTM neural network model to learn the development pattern of refractive errors; comparing and analyzing the temporal dependency and the development pattern with baseline data of populations grouped by age and gender, respectively, and outputting the subject's refractive error risk level within the preset future time period.

[0008] In one example, the contribution of standardized multimodal data features to the refractive error risk level is quantified using the SHAP value analysis method, and core risk factors are output. Specifically, this includes: obtaining the input features used by the LSTM neural network model to determine the refractive error risk level; the input features include the rate of change of axial length, annual increase in refractive power, pupillary response sensitivity, and trend of change in eye movement stability; using the SHAP value analysis method to calculate the contribution value of each input feature to the refractive error risk level; and sorting all input features from high to low contribution values ​​to select input features whose cumulative contribution reaches a preset threshold as core risk factors.

[0009] In one example, an intervention plan is generated using a pre-defined knowledge graph engine based on the visual acuity parameters, the refractive error risk level, and the core risk factors. Specifically, this includes: constructing a knowledge graph; the knowledge graph includes ophthalmological clinical guidelines, a drug database, and an intervention plan library; it is a fourth-order logical association graph of symptoms, causes, risk factors, and intervention measures; based on the core risk factors, preliminary intervention measures with a matching relationship to the core risk factors are selected from the knowledge graph; the individual characteristics of the examinee are obtained; the individual characteristics include age, lifestyle habits, previous intervention history, and local medical resources; the preliminary intervention measures are adapted and optimized according to the individual characteristics to generate an intervention plan; the intervention plan includes at least one of eye habit correction, vision training, and orthodontic appliance adaptation; and a pre-defined reinforcement learning algorithm is used to collect feedback on the execution effect of historical intervention plans from individuals with the same refractive error risk level and individual characteristics, and subsequent intervention plans are adjusted based on the execution effect feedback to form a closed-loop optimization mechanism.

[0010] In one example, the standardized multimodal data is judged based on a preset threshold. If there is data that does not meet the preset threshold, a terminal alarm is triggered. Specifically, this includes: extracting the vision test data to determine the pupillary reaction time; extracting the ocular physiological data to determine the blink frequency; comparing the pupillary reaction time with a preset reaction time threshold, and comparing the blink frequency with a preset frequency threshold; if the pupillary reaction time is greater than the reaction time threshold, and / or the blink frequency is greater than the frequency threshold, it is determined that there is a physiological abnormality in the standardized multimodal data, and an alarm is triggered.

[0011] In one example, acquiring the subject's multimodal vision data specifically includes: presenting the subject with random sequence optotype images generated by a digital vision chart, recording the subject's recognition results and reaction time, and generating vision test data; collecting the subject's eye movement trajectory, pupil diameter changes, and blink frequency by integrating an infrared camera and a pupil detector, and generating ocular physiological data; recording the ambient temperature, humidity, and light intensity during the subject's test by a temperature and humidity sensor and a light sensor, and generating test environment data; obtaining the subject's identity information through facial recognition, and retrieving the corresponding electronic health record based on the identity information.

[0012] In one example, the multimodal vision data is preprocessed to generate standardized multimodal data. Specifically, this includes: cleaning the vision test data, ocular physiological data, and detection environment data using a preset outlier removal algorithm; and uniformly converting the cleaned vision test data, ocular physiological data, and detection environment data according to the HL7 FHIR standard format to generate standardized multimodal data.

[0013] On the other hand, embodiments of this application provide a vision detection and management device for multimodal data fusion, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform any of the above-mentioned vision detection and management methods for multimodal data fusion.

[0014] On the other hand, embodiments of this application provide a non-volatile computer storage medium for vision detection and management based on multimodal data fusion, which stores computer-executable instructions that can execute any of the aforementioned methods for vision detection and management based on multimodal data fusion.

[0015] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: This solution deeply integrates multimodal data with artificial intelligence technology, achieving a full-process intelligent upgrade of vision testing and health management. First, it provides comprehensive data collection dimensions, integrating vision test, ocular physiology, and environmental data, improving detection accuracy compared to traditional methods. Second, the deep learning model based on attention mechanisms outputs vision parameters with an accuracy rate of no less than 98%. Combined with LSTM and SHAP value analysis, it can predict future refractive error risks and quantify core risk factors, further improving prediction accuracy. Third, it generates personalized intervention plans using a three-dimensional mechanism of risk targeting, individual adaptation, and reinforcement learning optimization. Based on knowledge graph screening measures and combined with individual characteristic adaptation, it improves plan execution rate. Fourth, it enables cross-institutional collaboration in management processes, achieving seamless integration of the entire chain from testing to diagnosis, intervention, and supervision through HL7 FHIR standardization and multi-terminal interaction, breaking down information silos. Attached Figure Description

[0016] To more clearly illustrate the technical solution of this application, some embodiments of this application will be described in detail below with reference to the accompanying drawings, in which: Figure 1 A flowchart illustrating a vision detection and management method based on multimodal data fusion, provided as an embodiment of this application; Figure 2 A logic diagram of LSTM risk prediction and contribution analysis for a vision detection and management method based on multimodal data fusion provided in this application embodiment; Figure 3 A flowchart illustrating a vision detection and management method based on multimodal data fusion, provided in an embodiment of this application; Figure 4 A system architecture diagram of a vision detection and management method based on multimodal data fusion provided in this application embodiment; Figure 5This is a schematic diagram of the structure of a vision detection and management device based on multimodal data fusion, provided in an embodiment of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0018] Some embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0019] Figure 1 This is a flowchart illustrating a vision detection and management method based on multimodal data fusion, provided as an embodiment of this application. This method can be applied to various business domains. Certain input parameters or intermediate results in this process allow for manual intervention and adjustment to help improve accuracy.

[0020] The analysis method involved in the embodiments of this application can be implemented by a terminal device or a server, and this application does not impose any special limitations on it. For ease of understanding and description, the following embodiments are all described in detail using a server as an example.

[0021] First, it should be noted that the architecture of this solution has four layers, including the detection terminal layer, the edge computing layer, the cloud intelligence layer, and the multi-terminal interaction layer.

[0022] Based on this Figure 1 The process may include the following steps: S101: Acquire the subject's multimodal visual acuity data; the multimodal visual acuity data includes visual acuity test data, ocular physiological data, and testing environment data.

[0023] In some embodiments of this application, this step achieves real-time acquisition of multimodal data through various dedicated modules deployed on the detection terminal layer.

[0024] The testing terminal layer uses an industrial-grade touchscreen and integrates a vision testing module, a physiological sensing module, an environmental perception module, and an identity recognition module. Specifically, the vision testing module presents a random sequence of optotype images to the test subject through a dynamically adjustable digital vision chart and accurately records the test subject's recognition results and reaction time, thereby generating vision test data.

[0025] The physiological sensing module integrates a 20-megapixel infrared camera and a high-precision pupil detector (accuracy ±0.1mm). It continuously collects the subject's eye movement trajectory, pupil diameter changes, and blink frequency at a sampling frequency of no less than 30Hz, generating ocular physiological data including eye rotation stability, pupil contraction / dilation rate, and blink frequency variation coefficient.

[0026] The environmental sensing module has built-in temperature and humidity sensors and light sensors to record the ambient temperature, humidity and light intensity in real time during detection, generating detection environment data for subsequent data correction.

[0027] The identity recognition module obtains the examinee's identity information through facial recognition or NFC technology, and automatically retrieves or associates the corresponding electronic health record based on the identity information, achieving seamless connection of historical data.

[0028] After being collected, the aforementioned multimodal data is transmitted in real time to the edge computing layer via an encrypted communication protocol for further processing.

[0029] S102: Preprocess the multimodal vision data to generate standardized multimodal data, and judge the standardized multimodal data based on a preset threshold. If there is data that does not meet the preset threshold, trigger a terminal alarm.

[0030] In some embodiments of this application, this step is performed by an edge computing layer deployed locally on the detection terminal.

[0031] The edge computing unit uses an ARM Cortex-A72 processor with 4GB of memory and supports real-time data processing. First, the data cleaning module uses outlier removal algorithms (such as the 3σ criterion) to clean the ocular physiological data, automatically filtering out noise data caused by environmental interference or the subject's blinking, ensuring the accuracy of physiological parameters.

[0032] Furthermore, the format standardization module converts the cleaned vision test data, ocular physiological data, and testing environment data into the HL7 FHIR standard format, generating standardized multimodal data, thereby ensuring cross-system and cross-institutional compatibility and interoperability.

[0033] Meanwhile, the real-time feedback module performs linked analysis on the standardized data based on preset thresholds. Specifically, it extracts pupil reaction time and blink frequency from the standardized multimodal data, compares the pupil reaction time with a preset reaction time threshold (500ms), and compares the blink frequency with a preset frequency threshold (30 times / minute). If the pupil reaction time is greater than 500ms and / or the blink frequency is greater than 30 times / minute, it is determined that there is a significant physiological abnormality in the standardized multimodal data. An alarm is immediately triggered through the display interface or sound module of the detection terminal so that on-site personnel can intervene in a timely manner or guide the examinee to retest.

[0034] S103: Using a preset attention mechanism deep learning model, feature extraction and fusion are performed on the standardized multimodal data to generate a fused feature vector, and visual parameters are output based on the fused feature vector and a preset activation function.

[0035] In some embodiments of this application, this step is performed by the multimodal fusion analysis module of the cloud intelligence layer.

[0036] This module is deployed on a cloud server cluster (8-core 16GB cloud host, using the TensorFlow 2.15 deep learning framework). It performs deep analysis on the uploaded standardized multimodal data based on an attention mechanism deep learning model trained on 200,000 clinical samples.

[0037] Specifically, feature extraction is performed first. Eight key features, such as optotype recognition accuracy, average reaction time, and average sharpness feedback, are extracted from the visual acuity test data. Twelve key features, such as eyeball rotation stability, pupil constriction rate, pupil dilation rate, blink frequency variation coefficient, and normalized value of axial length, are extracted from the ocular physiological data.

[0038] Furthermore, by combining the light intensity and temperature / humidity parameters from the detection environment data, environmental interference correction is applied to the aforementioned 20 types of features (e.g., automatically adjusting the weight of optotype recognition accuracy when there is insufficient light) to eliminate the influence of environmental factors on the detection results. Next, the corrected features are input into an attention mechanism model, which automatically assigns weight coefficients to each feature parameter (e.g., the weights of core features such as axial length and pupillary reaction time can reach 0.25 and 0.20 respectively), achieving adaptive weighted fusion of features. The weighted features are then mapped to a unified feature space, generating a 512-dimensional fused feature vector. Finally, the fused feature vector is input into a fully connected layer, processed by the Softmax activation function, and outputs accurate visual acuity parameters, including refractive power, astigmatism degree, and astigmatic axis, with a model accuracy of no less than 98%.

[0039] The core of this fusion logic lies in the fact that visual acuity test data reflects the external manifestation of visual function, while ocular physiological data reveals the internal characteristics of eye structure and function (such as the direct correlation between axial length and refractive power, and pupillary response reflecting visual sensitivity). The fusion of the two achieves dual support of "external manifestation + internal mechanism", which significantly improves the accuracy of visual acuity parameter calculation.

[0040] S104: The subject's historical multimodal data is used as a time series and input into a pre-trained LSTM neural network model to determine the subject's refractive error risk level within a preset future time period.

[0041] In some embodiments of this application, this step is performed by the risk prediction module of the cloud-based intelligent layer.

[0042] This module uses a time-series prediction model based on LSTM (Long Short-Term Memory) neural network. The model was trained using 50,000 longitudinal data sets of subjects. Each data set contains more than 3 years of continuous test records, covering population samples of different ages, genders, and regions. Each record contains multimodal fusion features, clinical diagnosis results, and the trajectory of refractive abnormality development.

[0043] During prediction, the system first obtains the subject's historical test records sorted by test timestamp within the past 1 to 3 years. These historical records are then input as time series data into a pre-trained LSTM neural network model. The LSTM model automatically captures the temporal dependencies of visual parameters over time through its gating units, learns the development patterns of refractive errors, and compares the learned temporal dependencies and development patterns with baseline data of populations grouped by age and gender. Finally, it outputs the subject's refractive error risk level (divided into three levels: low risk, medium risk, or high risk) for the next 12 months, thus achieving a leap from static assessment to dynamic prediction.

[0044] S105: Quantify the contribution of standardized multimodal data features to the refractive error risk level using the SHAP value analysis method, and output the core risk factor.

[0045] In some embodiments of this application, this step is tightly coupled with the risk prediction module, and the SHAP (SHapley Additive exPlanations) value analysis method is used to quantify the interpretability of the prediction results of the LSTM model.

[0046] Specifically, all input features used by the LSTM neural network model to determine the risk level of refractive errors are obtained, including the rate of change of axial length, annual increase in refractive power, pupillary response sensitivity, and trend of change in eye movement stability. The contribution value of each input feature to the risk prediction result is calculated using the SHAP value analysis method (with a value range of 0-1). The contribution value directly reflects the degree of influence of the feature on the occurrence of refractive errors.

[0047] Furthermore, all input features are sorted from highest to lowest contribution value, and features with a cumulative contribution percentage reaching a preset threshold (≥60%) are selected as core risk factors. Quantitative results are output, such as "High risk: Annual increase in axial length greater than 0.3mm, contribution 0.72; Annual increase in refractive error of 0.5D, contribution 0.65; Decrease in pupillary sensitivity of 15%, contribution 0.58." This transparent contribution analysis not only solves the "black box" problem of traditional deep learning models but also provides precise targeting basis for the subsequent development of personalized intervention plans.

[0048] S106: Using a pre-set knowledge graph engine, an intervention plan is generated based on the visual acuity parameters, the refractive error risk level, and the core risk factors.

[0049] In some embodiments of this application, this step is completed collaboratively by the knowledge graph engine and the solution generation module of the cloud-based intelligent layer.

[0050] First, a knowledge graph containing ophthalmology clinical guidelines, drug databases, and intervention program libraries was constructed. This graph is a fourth-order logical association graph of "symptom-cause-risk factor-intervention measures". The graph nodes contain more than 1,200 entities, and the edge relationships include 6 types of relationships such as causality, adaptation, and contraindication.

[0051] The solution generation adopts an innovative dual-drive mechanism of "risk factor targeted intervention + individual characteristic adaptation": Based on the core risk factors output by S105, one or more preliminary intervention measures with an adaptation relationship with the core risk factors are selected from the knowledge graph (e.g., excessive axial length → recommending orthokeratology lens adaptation combined with axial length monitoring); further, the individual characteristics of the examinee are obtained (including age, lifestyle habits, previous intervention history, and local medical resources), and the preliminary intervention measures are adapted and optimized (e.g., increasing outdoor activity guidance for children, optimizing the allocation of near-field eye use time for office workers), generating a personalized intervention plan that includes at least one of the measures of eye habit correction, vision training, and orthodontic device adaptation; further, through reinforcement learning algorithms, feedback on the implementation effect of historical intervention plans is collected from people with the same risk level and similar individual characteristics, and the follow-up examination cycle, training intensity, and other detailed parameters in the subsequently generated intervention plans are dynamically adjusted according to the feedback, forming a closed-loop optimization intelligent generation mechanism.

[0052] This mechanism enables intervention programs to be "targeted, adaptable, and dynamically optimized," significantly improving the relevance and adherence of health management.

[0053] S107: The intervention plan is pushed to the doctor's terminal and the examinee's terminal for visualization through a preset multi-terminal interaction module.

[0054] In some embodiments of this application, this step is performed by a multi-terminal interaction layer to achieve data visualization and multi-role collaboration. The multi-terminal interaction layer includes a doctor terminal, a patient terminal, a monitoring terminal, and a data sharing interface.

[0055] After the analysis results and intervention plans are generated, the system synchronously pushes them to various terminals via HTTPS+AES-256 encryption protocol: the doctor's terminal provides ophthalmologists with data query, risk assessment result review, and intervention plan adjustment functions, and supports the generation and digital signature of electronic diagnostic reports to facilitate clinical decision-making; the examinee's terminal displays vision test reports, refractive error risk levels, core risk factors, and personalized intervention suggestions in intuitive charts and text, while also supporting functions such as follow-up reminders and online consultation appointments to enhance patient participation and self-management capabilities; the monitoring terminal is for health management institutions to realize statistical analysis of regional vision health data and generate public health decision-making reference reports; the data sharing interface follows medical data security standards, providing standard interfaces with hospital information systems and electronic health record systems to achieve cross-institutional interoperability of test data and break down information silos.

[0056] In addition, the system has a built-in quality assessment module that regularly compiles statistics on the accuracy of detection data and the execution rate of intervention plans, generates system performance assessment reports, and feeds them back to the cloud platform for continuous optimization of models and knowledge graphs.

[0057] It should be noted that, although the embodiments in this application are based on... Figure 1 Steps S101 to S107 will be described sequentially, but this does not mean that steps S101 and S107 must be performed in a strict order. The reason this embodiment follows this order is... Figure 1 The order in which steps S101 to S107 are described is provided to facilitate understanding of the technical solutions of the embodiments of this application by those skilled in the art. In other words, in the embodiments of this application, the order of steps S101 to S107 can be appropriately adjusted according to actual needs.

[0058] pass Figure 1This approach deeply integrates multimodal data with artificial intelligence technology, achieving a fully intelligent upgrade of the entire vision testing and health management process. First, it provides comprehensive data collection dimensions, integrating vision test, ocular physiology, and environmental data, thus improving detection accuracy compared to traditional methods. Second, the deep learning model based on attention mechanisms outputs vision parameters with an accuracy rate of no less than 98%. Combined with LSTM and SHAP value analysis, it can predict future refractive error risks and quantify core risk factors, further improving prediction accuracy. Third, it generates personalized intervention plans using a three-dimensional mechanism of risk targeting, individual adaptation, and reinforcement learning optimization. Based on knowledge graph screening measures and combined with individual characteristic adaptation, it improves plan execution rate. Fourth, it enables cross-institutional collaboration in management processes, achieving seamless integration of the entire chain from testing to diagnosis, intervention, and supervision through HL7 FHIR standardization and multi-terminal interaction, breaking down information silos.

[0059] Figure 2 The following is a logic diagram of LSTM risk prediction and contribution analysis for a vision detection and management method based on multimodal data fusion provided in the embodiments of this application.

[0060] Figure 3 This is a flowchart illustrating a vision detection and management method based on multimodal data fusion, as provided in an embodiment of this application.

[0061] exist Figure 3 The document marks the six key steps from identity authentication to data archiving, as well as the participating nodes of each module.

[0062] Figure 4 This is a system architecture diagram of a vision detection and management method based on multimodal data fusion, provided in an embodiment of this application.

[0063] exist Figure 4 The text describes the composition and data flow of the detection terminal layer, edge computing layer, cloud intelligence layer, and multi-terminal interaction layer.

[0064] Figure 5 A schematic diagram of a vision detection and management device based on multimodal data fusion provided in this application embodiment includes: At least one processor; and, A memory that is communicatively connected to at least one processor; wherein, A multimodal data fusion vision detection and management method that stores instructions executable by at least one processor, such that the at least one processor is able to perform any of the above-mentioned functions.

[0065] Some embodiments of this application provide a non-volatile computer storage medium for vision detection and management based on multimodal data fusion, which stores computer-executable instructions capable of executing any of the aforementioned multimodal data fusion vision detection and management methods.

[0066] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and medium embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the description of the method embodiments.

[0067] The devices and media provided in this application are one-to-one with the methods. Therefore, the devices and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0068] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0069] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0070] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.

[0071] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0072] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0073] Memory may include non-persistent storage in computer-readable media, random access memory (RAM), and non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0074] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0075] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0076] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the technical principles of this application should fall within the protection scope of this application.

Claims

1. A multi-modal data fusion based vision detection and management method, characterized in that, The method includes: Acquire multimodal visual acuity data of the examinee; the multimodal visual acuity data includes visual acuity test data, ocular physiological data, and testing environment data; The multimodal vision data is preprocessed to generate standardized multimodal data, and the standardized multimodal data is judged based on a preset threshold. If there is data that does not meet the preset threshold, a terminal alarm is triggered. The standardized multimodal data is subjected to feature extraction and fusion using a pre-defined attention mechanism deep learning model to generate a fused feature vector. Based on the fused feature vector and a pre-defined activation function, visual parameters are output. The subject's historical multimodal data is used as a time series and input into a pre-trained LSTM neural network model to determine the subject's refractive error risk level within a preset future time period. The contribution of standardized multimodal data features to the refractive error risk level is quantified using the SHAP value analysis method, and the core risk factor is output. An intervention plan is generated based on the visual acuity parameters, the refractive error risk level, and the core risk factors using a pre-set knowledge graph engine. The intervention plan is pushed to the doctor's terminal and the examinee's terminal for visualization through a preset multi-terminal interaction module.

2. The method of claim 1, wherein, The process involves using a pre-defined attention mechanism deep learning model to extract and fuse features from the standardized multimodal data, generating a fused feature vector, and outputting visual acuity parameters based on the fused feature vector and a pre-defined activation function. Specifically, this includes: Visual function features are extracted from the vision test data in the standardized multimodal data; the visual function features include at least optotype recognition accuracy and average pupillary reaction time. Extract ocular physiological features from the standardized multimodal data; the ocular physiological features include at least ocular rotation stability, pupil constriction rate, pupil dilation rate, blink frequency variation coefficient, and normalized axial length. By combining the light intensity and temperature and humidity parameters in the standardized multimodal data, environmental interference correction is performed on the extracted visual functional features and ocular physiological features; The corrected visual function features and eye physiological features are input into a pre-defined attention mechanism deep learning model. The attention mechanism is used to assign weight coefficients to each feature parameter to perform feature weighting. The weighted features are mapped to a unified feature space to generate a fused feature vector. The fused feature vector is input into a fully connected layer, and visual acuity parameters are output through a preset Softmax activation function; the visual acuity parameters include refractive power, astigmatism degree, and astigmatism axis.

3. The method of claim 1, wherein, The step of using the subject's historical multimodal data as a time series and inputting it into a pre-trained LSTM neural network model to determine the subject's refractive error risk level within a preset future time period specifically includes: Obtain historical test records of the examinee within a preset historical time period, sorted by test timestamp; the historical test records include standardized multimodal data; The historical detection records are used as time series data and input into a pre-trained LSTM neural network model. By using the gating unit of the LSTM neural network model to capture the temporal dependence of visual parameters over time in the time series data, the development pattern of refractive abnormalities can be learned. The temporal dependence and developmental pattern are compared and analyzed with the baseline data of the population grouped by age and gender, respectively, and the refractive error risk level of the examinee within a preset time period is output.

4. The method of claim 1, wherein, The contribution of standardized multimodal data features to the refractive error risk level is quantified using the SHAP value analysis method, and core risk factors are output, specifically including: The input features used by the LSTM neural network model to determine the risk level of refractive errors are obtained; the input features include the rate of change of axial length, the annual increase in refractive power, pupillary response sensitivity, and the trend of change in eye movement stability. The SHAP value analysis method was used to calculate the contribution of each input feature to the risk level of refractive abnormality. All input features are sorted from highest to lowest according to their contribution value, and the input features whose cumulative contribution ratio reaches the preset ratio threshold are selected as the core risk factors.

5. The method of claim 1, wherein, The step involves generating an intervention plan using a pre-set knowledge graph engine, based on the visual acuity parameters, the refractive error risk level, and the core risk factors. Specifically, this includes: Construct a knowledge graph; the knowledge graph includes ophthalmology clinical guidelines, drug databases, and intervention program libraries; it is a fourth-order logical association graph of symptoms, etiologies, risk factors, and intervention measures; Based on the core risk factors, preliminary intervention measures that are compatible with the core risk factors are selected from the knowledge graph; Obtain the individual characteristics of the examinee; these characteristics include age, lifestyle habits, previous intervention history, and local medical resources. Based on the individual characteristics, the preliminary intervention measures are adapted and optimized to generate an intervention plan; the intervention plan includes at least one of eye habit correction, vision training, and orthodontic device adaptation; By using a pre-set reinforcement learning algorithm, feedback on the implementation effect of historical intervention programs is collected from individuals with the same refractive error risk level and individual characteristics. Subsequent intervention programs are then adjusted based on this feedback to form a closed-loop optimization mechanism.

6. The method of claim 1, wherein, The process of judging the standardized multimodal data based on a preset threshold, and triggering a terminal alarm if data does not meet the preset threshold, specifically includes: The visual acuity test data is extracted to determine the pupillary reaction time; The eye physiological data are extracted to determine the blinking frequency; The pupillary reaction time is compared with a preset reaction time threshold, and the blinking frequency is compared with a preset frequency threshold. If the pupillary reaction time is greater than the reaction time threshold, and / or the blinking frequency is greater than the frequency threshold, it is determined that there is a physiological abnormality in the standardized multimodal data, and an alarm is triggered.

7. The method of claim 1, wherein, The acquisition of the subject's multimodal visual acuity data specifically includes: The system presents the subject with a random sequence of optotypes generated by a digital vision chart, records the subject's recognition results and reaction time, and generates vision test data. By integrating an infrared camera and a pupil detector, the eye movement trajectory, pupil diameter changes and blinking frequency of the examinee are collected to generate eye physiological data. The ambient temperature, humidity, and light intensity during the test are recorded by temperature and humidity sensors and light sensors to generate test environment data. The examinee's identity information is obtained through facial recognition, and the corresponding electronic health record is retrieved based on the identity information.

8. The method of claim 1, wherein, The preprocessing of the multimodal vision data to generate standardized multimodal data specifically includes: The vision test data, ocular physiological data, and detection environment data are cleaned using a preset outlier removal algorithm. The cleaned vision test data, ocular physiological data, and testing environment data are uniformly converted according to the HL7 FHIR standard format to generate standardized multimodal data.

9. A multi-modal data fusion based vision detection and management device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform a vision detection and management method based on multimodal data fusion as described in any one of claims 1-8.

10. A multi-modal data fusion vision detection and management storage medium storing computer executable instructions, the method comprising: The computer-executable instructions are capable of executing the vision detection and management method based on multimodal data fusion as described in any one of claims 1-8.