Intelligent vision trend chart generation method and system based on multi-modal data fusion
By collecting and processing multimodal vision data, using attention mechanism and two-way LSTM network to model vision trends, we generate interactive stereoscopic vision trend charts, solving the problem of insufficient reliability and insufficient visualization of vision trend prediction in the prior art, and achieving high-precision and personalized vision health management.
Patent Information
- Application Number
- CN202510342856.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
AI Technical Summary
Existing visual trend analysis methods usually rely on a single data source, fail to fully consider the complex influencing factors of vision changes, and lack of space-time alignment mechanisms, resulting in insufficient prediction accuracy and does not provide interactive visual results, making it difficult to intuitively understand the changes in one's own vision.
Multimodal vision data is collected through intelligent terminal devices, space-time alignment processing is performed, and the feature weights of each modal data are dynamically allocated by the attention mechanism to generate a comprehensive feature vector. Then, a residual-connected bidirectional LSTM network is used to dynamically model the visual acuity change trend, output the predicted trajectory including confidence intervals, and an interactive stereoscopic vision trend chart is generated by an adaptive smoothing algorithm.
Improve the reliability and accuracy of vision trend prediction, and enable users to intuitively understand their own vision changes through interactive stereo vision trend charts and provide personalized health management advice.
Smart Images

Figure CN120167879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image generation, and particularly to an intelligent visual acuity trend graph generation method and system based on multi-modal data fusion. Background Art
[0002] With the development of digital medicine and artificial intelligence technologies, personalized visual health management has become an important direction in ophthalmic diagnosis and treatment. Traditional visual acuity assessment methods mainly rely on regular ophthalmic examinations and visual acuity tests. However, due to the low measurement frequency and single data dimension, it is difficult to comprehensively reflect the individual visual acuity change trend. In recent years, with the popularization of intelligent terminal devices, personalized health monitoring technologies based on multi-modal data have gradually emerged. By integrating electronic visual acuity detection, eye usage behavior analysis, environmental light intensity monitoring, and medical image evaluation, a high-precision visual acuity change model can be constructed.
[0003] Existing visual acuity trend analysis methods usually only rely on a single data source and do not fully consider the complex influencing factors of visual acuity changes. In addition, existing methods generally lack a spatio-temporal alignment mechanism and it is difficult to establish accurate temporal correlations in multi-source heterogeneous data, resulting in insufficient prediction accuracy. At the same time, current trend analysis is mostly static evaluation and does not provide interactive visualization results, making it difficult for users to intuitively understand their own visual acuity changes. Therefore, there is an urgent need for an intelligent visual acuity trend graph generation method based on multi-modal data fusion that can integrate multiple data sources, establish a high-precision visual acuity change prediction model, and present the individual visual acuity development trend in an intuitive and interactive manner, so as to provide support for precise visual health management. Summary of the Invention
[0004] Based on the above purpose, the present invention provides an intelligent visual acuity trend graph generation method based on multi-modal data fusion to solve the problem that visual acuity trend prediction is not reliable enough.
[0005] The intelligent visual acuity trend graph generation method based on multi-modal data fusion includes the following steps:
[0006] S1: Collect multi-modal visual acuity data of a user through an intelligent terminal device, where the multi-modal visual acuity data includes electronic visual acuity detection data, eye usage behavior trajectory data, environmental light intensity data, and medical image feature data;
[0007] S2: Perform spatio-temporal alignment processing on the multi-modal visual acuity data to generate a standardized data set with a unified timestamp, where the eye usage behavior trajectory data and the environmental light intensity data are multi-dimensionally matched through a dynamic sliding window;
[0008] S3: Input the standardized data set into a multi-modal feature fusion module, and dynamically allocate feature weights to each modal data through an attention mechanism to generate a comprehensive feature vector including spatio-temporal correlation information;
[0009] S4: Based on the comprehensive feature vector, construct a time series prediction model, and use a bidirectional LSTM network with residual connections to dynamically model the vision change trend, and output a prediction trajectory including a confidence interval;
[0010] S5: Superimpose the prediction trajectory on the historical data, and generate an interactive three-dimensional vision trend map through an adaptive smoothing algorithm. The three-dimensional vision trend map includes a pathological factor association layer and a behavior intervention suggestion layer.
[0011] Further, the S1 includes:
[0012] S11, Electronic vision detection data collection: Obtain the visual acuity threshold value of the user through the electronic vision test module embedded in the intelligent terminal device, and measure the vision using the standard contrast sensitivity function;
[0013] S12, Eye use behavior trajectory data collection: Use the eye movement tracking sensor of the intelligent terminal device to obtain the eye movement trajectory of the user;
[0014] S13, Ambient light intensity data collection: Real-time monitor the ambient light intensity through the light sensor of the intelligent terminal device, and calculate the light intensity mean and volatility;
[0015] S14, Medical image feature data collection: Combine an optical coherence tomography device to obtain a retinal image, and extract a pathological feature vector.
[0016] Further, the S2 includes:
[0017] S21, Perform time synchronization on the electronic vision detection data, eye use behavior trajectory data, ambient light intensity data, and medical image feature data, and use the linear interpolation method to map all data to a unified timestamp set;
[0018] S22, Dynamic sliding window matching: Use the dynamic sliding window method to align different modality data in the time and feature spaces;
[0019] S23, Normalization processing: Normalize all modality data to make their numerical distributions consistent.
[0020] Further, the S3 includes:
[0021] S31, Multimodal feature encoding: Perform feature encoding on different modality data in the standardized dataset, and define the input modality dataset;
[0022] S32, Feature embedding and spatio-temporal encoding: Use a trainable linear transformation to project different modality data into the same feature space;
[0023] S33, Attention mechanism dynamic feature fusion: Calculate the attention weights of each modality data for the comprehensive feature vector to obtain the final comprehensive feature vector.
[0024] Further, the S4 includes:
[0025] S41, Construct time series input data: Extract data segments of consecutive time windows from the comprehensive feature vector to construct a time series input.
[0026] S42, Train a bidirectional LSTM prediction model: Use a bidirectional LSTM network to model the time series data and process the forward and backward information of the time series simultaneously.
[0027] S43, Generate prediction trajectories and confidence intervals: Predict the future vision change trend and calculate the confidence interval of the prediction result using an uncertainty estimation algorithm.
[0028] Further, the S41 extracts data segments of consecutive time windows from the comprehensive feature vector to construct a time series input dataset and standardizes the input data.
[0029] Further, the S42 includes:
[0030] S421, Calculate LSTM cells;
[0031] S422, Calculate the bidirectional LSTM and predict the vision value through a fully connected layer.
[0032] Further, the S43 includes:
[0033] S431, Use the Monte Carlo Dropout method to model the uncertainty of the prediction result, perform multiple forward propagations, and obtain the sample distribution;
[0034] S432, Calculate Confidence interval.
[0035] Further, the S5 includes:
[0036] S51, Overlay the prediction trajectory with historical data: Align the vision change trajectory predicted by the LSTM with the historical real measurement data to form a complete time series.
[0037] S52, Adaptive smoothing algorithm: Use the exponentially weighted moving average method to make the vision change trend smoother while retaining the long-term trend.
[0038] S53, Generate a stereoscopic vision trend map: Add a pathological factor association layer and a behavioral intervention suggestion layer to the interactive three-dimensional vision trend map.
[0039] An intelligent visual acuity trend graph generation system based on multi-modal data fusion, which is used to implement the above-mentioned intelligent visual acuity trend graph generation method based on multi-modal data fusion, includes the following modules:
[0040] Data acquisition module: Collect multi-modal visual acuity data of users through intelligent terminal devices, including electronic visual acuity detection data, eye usage behavior trajectory data, environmental light intensity data, and medical image feature data;
[0041] Data processing module: Perform spatio-temporal alignment processing on the collected multi-modal visual acuity data to generate a standardized data set with a unified time stamp, and use a dynamic sliding window to achieve multi-dimensional matching of eye usage behavior trajectory data and environmental light intensity data;
[0042] Feature fusion module: Input the standardized data set into a feature fusion model based on the attention mechanism, dynamically allocate the feature weights of each modal data, and generate a comprehensive feature vector including spatio-temporal correlation information;
[0043] Prediction modeling module: Build a time series prediction model based on the comprehensive feature vector, use a bidirectional LSTM network with residual connections to dynamically model the visual acuity change trend, and output a prediction trajectory including a confidence interval;
[0044] Trend graph generation module: Overlay the prediction trajectory with historical data, and generate an interactive three-dimensional visual acuity trend graph through an adaptive smoothing algorithm. The three-dimensional visual acuity trend graph includes a pathological factor association layer and a behavior intervention suggestion layer.
[0045] Advantages of the present invention:
[0046] The present invention provides an intelligent visual acuity trend graph generation method based on multi-modal data fusion, which can comprehensively integrate electronic visual acuity detection data, eye usage behavior trajectory data, environmental light intensity data, and medical image feature data. Through spatio-temporal alignment and multi-dimensional matching, it improves the temporal consistency and feature expression ability of the data. By dynamically allocating the feature weights of each modal data through the attention mechanism, the information from different sources is reasonably fused to generate a comprehensive feature vector including spatio-temporal correlation information. Further, a bidirectional LSTM network with residual connections is used to model the visual acuity change trend, which not only improves the long-term stability of the prediction, but also can evaluate the uncertainty of the prediction result through the confidence interval, thereby improving the reliability of personalized visual acuity trend prediction.
[0047] The present invention can fuse the predicted trajectory with historical data and optimize the trend curve through an adaptive smoothing algorithm, making the visualization of vision changes more intuitive, smooth, and interactive. By constructing a three-dimensional vision trend graph, while showing the vision development trend, the present invention further superimposes a pathological factor correlation layer and a behavior intervention suggestion layer, enabling users to clearly understand the potential influencing factors of their own vision changes and obtain personalized health management suggestions. Compared with traditional single-data-source analysis methods, the present invention has higher accuracy and applicability in aspects such as data fusion, trend prediction, and visual interaction, providing technical support for precise and personalized vision health management. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0049] Figure 1 is the method flow chart of the embodiment of the present invention;
[0050] Figure 2 is the system module diagram of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following further elaborates on the present invention in combination with specific embodiments.
[0052] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those with ordinary skills in the field to which the present invention belongs. The "first", "second", and similar terms used in the present invention do not indicate any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0053] As Figure 1 shown, the intelligent vision trend graph generation method based on multi-modal data fusion includes the following steps:
[0054] S1: Collect the multi-modal vision data of the user through the intelligent terminal device. The multi-modal vision data includes electronic vision detection data, eye usage behavior trajectory data, environmental light intensity data, and medical image feature data;
[0055] S2: Perform spatio-temporal alignment processing on the multi-modal vision data to generate a standardized data set with a unified timestamp. Among them, the eye usage behavior trajectory data and the environmental light intensity data are multi-dimensionally matched through a dynamic sliding window;
[0056] S3: Input the standardized data set into the multi-modal feature fusion module, and dynamically allocate the feature weights of each modal data through the attention mechanism to generate a comprehensive feature vector including spatio-temporal correlation information;
[0057] S4: Build a time series prediction model based on the comprehensive feature vector, and use a residual-connected bidirectional LSTM network to dynamically model the vision change trend, and output a prediction trajectory including a confidence interval;
[0058] S5: Superimpose the prediction trajectory on the historical data, and generate an interactive three-dimensional vision trend map through an adaptive smoothing algorithm. The three-dimensional vision trend map includes a pathological factor association layer and a behavior intervention suggestion layer.
[0059] S1 includes:
[0060] S11, Electronic vision detection data collection: Obtain the visual acuity threshold of the user through the electronic vision test module embedded in the intelligent terminal device, and measure the vision using the standard contrast sensitivity function (CSF). The vision value is expressed as:
[0061] ;;
[0062] Among them, represents the minimum resolvable angle that the user can identify, and the vision value is expressed in the standard logarithmic minimum angle of resolution (LogMAR);
[0063] S12, Eye usage behavior trajectory data collection: Use the eye movement tracking sensor of the intelligent terminal device to obtain the eye movement trajectory of the user , which is expressed as:
[0064] ;
[0065] Among them, represents the line-of-sight focus coordinates of the th sampling, is the sampling timestamp, is the total number of sampling points;
[0066] Calculate the eye movement stability based on the line-of-sight focus density, which is expressed as:
[0067] ;
[0068] Among them, is the central coordinate of all sampling points;
[0069] S13, Environmental light intensity data acquisition: The environmental light intensity is monitored in real time through the light sensor of the intelligent terminal device , and the average light intensity and volatility are calculated, expressed as:
[0070] ;
[0071] ;
[0072] Among them, represents the sampling duration, represents the average light level, represents the degree of light change;
[0073] S14, Medical image feature data acquisition: Combining optical coherence tomography (OCT) equipment to obtain retinal images and extract pathological feature vectors, expressed as:
[0074] ;
[0075] Among them, is the dimension of the extracted image features, and each feature is automatically segmented and calculated by the U-Net segmentation model, specifically including:
[0076] (1) Input image normalization: Normalize to the range to improve the stability of the U-Net segmentation model;
[0077] (2) U-Net segmentation: Input the normalized image, extract features through the U-Net segmentation model and generate a segmentation mask, expressed as:
[0078] ;
[0079] Among them, is the probability distribution of the lesion area, and the value range is , indicating the possibility that the pixel point belongs to the lesion, is the pixel activation value generated by the U-Net segmentation model, which is converted into a probability distribution through the Sigmoid activation function, represents the exponential transformation of the pixel activation value;
[0080] (3) Feature calculation:
[0081] Macular thickness:
[0082] ;
[0083] Optic disc cup-to-disc ratio:
[0084] ;
[0085] Vascular density:
[0086] ;
[0087] Among them, is the average thickness of the retinal macula area, is the height of the inner retinal interface at the th position (close to the vitreous body), is the height of the outer retinal interface at the th position (close to the choroid), is the total number of pixel points in the macular area, is the area of the optic cup region, is the total area of the optic disc region, is the vascular marker value of the pixel in the binary vascular segmentation image, is the total number of pixels.
[0088] S2 includes:
[0089] S21, perform time synchronization on the electronic vision detection data, eye usage behavior trajectory data, environmental light intensity data, and medical image feature data, and map all data to a unified set of timestamps using linear interpolation. The data values of the timestamp set are expressed as:
[0090] ;
[0091] Among them, is the data value at time moment, is the adjacent sampled data point, is the corresponding data timestamp;
[0092] S22, dynamic sliding window matching: To match the eye usage behavior trajectory data and environmental light intensity data, use the dynamic sliding window method to align different modality data in the time and feature spaces. The window matching rule is expressed as:
[0093] ;
[0094] Among them, is the sliding window at time position, is the line-of-sight focus coordinate in the eye usage behavior trajectory data, is the environmental light intensity, is the timestamp of the data point, is the time span of the sliding window, which is adaptively adjusted according to the data change rate, , is the adjustment coefficient, is the number of data points within the window, is the mean of the line-of-sight focus within the window;
[0095] S23, normalization: Normalize all modal data to make their numerical distributions consistent, expressed as:
[0096] ;
[0097] where, is the normalized data value, is the original data value, is the minimum and maximum values of the data.
[0098] S3 includes:
[0099] S31, multi-modal feature encoding: Encode the features of different modal data in the standardized dataset, define the input modal dataset, expressed as:
[0100] ;
[0101] where, is the electronic vision detection feature vector, including LogMAR visual acuity value and contrast sensitivity, is the eye usage behavior trajectory feature vector, including line-of-sight focus density, eye movement speed, and fixation duration, is the environmental light feature vector, including average light intensity and light change rate, is the medical image feature vector, including optic disc cup-to-disc ratio, macular thickness, and vascular density;
[0102] S32, feature embedding and spatio-temporal encoding: Project different modal data into the same feature space using a trainable linear transformation, expressed as:
[0103] ;
[0104] where, is the embedded feature of modality , is the linear transformation weight matrix of modality , is the bias vector;
[0105] Combine the timestamp information and enhance the temporal features using positional encoding, expressed as:
[0106] ;
[0107] Among them, is the time position encoding, is the data timestamp, is the feature dimension, is the current dimension index;
[0108] S33, attention mechanism dynamic feature fusion: Calculate the attention weights of each modality data pair for the comprehensive feature vector, expressed as:
[0109] ;
[0110] Among them, is the attention weight of modality , is the trainable parameter for attention calculation, and the denominator is the normalized distribution of all modalities to ensure that the sum of weights is 1, is the embedded feature of modality ;
[0111] Obtain the final comprehensive feature vector, expressed as:
[0112] ;
[0113] Among them, is the fused comprehensive feature vector, including the spatio-temporal correlation information of all modalities.
[0114] S4 includes:
[0115] S41, construct time series input data: Extract data segments of continuous time windows from the comprehensive feature vector to construct a time series input;
[0116] S42, train a bidirectional LSTM prediction model: Use a bidirectional LSTM (Bi-LSTM) network to model the time series data, and process the forward and backward information of the time series simultaneously to improve the prediction accuracy;
[0117] S43, generate prediction trajectories and confidence intervals: Predict the future vision change trend, and use an uncertainty estimation algorithm to calculate the confidence interval of the prediction result.
[0118] S41 extracts data segments of continuous time windows from the comprehensive feature vector to construct a time series input dataset, expressed as:
[0119] ;
[0120] Among them, is time The feature input sequence at '[['END]] is the time The comprehensive feature vector at '[['END]] is the time window size, representing the length of historical data for prediction;
[0121] Normalize the input data to improve data stability and trainability, expressed as:
[0122] ;
[0123] Among them, is the normalized feature vector, is the mean of historical data, is the standard deviation of historical data.
[0124] S42 includes:
[0125] S421, calculate the LSTM cell, expressed as:
[0126] ;
[0127] ;
[0128] ;
[0129] ;
[0130] ;
[0131] ;
[0132] Among them, are the activation values of the input gate, forget gate, and output gate respectively, is the cell state at the current time step, is the hidden state at the current time step (i.e., the LSTM output), are the weight matrices of the input gate, forget gate, output gate, and cell state respectively, are the corresponding bias terms, is the Sigmoid activation function, controlling the information flow, is the hyperbolic tangent activation function, controlling the data range, represents element-wise multiplication;
[0133] S422, calculate the bidirectional LSTM, expressed as:
[0134] ;
[0135] ;
[0136] ;
[0137] Among them, is the forward LSTM processing the time series from the past to the present, is the backward LSTM processing the time series from the present to the past, are the hidden states of the forward and backward LSTMs respectively, is the final hidden state combined with the residual connection;
[0138] The visual acuity value is predicted through a fully connected layer, expressed as:
[0139] ;
[0140] Among them, is the time of the visual acuity prediction value, are the weights and biases of the fully connected layer.
[0141] S43 includes:
[0142] S431, using the Monte Carlo Dropout method to model the uncertainty of the prediction results, performing multiple forward propagations to obtain the sample distribution, expressed as:
[0143] ;
[0144] Among them, is the prediction value of the th forward propagation, mechanism is used to simulate the model uncertainty, randomly ignoring some neurons during each prediction to improve the generalization ability;
[0145] S432, calculating the confidence interval, expressed as:
[0146] ;
[0147] ;
[0148] ;
[0149] Among them, is the mean of all prediction values, is the variance of all prediction values, is the confidence interval range, representing the uncertainty of the prediction results.
[0150] S5 includes:
[0151] S51, Superposition of predicted trajectory and historical data: Align the visual acuity change trajectory predicted by LSTM with the historical real measurement data to form a complete time series, expressed as:
[0152] ;
[0153] where, is the superimposed visual acuity trajectory value at time , is the historical visual acuity measurement value at time , is the LSTM predicted visual acuity value at time , is the historical data weight factor, which is adaptively adjusted to make the predicted trajectory smoothly integrate with the historical trend, , where, controls the smoothness degree, is the set smooth starting point;
[0154] S52, Adaptive smoothing algorithm: Adopt the exponentially weighted moving average (EWMA) method to make the visual acuity change trend smoother while retaining the long-term trend, expressed as:
[0155] ;
[0156] where, is the smoothed visual acuity value at time , , is the smoothing coefficient, expressed as:
[0157] ;
[0158] where, controls the smoothness degree, and is dynamically adjusted according to the visual acuity change rate to make the mutated data smoothly transition;
[0159] S53, Generation of stereoscopic visual acuity trend diagram: Add a pathological factor association layer and a behavior intervention suggestion layer in the interactive three-dimensional visual acuity trend diagram to assist in personalized visual acuity management, specifically including:
[0160] (1) Pathological factor association layer:
[0161] Calculate the visual acuity decline rate, expressed as:
[0162] ;
[0163] Set the warning threshold , when exceeds the threshold, mark the pathological association area;
[0164] (2) Behavior intervention suggestion layer:
[0165] Calculate the environmental impact factor, expressed as:
[0166] ;
[0167] wherein, is the environmental light intensity, is the eye-using behavior (screen usage duration), is the near-eye-using time, is the behavior impact weight, obtained from expert experience or data training;
[0168] When exceeds the set threshold, generate behavior intervention suggestions, such as "increase outdoor activity time" or "reduce screen brightness".
[0169] As Figure 2 shown, an intelligent visual acuity trend graph generation system based on multimodal data fusion is used to implement the above-mentioned intelligent visual acuity trend graph generation method based on multimodal data fusion, including the following modules:
[0170] Data acquisition module: Collect multimodal visual acuity data of users through intelligent terminal devices, including electronic visual acuity detection data, eye-using behavior trajectory data, environmental light intensity data, and medical image feature data;
[0171] Data processing module: Perform spatio-temporal alignment processing on the collected multimodal visual acuity data to generate a standardized data set with a unified time stamp, and use a dynamic sliding window to achieve multi-dimensional matching of eye-using behavior trajectory data and environmental light intensity data;
[0172] Feature fusion module: Input the standardized data set into a feature fusion model based on the attention mechanism, dynamically allocate the feature weights of each modal data, and generate a comprehensive feature vector including spatio-temporal correlation information;
[0173] Prediction modeling module: Construct a time series prediction model based on the comprehensive feature vector, dynamically model the visual acuity change trend using a bidirectional LSTM network with residual connections, and output a prediction trajectory including a confidence interval;
[0174] Trend graph generation module: Overlay the prediction trajectory with historical data, and generate an interactive three-dimensional visual acuity trend graph through an adaptive smoothing algorithm. The three-dimensional visual acuity trend graph includes a pathological factor association layer and a behavior intervention suggestion layer
[0175] Those of ordinary skill in the art should understand that any discussion of the above embodiments is merely exemplary and is not intended to imply that the scope of the present invention is limited to these examples; under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present invention as described above, and they are not provided in detail for the sake of brevity.
Claims
1. An intelligent vision trend graph generation method based on multimodal data fusion, characterized in that: The following steps are involved: S1: Collecting multimodal vision data of the user through a smart terminal device, wherein the multimodal vision data includes electronic vision detection data, eye behavior trajectory data, ambient light intensity data, and medical imaging feature data; S2: Perform spatiotemporal alignment processing on multimodal vision data to generate a standardized data set with a unified timestamp, in which the eye behavior trajectory data and the ambient light intensity data are matched in multiple dimensions through a dynamic sliding window; S3: Input the standardized data set into the multimodal feature fusion module, dynamically assign the feature weights of each modality data through the attention mechanism, and generate a comprehensive feature vector including spatiotemporal correlation information; S4: constructing a time series prediction model based on the comprehensive feature vector, using a residual-connected bidirectional LSTM network to dynamically model the trend of vision changes, and outputting a prediction trajectory including a confidence interval; S5: Superimpose the predicted trajectory with the historical data, and generate an interactive stereoscopic vision trend map through an adaptive smoothing algorithm, wherein the stereoscopic vision trend map includes a pathological factor association layer and a behavioral intervention suggestion layer.
2. The method for generating an intelligent vision trend graph based on multimodal data fusion according to claim 1, characterized in that: The S1 includes: S11, electronic vision test data collection: obtaining the user's visual acuity value through the electronic vision test module embedded in the smart terminal device, and using the standard contrast sensitivity function to measure vision; S12, eye behavior trajectory data collection: using the eye tracking sensor of the smart terminal device to obtain the user's eye movement trajectory; S13, ambient light intensity data collection: real-time monitoring of ambient light intensity through the light sensor of the smart terminal device, and calculation of the light mean and volatility; S14, medical image feature data acquisition: combine optical coherence tomography equipment to obtain retinal images and extract pathological feature vectors.
3. The method for generating an intelligent vision trend graph based on multimodal data fusion according to claim 1, characterized in that: The S2 includes: S21, time synchronization of electronic vision test data, eye behavior trajectory data, ambient light intensity data, and medical imaging feature data, and mapping all data to a unified timestamp set using a linear interpolation method; S22, dynamic sliding window matching: a dynamic sliding window method is used to align different modal data in time and feature space; S23, normalization processing: normalize all modal data to make their numerical distribution consistent.
4. The method for generating an intelligent vision trend graph based on multimodal data fusion according to claim 3, characterized in that: The S3 includes: S31, multimodal feature encoding: feature encoding of different modal data in the standardized data set to define the input modal data set; S32, Feature Embedding and Spatiotemporal Coding: Use trainable linear transformation to project different modal data into the same feature space; S33, dynamic feature fusion with attention mechanism: calculate the attention weight of each modal data on the comprehensive feature vector to obtain the final comprehensive feature vector.
5. The method for generating an intelligent vision trend graph based on multimodal data fusion according to claim 4, characterized in that: The S4 includes: S41, constructing time series input data: extracting data segments of continuous time windows from the comprehensive feature vector to construct time series input; S42, training a bidirectional LSTM prediction model: using a bidirectional LSTM network to model time series data and process both the forward and backward information of the time series; S43, generate prediction trajectory and confidence interval: predict the future trend of vision change, and use uncertainty estimation algorithm to calculate the confidence interval of the prediction result.
6. The method for generating an intelligent vision trend graph based on multimodal data fusion according to claim 5, characterized in that: The S41 extracts data segments of continuous time windows from the comprehensive feature vector, constructs a time series input data set, and standardizes the input data.
7. The method for generating an intelligent vision trend graph based on multimodal data fusion according to claim 5, characterized in that: The S42 includes: S421, calculating LSTM unit; S422, calculate the bidirectional LSTM and predict the visual acuity value through the fully connected layer.
8. The method for generating an intelligent vision trend graph based on multimodal data fusion according to claim 5, characterized in that: The S43 includes: S431, Monte Carlo Dropout method is used to model the uncertainty of the prediction results, and multiple forward propagations are performed to obtain the sample distribution; S432, calculation Confidence interval.
9. The method for generating an intelligent vision trend graph based on multimodal data fusion according to claim 8, characterized in that: The S5 includes: S51, superposition of predicted trajectory and historical data: aligning the LSTM predicted vision change trajectory with the historical real measurement data to form a complete time series; S52, adaptive smoothing algorithm: uses an exponentially weighted moving average method to make the visual acuity change trend smoother while retaining the long-term trend; S53, Stereoscopic Vision Trend Chart Generation: Add pathological factor association layer and behavioral intervention suggestion layer to the interactive three-dimensional vision trend chart.
10. An intelligent vision trend graph generation system based on multimodal data fusion, used to implement the intelligent vision trend graph generation method based on multimodal data fusion as described in any one of claims 1 to 9, characterized in that: Includes the following modules: Data collection module: collects the user's multimodal vision data through smart terminal devices, including electronic vision test data, eye behavior trajectory data, ambient light intensity data and medical imaging feature data; Data processing module: Performs spatiotemporal alignment processing on the collected multimodal vision data to generate a standardized data set with a unified timestamp, and uses a dynamic sliding window to achieve multi-dimensional matching of eye behavior trajectory data and ambient light intensity data; Feature fusion module: inputs the standardized data set into the feature fusion model based on the attention mechanism, dynamically assigns the feature weights of each modality data, and generates a comprehensive feature vector including spatiotemporal correlation information; Prediction modeling module: builds a time series prediction model based on the comprehensive feature vector, uses a residual-connected bidirectional LSTM network to dynamically model the trend of vision changes, and outputs the prediction trajectory including the confidence interval; Trend chart generation module: superimposes the predicted trajectory with historical data, and generates an interactive stereoscopic vision trend chart through an adaptive smoothing algorithm. The stereoscopic vision trend chart includes a pathological factor association layer and a behavioral intervention suggestion layer.
Citation Information
Cited By
Urban building fire safety intelligent management platform based on BIM and GIS multi-source fusion
CN120634394A
Myopia occurrence risk assessment method based on multi-source data
CN120727301A
A myopia occurrence risk assessment method based on multi-source data
CN120727301B