Teenager vision development trend prediction method and system based on multi-modal data fusion
By using multimodal data fusion and deep learning models, the problem of insufficient multimodal data fusion in existing adolescent vision prediction technologies has been solved, improving the accuracy and adaptability of the prediction model, especially the prediction effect under irregular time series data.
Patent Information
- Application Number
- CN202511169165.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-12-12
AI Technical Summary
Existing vision prediction technologies for adolescents mainly rely on single-modal data, lack multimodal data fusion, have limited ability to model complex dynamic processes, and are deficient in processing irregular time series data. As a result, the prediction results are difficult to fully reflect the multidimensional characteristics of adolescent vision development. In particular, when faced with irregularly collected data, the generalization ability and accuracy of the model are limited.
The system acquires eye image data, physiological parameter data, and behavioral habit data of adolescents through a multi-source data acquisition module. After preprocessing, it uses a multimodal feature association analysis module to mine deep relationships. It then combines dynamic time warping algorithm and deep learning model to generate prediction results and optimizes model performance through adaptive weight adjustment and anomaly detection.
A more comprehensive and accurate predictive model for adolescent vision development trends was constructed, which improved the ability to fuse multimodal data, model complex dynamic processes, and process irregular time series data, thereby enhancing the comprehensiveness and accuracy of the prediction results and providing reliable support for the scientific assessment and intervention of adolescent vision development trends.
Smart Images

Figure CN121122701A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medical health and artificial intelligence technology, specifically to a method and system for predicting the development trend of adolescent vision based on multimodal data fusion. Background Technology
[0002] With the increasing prominence of myopia among teenagers, how to predict and intervene in the development trend of adolescent vision through scientific means has become a research hotspot. Existing technical solutions mainly focus on the processing and analysis of single-modal data (such as image data or numerical data). Although these solutions have improved prediction accuracy to some extent, they still have problems such as insufficient multimodal data fusion and limited generalization ability of prediction models, making it difficult to fully reflect the complex dynamic process of adolescent vision development.
[0003] Patent CN118512149B proposes a method based on the MTSformer prediction model, including an MCAM module, a TKMM module, and a fully connected layer. This method can handle the relationships between feature dimensions and time dimensions in time-series data of adolescent myopia vision and enhances the model's expressive power. However, this technical solution mainly relies on single-modal time-series data and lacks comprehensive analysis of multimodal data (such as images and physiological parameters), which may result in prediction results that fail to fully reflect the multidimensional characteristics of adolescent vision development. Furthermore, the model's insufficient exploration of correlations between data from different sources may affect its prediction accuracy in complex scenarios.
[0004] The patent with announcement number CN117153407B employs a data preprocessing method combining one-hot encoding and data standardization. It combines deep convolutional neural networks and temporal long short-term memory neural networks to predict image and numerical data respectively, improving data quality and prediction dimensionality. However, this technical solution still falls short in multimodal data fusion, failing to fully explore the deep correlations between different modalities, which may limit the comprehensiveness and accuracy of the prediction results. Furthermore, the solution's ability to process irregular time-series data needs improvement, limiting its adaptability to the irregular data collection scenarios during the visual development of adolescents.
[0005] The aforementioned problems indicate that existing technologies for predicting adolescent vision still have room for improvement in areas such as multimodal data fusion, complex dynamic process modeling, and processing irregular time series data. Therefore, this invention provides a method and system for predicting adolescent vision development trends based on multimodal data fusion. The aim is to integrate multiple data sources (such as images, physiological parameters, and behavioral data) to construct a more comprehensive and accurate prediction model, thereby better supporting the scientific assessment and intervention of adolescent vision development trends. Summary of the Invention
[0006] The purpose of this invention is to provide a method and system for predicting the development trend of adolescent vision based on multimodal data fusion, addressing the shortcomings of existing adolescent vision prediction technologies mentioned in the background, which primarily rely on single-modal data (such as image data or numerical data) for processing and analysis. While these technologies improve prediction accuracy to some extent, they suffer from insufficient multimodal data fusion, limited ability to model complex dynamic processes, and a lack of ability to process irregular time series data. These problems make it difficult for the prediction results to fully reflect the multidimensional characteristics of adolescent vision development, especially when faced with irregularly collected data, where the model's generalization ability and accuracy are limited.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a method for predicting the development trend of adolescent vision based on multimodal data fusion, characterized by comprising the following steps:
[0008] The system acquires eye image data, physiological parameter data, and behavioral habit data of adolescents through a multi-source data acquisition module.
[0009] The multimodal data is preprocessed to extract key features from each modality and map these features to a unified feature space.
[0010] The multimodal feature association analysis module is used to mine the deep relationships between different modal data and construct a multimodal feature association map;
[0011] Based on the multimodal feature association map, a prediction result of the vision development trend of adolescents is generated by combining the dynamic time warping algorithm and the deep learning model.
[0012] Preferably, the following steps are also included:
[0013] The time axis of each modality data is aligned by the time synchronization unit to ensure that the time labels of all modality data are consistent.
[0014] Preferably, the following steps are also included:
[0015] Multimodal data is divided into training, validation, and test sets, and an adaptive weight adjustment mechanism is introduced during the training phase. By assigning dynamic weights to different modal data, the model's comprehensive fitting ability to multimodal data is optimized.
[0016] Preferably, the following steps are also included:
[0017] For irregular time series data, linear interpolation is used to supplement missing values, and sliding window technology is used to segment the data into fixed-length segments.
[0018] Preferably, the following steps are also included:
[0019] When the feature strength of a certain modality data is lower than a preset threshold, the feature enhancement unit reconstructs the features of the modality data to improve its contribution in the multimodal feature association map.
[0020] Preferably, the following steps are also included:
[0021] By calculating the correlation coefficients between different modalities, samples with significantly deviated correlations from the normal range were screened out and marked as anomalous data.
[0022] The adolescent vision development trend prediction system based on multimodal data fusion includes the following components:
[0023] The multi-source data acquisition module is used to acquire eye image data, physiological parameter data, and behavioral habit data of adolescents;
[0024] The data preprocessing module, connected to the multi-source data acquisition module, is used to clean, standardize, and extract features from multimodal data;
[0025] A multimodal feature association analysis module, connected to the data preprocessing module, is used to mine the deep relationships between different modal data and construct a multimodal feature association map;
[0026] The prediction model module is connected to the multimodal feature association analysis module and is used to generate prediction results of adolescent vision development trends based on the multimodal feature association map.
[0027] The results display module is connected to the prediction model module and is used to present the prediction results to the user in a visual form.
[0028] Preferred options also include:
[0029] The time synchronization module is used to align the timelines of multimodal data to ensure that the time stamps of all modal data are consistent.
[0030] Preferred options also include:
[0031] The adaptive weight adjustment module is used to dynamically adjust the weight allocation of each modality during the training phase to optimize the model's overall fitting ability to multimodal data.
[0032] Preferred options also include:
[0033] The anomaly detection module is used to identify and remove samples whose correlation significantly deviates from the normal range by calculating the correlation coefficient between data from different modalities.
[0034] Compared with existing technologies, the beneficial effects of this invention are as follows: This method and system for predicting adolescent vision development trends based on multimodal data fusion integrates multiple data sources (such as images, physiological parameters, behavioral data, etc.) to construct a more comprehensive and accurate prediction model. In particular, it proposes targeted technical means in multimodal data fusion, complex dynamic process modeling, and irregular time series data processing, effectively improving the comprehensiveness and accuracy of the prediction results and providing reliable support for the scientific assessment and intervention of adolescent vision development trends. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of the process flow of the present invention;
[0036] Figure 2 This is a schematic diagram of the system framework structure of the present invention. Detailed Implementation
[0037] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0038] Please see Figure 1-2 The present invention provides a technical solution: a method and system for predicting the development trend of adolescent vision based on multimodal data fusion.
[0039] In this embodiment, the system includes a multi-source data acquisition module, a data preprocessing module, a multimodal feature association analysis module, a prediction model module, and a result display module. The multi-source data acquisition module is responsible for acquiring eye image data, physiological parameter data, and behavioral habit data of adolescents, and transmitting this data to the data preprocessing module. The data preprocessing module cleans, standardizes, and extracts features from the received data, and then transmits the processed data to the multimodal feature association analysis module. The multimodal feature association analysis module mines the deep relationships between different modalities of data and constructs a multimodal feature association map. Finally, the prediction model module generates prediction results based on this map and presents them to the user through the result display module.
[0040] In practice, the multi-source data acquisition module utilizes a variety of devices working collaboratively to complete data collection tasks. Eye image data is acquired by a high-resolution camera mounted on a fixed bracket and connected to a computer, ensuring real-time transmission of image data to the data preprocessing module. Physiological parameter data is collected through wearable devices such as smart bracelets or smart glasses. These devices have built-in sensors to monitor vision-related indicators such as heart rate and blood pressure, and transmit the data to the data preprocessing module via Bluetooth or wireless networks. Behavioral habit data primarily comes from a questionnaire survey system and mobile applications. The questionnaire survey system collects information on adolescents' eye habits through an online platform, while mobile applications record behavioral data such as screen time and reading posture. This data is transmitted to the data preprocessing module via a local area network or the internet.
[0041] After receiving multi-source data, the data preprocessing module first cleans the data to remove noise and invalid values. For example, for eye image data, edge detection algorithms are used to remove blurry or incomplete images; for physiological parameter data, a sliding window technique is used to filter abnormal fluctuations; and for behavioral habit data, a rule engine is used to screen for logical errors or contradictory information. After cleaning, the data preprocessing module standardizes the data of each modality to conform to the unified format required for subsequent analysis. For example, eye image data is converted to grayscale images and the resolution is adjusted to a unified standard; physiological parameter data is normalized to between 0 and 1; and behavioral habit data is encoded as numerical variables for easy calculation. After standardization, the data preprocessing module calls the feature extraction unit to separate key features from each modality. For example, features such as corneal curvature and pupil diameter are extracted from eye image data; features such as heart rate variability and blood pressure trends are extracted from physiological parameter data; and features such as daily screen usage time and continuous eye use time are extracted from behavioral habit data. These features are then mapped to a unified feature space to form feature vectors that can be used for subsequent analysis.
[0042] After receiving feature vectors from the data preprocessing module, the multimodal feature association analysis module first performs temporal and spatial alignment operations on each modality feature through the feature alignment submodule. Temporal alignment is achieved through a time synchronization unit, which calibrates the time axis based on the timestamp information of each modality data to ensure consistency in time labels across all modality data. For example, if the timestamp of a certain eye image data is 10:00:00, but the corresponding physiological parameter data has a 1-second delay, linear interpolation is used to fill in the missing values and eliminate the time discrepancy. Spatial alignment is achieved through coordinate transformation techniques, such as matching the pixel coordinates in the eye image data with the sensor positions in the physiological parameter data to ensure consistency in the spatial dimension. After alignment, the feature fusion submodule generates a unified multimodal feature representation using a weighted fusion algorithm. The weighted fusion algorithm dynamically assigns weights based on the importance of each modality data; for example, the weight of eye image data may be higher than that of behavioral habit data because the former has a more direct impact on vision prediction. The final generated multimodal feature representation is then passed to the multimodal feature association map construction unit.
[0043] The multimodal feature association graph construction unit mines deep relationships between different modalities using a deep learning model. Specifically, this unit employs graph neural network technology to construct a multimodal feature association graph, where each node represents a modal feature, and the edge weights represent the correlation strength between features. To improve the model's prediction accuracy, an attention mechanism is embedded in the graph, dynamically adjusting weight allocation by calculating the importance score of each modal feature. For example, when a modal feature has a high importance score, the model prioritizes that feature and assigns it a higher weight, thereby improving the accuracy of the prediction results. Furthermore, the multimodal feature enhancement module plays a role in the graph construction process. When the feature strength of a modal data falls below a preset threshold, the feature enhancement unit is triggered to reconstruct the features of that modal data. For example, if the corneal curvature feature in eye image data has a low strength, it is enhanced using a convolutional neural network to increase its contribution to the graph.
[0044] After receiving the multimodal feature association map, the prediction model module combines the dynamic time warping algorithm and the deep learning model to generate prediction results for adolescent vision development trends. The dynamic time warping algorithm is mainly used to process irregular time series data. It fills in missing values through linear interpolation and uses a sliding window technique to segment the data into fixed-length segments. For example, for behavioral habit data with uneven sampling intervals, missing values are first filled in through linear interpolation, and then the data is segmented into segments every 5 minutes to facilitate model processing. The deep learning model uses a long short-term memory network to model the multimodal feature association map, capturing the long-term dependencies in vision development. During the training phase, the adaptive weight adjustment module dynamically adjusts the weight allocation of each modality of data, optimizing the model's overall fitting ability to multimodal data by dynamically assigning weights to different modalities. For example, in the early stages of training, the weight of eye image data may be higher, while in the later stages of training, the weight of behavioral habit data is gradually increased to balance the model's generalization ability.
[0045] The results display module receives the prediction results generated by the prediction model module and presents them to the user in a visual format. Specifically, the results display module uses a line graph to show the trend of adolescent vision development, with the horizontal axis representing time and the vertical axis representing changes in vision indicators such as refractive error. Simultaneously, the results display module also provides a bar chart to show the contribution of each modality of data to the prediction results; for example, the contribution of ocular image data is 40%, physiological parameter data is 30%, and behavioral habit data is 30%. Furthermore, the results display module includes an anomaly detection module, which identifies and removes abnormal samples by calculating the correlation coefficients between the data of each modality. For example, if the correlation coefficient between a certain ocular image data and the corresponding physiological parameter data deviates significantly from the normal range, it is marked as abnormal data and removed from the prediction results.
[0046] Throughout the system's operation, modules interact via a local area network (LAN) or the internet to ensure real-time and reliable data transmission. For example, the multi-source data acquisition module transmits data to the data preprocessing module via a wireless network. The data preprocessing module then transmits the processed data to the multimodal feature association analysis module via the LAN, and so on, until the prediction result is generated. Furthermore, the system includes a time synchronization module and an adaptive weight adjustment module, used to align the time axis of the multimodal data and dynamically adjust the weight allocation of each modality, respectively, thereby further improving the system's prediction accuracy and stability.
[0047] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for predicting the development trend of adolescent vision based on multimodal data fusion, characterized in that: Includes the following steps, The system acquires eye image data, physiological parameter data, and behavioral habit data of adolescents through a multi-source data acquisition module. The multimodal data is preprocessed to extract key features from each modality and map these features to a unified feature space. The multimodal feature association analysis module is used to mine the deep relationships between different modal data and construct a multimodal feature association map; Based on the multimodal feature association map, a prediction result of the vision development trend of adolescents is generated by combining the dynamic time warping algorithm and the deep learning model.
2. The method for predicting the development trend of adolescent vision based on multimodal data fusion according to claim 1, characterized in that, It also includes the following steps: The time axis of each modality data is aligned by the time synchronization unit to ensure that the time labels of all modality data are consistent.
3. The method for predicting the development trend of adolescent vision based on multimodal data fusion according to claim 1, characterized in that, It also includes the following steps: Multimodal data is divided into training, validation, and test sets, and an adaptive weight adjustment mechanism is introduced during the training phase. By assigning dynamic weights to different modal data, the model's comprehensive fitting ability to multimodal data is optimized.
4. The method for predicting the development trend of adolescent vision based on multimodal data fusion according to claim 1, characterized in that, It also includes the following steps: For irregular time series data, linear interpolation is used to supplement missing values, and sliding window technology is used to segment the data into fixed-length segments.
5. The method for predicting the development trend of adolescent vision based on multimodal data fusion according to claim 1, characterized in that, It also includes the following steps: When the feature strength of a certain modality data is lower than a preset threshold, the feature enhancement unit reconstructs the features of the modality data to improve its contribution in the multimodal feature association map.
6. The method for predicting the development trend of adolescent vision based on multimodal data fusion according to claim 1, characterized in that, It also includes the following steps: By calculating the correlation coefficients between different modalities, samples with significantly deviated correlations from the normal range were screened out and marked as anomalous data.
7. A predictive system for adolescent vision development trends based on multimodal data fusion, characterized in that, It includes the following components: The multi-source data acquisition module is used to acquire eye image data, physiological parameter data, and behavioral habit data of adolescents; The data preprocessing module, connected to the multi-source data acquisition module, is used to clean, standardize, and extract features from multimodal data; A multimodal feature association analysis module, connected to the data preprocessing module, is used to mine the deep relationships between different modal data and construct a multimodal feature association map; The prediction model module is connected to the multimodal feature association analysis module and is used to generate prediction results of adolescent vision development trends based on the multimodal feature association map. The results display module is connected to the prediction model module and is used to present the prediction results to the user in a visual form.
8. The adolescent vision development trend prediction system based on multimodal data fusion according to claim 7, characterized in that, Also includes: The time synchronization module is used to align the timelines of multimodal data to ensure that the time stamps of all modal data are consistent.
9. The adolescent vision development trend prediction system based on multimodal data fusion according to claim 7, characterized in that, Also includes: The adaptive weight adjustment module is used to dynamically adjust the weight allocation of each modality during the training phase to optimize the model's overall fitting ability to multimodal data.
10. The adolescent vision development trend prediction system based on multimodal data fusion according to claim 7, characterized in that, Also includes: The anomaly detection module is used to identify and remove samples whose correlation significantly deviates from the normal range by calculating the correlation coefficient between data from different modalities.
Citation Information
Patent Citations
A method and system for predicting myopia in adolescents for vision correction.
CN117153407B
A method for predicting myopia vision for adolescent vision training
CN118512149B