Ocular surface state fusion algorithm based on multi-modal data
By constructing a multimodal data fusion algorithm, the problems of collaborative acquisition and feature alignment of multimodal data in ocular surface condition assessment were solved, enabling accurate quantitative assessment of key pathological indicators of the ocular surface and interpretable diagnostic results, thereby improving the accuracy and clinical applicability of ocular surface condition assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HENAN JINRUI MEDICAL EQUIPMENT CO LTD
- Filing Date
- 2026-01-12
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies suffer from inconsistencies in the collaborative acquisition of multimodal data, lack of cross-modal feature alignment, coarse modeling of ocular surface-specific indicators, and lack of interpretability in fusion decisions, resulting in insufficient accuracy and clinical applicability in ocular surface condition assessment.
A multimodal data ocular surface state fusion algorithm is constructed, including a multimodal synchronous acquisition module, an ocular surface specific feature extraction subnetwork, a cross-modal spatiotemporal dynamic alignment unit, a multi-scale fusion inference engine, and an interpretable output layer. Through high-dimensional feature collaborative extraction, cross-modal dynamic alignment, and intelligent fusion inference, the algorithm enables accurate quantitative assessment of key ocular surface pathological indicators such as tear film stability, epithelial integrity, and inflammation degree.
It achieves feature fusion of multi-source heterogeneous data under a unified spatiotemporal framework, improves the consistency and accuracy of ocular surface condition assessment, generates clinically interpretable diagnostic results, and breaks through the limitations of existing technologies.
Smart Images

Figure CN121867679A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of ophthalmic intelligent diagnostic technology, specifically an ocular surface state fusion algorithm based on multimodal data. Background Technology
[0002] With the continuous development of ophthalmic diagnostic and treatment technologies, accurate assessment of ocular surface condition is of great significance for the early screening and intervention of various eye diseases such as dry eye, corneal injury, and ocular surface inflammation. In recent years, multimodal data fusion technology has provided a new approach for the comprehensive analysis of ocular surface condition. By integrating multi-source information such as images, physiological signals, and tear film parameters, it is expected to achieve more accurate and personalized health assessments. However, existing technologies still have significant shortcomings in the collaborative acquisition, feature alignment, and fusion modeling of multimodal data, limiting the accuracy and clinical applicability of ocular surface condition assessment.
[0003] A search revealed a patent with publication number CN118942696B that proposes a multimodal recognition device for health screening based on eye imaging and eye movement tracking. This device can simultaneously acquire iris images, ocular surface images, fundus images, eye position images, and eye movement information, and determine the user's health status based on multimodal data. However, this solution mainly focuses on identity authentication and overall health screening. Its modeling of the ocular surface condition itself is relatively coarse, lacking fine-grained feature extraction for ocular surface-specific indicators (such as tear film stability, epithelial integrity, and degree of inflammation), and lacking a dynamic alignment mechanism for different modal data in the temporal and spatial dimensions, making it difficult to achieve high-precision identification of ocular surface pathological conditions.
[0004] A search revealed a patent with publication number CN120477690B, which discloses a wireless ocular surface pressure monitoring system based on a multi-frequency phase-shift decoupling self-calibration algorithm. This system integrates eyelid pressure and intraocular pressure sensors through a wearable contact lens to achieve real-time sensing of two parameters in a closed-eye state. Although this technology introduces simultaneous monitoring of multi-parameter physiological signals, its data modality is limited to mechanical sensing signals and does not integrate information from other key dimensions such as optical imaging, tear biochemistry, or eye movement behavior. This results in an incomplete depiction of the overall microenvironment of the ocular surface and an inability to fully reflect complex pathological changes such as dry eye, inflammation, or repair processes.
[0005] The aforementioned problems indicate that while existing technologies have made initial explorations into the application of multimodal data in eye health assessment, significant shortcomings remain in areas such as in-depth mining of ocular surface-specific features, spatiotemporal alignment of multi-source heterogeneous data, and the construction of fusion decision models tailored to clinical needs. Therefore, this invention proposes an ocular surface state fusion algorithm based on multimodal data. This algorithm aims to achieve accurate, comprehensive, and interpretable quantitative assessment of ocular surface health status through high-dimensional feature collaborative extraction, cross-modal dynamic alignment, and intelligent fusion reasoning, providing technical support for the early diagnosis and personalized intervention of diseases such as dry eye and corneal injury. Summary of the Invention
[0006] To address the problems in existing technologies, this invention provides an ocular surface state fusion algorithm based on multimodal data, solving technical deficiencies in current ocular surface state assessment methods such as inconsistent multimodal data collaborative acquisition, lack of cross-modal feature alignment, coarse modeling of ocular surface-specific indicators, and lack of interpretability in fusion decisions. This invention aims to achieve accurate quantitative assessment of key ocular surface pathological indicators such as tear film stability, epithelial integrity, and inflammation degree by constructing an algorithm system with spatiotemporal dynamic alignment capabilities, a high-dimensional fine-grained feature extraction mechanism, and an interpretable fusion inference structure.
[0007] The technical solution adopted by this invention to solve the above-mentioned technical problems is: an ocular surface state fusion algorithm based on multimodal data, including a multimodal synchronous acquisition module, an ocular surface specific feature extraction subnetwork, a cross-modal spatiotemporal dynamic alignment unit, a multi-scale fusion inference engine, and an interpretable output layer; the multimodal synchronous acquisition module includes a high-resolution slit-lamp imaging unit, a non-contact tear evaporation rate detection unit, a miniature eye-tracking camera, a wireless tear biochemical sensing patch, and a flexible eyelid mechanical sensing ring; the high-resolution slit-lamp imaging unit is installed in front of the examination table and is used to acquire a sequence of corneal anterior surface images at a frame rate of 30 frames per second; the non-contact tear evaporation rate detection unit integrates infrared thermal imaging... The sensor and humidity gradient array are installed 15 cm to the side of the slit lamp path, sharing the same trigger clock signal with the imaging unit; the miniature eye-tracking camera is embedded in the front edge of the lightweight headband worn by the subject, with a sampling frequency of 200 Hz; the wireless tear biochemical sensing patch is attached to the inner side of the lower eyelid conjunctival sac, and contains a three-channel electrochemical sensor for glucose, lactoferrin, and lysozyme, with a sampling interval of 5 seconds; the flexible eyelid mechanical sensing ring is embedded in the inner edge of a custom silicone eye mask, containing 6 distributed piezoresistive strain gauges, with a sampling frequency of 100 Hz; all acquisition units are connected to the central data buffer through hardware synchronous triggers to ensure that the timestamp error of each modality data does not exceed ±2 milliseconds.
[0008] The ocular surface-specific feature extraction subnetwork includes a tear film breakup time (TBUT) dynamic segmentation module, a corneal epithelial micro-damage detection module, an inflammatory factor concentration mapping module, and an eyelid closure dynamics modeling module. The tear film breakup time dynamic segmentation module adopts the U-Net++ architecture, taking a slit-lamp image sequence as input. It extracts the tear film edge contour through multi-scale skip connections, calculates the tear film area change rate between adjacent frames, and outputs a tear film stability index. The corneal epithelial micro-damage detection module is based on attention-enhanced ResNet-34 and processes slit-lamp fluorescein-stained images... Pixel-level classification is performed to identify punctate epithelial defect areas with a diameter of less than 50 micrometers, and an epithelial integrity score is output. The inflammatory factor concentration mapping module maps the raw current signal output from the wireless tear biochemical sensor patch to the molar concentration values of lactoferrin and lysozyme through a pre-trained multilayer perceptron after wavelet denoising, thus forming an inflammation degree vector. The eyelid closure dynamics modeling module receives 6-channel strain data from the flexible eyelid mechanical sensing ring, uses a long short-term memory network (LSTM) to model the temporal characteristics of pressure distribution during eyelid closure, and outputs blink integrity parameters.
[0009] The cross-modal spatiotemporal dynamic alignment unit includes a spatial coordinate normalization layer, a timestamp resampling buffer, and a cross-modal attention alignment matrix. The spatial coordinate normalization layer projects the slit-lamp image, eye-tracking point cloud, and mechanical sensor positions onto a spherical coordinate system with the corneal center as the origin. The slit-lamp image is converted into a polar coordinate grid after lens distortion correction by a calibration plate. The eye-tracking point cloud is mapped to the same sphere after being solved by the PnP algorithm. The positions of the six strain gauges of the mechanical sensor loop are preset with fixed coordinates based on the eye mask geometry model. The timestamp resampling buffer uses the frame rate of the slit-lamp image as a reference and performs cubic spline interpolation resampling on the tear biochemical data and mechanical sensor data to align all modal data to 30Hz in the time dimension. The cross-modal attention alignment matrix consists of a three-layer Transformer encoder. The input is the modal feature vectors after spatial normalization and temporal resampling. The correlation weights between modalities are calculated through a query-key-value mechanism to generate an aligned joint feature tensor.
[0010] The multi-scale fusion inference engine includes a low-order feature splicing layer, a mid-order graph neural network aggregation layer, and a high-order Bayesian inference layer. The low-order feature splicing layer directly splices the tear film stability index, epithelial integrity score, inflammation degree vector, and blink integrity parameter into a 128-dimensional vector. The mid-order graph neural network aggregation layer constructs a heterogeneous graph with ocular surface regions as nodes and physiological associations as edges. Node features include tear film thickness, epithelial damage density, and local inflammation concentration in each region. Edge weights are dynamically updated based on eyelid biomechanical distribution and blink frequency, and neighborhood information is aggregated through two-layer graph convolution operations. The high-order Bayesian inference layer pre-determines the prior probability distributions of three pathological states: dry eye, corneal damage, and ocular surface inflammation. Combined with the post-validation evidence output by the mid-order graph neural network, variational inference is used to calculate the confidence level of each pathological state.
[0011] The interpretable output layer includes a feature contribution heatmap generation module and a clinical indicator mapping table. The feature contribution heatmap generation module uses the gradient-weighted class activation mapping (Grad-CAM) method to backpropagate the gradients of each intermediate layer in the multi-scale fusion inference engine, generating a visualization image of the contribution of each ocular surface region to the final diagnostic result. The clinical indicator mapping table linearly maps the tear film stability index, epithelial integrity score, and inflammation degree vector output by the algorithm to the TBUT grade, Oxford staining score, and inflammation activity level in the International Dry Eye Workshop (DEWS II) standard, respectively.
[0012] Preferably, the cross-modal spatiotemporal dynamic alignment unit further includes an eye-tracking compensation displacement correction submodule. The eye-tracking compensation displacement correction submodule receives the pupil center coordinate sequence output by the miniature eye-tracking camera, calculates the translation and rotation offset between adjacent slit lamp image frames, and performs inverse transformation correction on subsequent frames through bilinear interpolation to eliminate image misalignment caused by micro-movements of the eyeball. This correction submodule is connected in series with the spatial coordinate normalization layer, and its output serves as one of the inputs of the normalization layer.
[0013] Preferably, the ocular surface-specific feature extraction subnetwork also includes a tear river height dynamic measurement module. The tear river height dynamic measurement module uses an active contour model to fit the tear river boundary based on the gray-level gradient abrupt change points between the lower eyelid margin and the tear river meniscus in the slit-lamp image sequence, calculates the vertical height of the tear river in each frame, and outputs a dynamic curve of tear secretion. This module shares the underlying convolutional feature map with the tear film breakup time dynamic segmentation module and outputs the tear river height sequence through an independent fully connected head.
[0014] Preferably, the multi-scale fusion inference engine also includes an uncertainty quantification module. The uncertainty quantification module uses the Monte Carlo Dropout method to perform 50 random samplings on the higher-order Bayesian inference layer during the inference phase, and calculates the standard deviation of the confidence of each pathological state as a reliability index of the diagnostic result. The output of this module is connected in parallel with the interpretability output layer, and its standard deviation is embedded in the clinical indicator mapping table as a confidence interval label.
[0015] The structural composition, implementation method, and operating principle of this invention are as follows: The multimodal synchronous acquisition module ensures strict alignment of five types of data—slit-lamp imaging, tear evaporation detection, eye tracking, tear biochemical sensing, and eyelid biomechanical sensing—on the time axis through hardware synchronization triggers, with timestamp errors controlled within ±2 milliseconds. The acquired raw data are input to corresponding modules of the ocular surface-specific feature extraction subnetwork. The slit-lamp image sequence is processed by the tear film breakup time dynamic segmentation module to output a tear film stability index, while the corneal epithelial micro-damage detection module outputs an epithelial integrity score. Tear biochemical sensing patch data is converted into an inflammation degree vector by an inflammatory factor concentration mapping module, and flexible eyelid biomechanical sensing ring data generates blink integrity parameters through an eyelid closure dynamics modeling module. The cross-modal spatiotemporal dynamic alignment unit first... The spatial coordinate normalization layer unifies modal data from different physical spaces into a spherical coordinate system with the corneal center as the origin. Then, the non-image modal data is interpolated to a 30Hz frame rate through a timestamp resampling buffer. Finally, the cross-modal attention alignment matrix calculates the intermodal correlation weights to generate a joint feature tensor. The multi-scale fusion inference engine sequentially executes low-order feature stitching, mid-order graph neural network aggregation, and high-order Bayesian inference. The graph neural network constructs a heterogeneous graph with ocular surface regions as nodes, and the edge weights are dynamically updated by the eyelid biomechanical distribution. The Bayesian inference layer combines prior pathological distribution and subsequent validation evidence to calculate the final diagnostic confidence. The interpretable output layer generates a feature contribution heatmap using the Grad-CAM method and linearly maps the algorithm output indicators to the DEWS II clinical standard, forming a diagnostic report that can be directly interpreted by ophthalmologists. Throughout the process, the eye movement compensation displacement correction submodule corrects image misalignment caused by micro-movements of the eyeball in real time, the tear river height dynamic measurement module supplements tear secretion dimension information, and the uncertainty quantification module provides diagnostic reliability indicators through Monte Carlo Dropout, together forming a complete closed-loop evaluation system.
[0016] The beneficial effects of this invention are as follows: 1. A cross-modal spatiotemporal dynamic alignment unit, through hardware-level time synchronization and spherical coordinate space normalization, solves the problem of misalignment in the spatiotemporal dimension of multi-source heterogeneous data, enabling feature fusion of modalities such as tear film images, mechanical sensing, and biochemical indicators within a unified spatiotemporal framework, significantly improving the consistency of ocular surface condition assessment. 2. An ocular surface-specific feature extraction sub-network, with dedicated deep learning modules designed for key clinical indicators such as tear film stability, epithelial integrity, and inflammation degree, achieves micron-level epithelial damage identification and molar-level inflammatory factor concentration mapping, overcoming the limitations of existing technologies in the coarse modeling of ocular surface pathological states. 3. A multi-scale fusion inference engine, combining graph neural networks and Bayesian inference, introduces physiological association priors between ocular surface regions while preserving the original semantics of each modality, generating clinically interpretable diagnostic results, which is superior to the decision transparency of existing black-box fusion models. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the overall system architecture of the ocular surface state fusion algorithm based on multimodal data of this invention;
[0018] Figure 2 This is a schematic diagram showing the structural layout and hardware connection relationship of the multimodal synchronous acquisition module;
[0019] Figure 3 This is a schematic diagram of the internal module composition and data flow of the ocular surface specific feature extraction subnetwork;
[0020] Figure 4 This is a schematic diagram of the processing flow and submodule connection relationship of the cross-modal spatiotemporal dynamic alignment unit;
[0021] Figure 5 This is a schematic diagram of the hierarchical structure and information fusion mechanism of a multi-scale fusion inference engine;
[0022] Figure 6 This is a schematic diagram of the functional modules of the interpretable output layer and their mapping relationship with clinical standards.
[0023] In the figure: 1. Multimodal synchronous acquisition module; 11. High-resolution slit-lamp imaging unit; 12. Non-contact tear evaporation rate detection unit; 121. Infrared thermal imaging sensor; 122. Humidity gradient array; 13. Miniature eye-tracking camera; 14. Wireless tear biochemical sensing patch; 141. Glucose electrochemical sensor; 142. Lactoferrin electrochemical sensor; 143. Lysozyme electrochemical sensor; 15. Flexible eyelid mechanical sensing ring; 151. Piezoresistive strain gauge; 16. Hardware synchronous trigger; 17. Central data buffer; 2. Ocular surface specific feature extraction subnetwork; 21. Tear film breakup time dynamic segmentation module; 22. 1. Corneal epithelial micro-damage detection module; 23. Inflammatory factor concentration mapping module; 24. Eyelid closure dynamic modeling module; 25. Tear river height dynamic measurement module; 3. Cross-modal spatiotemporal dynamic alignment unit; 31. Spatial coordinate normalization layer; 32. Timestamp resampling buffer; 33. Cross-modal attention alignment matrix; 34. Eye movement compensation displacement correction submodule; 4. Multi-scale fusion inference engine; 41. Low-order feature splicing layer; 42. Mid-order graph neural network aggregation layer; 43. High-order Bayesian inference layer; 44. Uncertainty quantification module; 5. Interpretability output layer; 51. Feature contribution heatmap generation module; 52. Clinical indicator mapping table. Detailed Implementation
[0024] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.
[0025] This invention provides an ocular surface state fusion algorithm based on multimodal data, addressing the technical shortcomings of existing ocular surface state assessment methods, such as inconsistent multimodal data acquisition, lack of cross-modal feature alignment, coarse modeling of ocular surface-specific indicators, and lack of interpretability in fusion decisions. This invention aims to achieve accurate quantitative assessment of key ocular surface pathological indicators such as tear film stability, epithelial integrity, and inflammation degree by constructing an algorithm system with spatiotemporal dynamic alignment capabilities, a high-dimensional fine-grained feature extraction mechanism, and an interpretable fusion inference structure.
[0026] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0027] See Figures 1 to 6 An ocular surface state fusion algorithm based on multimodal data includes 1, 2, 3, 4 and 5; wherein 1 includes 11, 12, 13, 14, 15, 16 and 17; 2 includes 21, 22, 23, 24 and 25; 3 includes 31, 32, 33 and 34; 4 includes 41, 42, 43 and 44; and 5 includes 51 and 52.
[0028] See Figure 2 11 is installed in front of the examination table to acquire a sequence of images of the anterior corneal surface at a frame rate of 30 frames per second; 12 integrates 121 and 122 and is installed 15 cm to the side of the optical path of 11, sharing the same trigger clock signal with 11; 13 is embedded in the front edge of the lightweight headband worn by the examinee, with a sampling frequency of 200 Hz; 14 is attached to the inner side of the lower eyelid conjunctival sac, containing 141, 142 and 143, with a sampling interval of 5 seconds; 15 is embedded in the inner edge of the custom silicone eye mask, containing 6 distributed 151, with a sampling frequency of 100 Hz; 16 connects 11, 12, 13, 14 and 15 and is connected to 17 to ensure that the timestamp error of each modality data does not exceed ±2 milliseconds.
[0029] See Figure 3 21. Using the U-Net++ architecture, the input is the image sequence output by 11. The tear film edge contour is extracted through multi-scale skip connections, and the tear film area change rate between adjacent frames is calculated to output the tear film stability index. 22. Based on attention-enhanced ResNet-34, the fluorescein-stained images acquired by 11 are classified at the pixel level to identify punctate epithelial defect areas with a diameter of less than 50 micrometers and output the epithelial integrity score. 23. The raw current signal output by 14 is denoising by wavelet and then mapped to the molar concentration values of lactoferrin and lysozyme through a pre-trained multilayer perceptron to form an inflammation degree vector. 24. The 6-channel strain data output by 15 is received, and the temporal features of pressure distribution during eyelid closure are modeled using LSTM to output the blink integrity parameter. 25. Based on the gray-level gradient abrupt change points of the lower eyelid margin and the meniscus of the tear duct in the image sequence of 11, an active contour model is used to fit the tear duct boundary. The vertical height of the tear duct is calculated for each frame, and the dynamic curve of tear secretion is output. 25 and 21 share the bottom-level convolutional feature map and output the tear duct height sequence through independent fully connected heads.
[0030] See Figure 434 receives the pupil center coordinate sequence output by 13, calculates the translation and rotation offset between adjacent image frames 11, and performs inverse transformation correction on subsequent frames through bilinear interpolation to eliminate image misalignment caused by micro-movements of the eyeball. The output of 34 is connected to 31. 31 projects the image 11, the trajectory point cloud of 13, and the position of 15 onto a spherical coordinate system with the corneal center as the origin. The image 11 is converted into a polar coordinate grid after lens distortion is corrected by a calibration plate. The trajectory point cloud of 13 is mapped to the same sphere after being solved by the PnP algorithm. The six positions of 151 of 15 are preset with fixed coordinates according to the eye mask geometric model. 32 uses the frame rate of 11 of 30Hz as a reference and performs cubic spline interpolation resampling on the data output by 14 and 15 to align all modal data to 30Hz in the time dimension. 33 consists of a three-layer Transformer encoder. The input is the modal feature vectors after temporal resampling by 31. The correlation weight between modalities is calculated through a query-key-value mechanism to generate an aligned joint feature tensor.
[0031] See Figure 5 41. The tear film stability index output from 21, the epithelial integrity score output from 22, the inflammation degree vector output from 23, and the blink integrity parameter output from 24 are directly concatenated into a 128-dimensional vector. 42. A heterogeneous graph is constructed with ocular surface regions as nodes and physiological associations as edges. The node features include the tear film thickness, epithelial damage density, and local inflammation concentration of each region. The edge weights are dynamically updated by the eyelid mechanical distribution and blink frequency output from 15. Neighborhood information is aggregated through two-layer graph convolution operations. 43. The prior probability distributions of three pathological states—dry eye, corneal damage, and ocular surface inflammation—are preset. Combined with the post-validation evidence output from 42, variational inference is used to calculate the confidence of each pathological state. 44. The Monte Carlo Dropout method is used to perform 50 random samplings on 43 during the inference phase. The standard deviation of the confidence of each pathological state is calculated as a reliability index of the diagnostic results. The output of 44 is connected in parallel with 5.
[0032] See Figure 6 51 Based on the Grad-CAM method, backpropagation is performed on the gradients of each intermediate layer in 4 to generate a visualization image of the contribution of each ocular surface region to the final diagnostic result; 52 The tear film stability index output by 21, the epithelial integrity score output by 22, and the inflammation degree vector output by 23 are linearly mapped to the TBUT grade, Oxford staining score and inflammation activity level in the DEWS II standard, respectively, and the standard deviation output by 44 is embedded in 52 as a confidence interval label.
[0033] During operation, the subject wears a customized headpiece containing components 13 and 15, with component 14 attached to the lower eyelid, and sits facing component 11. Activation 16 simultaneously triggers components 11, 12, 13, 14, and 15. Component 11 acquires corneal images at 30Hz, 12 records the tear evaporation rate, 13 tracks the pupil center at 200Hz, 14 outputs a biochemical current signal every 5 seconds, and 15 records eyelid mechanical strain at 100Hz. All raw data is synchronized by 16 and stored in component 17. Component 17 outputs data to component 2. Component 21 processes the image from component 11 to output the tear film stability index, 22 processes the fluorescein-stained image to output the epithelial integrity score, and 23 processes the signal from component 14 to output the inflammation degree vector. 4. Process data from step 15 to output blink integrity parameters, and 25. Simultaneously output tear river height sequence; 2. Output to step 3, 34. First, perform eye movement compensation correction on image 11, 31. Unify the corrected image, point cloud of step 13, and position of step 15 to spherical coordinate system, 32. Resample data from steps 14 and 15 to 30Hz, 33. Generate a cross-modal aligned joint feature tensor; This tensor is input to step 4, 41. Concatenate low-order features, 42. Construct ocular surface heterogeneity map and aggregate neighborhood information, 43. Calculate pathological confidence based on prior distribution, 44. Simultaneously output uncertainty standard deviation; Finally, 51. Generate heat map, 52. Map indicators to clinical standards and label confidence intervals to form a complete diagnostic report.
[0034] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of protection claimed by the present invention. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. An ocular surface state fusion algorithm based on multi-modal data, characterized in that, The system includes a multimodal synchronous acquisition module (1), an ocular surface-specific feature extraction subnetwork (2), a cross-modal spatiotemporal dynamic alignment unit (3), a multi-scale fusion inference engine (4), and an interpretable output layer (5). The multimodal synchronous acquisition module (1) includes a high-resolution slit-lamp imaging unit (11), a non-contact tear evaporation rate detection unit (12), a miniature eye-tracking camera (13), a wireless tear biochemical sensing patch (14), and a flexible eyelid mechanical sensing ring (15). The high-resolution slit-lamp imaging unit (11) acquires a sequence of corneal anterior surface images at a frame rate of 30 frames per second. The non-contact tear evaporation rate detection unit (12) integrates an infrared thermal imaging sensor (121) and a humidity gradient array (122). The device is installed 15 cm to the side of the optical path of the high-resolution slit lamp imaging unit (11) and shares the same trigger clock signal with the high-resolution slit lamp imaging unit (11); the sampling frequency of the miniature eye-tracking camera (13) is 200 Hz; the wireless tear biochemical sensing patch (14) has a built-in glucose electrochemical sensor (141), lactoferrin electrochemical sensor (142) and lysozyme electrochemical sensor (143), with a sampling interval of 5 seconds; the flexible eyelid mechanical sensing ring (15) contains 6 distributed piezoresistive strain gauges (151), with a sampling frequency of 100 Hz; all acquisition units are connected to the central data buffer (17) through a hardware synchronization trigger (16), and the timestamp error of each modal data does not exceed ±2 milliseconds.
2. The ocular surface state fusion algorithm based on multi-modal data according to claim 1, characterized in that, The ocular surface-specific feature extraction subnetwork (2) includes a tear film breakup time dynamic segmentation module (21), a corneal epithelial micro-damage detection module (22), an inflammatory factor concentration mapping module (23), an eyelid closure dynamic modeling module (24), and a tear river height dynamic measurement module (25). The tear film breakup time dynamic segmentation module (21) adopts the U-Net++ architecture, and the input is the image sequence output by the high-resolution slit-lamp imaging unit (11). The corneal epithelial micro-damage detection module (22) is based on attention-enhanced ResNet-34 and processes high-resolution slit-lamp imaging. The unit (11) collects fluorescein staining images; the inflammatory factor concentration mapping module (23) receives the raw current signal output by the wireless tear biochemical sensing patch (14); the eyelid closure dynamic modeling module (24) receives 6-channel strain data output by the flexible eyelid mechanical sensing ring (15); the tear river height dynamic measurement module (25) uses an active contour model to fit the tear river boundary based on the gray-scale gradient abrupt change points between the lower eyelid margin and the tear river meniscus in the image sequence of the high-resolution slit lamp imaging unit (11), and shares the underlying convolutional feature map with the tear film breakup time dynamic segmentation module (21).
3. The ocular surface state fusion algorithm based on multi-modal data according to claim 1, wherein, The cross-modal spatiotemporal dynamic alignment unit (3) includes an eye-tracking compensation displacement correction submodule (34), a spatial coordinate normalization layer (31), a timestamp resampling buffer (32), and a cross-modal attention alignment matrix (33); the eye-tracking compensation displacement correction submodule (34) receives the pupil center coordinate sequence output by the miniature eye-tracking camera (13), calculates the translation and rotation offset between adjacent high-resolution slit-lamp imaging units (11) image frames, and performs inverse transformation correction on the image frames through bilinear interpolation; the spatial coordinate normalization layer ( 31) Project the corrected high-resolution slit lamp imaging unit (11) image, the trajectory point cloud of the miniature eye-tracking camera (13) and the position of the flexible eyelid mechanical sensing ring (15) onto a spherical coordinate system with the corneal center as the origin; the timestamp resampling buffer (32) performs cubic spline interpolation resampling on the data output by the wireless tear biochemical sensing patch (14) and the flexible eyelid mechanical sensing ring (15) with 30Hz as the reference; the cross-modal attention alignment matrix (33) is composed of three layers of Transformer encoders.
4. The ocular surface state fusion algorithm based on multi-modal data according to claim 3, characterized in that, The spatial coordinate normalization layer (31) converts the image of the high-resolution slit lamp imaging unit (11) into a polar coordinate grid after the lens distortion is corrected by the calibration plate. The trajectory point cloud of the miniature eye-tracking camera (13) is mapped to the spherical coordinate system after being solved by the PnP algorithm. The positions of the six piezoresistive strain gauges (151) of the flexible eyelid mechanical sensing ring (15) are fixed according to the preset coordinates of the customized silicone eye mask geometric model.
5. The ocular surface state fusion algorithm based on multi-modal data according to claim 1, wherein, The multi-scale fusion inference engine (4) includes a low-order feature splicing layer (41), a mid-order graph neural network aggregation layer (42), a high-order Bayesian inference layer (43), and an uncertainty quantification module (44). The low-order feature splicing layer (41) splices the tear film stability index, epithelial integrity score, inflammation degree vector, and blink integrity parameter into a 128-dimensional vector. The mid-order graph neural network aggregation layer (42) constructs a heterogeneous graph with ocular surface regions as nodes and physiological associations as edges. The node features include tear film thickness, epithelial damage density, and local inflammation concentration. The edge weights are dynamically updated by the eyelid mechanical distribution and blink frequency output by the flexible eyelid mechanical sensing ring (15). The high-order Bayesian inference layer (43) presets the prior probability distribution of three pathological states: dry eye, corneal damage, and ocular surface inflammation. The uncertainty quantification module (44) uses the Monte Carlo Dropout method to perform 50 random samplings on the high-order Bayesian inference layer (43) during the inference stage.
6. The ocular surface state fusion algorithm based on multimodal data according to claim 5, characterized in that, The intermediate-order graph neural network aggregation layer (42) aggregates neighborhood information through two layers of graph convolution operations.
7. The ocular surface state fusion algorithm based on multimodal data according to claim 5, characterized in that, The uncertainty quantification module (44) outputs the standard deviation of the confidence level of each pathological state, which is connected in parallel with the interpretability output layer (5).
8. The ocular surface state fusion algorithm based on multimodal data according to claim 1, characterized in that, The interpretability output layer (5) includes a feature contribution heatmap generation module (51) and a clinical indicator mapping table (52). The feature contribution heatmap generation module (51) performs backpropagation on the gradient of each intermediate layer in the multi-scale fusion inference engine (4) based on the gradient weighted class activation mapping method. The clinical indicator mapping table (52) linearly maps the tear film stability index, epithelial integrity score and inflammation degree vector to the TBUT grade, Oxford staining score and inflammation activity level in the International Dry Eye Workshop (DEWS II) standard, respectively.
9. The ocular surface state fusion algorithm based on multimodal data according to claim 8, characterized in that, The clinical indicator mapping table (52) embeds the standard deviation output by the uncertainty quantification module (44) as the confidence interval label.
10. The ocular surface state fusion algorithm based on multimodal data according to claim 2, characterized in that, The tear river height dynamic measurement module (25) outputs the tear river height sequence through an independent full connector.
Citation Information
Patent Citations
Multimodal recognition device for health screening based on eye imaging and eye tracking
CN118942696B
Wireless ocular surface pressure monitoring system based on multi-frequency phase shift decoupling self-calibration algorithm
CN120477690B