Ophthalmic image information processing system and ocular light exposure device
Patent Information
- Application Number
- CN202610965574.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
AI Technical Summary
相关技术中,主要依赖临床医师经验判断,缺乏客观量化的辅助手段,导致评估结果主观性强、一致性差,难以准确反映用户真实的眼部状态
[0015]本申请实施例中,对于眼科图像信息处理系统,通过获取模块采集第一用户的多模态数据,利用预先训练的多模态神经网络模型对眼科图像、生理参数及光生物调节参数进行联合处理,生成与视觉功能相关的量化参数,实现了对用户眼部状态的客观量化评估,以数据驱动方式替代传统经验判断,能够提高视觉功能相关指标评估的准确度和一致性。
Smart Images

Figure CN122822243A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of medical image processing, specifically relating to an ophthalmic image information processing system and an ophthalmic illumination device. Background Technology
[0002] Photobiomodulation (PBM) technology uses specific wavelengths of light to irradiate retinal tissue, showing potential applications in ophthalmology fields such as myopia control and retinal disease treatment. However, there are significant differences in individual responses to PBM irradiation.
[0003] Visual function indicators (including visual acuity, axial length, and choroidal thickness) are core criteria for assessing ocular condition. Accurately extracting the quantitative characteristics of these indicators from ophthalmic image data is of significant reference value for clinical assessment. However, current techniques primarily rely on the experience of clinicians, lacking objective quantitative aids. This leads to highly subjective and inconsistent assessment results, making it difficult to accurately reflect the user's true ocular condition. Summary of the Invention
[0004] The purpose of this application is to provide an ophthalmic image information processing system and an ophthalmic illumination device that can improve the accuracy and consistency of visual function-related index assessment.
[0005] According to a first aspect of the embodiments of this application, an ophthalmic image information processing system is provided, the system comprising: The acquisition module is used to acquire multimodal data of the first user; wherein, the multimodal data includes ophthalmic images, physiological parameters, and photobiological regulation parameters; The processing module is used to process the multimodal data through a pre-trained multimodal neural network model to obtain the quantitative parameters related to the visual function of the first user. The output module is used to output the quantization parameters; The multimodal neural network model includes: a feature extraction network for extracting features of multiple modalities from the multimodal data; an attention fusion module for fusing the features of the multiple modalities through an attention mechanism to obtain fused features; and a prediction head for predicting the quantization parameters based on the fused features.
[0006] Optionally, as an embodiment, the ophthalmic images include: tomographic images and fundus images; The physiological parameters include at least one of the following: age, sex, refractive error, intraocular pressure, and axial length; The photobiological regulation parameters include at least one of the following: wavelength, power density, irradiation duration, and number of irradiations.
[0007] Optionally, as an embodiment, the quantification parameters include at least one of the following: visual acuity value, axial length value, and choroid thickness value.
[0008] Optionally, as an embodiment, the multimodal data further includes: visual field inspection results; The visual field inspection results include at least one of the following: light sensitivity value matrix, average deviation, pattern deviation, and visual field index.
[0009] Optionally, as an embodiment, the feature extraction network includes: The first feature extraction branch uses a 3D convolutional neural network to extract the first feature from the tomographic scan image; The second feature extraction branch uses a 2D convolutional neural network to extract the second feature from the fundus image; The third feature extraction branch uses a multilayer perceptron to extract the third feature from the physiological parameters; The fourth feature extraction branch uses a multilayer perceptron to extract the fourth feature from the photobiological regulatory parameters.
[0010] Optionally, as an embodiment, the multimodal neural network model further includes: a temporal encoder connected between the attention fusion module and the prediction head; The temporal encoder is used to construct a temporal sequence from multiple fused features corresponding to the baseline time point and / or follow-up time point output by the attention fusion module, and to obtain a temporal feature vector by performing temporal feature encoding processing through an encoder based on a self-attention mechanism. The prediction head is specifically used to predict quantization parameters at multiple future time points based on the time-series feature vector.
[0011] Optionally, as an embodiment, the system further includes: An uncertainty quantization module, connected to the processing module, is used to perform multiple forward propagations on the prediction results of the multimodal neural network model using the Monte Carlo Dropout algorithm or a deep ensemble algorithm, and to calculate the confidence interval of the quantization parameter based on the statistical distribution information of the multiple forward propagation results.
[0012] Optionally, as an embodiment, the system further includes: A visualization and interpretation module, connected to the processing module, is used for at least one of the following: Based on the intermediate layer feature maps and gradient information of the prediction results of the multimodal neural network model, a heatmap for the ophthalmic image is generated, and the image region that contributes the most to the quantization parameter in the heatmap is marked. Calculate the sensitivity of each input feature in the multimodal data to the quantization parameter; Generate natural language interpretation reports; The quantification parameters, heatmaps, sensitivity, and natural language interpretation reports are displayed graphically.
[0013] According to a second aspect of the embodiments of this application, an eye illumination device is provided, the device comprising: The light source module is used to emit illumination light. The ophthalmic image information processing system described in the first aspect; The control module is connected to both the ophthalmic image information processing system and the light source module, and is used to adjust the illumination parameters of the light source module according to the quantization parameters output by the ophthalmic image information processing system.
[0014] According to a third aspect of the embodiments of this application, a computer-readable storage medium is provided, storing a computer program that, when executed by a processor, implements the functions of the ophthalmic image information processing system described in the first aspect, or the functions of the ophthalmic illumination device described in the second aspect.
[0015] In this embodiment of the application, for the ophthalmic image information processing system, the acquisition module collects multimodal data of the first user, and uses a pre-trained multimodal neural network model to jointly process ophthalmic images, physiological parameters and photobiological regulation parameters to generate quantitative parameters related to visual function. This realizes an objective quantitative assessment of the user's eye condition, and replaces traditional experience-based judgment with a data-driven approach, which can improve the accuracy and consistency of the assessment of visual function-related indicators.
[0016] In this embodiment of the application, for the ophthalmic illumination device, by integrating the ophthalmic image information processing system and the light source module into the same device, the control module can adjust the illumination parameters of the light source module in real time according to the quantitative parameters output by the ophthalmic image information processing system, thereby realizing the automated connection from ophthalmic condition assessment to illumination scheme adjustment, and providing an objective adjustment basis for personalized illumination. Attached Figure Description
[0017] Figure 1 This is one of the structural schematic diagrams of an ophthalmic image information processing system provided in some embodiments of this application; Figure 2 This is one of the network structure diagrams of a multimodal neural network model provided in some embodiments of this application; Figure 3 This is the second network structure diagram of a multimodal neural network model provided in some embodiments of this application; Figure 4This is a second schematic diagram of the structure of an ophthalmic image information processing system provided in some embodiments of this application; Figure 5 These are schematic diagrams of the structure of an eye-illuminating device provided in some embodiments of this application; Figure 6 This is a flowchart of a control method for an ophthalmic image information processing system provided in some embodiments of this application; Figure 7 This is a flowchart of an eye illumination device control method provided in some embodiments of this application. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] It should be understood that the terms "comprising" and "including" used in the specification and claims of this application indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0020] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this specification and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this specification and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.
[0021] As used in this specification and claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if the described condition or event is detected" may be interpreted, depending on the context, as meaning "once determined."
[0022] To facilitate understanding, the relevant concepts and application scenarios involved in the embodiments of this application will be introduced below.
[0023] I. Related Concepts Photobiomodulation (PBM) refers to a non-invasive intervention method that uses light of specific wavelengths (usually in the red to near-infrared range, such as 630nm to 680nm or 810nm to 850nm) to irradiate biological tissues. Through intracellular photobiomodulation effects, it promotes mitochondrial function, improves cell metabolism, and regulates local blood perfusion, thereby influencing tissue function. In ophthalmology, photobiomodulation is mainly used for myopia control and retinal function maintenance.
[0024] An optical coherence tomography (OCT) image is a three-dimensional structural image of ocular tissues obtained through optical coherence tomography. This technique utilizes the principle of low-coherence optical interference to perform high-resolution imaging of various layers of the retina and the morphology of the choroid, revealing the morphological characteristics of each layer, including the retinal nerve fiber layer, ganglion cell layer, inner nuclear layer, outer nuclear layer, photoreceptor layer, and retinal pigment epithelium, as well as the distribution of choroidal vessels.
[0025] Fundus images are two-dimensional images of the retinal surface acquired using a fundus camera. These images can display the distribution of the retinal vascular network, the morphology of the optic disc, the structure of the macular region, and the state of the pigment epithelium, serving as important evidence for assessing the surface structure and vascular condition of the retina.
[0026] II. Application Scenarios The ophthalmic image information processing system and ocular illumination device provided in this application are mainly used in the quantitative evaluation and personalized parameter adjustment of photobiological irradiation effects.
[0027] In clinical practice, significant differences exist in individual responses to photobiological irradiation. By collecting ophthalmic images, physiological parameters, and irradiation protocol parameters of users before or during irradiation, and inputting these parameters into a pre-trained multimodal neural network model, the system can output quantitative parameters related to visual function, providing an objective reference for assessing the user's eye condition. Based on these quantitative parameters, the ophthalmic irradiation device can adaptively adjust the irradiation parameters to achieve personalized irradiation protocols.
[0028] The embodiments of this application can be applied to the following specific scenarios: Eye condition assessment and irradiation parameter adjustment in myopia prevention and control; Quantitative assessment of the effectiveness of retinal function maintenance; Rapid assessment of the applicability of irradiation in primary screening; Dynamic monitoring and trend analysis during multi-time point follow-up.
[0029] The ophthalmic image information processing system and ocular illumination device provided in this application will now be described in further detail with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application.
[0030] Figure 1 This is one of the structural schematic diagrams of an ophthalmic image information processing system provided in some embodiments of this application, such as... Figure 1 As shown, the ophthalmic image information processing system 100 may include: an acquisition module 101, a processing module 102, and an output module 103.
[0031] The acquisition module 101 is used to acquire the multimodal data of the first user; wherein, the multimodal data includes ophthalmic images, physiological parameters and photobiological regulation parameters.
[0032] In this embodiment of the application, the first user refers to the user who receives eye image acquisition and photobiological modulation irradiation.
[0033] In this embodiment of the application, ophthalmic images are used to characterize the structural information of the eye tissue of the first user.
[0034] In this embodiment of the application, physiological parameters are used to reflect the individual characteristics and basic eye condition of the first user.
[0035] In this embodiment, the photobiological adjustment parameters are used to characterize the specific settings of the first user's irradiation scheme.
[0036] In this embodiment of the application, by jointly inputting ophthalmic images, physiological parameters and photobiological regulation parameters, the model can obtain complementary information in three dimensions: structure, individual state and irradiation conditions, thereby improving the comprehensiveness and accuracy of quantitative parameter evaluation.
[0037] In some embodiments, ophthalmic images may include tomographic images and fundus images. Tomographic images, acquired using optical coherence tomography (OCT), can visualize the various layers of the retina and the morphology of the choroid; fundus images, acquired using a fundus camera, are used to display the distribution of retinal vessels, the morphology of the optic disc, and the state of the pigment epithelium. The combination of these two images can complementaryly characterize ocular tissue features from both structural depth and surface morphology dimensions.
[0038] In some embodiments, physiological parameters may include at least one of the following: age, sex, refractive error, intraocular pressure, and axial length. Age and sex are demographic characteristics, refractive error reflects the degree of myopia or hyperopia, intraocular pressure characterizes the intraocular pressure state, and axial length reflects the developmental process of the eyeball. These parameters can provide the model with information on individual differences related to visual function, helping to improve the personalization of the assessment results.
[0039] In some embodiments, photobiological modulation parameters may include at least one of the following: wavelength, power density, irradiation duration, and number of irradiations. For example, wavelengths are typically in the red light band of 630 nm to 680 nm or the near-infrared band of 810 nm to 850 nm, and power densities are typically around 8 mW / cm². 2 Up to 50mW / cm 2 Adjustable within a certain range. Photobiological adjustment parameters are used to characterize the specific settings of the irradiation scheme, enabling the model to output corresponding quantitative parameters according to different irradiation conditions, and realizing the correlation modeling between the irradiation scheme and the evaluation results.
[0040] The processing module 102 is connected to the acquisition module 101 and is used to process the multimodal data of the first user through a pre-trained multimodal neural network model to obtain the quantitative parameters related to the visual function of the first user.
[0041] In this embodiment, the quantitative parameters related to visual function of the first user are used to objectively characterize the eye condition and can serve as a quantitative reference for clinical evaluation.
[0042] In some embodiments, the quantitative parameters related to visual function may include at least one of the following: visual acuity, axial length, and choroidal thickness. Visual acuity characterizes the state of visual function, axial length characterizes the state of ocular structural development, and choroidal thickness characterizes the state of choroidal blood perfusion. These three parameters together constitute a quantitative indicator system for ocular status from both functional and structural dimensions, providing multi-faceted reference for clinical assessment.
[0043] The output module 103 is connected to the processing module 102 and is used to output the quantization parameters related to the visual function of the first user.
[0044] In this embodiment, the output module can output the first user's visual function-related quantitative parameters in numerical form, and present the data through a visualization interface or transmit it to an external device in a structured data format.
[0045] In some embodiments, the multimodal data of the first user may further include: visual field test results of the first user, wherein the visual field test results are used to characterize the retinal functional state of the first user.
[0046] In this embodiment, the visual field examination results may include at least one of the following: a light sensitivity value matrix, an average deviation, a pattern deviation, and a visual field index. The light sensitivity value matrix is two-dimensional grid data reflecting the sensitivity of each test point to light stimuli; the average deviation is used to quantify the overall degree of visual field impairment; the pattern deviation is used to identify local visual field defect features; and the visual field index is used to characterize the percentage of central visual field function.
[0047] In this embodiment of the application, since the visual field examination results can provide information on the distribution of functional sensitivity in different areas of the retina, supplementing the local functional impairment features that cannot be directly reflected by structural images from a functional perspective, the model can obtain multi-dimensional data input that integrates structure and function.
[0048] As can be seen from the above embodiments, in this embodiment, the multimodal data of the first user is collected by the acquisition module, and the ophthalmic images, physiological parameters and photobiological regulation parameters are jointly processed by the pre-trained multimodal neural network model to generate quantitative parameters related to visual function. This realizes an objective quantitative assessment of the eye condition, and replaces traditional experience judgment with a data-driven approach, which can improve the accuracy and consistency of the assessment of visual function-related indicators.
[0049] Figure 2 This is one of the network structure diagrams of a multimodal neural network model provided in some embodiments of this application, such as... Figure 2 As shown, the multimodal neural network model 200 may include: a feature extraction network 201, an attention fusion module 202, and a prediction head 203.
[0050] Feature extraction network 201 is used to extract features of multiple modalities from the multimodal data of the first user.
[0051] In this embodiment, the feature extraction network can adopt an appropriate branch structure for different types of data in order to fully capture the deep information related to visual function in each modality of data.
[0052] In some embodiments, the feature extraction network may include a first feature extraction branch, a second feature extraction branch, a third feature extraction branch, and a fourth feature extraction branch.
[0053] The first feature extraction branch employs a 3D convolutional neural network to extract the first feature from the tomographic scan image. The tomographic scan image is three-dimensional volumetric data, containing morphological information of various retinal layers (such as the nerve fiber layer, ganglion cell layer, inner nuclear layer, outer nuclear layer, photoreceptor layer, and retinal pigment epithelium layer) and the spatial distribution of the choroidal structure. The 3D convolutional neural network simultaneously extracts structural features of the image in both spatial (horizontal and vertical) and depth (retinal layers) dimensions through three-dimensional convolutional kernels, outputting a high-dimensional feature map. Each feature vector corresponds to a retinal structural representation at a specific location and depth, such as changes in retinal layer thickness in a certain region, choroidal vessel morphology, and the presence or absence of subretinal fluid.
[0054] The second feature extraction branch uses a 2D convolutional neural network to extract second features from fundus images. Fundus images are two-dimensional planar images containing morphological information of the retinal vascular network, optic disc region, macula, and pigment epithelium. The 2D convolutional neural network extracts visual features from low to high order through layer-by-layer convolution and pooling operations. Low-order features correspond to local structures such as vessel edges, optic disc boundaries, and fovea, while high-order features correspond to global morphological features such as vessel distribution patterns, optic disc tilt direction, pigment deposition distribution, and macular morphological abnormalities. The output is a feature vector, where each dimension corresponds to a quantitative representation of a specific fundus structure.
[0055] The third feature extraction branch uses a multilayer perceptron to extract third features from physiological parameters. Physiological parameters are one-dimensional numerical vectors, including age, sex, refractive error, intraocular pressure, and axial length. The multilayer perceptron performs a nonlinear transformation on the original parameters through a multilayer fully connected network, mapping discrete or continuous original values to a high-dimensional continuous feature space. The output is a feature vector, where each dimension corresponds to an abstract representation of the combination of physiological parameters, such as the interaction effect between age and axial length, or the joint distribution characteristics of refractive error and intraocular pressure.
[0056] The fourth feature extraction branch employs a multilayer perceptron to extract a fourth feature from photobiological regulatory parameters. These parameters include wavelength, power density, irradiation duration, and number of irradiations. The multilayer perceptron encodes these parameters into conditional feature vectors, enabling the model to distinguish the expected response under different irradiation schemes. The output is a feature vector, where each dimension corresponds to an abstract representation of the combination of irradiation parameters, such as the synergistic effect of wavelength and power density, and the joint features of cumulative irradiation dose and irradiation frequency.
[0057] As can be seen, in this embodiment of the application, through the multi-branch structure of the feature extraction network, the model can extract complementary features from heterogeneous data, providing information-rich feature representations for subsequent fusion.
[0058] In some embodiments, the feature extraction network may employ a lightweight network structure, such as MobileNetV3, ShuffleNet, or SqueezeNet, to adapt to mobile deployment scenarios.
[0059] The attention fusion module 202 is connected to the feature extraction network 201 and is used to fuse features from multiple modalities through an attention mechanism to obtain fused features. In this embodiment, the attention mechanism of the attention fusion module can adaptively learn the correlation weights between features of different modalities, enabling the model to focus on feature dimensions that contribute more to the final prediction and suppress redundant or irrelevant information.
[0060] In some embodiments, the attention fusion module may employ a multi-head cross-attention mechanism to achieve the fusion of multimodal features.
[0061] Taking the fusion of tomographic image features and fundus image features as an example, the tomographic image features are used as the query, and the fundus image features as the key and value. After performing a dot product operation between the query and the key, the result is scaled and normalized using Softmax to obtain an attention weight matrix. This weight matrix reflects the correlation between each region in the tomographic image and each region in the fundus image. Multiplying the attention weights by the values yields a weighted feature representation, enabling the model to retrieve relevant complementary information from the fundus image based on the features of the tomographic image. The multi-head mechanism replicates this process as multiple parallel attention heads, each using an independent linear projection matrix to calculate attention weights from different feature subspaces. The outputs of each head are concatenated and restored to the original dimension through linear projection. The multi-head mechanism can capture various types of correlations between different modalities, such as the correspondence between structural features and vascular features, and the synergistic relationship between layer thickness and pigment distribution.
[0062] In this embodiment, the attention fusion module employs a similar mechanism when fusing physiological parameter features and photobiological regulatory parameter features. Physiological parameter features are used as queries, and irradiation parameter features are used as keys and values, enabling the model to learn the influence weights of irradiation parameters on quantified parameters under different physiological states. For example, the prediction results corresponding to the same set of irradiation parameters may differ for users of different ages or axial lengths; the attention mechanism can automatically capture this interaction effect.
[0063] In this embodiment, when there are three or more modalities, the attention fusion module can adopt a cascaded or parallel fusion method. The cascaded method refers to sequentially fusing the features of each modality pairwise to gradually generate a comprehensive feature; the parallel method refers to simultaneously inputting all modal features into the multi-head cross-attention module to calculate the correlation matrix between all modalities. Both methods can achieve deep fusion of multimodal information and output a comprehensive feature vector.
[0064] As can be seen, in the embodiments of this application, through the attention fusion module, the model can break through the linear limitation of traditional weighted fusion or splicing fusion, and dynamically adjust the contribution weight of each modality feature in a nonlinear manner, making the fused features more discriminative.
[0065] The prediction head 203 is connected to the attention fusion module 202 and is used to predict quantization parameters based on the fusion features.
[0066] In this embodiment, the prediction head can adopt a multi-layer fully connected network structure to gradually map high-dimensional fused features to the target output space.
[0067] In this embodiment, the fusion feature vector output by the attention fusion module is a high-dimensional continuous representation, typically with a dimension of 256 or 512. This feature vector encodes a comprehensive representation of multi-source information, including tomographic images, fundus images, physiological parameters, and photobiological regulatory parameters. The prediction head transforms this feature vector layer by layer through sequentially connected fully connected layers, achieving the conversion from abstract features to specific quantitative indicators.
[0068] In some embodiments, the prediction head can consist of two or three stacked fully connected layers. Taking a three-layer structure as an example, the first fully connected layer reduces the input feature vector from 256 dimensions to 128 dimensions and introduces nonlinear transformation capability using the ReLU activation function; the second fully connected layer reduces the 128-dimensional features to 64 dimensions, also using the ReLU activation function; the third fully connected layer reduces the 64-dimensional features to the output dimension, which corresponds to the number of target quantization parameters. For single-task prediction that only outputs visual acuity values, the output dimension is 1; for multi-task prediction that simultaneously outputs visual acuity values, axial length values, and choroidal thickness values, the output dimension is 3. The third layer typically does not use an activation function to keep the numerical range of the output values unconstrained, allowing the predicted values to cover the true continuous numerical range.
[0069] In this embodiment, during the training phase, the prediction head, feature extraction network, and attention fusion module jointly participate in end-to-end optimization. The loss function can be Mean Squared Error (MSE) or Mean Absolute Error (MAE) to measure the deviation between the predicted and true values. For multi-task prediction scenarios, the total loss function is a weighted sum of the losses from each task. The weights can be set as hyperparameters or automatically learned through an uncertainty-weighted method, allowing the model to adaptively allocate optimization resources across different prediction objectives, preventing the loss of any one task from dominating the overall optimization process.
[0070] In this embodiment of the application, the prediction head realizes a nonlinear mapping from multimodal fusion features to visual function quantification indicators through the above structure, providing an output interface for the end-to-end evaluation system.
[0071] As can be seen from the above embodiments, in this embodiment, the feature extraction network extracts deep features of heterogeneous modalities from ophthalmic images, physiological parameters and photobiological regulation parameters respectively, and then the attention fusion module performs adaptive weighted fusion of the features of each modality. Finally, the prediction head maps the fused features to the target quantization parameters. This end-to-end architecture can fully explore the complementary information in multi-source data, transform the original images and parameters into objective numerical indicators related to visual function, and realize the overall mapping from multimodal input to quantization output.
[0072] Figure 3This is a second network structure diagram of a multimodal neural network model provided in some embodiments of this application, such as... Figure 3 As shown, the multimodal neural network model 200 may further include: a temporal encoder 204, which is connected between the attention fusion module 202 and the prediction head 203.
[0073] The temporal encoder 204 is used to construct a temporal sequence from multiple fused features corresponding to the baseline time point and / or follow-up time point output by the attention fusion module 202, and to obtain a temporal feature vector by performing temporal feature encoding processing through an encoder based on a self-attention mechanism.
[0074] Considering that in clinical practice, the same user often undergoes multiple examinations at different time points, in this embodiment, the baseline time point reflects the initial state before irradiation, and the follow-up time point reflects the dynamic changes during the irradiation process. The temporal encoder organizes the fused features of these discrete time points into sequence data, uses a self-attention mechanism to capture the dependencies between time points, and extracts trend features.
[0075] In some embodiments, the temporal encoder may employ a Transformer encoder structure, wherein the Transformer encoder consists of multiple identical encoder layers stacked together, each encoder layer containing two sub-layers: a multi-head self-attention sub-layer and a feedforward neural network sub-layer. Each sub-layer is followed by a residual connection and a layer normalization operation.
[0076] For example, let f1, f2, ..., f be the fusion features corresponding to a user at T time points (e.g., baseline, month 1 follow-up, month 3 follow-up, month 6 follow-up). T Each fused feature has a dimension of d. The temporal encoder first performs positional encoding on these features, embedding the sequence information or actual time interval information of each time point into the feature vector. Positional encoding can use sine and cosine functions to generate fixed codes, or it can use a learnable positional embedding layer.
[0077] The position-encoded feature sequence is input into a multi-head self-attention sublayer. The self-attention mechanism calculates the correlation weights between any two positions in the sequence, enabling the model to capture dependencies between different time points. For example, when encoding features from the third month follow-up, it can simultaneously focus on the baseline state and the state from the first month follow-up, extracting information on changing trends. The multi-head mechanism replicates the attention computation as multiple parallel heads, each learning temporal dependencies from a different subspace. The outputs of all heads are concatenated and fused using linear projection.
[0078] The output of the multi-head self-attention sublayer is fed into the feedforward neural network sublayer after residual connections and layer normalization. The feedforward neural network sublayer consists of two fully connected layers. The first layer expands the dimension from d to 4d, activates it with ReLU, and then compresses it back to d by the second layer. This sublayer introduces non-linear transformation capability into the encoder, enhancing its feature representation ability.
[0079] After layer-by-layer transformation through multiple encoder layers (typically 3 to 6 layers), the temporal encoder outputs the encoded feature sequence z1, z2, ..., z T In some embodiments, the output z corresponding to the last time point can be... T As a temporal feature vector, this vector integrates information from all historical time points; in other embodiments, the outputs at each time point can be averaged or max-pooled to obtain a global temporal feature vector.
[0080] Prediction head 203 is specifically used to predict quantization parameters for multiple future time points based on time-series feature vectors.
[0081] In this embodiment, with the introduction of a temporal encoder, the input to the prediction head is replaced by the fused features at a single time point, which are then replaced by the temporal feature vector output by the temporal encoder. The prediction head can still employ a multi-layer fully connected network structure, with its output dimension expanded to the product of multiple future time points and multiple quantization parameters.
[0082] Specifically, if we need to predict K quantized parameters (such as visual acuity, axial length, and choroidal thickness) at M future time points (e.g., 3 months, 6 months, and 12 months after irradiation), then the output dimension of the prediction head is M×K. The prediction head maps the temporal feature vector to this multi-dimensional output space through a fully connected network. The output values are arranged in a predetermined order; for example, the first K outputs correspond to the K quantized parameters at the first time point, the middle K outputs correspond to the K quantized parameters at the second time point, and so on.
[0083] In this embodiment, during training, the loss function simultaneously considers the prediction errors of all prediction time points and all quantization parameters. The total loss function can be expressed as a weighted sum of the prediction errors of each time point and each quantization parameter, such as the sum of mean squared error or mean absolute error. Through end-to-end optimization, the time encoder and the prediction head jointly learn how to extract trends from historical data and extrapolate them to future time points.
[0084] As can be seen, in this embodiment of the application, by introducing a time encoder into the model, the time encoder transforms the fusion features of discrete time points into a continuous trend representation. The prediction head maps this representation to the quantization parameter space of multiple future time points. This time-series modeling capability enables the model to improve prediction accuracy by utilizing historical information, while outputting prediction results for multiple time points, providing a basis for dynamic monitoring.
[0085] Figure 4 This is a second schematic diagram of the structure of an ophthalmic image information processing system provided in some embodiments of this application, such as... Figure 4 As shown, the ophthalmic image information processing system 100 may further include: an uncertainty quantification module 104 and a visualization and interpretation module 105.
[0086] The uncertainty quantization module 104 is connected to the processing module 102 and is used to perform multiple forward propagations on the prediction results of the multimodal neural network model using the Monte Carlo Dropout algorithm or the deep ensemble algorithm, and calculate the confidence interval of the quantization parameters based on the statistical distribution information of the multiple forward propagation results.
[0087] In some embodiments, the uncertainty quantification module may employ the Monte Carlo Dropout algorithm. During model training, the Dropout layer randomly deactivates some neurons with a certain probability to prevent overfitting. During inference, the Dropout layer is typically disabled to obtain deterministic output. Monte Carlo Dropout keeps the Dropout layer enabled during inference, performing multiple forward propagations (e.g., 50 or 100 times) on the same input data. Each forward propagation yields slightly different predictions due to the random deactivation of neurons. These multiple predictions constitute a sample distribution. The uncertainty quantification module calculates the mean of this distribution as the final prediction and the standard deviation or quantile interval (e.g., from the 2.5% quantile to the 97.5% quantile) as the 95% confidence interval. This confidence interval reflects the cognitive uncertainty of the model parameters, i.e., the range of prediction fluctuations caused by limited training data or insufficient knowledge.
[0088] In some embodiments, the uncertainty quantification module can employ a deep ensemble algorithm. During the training phase, multiple (e.g., 5 or 10) multimodal neural network models with identical structures but different initialization parameters are trained independently. During the inference phase, the same input data is simultaneously fed into all models, and each model outputs a predicted value. The predictions from the multiple models constitute a sample distribution. The uncertainty quantification module calculates the statistical characteristics of this distribution, obtaining the mean and confidence interval. Deep ensemble algorithms can capture the diversity of models under different initial states and convergence paths.
[0089] In this embodiment, the uncertainty quantification module outputs the confidence interval along with the predicted value of the quantification parameter. For example, for visual acuity prediction, it outputs "0.10 logMAR (95% confidence interval: 0.05-0.15)". The width of the confidence interval reflects the reliability of the prediction result: the narrower the interval, the more confident the model is in the prediction; the wider the interval, the higher the uncertainty, and caution should be exercised when using it. This information provides a reference for the credibility of the prediction result in clinical evaluation, avoiding the potential misleading effect of a single numerical value.
[0090] The visualization and interpretation module 105 is connected to the processing module 102 and is used to interpret the prediction basis of the model from multiple dimensions, thereby enhancing the understandability and verifiability of the prediction process.
[0091] In this embodiment, the visualization and interpretation module can have a heatmap generation function. Accordingly, the visualization and interpretation module generates a heatmap for ophthalmic images based on the intermediate layer feature maps of the multimodal neural network model and the gradient information of the prediction results, and marks the image regions in the heatmap that contribute the most to the quantization parameters.
[0092] The heatmap generation function can be implemented based on gradient-weighted class activation mapping (JEM). Specifically, the last convolutional feature map of the model contains multiple channels, each corresponding to the response intensity of different semantic features in the image. For a given prediction result, this technique calculates the gradient of the prediction result with respect to each feature channel to obtain the importance weight of each channel to the prediction result. The larger the gradient value, the more significant the impact of the feature change of that channel on the prediction result. By weighting and superimposing all feature channels according to their respective weights, a heatmap corresponding to the size of the input image can be generated. Darker areas in the heatmap indicate that the area contributes more to the prediction result. Overlaying the heatmap onto the original ophthalmic image can visually display the retinal areas that the model focuses on during prediction. For example, for visual acuity prediction, the heatmap may highlight the fovea region; for axial length prediction, the heatmap may highlight the peridiscal region or the choroid region.
[0093] In this embodiment, the visualization and interpretation module may have a feature sensitivity calculation function. Accordingly, the visualization and interpretation module calculates the sensitivity of each input feature in the multimodal data to the quantization parameters, which is used to identify key factors affecting the prediction results.
[0094] For numerical input features (such as age and axial length in physiological parameters, or wavelength and power density in photobiological regulation parameters), the sensitivity calculation adopts the feature perturbation method: keeping other inputs constant, a certain feature is perturbed slightly within a certain range (e.g., increased or decreased by one unit), and the magnitude of change in the prediction result is observed. The larger the magnitude of change, the higher the sensitivity of the feature to the prediction result.
[0095] For image-based input features (such as tomographic images and fundus images), sensitivity calculation employs either the gradient integral method or the occlusion experiment method. The gradient integral method calculates the gradient of the prediction result with respect to the input pixels, accumulating the gradient values along a straight path to obtain the contribution of each pixel. The occlusion experiment method replaces a region in the image with background values and observes the change in the prediction result; the greater the change, the more important the occluded region is to the prediction result.
[0096] In this embodiment, the visualization and interpretation module can generate a natural language interpretation report. Accordingly, the visualization and interpretation module converts the heatmap analysis results and sensitivity analysis results into structured text descriptions, generating a natural language interpretation report. This report uses a preset text template, filling the template with quantitative analysis results to form a readable natural language expression.
[0097] For example, based on the location and extent of the highlighted areas in the heatmap, the report generates "The prediction results are mainly based on the structural characteristics of the retina in the macular region, which shows a trend of choroidal thickening"; based on the sensitivity analysis results, the report generates "Changes in axial length have the greatest impact on the user's vision prediction, contributing 42%". The natural language report reduces the difficulty of understanding, enabling non-professional users to understand the basis of the prediction.
[0098] In this embodiment, the visualization and interpretation module can have graphical display capabilities. Accordingly, the visualization and interpretation module graphically displays quantification parameters, confidence intervals, heatmaps, sensitivity, and natural language interpretation reports. The display interface typically adopts a segmented layout: the left area displays the original ophthalmic image and overlaid heatmap; the right area displays the quantification parameter values and confidence interval bar charts; and the bottom area displays the feature sensitivity histogram and natural language interpretation text. Through a unified visualization interface, users can quickly obtain prediction results, understand the basis for prediction, and assess the reliability of predictions, providing a comprehensive reference for subsequent decision-making.
[0099] As can be seen from the above embodiments, in this embodiment, the visualization and interpretation module can transform the feature extraction and decision-making process inside the model into understandable visual information and text descriptions, thereby enhancing the interpretability and clinical usability of the model.
[0100] Figure 5 These are schematic diagrams of the structure of an eye-illuminating device provided in some embodiments of this application, such as... Figure 5 As shown, the eye illumination device 500 may include: a light source module 501, an ophthalmic image information processing system 502, and a control module 503.
[0101] The light source module 501 is used to emit illumination light.
[0102] In some embodiments, the light source module may include at least one light-emitting unit capable of emitting light of a specific wavelength. The light-emitting unit may be implemented using a light-emitting diode or a laser diode, and the emitted light is uniformly projected onto the user's eye through an optical lens group. The light source module may support multiple operating modes, including continuous illumination mode and pulse illumination mode, and the pulse frequency can be adjusted within a certain range.
[0103] In some embodiments, the light source module may use red light of 630nm to 680nm and / or near-infrared light of 810nm to 850nm, and the output power density may be 8mW / cm² to 50mW / cm².
[0104] The ophthalmic image information processing system 502 is the ophthalmic image information processing system described in any of the foregoing embodiments. This system acquires multimodal data from the user through an acquisition module, including ophthalmic images (computed tomographic images, fundus images), physiological parameters (age, gender, refractive error, intraocular pressure, axial length), and photobiological modulation parameters (wavelength, power density, irradiation duration, number of irradiations). After analysis by the multimodal neural network model in the processing module, it outputs quantitative parameters related to visual function, such as visual acuity, axial length, or choroidal thickness.
[0105] The control module 503 is connected to the ophthalmic image information processing system 502 and the light source module 501 respectively, and is used to adjust the illumination parameters of the light source module 501 according to the quantization parameters output by the ophthalmic image information processing system 502.
[0106] In this application embodiment, the irradiation parameters may include at least one of the following: wavelength, power density, pulse frequency, duty cycle, and irradiation duration.
[0107] For example, taking user A as an example, when user A uses an eye-illuminating device for the first time, the ophthalmic image information processing system outputs quantitative parameters based on user A's baseline data (computed tomographic images showing thin choroidal thickness, fundus images showing low vascular density, physiological parameters showing longer axial length, and age 12 years): current visual acuity is 0.30 logMAR, predicting that the device will work best with a wavelength of 660nm and a power density of 25mW / cm². 2 Under the prescribed irradiation scheme, visual acuity improved to 0.20 logMAR after 3 months, and choroidal thickness increased by approximately 15 μm. Upon receiving this quantification parameter, the control module determined that the predicted improvement was within the expected range and set the irradiation parameters of the light source module to a wavelength of 660 nm and a power density of 25 mW / cm². 2 The duration of a single irradiation session is 3 minutes, and this setting is recorded for future use.
[0108] For example, taking user B as an example, their baseline data is similar to user A, but the quantitative parameters output by the ophthalmic image information processing system show that the predicted improvement is small under the standard irradiation protocol, while the improvement is smaller under the standard irradiation protocol with a wavelength of 830nm and a power density of 30mW / cm². 2 When using a near-infrared irradiation scheme, the predicted improvement in choroidal thickness is more significant. After receiving this quantified parameter, the control module automatically adjusts the light source module to near-infrared irradiation mode and sets the corresponding power density and irradiation duration based on the prediction results. This eye illumination device, through a linkage mechanism of real-time quantified evaluation and parameter adjustment, enables the irradiation scheme to be dynamically adapted to the user's individual characteristics.
[0109] As can be seen from the above embodiments, in this embodiment, by integrating the ophthalmic image information processing system and the light source module into the same device, the control module can adjust the illumination parameters of the light source module in real time according to the quantitative parameters output by the system, realizing the automated connection from eye condition assessment to illumination scheme adjustment, and providing an objective adjustment basis for personalized illumination.
[0110] Corresponding to the ophthalmic image information processing system and ocular illumination device in the foregoing embodiments, this application also provides corresponding control methods. The control method for the ophthalmic image information processing system includes control steps executed by the processing module, used to implement the corresponding functions of the acquisition module, processing module, and output module in any of the foregoing embodiments. The control method for the ocular illumination device includes control steps executed by the control module, used to implement the corresponding functions of the light source module in any of the foregoing embodiments.
[0111] Figure 6 This is a flowchart illustrating a control method for an ophthalmic image information processing system according to some embodiments of this application, such as... Figure 6 As shown, the method includes the following steps: step 601, step 602 and step 603.
[0112] In step 601, the control acquisition module acquires the multimodal data of the first user; wherein, the multimodal data includes ophthalmic images, physiological parameters and photobiological regulation parameters.
[0113] In step 602, the multimodal data is processed by a pre-trained multimodal neural network model to obtain the quantization parameters related to the visual function of the first user. The multimodal neural network model includes: a feature extraction network for extracting features of multiple modalities from the multimodal data; an attention fusion module for fusing the features of multiple modalities through an attention mechanism to obtain fused features; and a prediction head for predicting the quantization parameters based on the fused features.
[0114] In step 603, the control output module outputs quantization parameters.
[0115] As can be seen from the above embodiments, in this embodiment, the multimodal data of the first user is collected by the acquisition module, and the ophthalmic images, physiological parameters and photobiological regulation parameters are jointly processed by the pre-trained multimodal neural network model to generate quantitative parameters related to visual function. This realizes an objective quantitative assessment of the user's eye condition, and replaces traditional experience judgment with a data-driven approach, which can improve the accuracy and consistency of the assessment of visual function-related indicators.
[0116] Optionally, as an embodiment, the ophthalmic images include: tomographic images and fundus images; The physiological parameters include at least one of the following: age, sex, refractive error, intraocular pressure, and axial length; The photobiological regulation parameters include at least one of the following: wavelength, power density, irradiation duration, and number of irradiations.
[0117] Optionally, as an embodiment, the quantification parameters include at least one of the following: visual acuity value, axial length value, and choroid thickness value.
[0118] Optionally, as an embodiment, the multimodal data further includes: visual field inspection results; wherein the visual field inspection results include at least one of the following: light sensitivity value matrix, average deviation, mode deviation, and visual field index.
[0119] Optionally, as an embodiment, the feature extraction network includes: The first feature extraction branch uses a 3D convolutional neural network to extract the first feature from the tomographic scan image; The second feature extraction branch uses a 2D convolutional neural network to extract the second feature from the fundus image; The third feature extraction branch uses a multilayer perceptron to extract the third feature from the physiological parameters; The fourth feature extraction branch uses a multilayer perceptron to extract the fourth feature from the photobiological regulatory parameters.
[0120] Optionally, as an embodiment, the multimodal neural network model further includes: a temporal encoder connected between the attention fusion module and the prediction head; The temporal encoder is used to construct a temporal sequence from multiple fused features corresponding to the baseline time point and / or follow-up time point output by the attention fusion module, and to obtain a temporal feature vector by performing temporal feature encoding processing through an encoder based on a self-attention mechanism. The prediction head is specifically used to predict quantization parameters at multiple future time points based on the time-series feature vector.
[0121] Optionally, as an embodiment, the method further includes: The uncertainty quantization module uses the Monte Carlo Dropout algorithm or the deep ensemble algorithm to perform multiple forward propagations on the prediction results of the multimodal neural network model, and calculates the confidence interval of the quantization parameters based on the statistical distribution information of the multiple forward propagation results.
[0122] Optionally, as an embodiment, the method further includes: The visualization and interpretation module is controlled to perform at least one of the following operations: Based on the intermediate layer feature maps and gradient information of the prediction results of the multimodal neural network model, a heatmap for the ophthalmic image is generated, and the image region that contributes the most to the quantization parameter in the heatmap is marked. Calculate the sensitivity of each input feature in the multimodal data to the quantization parameter; Generate natural language interpretation reports; The quantification parameters, heatmaps, sensitivity, and natural language interpretation reports are displayed graphically.
[0123] Figure 7 This is a flowchart of a method for controlling an eye-illuminating device according to some embodiments of this application, such as... Figure 7 As shown, the method may include the following steps: Step 701.
[0124] In step 701, the control module acquires the quantization parameters output by the ophthalmic image information processing system and adjusts the illumination parameters of the light source module according to the quantization parameters; wherein, the light source module is used to emit illumination light, and the illumination parameters include at least one of wavelength, power density, pulse frequency, duty cycle, and illumination duration.
[0125] As can be seen from the above embodiments, in this embodiment, by integrating the ophthalmic image information processing system and the light source module into the same device, the control module can adjust the illumination parameters of the light source module in real time according to the quantitative parameters output by the ophthalmic image information processing system, thereby realizing the automated connection from eye condition assessment to illumination scheme adjustment, and providing an objective adjustment basis for personalized illumination.
[0126] In summary, through the above control methods, the ophthalmic image information processing system can automatically complete the processing flow from data acquisition to quantitative parameter output, and the ophthalmic illumination device can automatically adjust the illumination parameters of the light source module according to the quantitative evaluation results.
[0127] Since the method embodiments correspond to the system and device embodiments in terms of technical concept and implementation principle, their specific steps, technical features, and beneficial effects can be directly and clearly derived from the understanding of the aforementioned system and device embodiments. Therefore, the content of the method embodiments will not be repeated here. After reading the system and device embodiments of this application, those skilled in the art will be able to understand and implement the corresponding control methods without any doubt.
[0128] Additionally or optionally, embodiments of this application can also be implemented as a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, causes the processor to perform some or all of the steps in the above-described method embodiments of this application. The computer-readable storage medium can be any tangible medium containing or storing a program, such as a hard disk, solid-state drive, random access memory, read-only memory, flash memory or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the required program code and is accessible by a computer.
[0129] While numerous embodiments of this application have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will arise for those skilled in the art without departing from the spirit and intent of this application. It should be understood that various alternatives to the embodiments of this application described herein may be employed in the practice of this application. The appended claims are intended to define the scope of protection of this application and therefore cover equivalents or alternatives within the scope of these claims.
[0130] The collection and acquisition of various data in this application comply with relevant laws and regulations and are authorized by the data providers. Any organization or individual that needs to obtain external data shall obtain authorization in accordance with the law and ensure data security, and shall not illegally collect, use, process, or transmit unauthorized or unprotected data, nor shall it illegally buy, sell, provide, or disclose unauthorized or unprotected data.
Claims
1. An ophthalmic image information processing system, characterized in that, The system includes: The acquisition module is used to acquire multimodal data of the first user; wherein, the multimodal data includes ophthalmic images, physiological parameters, and photobiological regulation parameters; The processing module is used to process the multimodal data through a pre-trained multimodal neural network model to obtain the quantitative parameters related to the visual function of the first user. The output module is used to output the quantization parameters; The multimodal neural network model includes: a feature extraction network for extracting features of multiple modalities from the multimodal data; an attention fusion module for fusing the features of the multiple modalities through an attention mechanism to obtain fused features; and a prediction head for predicting the quantization parameters based on the fused features.
2. The system according to claim 1, characterized in that, The ophthalmic images include: tomographic images and fundus images; The physiological parameters include at least one of the following: age, sex, refractive error, intraocular pressure, and axial length; The photobiological regulation parameters include at least one of the following: wavelength, power density, irradiation duration, and number of irradiations.
3. The system according to claim 1, characterized in that, The quantitative parameters include at least one of the following: visual acuity value, axial length value, and choroid thickness value.
4. The system according to claim 1, characterized in that, The multimodal data also includes: visual field inspection results; The visual field inspection results include at least one of the following: light sensitivity value matrix, average deviation, pattern deviation, and visual field index.
5. The system according to claim 1, characterized in that, The feature extraction network includes: The first feature extraction branch uses a 3D convolutional neural network to extract the first feature from the tomographic scan image; The second feature extraction branch uses a 2D convolutional neural network to extract the second feature from the fundus image; The third feature extraction branch uses a multilayer perceptron to extract the third feature from the physiological parameters; The fourth feature extraction branch uses a multilayer perceptron to extract the fourth feature from the photobiological regulatory parameters.
6. The system according to claim 1, characterized in that, The multimodal neural network model further includes a temporal encoder, which is connected between the attention fusion module and the prediction head; The temporal encoder is used to construct a temporal sequence from multiple fused features corresponding to the baseline time point and / or follow-up time point output by the attention fusion module, and to obtain a temporal feature vector by performing temporal feature encoding processing through an encoder based on a self-attention mechanism. The prediction head is specifically used to predict quantization parameters at multiple future time points based on the time-series feature vector.
7. The system according to claim 1, characterized in that, The system also includes: An uncertainty quantization module, connected to the processing module, is used to perform multiple forward propagations on the prediction results of the multimodal neural network model using the Monte Carlo Dropout algorithm or a deep ensemble algorithm, and to calculate the confidence interval of the quantization parameter based on the statistical distribution information of the multiple forward propagation results.
8. The system according to claim 1, characterized in that, The system also includes: A visualization and interpretation module, connected to the processing module, is used for at least one of the following: Based on the intermediate layer feature maps and gradient information of the prediction results of the multimodal neural network model, a heatmap for the ophthalmic image is generated, and the image region that contributes the most to the quantization parameter in the heatmap is marked. Calculate the sensitivity of each input feature in the multimodal data to the quantization parameter; Generate natural language interpretation reports; The quantification parameters, heatmaps, sensitivity, and natural language interpretation reports are displayed graphically.
9. An eye illumination device, characterized in that, The device includes: The light source module is used to emit illumination light. The ophthalmic image information processing system according to any one of claims 1-8; The control module is connected to both the ophthalmic image information processing system and the light source module, and is used to adjust the illumination parameters of the light source module according to the quantization parameters output by the ophthalmic image information processing system.
10. A computer-readable storage medium storing a computer program, characterized in that, When the program is executed by the processor, it implements the function of the ophthalmic image information processing system according to any one of claims 1-8, or the function of the ophthalmic illumination device according to claim 9.