A method and system for predicting the risk of coronary heart disease based on multimodal fundus images

Through multimodal fundus image technology and deep learning model, the problem of non-invasive screening in early coronary heart disease is solved, and non-invasive, fast and accurate coronary heart disease risk prediction is achieved, which improves the sensitivity and specificity of the detection, and is suitable for community hospitals and primary medical institutions.

CN119400403BActive Publication Date: 2025-07-25FUWAI HOSPITAL CHINESE ACAD OF MEDICAL SCI & PEKING UNION MEDICAL COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411265435.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2025-07-25
Estimated Expiration
2044-09-10

AI Technical Summary

Technical Problem

The diagnosis and prognostic evaluation of coronary heart disease in the prior art relies on high-cost, time-consuming and invasive imaging examinations, making it difficult to achieve early non-invasive screening and diagnosis, resulting in patients seeking medical treatment in the middle and late stages, increasing the complexity of treatment and poor prognosis.

Method used

Multimodal fundus image technology is used to obtain fundus plane and stereoscopic features through color fundus photography and optical coherence tomography. Combining multi-scale attention model and area-guided attention model, multimodal fusion deep learning model is used to predict coronary heart disease risk.

Benefits of technology

It has achieved non-invasive, fast and accurate risk prediction of coronary heart disease, improved the sensitivity and specificity of early detection, reduced medical costs, and was suitable for widespread applications in community hospitals and primary medical institutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119400403B_ABST
    Figure CN119400403B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for predicting the risk of coronary heart disease based on multimodal fundus images, which relates to the technical field of intelligent medical systems. The method includes: obtaining a multimodal fundus image group corresponding to a target user; extracting a fundus plane feature group corresponding to a first fundus image based on a multi-scale attention model; extracting a fundus three-dimensional feature group corresponding to a second fundus image based on a region-guided attention model; and inputting the fundus plane feature group and the fundus three-dimensional feature group into a risk prediction model to determine a corresponding risk prediction result of coronary heart disease, where the risk prediction model adopts a multimodal fusion deep learning model. Thus, through big data analysis of multimodal and multi-dimensional fundus images, the prevention and control of coronary heart disease is advanced to an earlier stage, realizing non-invasive, rapid and accurate risk prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent medical systems, and in particular, to a method and system for predicting the risk of coronary heart disease based on multimodal fundus images. Background Technique

[0002] Coronary Heart Disease (CHD) is a heart disease caused by the stenosis or occlusion of blood vessels due to coronary atherosclerosis, which in turn leads to myocardial ischemia, hypoxia or necrosis. Coronary heart disease is not only one of the main causes of death and disability globally, but also a disease with the highest incidence and mortality rate among cardiovascular diseases. Therefore, the early detection and risk prediction of coronary heart disease are of crucial significance for reducing the incidence and mortality rate.

[0003] Currently, the diagnosis and prognosis evaluation of coronary heart disease mainly rely on imaging examinations such as coronary CTA (Coronary Computed Tomography Angiography), cardiac magnetic resonance imaging, and coronary angiography. These methods are expensive, time-consuming, partially invasive, and have high operation requirements, making it difficult to achieve early non-invasive screening, diagnosis, prognosis evaluation, and long-term follow-up. Most patients are already in the middle and late stages when they are clearly diagnosed, resulting in complex coronary artery lesions, increased difficulty in interventional and surgical treatments, and poor prognosis.

[0004] In response to the above problems, the industry has not yet proposed a better technical solution. Summary of the Invention

[0005] Embodiments of the present invention provide a method and system for predicting the risk of coronary heart disease based on multimodal fundus images, which are used to at least solve the problems of high cost, long time consumption, and strong invasiveness existing in current imaging examination methods.

[0006] In a first aspect, embodiments of the present invention provide a method for predicting the risk of coronary heart disease based on multimodal fundus images. The method includes: obtaining a multimodal fundus image group corresponding to a target user; the multimodal fundus image group includes a first fundus image and a second fundus image; the first fundus image is obtained by color fundus photography, and the second fundus image is obtained by optical coherence tomography angiography; based on a multi-scale attention model, extracting a fundus plane feature group corresponding to the first fundus image; based on a region-guided attention model, extracting a fundus three-dimensional feature group corresponding to the second fundus image; inputting the fundus plane feature group and the fundus three-dimensional feature group into a risk prediction model to determine a corresponding coronary heart disease risk prediction result, and the risk prediction model uses a multimodal fusion deep learning model.

[0007] In a second aspect, an embodiment of the present invention provides a coronary heart disease risk prediction system based on multimodal fundus images. The system includes: a multimodal data acquisition unit configured to acquire a group of multimodal fundus images corresponding to a target user; the group of multimodal fundus images includes a first fundus image and a second fundus image; the first fundus image is obtained by color fundus photography, and the second fundus image is obtained by optical coherence tomography angiography; a planar feature extraction unit configured to extract a group of fundus planar features corresponding to the first fundus image based on a multi-scale attention model; a three-dimensional feature extraction unit configured to extract a group of fundus three-dimensional features corresponding to the second fundus image based on a region-guided attention model; a risk prediction unit configured to input the group of fundus planar features and the group of fundus three-dimensional features into a risk prediction model to determine a corresponding coronary heart disease risk prediction result, and the risk prediction model adopts a multimodal fusion deep learning model.

[0008] In a third aspect, an embodiment of the present invention provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the steps of the above method.

[0009] In a fourth aspect, an embodiment of the present invention provides a storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute the steps of the above method of the present invention.

[0010] In a fifth aspect, an embodiment of the present invention further provides a computer program product, which includes a computer program stored on a storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer is enabled to execute the steps of the above method.

[0011] Through a coronary heart disease risk prediction method, system, electronic device, and non-transitory computer-readable storage medium provided by the present invention, the following technical effects can be achieved at least:

[0012] (1) By adopting the method of multi-modal fundus images for coronary heart disease risk prediction, it is completely non-invasive, reducing the pain of patients and the consumption of medical resources. At the same time, it also avoids the problems of high cost, long time consumption, and strong invasiveness in traditional imaging examination methods. Through this technical solution, by using two methods, namely Color Fundus Photography (CFP) and Optical Coherence Tomography Angiography (OCTA), the planar and three-dimensional features of the fundus are comprehensively captured and analyzed through a multi-modal fusion deep learning model, improving the sensitivity and specificity of early detection of coronary heart disease, and enabling screening in the early stage when patients have no obvious symptoms, thus achieving early detection and intervention of coronary heart disease.

[0013] (2) In this technical solution, a multi-scale attention model is adopted to extract the planar features of color fundus images. The multi-scale attention model can capture the detailed features of fundus images at different scales, and a region-guided attention model is used to extract the three-dimensional features of optical coherence tomography angiography images. The region-guided attention model can focus on the regional features with important clinical significance, thereby improving the accuracy and efficiency of feature extraction.

[0014] (3) By inputting the feature groups of different-modal fundus images into the multi-modal fusion deep learning model, comprehensive analysis of different-modal features is achieved. Compared with the analysis and prediction method of single-modal images, the multi-modal fusion model can more comprehensively reflect the complex pathological features of coronary heart disease, improving the accuracy and reliability of risk prediction.

[0015] Through this technical solution, fundus imaging technology is used to predict coronary heart disease before its onset, extracting the important features in the two fundus imaging methods respectively and performing feature fusion. Through big data analysis of multi-modal and multi-dimensional fundus images, the prevention and control of coronary heart disease is advanced to the front end, achieving non-invasive, fast, and accurate prediction, making this platform more valuable for disease prevention, greatly facilitating clinical applications, especially suitable for community hospitals and primary medical institutions, and improving the coverage and popularity of coronary heart disease screening. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1Shows a flowchart of an example of a coronary heart disease risk prediction method based on multimodal fundus images according to an embodiment of the present invention;

[0018] Figure 2 Shows a schematic structural connection diagram of an example of a risk prediction model according to an embodiment of the present invention;

[0019] Figure 3 Shows a schematic structural connection diagram of an example of a multi-scale attention model according to an embodiment of the present invention;

[0020] Figure 4 Shows a schematic structural connection diagram of an example of a region-guided attention model according to an embodiment of the present invention;

[0021] Figure 5 Shows an operation flowchart of an example of determining a membrane layer region image by an adaptive threshold segmentation method according to an embodiment of the present invention;

[0022] Figure 6 Shows a schematic block diagram of an example of a coronary heart disease risk prediction system based on multimodal fundus images according to an embodiment of the present invention;

[0023] Figure 7 Is a schematic structural diagram of an embodiment of an electronic device of the present invention. Detailed implementation manners

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0025] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the art to which the present invention pertains. The "first", "second", and similar terms used in the present invention do not denote any order, quantity, or importance, but are only used to distinguish different components. Similarly, terms such as "a", "one", or "the" do not denote a quantity limitation, but mean that there is at least one. The terms "including" or "comprising" and the like mean that the elements or items appearing before the term cover the elements or items listed after the term and their equivalents, without excluding other elements or items. The terms "connected" or "coupled" and the like are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.

[0026] It should be noted that the terms "up", "down", "left", "right", "front", "back", etc. used in the present invention are only used to indicate relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.

[0027] Fundus microcirculation has similar pathophysiological characteristics to cardiovascular circulation. Many studies have discovered fundus imaging markers related to cardiovascular disease, such as retinal vessel diameter, retinal and choroidal blood flow density, thickness, etc., which are related to the incidence of cardiovascular events, and pupil area is related to the prognosis of heart failure. Fundus images can generally be obtained from opticians or ophthalmology clinics. The collection method is safe, non-invasive, convenient and inexpensive, with high health economic benefits and various imaging methods. Therefore, artificial intelligence is used to integrate multimodal information of fundus imaging, and then non-invasively and accurately predict the risk of coronary heart disease, and achieve early warning and intervention. The development of an efficient, rapid, economical, non-invasive, popular and scalable coronary heart disease prediction and evaluation model is crucial for early warning and prevention for people with coronary heart disease.

[0028] Figure 1 A flowchart of an example of a method for predicting the risk of coronary heart disease based on multimodal fundus images according to an embodiment of the present invention is shown.

[0029] Regarding the executor of the method of the embodiment of the present invention, it can be any controller or processor with computing or processing capabilities, which can be integrated or partially integrated in a server or a coronary heart disease risk prediction platform, so as to achieve non-invasive, early, high-precision, rapid and economical coronary heart disease risk prediction by combining multimodal fundus images and deep learning models, thereby greatly optimizing the existing detection and evaluation methods of coronary heart disease.

[0030] In some examples, it can be integrated into the configuration server by software, hardware, or a combination of software and hardware, which should not be limited here. It should be understood that one or more steps involved in the following process can be implemented by one or more controllers or software installed and deployed in the client or server.

[0031] like Figure 1 As shown, in step S110, a multimodal fundus image group corresponding to the target user is acquired, and the multimodal fundus image group includes a first fundus image and a second fundus image.

[0032] Here, the first fundus image is obtained by color fundus photography, and the second fundus image is obtained by optical coherence tomography blood flow imaging.

[0033] In some embodiments, a high-resolution color fundus camera is used to photograph the fundus of a target user to obtain a color fundus image. Since the photographing can be completed only by having the patient fixate on a certain direction, the whole process is non-invasive and fast. Additionally, an optical coherence tomography (OCT) device is used to scan the retina of the user to obtain a high-resolution three-dimensional fundus image, especially a blood flow imaging image. Similarly, the sampling process is also non-invasive and only requires the user to keep their eyes fixed and look in a specific direction for a few seconds. Thus, a non-invasive and fast acquisition method is achieved, improving the comfort and compliance of the patient and facilitating large-scale screening.

[0034] In step S120, based on the multi-scale attention model, the fundus plane feature group corresponding to the first fundus image is extracted.

[0035] In some embodiments, preprocessing operations such as denoising, contrast enhancement, and normalization are performed on the first fundus image, and multi-level feature extraction of the color fundus image is carried out through the multi-scale attention mechanism of deep learning. Through this embodiment, the plane features of the fundus are accurately extracted by the multi-scale attention mechanism, which are highly correlated with the risk of coronary heart disease and can effectively capture important feature details at different scales, improving the accuracy and integrity of feature extraction.

[0036] In some examples of the embodiments of the present invention, the fundus plane feature group includes at least one of the following: cup-to-disc ratio of the left and right eyes, vertical cup-to-disc ratio, cup area, cup volume, disc area, average thickness of the retinal nerve fiber layer, area of the foveal avascular zone, perimeter of the foveal avascular zone, and near-circularity index.

[0037] It should be noted that the cup-to-disc ratio (CDR) of the left and right eyes is the ratio of the cup diameter to the disc diameter, which is used to evaluate the morphology of the optic disc. A higher CDR may be associated with optic nerve damage, and there is a certain correlation between optic nerve health and overall cardiovascular health. The vertical cup-to-disc ratio is the ratio of the vertical diameter of the optic cup to the vertical diameter of the optic disc. Similar to the horizontal CDR, the vertical cup-to-disc ratio provides another dimension of information on optic nerve health, helps to comprehensively evaluate the changes in the optic disc, and increases the accuracy of coronary heart disease risk prediction. The optic cup area is the area of the central depression of the optic disc. Changes in the optic cup area may reflect damage to the optic nerve fibers, which is related to the vascular health status and can indirectly reflect the risk of coronary heart disease. The optic cup volume is the three-dimensional volume of the optic cup. The optic cup volume can more stereoscopically reflect the changes in the size and shape of the optic cup, helping to more comprehensively evaluate the health status of the optic nerve and blood vessels. The optic disc area is the area of the region where the optic nerve enters the retina. Changes in the optic disc area may be related to the blood flow and vascular conditions in the fundus. Monitoring the changes in the optic disc area can provide an important reference for coronary heart disease risk assessment. The average thickness of the retinal nerve fiber layer is the average thickness value of the retinal nerve fiber layer. The thickness of the retinal nerve fiber layer is closely related to optic nerve health. Abnormal thickness may indicate vascular lesions and help to predict the risk of coronary heart disease at an early stage. The foveal avascular zone is the area in the center of the retina without blood vessel supply. Its area reflects the size of the avascular zone. Changes in the area of the foveal avascular zone may be related to abnormal retinal blood flow. Monitoring the changes in this area can provide an indirect basis for coronary heart disease risk. The perimeter of the foveal avascular zone can help to quantify the changes in the avascular zone and provide more information on the retinal blood flow status. The near-circularity index measures the degree to which the shape of the foveal avascular zone approaches a circle. Changes in the near-circularity index may reflect abnormalities in the retinal microcirculation state.

[0038] In step S130, based on the region-guided attention model, the fundus stereo feature group corresponding to the second fundus image is extracted.

[0039] In some embodiments, preprocessing operations such as denoising, contrast enhancement, and normalization are performed on the second fundus image. Through the region-guided attention mechanism of deep learning, in-depth analysis is carried out on the OCT blood flow imaging image to extract the features of the three-dimensional structure. Thus, through the region-guided attention model, the three-dimensional features of the fundus blood vessels can be accurately extracted, the key regions of the vascular structure can be focused on, the accuracy of three-dimensional feature extraction is improved, the features of arteries and veins can be effectively distinguished and analyzed, and data support is provided for more detailed risk assessment.

[0040] In some examples of the embodiments of the present invention, the fundus three-dimensional feature group includes at least one of the following: total blood vessel diameter, artery diameter, vein diameter, total blood vessel fractal dimension, artery fractal dimension, vein fractal dimension, total blood vessel curvature, artery curvature, vein curvature, total blood vessel curvature density, artery curvature density, vein curvature density, and total blood vessel branch angle.

[0041] It should be noted that the total blood vessel diameter is the total diameter of all major blood vessels on the retina, reflecting the overall health status of the retinal blood vessels. Abnormal blood vessel diameter may indicate systemic blood vessel health problems, including coronary heart disease. The artery diameter, that is, the diameter of the retinal artery, and changes in the artery diameter can directly reflect arteriosclerosis and stenosis, which are important risk factors for coronary heart disease. The vein diameter, that is, the diameter of the retinal vein, and changes in the vein diameter may reflect venous return disorders and hemodynamic abnormalities, helping to evaluate the overall blood vessel health status. The total blood vessel fractal dimension reflects the complexity and fractal characteristics of the retinal blood vessels. The fractal dimension of the blood vessels is related to their health status and function. A higher fractal dimension may indicate a healthier blood vessel network structure, helping to evaluate the risk of coronary heart disease. The artery fractal dimension is used to measure the fractal characteristics of the retinal artery, which can provide detailed information on the artery health status and help identify early risk factors for coronary heart disease. The vein fractal dimension is used to measure the fractal characteristics of the retinal vein, which reflects the complexity of the vein network. Abnormalities may indicate blood flow disorders and indirectly reflect the risk of coronary heart disease. The total blood vessel curvature represents the curvature of all blood vessels on the retina. Changes in the blood vessel curvature may reflect the elasticity and functional status of the blood vessels and have important reference value for predicting the risk of coronary heart disease. The artery curvature is the curvature of the retinal artery. Abnormal artery curvature may indicate arteriosclerosis and blood flow disorders, which are directly related to the risk of coronary heart disease. The vein curvature represents the curvature of the retinal vein, and its change degree may reflect venous return and hemodynamic abnormalities. The total blood vessel curvature density represents the curvature density of all blood vessels on the retina, reflecting the density and health status of the blood vessel network. The artery curvature density, that is, the curvature density of the retinal artery, and its change can indicate the health status of the artery and have important reference value for predicting the risk of coronary heart disease. Changes in the vein curvature density can reflect venous blood flow and blood vessel health status, helping to comprehensively evaluate the risk of coronary heart disease. The total blood vessel branch angle represents the angle of all blood vessel branches on the retina, and its change may reflect the growth and functional status of the blood vessels, providing quantitative analysis for predicting the risk of coronary heart disease.

[0042] Through this embodiment, by comprehensively extracting and analyzing these planar and three-dimensional features, the information of fundus images can be fully utilized to provide accurate and comprehensive prediction of the risk of coronary heart disease. Through the complementary and comprehensive analysis of these features, it helps to improve the accuracy and reliability of the prediction model and provides a scientific basis for early intervention and treatment.

[0043] In step S140, the fundus plane feature group and the fundus three-dimensional feature group are input into the risk prediction model to determine the corresponding coronary heart disease risk prediction result, and the risk prediction model adopts a multi-modal fusion deep learning model.

[0044] In some embodiments, the extracted fundus plane feature group and three-dimensional feature group are subjected to data fusion to construct a unified feature vector. By using a multi-modal fusion deep learning model, features of different modalities are effectively fused and comprehensively analyzed for coronary heart disease risk prediction, so as to improve the accuracy and robustness of the prediction result. Exemplarily, through supervised learning using a large amount of labeled data, training is carried out through a large-scale labeled data set to optimize model parameters and improve prediction accuracy.

[0045] Through this embodiment, by using a multi-modal fusion deep learning model, the advantages of plane and three-dimensional features can be comprehensively utilized to provide a more accurate and comprehensive risk prediction. Through the risk prediction result of the model, early screening and intervention can be realized, and the prognosis effect of patients can be improved. Thus, through the acquisition and analysis of multi-modal fundus images, combined with a multi-scale attention model and a region-guided attention model, efficient feature extraction and analysis are realized, significantly improving the accuracy and reliability of the prediction, and having the advantages of early non-invasive screening, strong real-time performance, low cost, etc., and can effectively reduce the incidence and mortality of coronary heart disease.

[0046] It should be noted that color fundus photography (CFP) and optical coherence tomography (OCT) are two fundus imaging methods. Optical coherence tomography angiography (OCTA) is a non-invasive blood flow detection technology developed recently based on OCT (optical coherence tomography). By using moving blood as an intrinsic contrast agent, the difference in the forward and backward movement of red blood cells in the retinal blood vessels detected is used to obtain a clear fundus image through computer algorithms. OCTA is three-dimensional stereoscopic imaging, which can observe blood vessel images at different layers and perform high-resolution imaging on the retinal membrane layer, but the observation range is limited. CFP can obtain a large-range color fundus two-dimensional image, but it cannot observe the depth structure and membrane layer structure of the retina. The image features of CFP and OCTA each have their own characteristics, and traditional deep learning cannot effectively extract complementary features.

[0047] Through the embodiment of the present application, a modality-specific attention deep learning network is proposed. Through the modality-specific attention module, important features in CFP and OCTA images are respectively extracted to provide information in different dimensions for coronary heart disease diagnosis, and clinical and biological features are incorporated to establish a multi-modal information fusion model to realize early diagnosis and prognosis evaluation of coronary heart disease.

[0048] Figure 2 A schematic structural connection diagram of an example of the risk prediction model according to an embodiment of the present invention is shown.

[0049] As shown Figure 2 in FIG. 2, the risk prediction model 200 includes a feature fusion layer 210, a deep residual network layer (ResNet) 220, a bidirectional long short-term memory network layer (Bidirectional Long Short-Term Memory, BiLSTM) 230, a multi-head self-attention layer 240, and a fully connected output layer 250.

[0050] The feature fusion layer 210 is used to fuse the fundus plane feature group and the fundus stereo feature group. Specifically, by performing a concatenation operation on the fundus plane feature group and the fundus stereo feature group, a fused feature representation is obtained:

[0051] F fusion = [F p , F s , Equation (1)

[0052] In the formula, F p represents the fundus plane feature group, F s represents the fundus stereo feature group, and F fusion represents the fundus fused feature group.

[0053] Through the feature fusion layer, the planar features and stereo features in the color fundus photography (CFP) and planar optical coherence tomography (OCT-A) images are integrated, ensuring the diversity and richness of the input features. By fusing the planar and stereo features of the fundus, different dimensional information of the fundus blood vessels, such as the cup-to-disc ratio, optic cup area, blood vessel diameter, blood vessel curvature, etc., can be comprehensively captured, improving the comprehensive assessment ability of the risk of coronary heart disease.

[0054] The deep residual network layer 220 includes multiple residual blocks, and each residual block includes multiple convolutional layers and short connections between different convolutional layers. For example, a short connection is constructed between every two convolutional layers. The deep residual network layer 220 is used to extract the deep convolutional features corresponding to the fundus fused feature group:

[0055] H = F fusion + F(F fusion , {W i}), Equation (2)

[0056] In the formula, F represents the convolutional operation in the residual block, W i represents the weight of the i-th convolutional layer in the residual block, and H represents the deep convolutional features.

[0057] Through the deep residual network (ResNet) layer, multi-level deep features can be effectively extracted through residual blocks. By increasing the depth of the model based on the ResNet layer, the feature representation ability is improved, and complex fundus vascular features can be better captured, thereby enhancing the model's discriminative ability for coronary heart disease risk.

[0058] The bidirectional long short-term memory network layer 230 is used to determine the bidirectional temporal dependence features corresponding to the deep convolutional features. Specifically, in order to capture the time series information and global dependence in the features, the forward and backward LSTM networks in the bidirectional long short-term memory network are used, which can consider the front and back dependencies of the sequence simultaneously.

[0059]

[0060] In the formula, H BiLSTM represents the bidirectional temporal dependence features, represents the output features of the forward LSTM of H, represents the output features of the backward LSTM of H.

[0061] Through the bidirectional long short-term memory network (BiLSTM) layer, the forward and backward dependencies of the feature sequence can be considered simultaneously, enhancing the time series information and global dependence of the feature representation. The BiLSTM layer improves the model's processing ability for time series data by capturing the time dynamic changes in the fundus image features, and can more accurately predict the risk of coronary heart disease.

[0062] The multi-head self-attention layer 240 is used to process the bidirectional temporal dependence features to obtain the corresponding global attention features. By using the multi-head self-attention mechanism, the global dependence relationship and important information in the features can be further captured.

[0063] H attention = MultiHead(Q, K, V) = [head1, head2, …, head h W O , Equation (4)

[0064]

[0065] Q = H BiLSTM W Q , Equation (6)

[0066] K = H BiLSTM W K , Equation (7)

[0067] V = H BiLSTM W V , Equation (8)

[0068] In the formula, WQ , W K and W V represent the query weight matrix, the key weight matrix, and the value weight matrix respectively, and Q, K, and V represent the query matrix, the key matrix, and the value matrix respectively; softmax represents the softmax function, represents the dimension of the key vector, and head i represents the output of the i-th attention head; W O represents the multi-head output weight matrix, h represents the total number of attention heads, and H attention represents the global attention feature.

[0069] Through the multi-head self-attention layer, the correlation between features is calculated based on multiple attention heads to capture the global dependencies between features. The multi-head self-attention mechanism can focus on the key information in the features, enhancing the global consistency and discriminability of the feature representation, thereby improving the model's ability to identify the risk of coronary heart disease.

[0070] The fully connected output layer 250 is used to process the global attention feature to obtain the corresponding coronary heart disease risk probability. Here, the global attention feature is further processed through the fully connected layer, and the coronary heart disease risk probability is output through the Sigmoid activation function.

[0071] H FC = ReLU(W FC H attention + b FC ), Equation (9)

[0072] y = Sigmoid(W y H FC + b y ), Equation (10)

[0073] In the formula, y represents the finally output coronary heart disease risk probability, W FC and b FC represent the weight matrix and the bias term of the fully connected layer respectively; ReLU represents the non-linear activation function, W y and b y represent the weight matrix and the bias term of the output layer respectively; Sigmoid represents the Sigmoid activation function.

[0074] The fused deep features are further processed through the fully connected layer and the output layer, and finally the coronary heart disease risk probability is output through the Sigmoid activation function. The fully connected layer optimizes the non-linear transformation ability of the feature representation through the calculation of multiple layers of neurons, and the output layer normalizes the prediction result into a probability value through the Sigmoid function, which is convenient for risk assessment and interpretation.

[0075] Through this embodiment, through the comprehensive application of multi-modal feature fusion, deep residual network, bidirectional long short-term memory network, and multi-head self-attention mechanism, multi-dimensional information in fundus images can be comprehensively and deeply extracted and fused. Furthermore, the combination of multi-modal feature fusion and advanced attention mechanism enables the model to have strong adaptability to different types and sources of data, improving the robustness and generalization ability of the model. Thus, when facing different individuals and diverse fundus image data, the model can maintain stable prediction performance, providing an effective technical means for the early prediction of coronary heart disease.

[0076] In some examples of the embodiments of the present invention, the loss function of the risk prediction model can adopt a comprehensive loss function that combines focal loss and Bayesian optimization loss. By introducing focal loss, the model can better handle the class imbalance problem, assign higher weights to difficult-to-classify samples, and improve the performance of the model on small-sample classes. By introducing Bayesian optimization loss and combining the uncertainty information of the model to guide parameter optimization, the prediction accuracy and robustness of the model can be effectively improved, and the prediction error can be reduced.

[0077] More specifically, the loss function L of the risk prediction model total is:

[0078] L total = λ1·L focal + λ2·L bayesian , Equation (11)

[0079]

[0080]

[0081] In the formula, L focal represents the focal loss term, L bayesian represents the Bayesian optimization loss term, λ1 and λ2 respectively represent the importance coefficients of the corresponding loss terms; N represents the total number of samples in the data sample set; β t represents the class weight of the t-th sample, p t is the prediction probability that the model assigns the t-th sample to the true class, γ represents the adjustment factor; σ t is the prediction uncertainty of the t-th sample, y t is the true label of the t-th sample, is the predicted value of the t-th sample.

[0082] In some embodiments, by adjusting the weight coefficients λ1 and λ2, the proportions of the focal loss and the Bayesian optimization loss in the total loss can be flexibly adjusted, achieving a better balance in dealing with class imbalance and optimizing model parameters, thereby realizing the optimal performance in different application scenarios.

[0083] It should be noted that for the Bayesian optimization loss term in this embodiment, in the coronary heart disease risk prediction model, determining the prediction uncertainty (σ i ) of the sample is a crucial step. Here, it can be achieved through a Bayesian neural network (BNN). By introducing a probability distribution over the model parameters in the Bayesian neural network, uncertainty is considered during the prediction process.

[0084] Specifically, the Bayesian neural network quantifies the model uncertainty by introducing a probability distribution over the weights and biases of the neural network. The prediction uncertainty can be divided into model uncertainty and aleatoric uncertainty.

[0085] Model uncertainty stems from the variation of model parameters, which is achieved by sampling the weights of the neural network and can be estimated using the Monte Carlo (MC) sampling method. In addition, aleatoric uncertainty comes from the noise and unpredictability of the input data and can be achieved by adding additional prediction distribution parameters (e.g., variance) to the output layer.

[0086] Exemplarily, a probability distribution is introduced over the weights and biases of the neural network, such as using a Gaussian distribution and then multiple forward propagations are performed through the Monte Carlo sampling method to calculate the mean and variance of the prediction. Suppose Q samplings are performed:

[0087]

[0088] where W q is the weight of the q-th sampling.

[0089] Furthermore, based on the sampling results, the prediction mean and the prediction uncertainty σ i are calculated through the following formula:

[0090]

[0091] In this way, based on the Bayesian neural network, the prediction uncertainty of the sample can be effectively estimated. In the coronary heart disease risk prediction model, multiple forward propagations are performed through the Monte Carlo sampling method to calculate the mean and variance of the prediction, thereby quantifying the prediction uncertainty. Thus, not only the uncertainty of the model parameters is considered, but also the uncertainty of the data itself is considered, providing the model with stronger robustness and higher prediction accuracy.

[0092] Through the comprehensive loss function provided by this embodiment, at least the following technical effects can be produced:

[0093] Based on the focal loss term, in medical image analysis, the positive samples (patients with coronary heart disease) are usually few. Through the focal loss, the model can effectively enhance the attention to these small sample categories, improving the recognition rate of positive samples. Thus, by reducing the weights of easy-to-classify samples, the model can better focus on difficult-to-classify samples, thereby improving the overall classification effect.

[0094] Based on the Bayesian optimization loss term, by introducing the uncertainty of weights using a Bayesian neural network, the uncertainty of the prediction results can be quantified, providing more reliable information for subsequent decision-making. Through multiple samplings and uncertainty evaluations, the model can more accurately adjust the parameters, reduce the prediction error, and improve the prediction accuracy.

[0095] By adjusting the weight coefficients in the loss function, the model can, according to the actual application requirements, adjust the weight coefficients of the focal loss and the Bayesian optimization loss, enabling the model to be widely applied to different medical image analysis tasks and improving the practicality and generality of the model.

[0096] Figure 3 The structural connection diagram of an example of the multi-scale attention model according to an embodiment of the present invention is shown.

[0097] As Figure 3 shown, the multi-scale attention model 300 includes a plurality (i.e., K) of cascaded attention modules 310 and a combination module 320. Each attention module 310 respectively includes a convolutional layer 311 and an attention layer 312 for extracting attention features at corresponding scales. Therefore, a corresponding attention module is assigned to each feature scale. The combination module 320 is used to determine the corresponding multi-scale attention feature map by combining the attention features of each scale, thereby obtaining the fundus plane feature group.

[0098] Through the convolutional layer 311 in each attention module 310, the multi-scale image features corresponding to the first fundus image are extracted:

[0099] F c =Conv c (I1), Equation (17)

[0100] In the formula, c = 1, 2,..., K, K represents the total number of attention modules, and the convolutional layers in each attention module respectively have corresponding convolutional kernel sizes; I1 represents the first fundus image; Conv c represents the convolutional operation of the convolutional layer in the c-th attention module, and F c represents the feature map extracted by the c-th convolutional layer.

[0101] Using multiple convolutional layers to extract feature maps of different sizes can capture the vascular and retinal structure information at different resolutions. The multi-scale representation helps to identify subtle features that are difficult to detect at a single scale.

[0102] Based on the first-scale attention module ranked first in each attention module 310, since there is no feature of the previous scale for comparison, the self-attention feature operation is performed to obtain the corresponding self-attention feature:

[0103]

[0104] G1 = [g 1,1 , g 1,2 ,..., g 1,C , Equation (19)

[0105] α1 = softmax(W a1 ·G1 + b a1 ), Equation (20)

[0106] A1 = α1·F1, Equation (21)

[0107] In the formula, G1 represents the global feature vector extracted by the first-scale attention module, W a1 and b a1 respectively represent the weight matrix and bias term of the first-scale attention module; α1 represents the attention weight corresponding to the first-scale attention module, and A1 represents the self-attention feature output by the first-scale attention module; f u,v,q represents the value of the q-th channel at the position (u, v) of the feature map F1, and g 1,q represents the value of the q-th channel after global average pooling. H, W, and C respectively represent the height, width, and number of channels of F1.

[0108] Here, first, through the Global Average Pooling (GAP) operation, by averaging the spatial dimension of the entire feature map, a global feature vector is generated, which can significantly reduce the number of parameters and effectively avoid overfitting. In addition, the self-attention mechanism of the first scale can assign different weights to each position of the feature map, highlighting the features of key regions and suppressing noise and irrelevant information.

[0109] Based on each attention module ranked non-first, the calculated attention features are respectively fused with the attention features output by the previous attention module for attention calculation, thereby establishing the connection between different-scale features:

[0110] Q e = F e W Qe , Ke-1 = A e-1 W Ke , V e-1 = A e-1 W Ve , Equation (22)

[0111] In the formula, Q e represents the query matrix output by the e-th attention module, K e-1 and V e-1 respectively represent the key matrix and value matrix output by the (e - 1)-th attention module; W Qe , W Ke and W Ve respectively represent the query weight matrix, key weight matrix, and value weight matrix of the e-th attention module; α e,g represents the attention weight of the e-th attention module to the g-th position, e > 1.

[0112] The attention weight is determined by calculating the cosine similarity:

[0113]

[0114] The attention feature is calculated according to the attention weight:

[0115]

[0116] In the formula, A e represents the attention feature of the e-th scale.

[0117] Through this embodiment, the subsequent scale establishes the connection between different scale features by calculating the similarity with the attention feature of the previous scale, enabling the model to better understand and integrate multi-scale information, and further enhancing the feature representation ability.

[0118] After pooling and concatenating the attention features of all scales, a complete multi-scale attention feature map is obtained:

[0119] P e = Pooling(A e ), Equation (25)

[0120] F MSA = Concat(P1, P2,..., P K ), Equation (26)

[0121] In the formula, P e represents the pooling result of the attention feature of the e-th scale, and F MSA represents the multi-scale attention feature map.

[0122] Through this embodiment, features of different scales are fused through the attention mechanism, enabling the model to consider both local and global information simultaneously, and improving the model's performance in processing complex and diverse fundus images. In addition, the attention mechanism weights according to the importance of the features, enabling the model to effectively suppress unimportant or interfering features while focusing on key features, achieving the ability of adaptive weighting, and enhancing the model's stability under different image qualities and shooting conditions.

[0123] Figure 4 FIG. shows a schematic structural connection diagram of an example of a region-guided attention model according to an embodiment of the present invention.

[0124] The Region-Guided Attention (RGA) model aims to effectively extract important features in optical coherence tomography angiography (OCT-A) images, especially features concentrated in the central macular area, while suppressing the interference of background noise on the model.

[0125] As Figure 4 shown, the region-guided attention model 400 includes a plurality (i.e., U) of region attention modules 410 and a fusion module 420. Each region attention module 410 includes a cascaded region convolutional layer 411 and a region-guided attention layer 412 for extracting region-guided attention features of the corresponding region. The fusion module 420 is used to fuse the region-guided attention features of each region to obtain the corresponding fundus stereo feature group.

[0126] Through the region convolutional layer 411 in each region attention module 410, the region convolutional feature maps of the second fundus image in each specific region are extracted:

[0127] R l = Conv l (I2), Equation (27)

[0128] In the formula, l = 1, 2,..., U, U is the total number of region attention modules, and each region attention module is respectively used to process the corresponding defined pixel region; I2 represents the second fundus image; R l represents the region convolutional feature map extracted by the region convolutional layer in the l-th region attention module, and Conv l represents the convolution operation of the region convolutional layer in the l-th region attention module.

[0129] By separately extracting the feature differences of different regions in the second fundus image through multiple region convolutional layers, the macular details at different resolutions can be captured, local feature calculations can be performed for the details of different regions, enabling the model to more carefully focus on the features of each region, and further enhancing the feature representation ability.

[0130] Through the region guiding attention layer 412 in each region attention module 410, the region convolutional feature map is weighted using the membrane layer region image to obtain the corresponding region guiding feature map:

[0131] R lM = R l ·M, Equation (28)

[0132] In the formula, M represents the membrane layer region image, and R lM represents the region guiding feature map determined by the region guiding attention layer in the l-th region attention module.

[0133] It should be noted that the membrane layer region image is extracted from the optical coherence tomography angiography (OCT-A) image through image segmentation technology. It is a binary image containing the membrane layer region, which represents the important membrane layer region in the OCT-A image. The value of 1 in the image represents the membrane layer region, and the value of 0 represents the background region. The membrane layer region image can be achieved through fixed threshold segmentation, global threshold segmentation, or k-means clustering method. More preferably, it can be achieved through adaptive threshold segmentation, and more details will be elaborated in combination with other examples below.

[0134] Through the region guiding attention mechanism, the feature map is weighted using the membrane layer region image, enabling the model to pay attention to the membrane layer region features while suppressing the influence of background noise, thereby increasing the robustness of the model.

[0135] The region guiding feature map R lM is divided into multiple sub-region guiding feature maps where s represents the s-th sub-region.

[0136] Here, further dividing the sub-regions based on the region guiding feature map can achieve image segmentation with differential feature performance within the region, effectively enhancing the feature representation ability of different sub-regions.

[0137] For each sub-region guiding feature map the corresponding local feature vector is obtained through local average pooling

[0138] The local feature vector is calculated through a fully connected layer to obtain the corresponding local attention weight

[0139]

[0140] In the formula, respectively represent the weight matrix and bias term for the attention weight calculation of the fully connected layer for the l-th region and the s-th sub-region, Indicates the attention weight of the s-th sub-region in the l-th region.

[0141] Through this embodiment, calculating the attention weight based on the local feature vector enables the model to adaptively adjust the degree of attention to different regions, improving the generalization ability of the model in different image scenarios.

[0142] For each sub-region guiding feature map Perform weighting to obtain the sub-region guiding attention feature:

[0143]

[0144] In the formula, Indicates the sub-region guiding attention feature of the s-th sub-region in the l-th region.

[0145] Through this embodiment, by using multi-scale feature extraction and local feature calculation, the model can capture the subtle features in the OCT-A image, improving the recognition ability of coronary heart disease-related lesions.

[0146] Concatenate the attention features of all sub-regions to obtain the region guiding attention feature map, and fuse it through the fusion module 420:

[0147]

[0148] F RGA = Concat(F RGA,1 , F RGA,2 ,..., F RGA,U ), Equation (32)

[0149] In the formula, F RGA,l Indicates the region guiding attention feature map of the l-th region, S represents the total number of sub-regions included in the l-th region, and F RGA Indicates the fusion feature based on the region guiding attention features of each region.

[0150] Through this embodiment, the attention weight calculation method based on the local feature vector enables the model to dynamically adjust the attention to different regions, improving the accuracy of feature representation and the accuracy of prediction. In addition, the local feature calculation and weighting operations can effectively reduce the consumption of computing resources by processing the features of each sub-region in parallel.

[0151] Figure 5 Shows an operation flowchart of an example for determining the membrane layer region image through the adaptive threshold segmentation method according to an embodiment of the present invention.

[0152] As Figure 5As shown, in step S510, for each pixel in the second fundus image, the local mean within the window centered on this pixel is calculated.

[0153] More specifically, it is calculated by the following formula:

[0154]

[0155] In the formula, (i,j) represents the position of the pixel in the image, μ(i,j) represents the local mean pixel value at (i,j), I2(i+m,j+n) represents the pixel value of the second fundus image I2 at the position (i+m,j+n), and w represents the window size.

[0156] In step S520, the local standard deviation is calculated.

[0157] More specifically, it is calculated by the following formula:

[0158]

[0159] In the formula, σ(i,j) represents the local standard deviation at (i,j).

[0160] In step S530, an adaptive threshold is calculated using the local mean and standard deviation.

[0161] More specifically, it is calculated in the following way:

[0162]

[0163] In the formula, r represents the adjustment parameter, and T(i,j) represents the adaptive threshold at (i,j).

[0164] In step S540, the second fundus image is binarized according to the adaptive threshold to obtain the membrane layer region image.

[0165] More specifically, it is processed in the following way:

[0166]

[0167] In the formula, M(i,j) represents the pixel value of the membrane layer region image M at (i,j), and I2(i,j) represents the pixel value of the image I2 at (i,j).

[0168] In the regional attention guidance model, the key steps of the adaptive threshold segmentation algorithm involve local standard deviation calculation, adaptive threshold calculation, and generation of the membrane layer region image. The local standard deviation σ(i,j) is used to measure the degree of dispersion of pixel values within a specific window (centered on pixel (i,j)), which reflects the variation of pixel values within this region, that is, the degree to which pixel values deviate from their local mean μ(i,j). The adaptive threshold T(i,j) is used to dynamically determine the threshold for segmenting the image based on local statistics (mean and standard deviation) to distinguish the membrane layer region from the background region, enabling this threshold to be adaptively adjusted to suit the image characteristics of different regions.

[0169] Through the embodiments of the present invention, by using the adaptive threshold segmentation algorithm, the threshold can be dynamically adjusted to adapt to the image characteristics of different regions, effectively separating the membrane layer region, reducing the interference of background noise, improving the accuracy of feature extraction, ensuring the accurate separation of the membrane layer region and the background region, thereby enhancing the accuracy of coronary heart disease risk prediction. In addition, by measuring the degree of dispersion of local pixel values and dynamically adjusting the segmentation threshold to adapt to the different regional characteristics of the image, it has stronger generalization ability.

[0170] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of actions combined. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention. In the above embodiments, each embodiment is described with emphasis. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0171] Figure 6 The structural block diagram of an example of a coronary heart disease risk prediction system based on multimodal fundus images according to an embodiment of the present invention is shown.

[0172] As Figure 6 shown, the coronary heart disease risk prediction system 600 based on multimodal fundus images includes a multimodal data acquisition unit 610, a planar feature extraction unit 620, a stereoscopic feature extraction unit 630, and a risk prediction unit 640.

[0173] The multimodal data acquisition unit 610 is used to acquire a group of multimodal fundus images corresponding to a target user; the group of multimodal fundus images includes a first fundus image and a second fundus image; the first fundus image is obtained by means of color fundus photography, and the second fundus image is obtained by means of optical coherence tomography angiography.

[0174] The planar feature extraction unit 620 is configured to extract the fundus planar feature group corresponding to the first fundus image based on a multi-scale attention model.

[0175] The three-dimensional feature extraction unit 630 is configured to extract the fundus three-dimensional feature group corresponding to the second fundus image based on a region-guided attention model.

[0176] The risk prediction unit 640 is configured to input the fundus planar feature group and the fundus three-dimensional feature group into a risk prediction model to determine the corresponding coronary heart disease risk prediction result, and the risk prediction model adopts a multi-modal fusion deep learning model.

[0177] In some embodiments, the embodiments of the present invention provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute the above-mentioned coronary heart disease risk prediction method based on multi-modal fundus images of the present invention.

[0178] In some embodiments, the embodiments of the present invention further provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer is enabled to execute the above-mentioned coronary heart disease risk prediction method based on multi-modal fundus images.

[0179] In some embodiments, the embodiments of the present invention further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the coronary heart disease risk prediction method based on multi-modal fundus images.

[0180] Figure 7 FIG. is a schematic hardware structure diagram of an electronic device for executing the coronary heart disease risk prediction method based on multi-modal fundus images provided by another embodiment of the present invention. As Figure 7 shown, the device includes:

[0181] One or more processors 710 and a memory 720, Figure 7 Taking one processor 710 as an example.

[0182] The device for executing the coronary heart disease risk prediction method based on multi-modal fundus images may further include: an input device 730 and an output device 740.

[0183] The processor 710, the memory 720, the input device 730, and the output device 740 can be connected through a bus or other means. Figure 7 Taking the connection through the bus as an example.

[0184] As a non-volatile computer-readable storage medium, the memory 720 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for predicting the risk of coronary heart disease based on multimodal fundus images in the embodiments of the present invention. By running the non-volatile software programs, instructions, and modules stored in the memory 720, the processor 710 executes various functional applications and data processing of the server, that is, implements the method for predicting the risk of coronary heart disease based on multimodal fundus images in the above method embodiments.

[0185] The memory 720 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 720 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 720 optionally includes a memory remotely set relative to the processor 710, and these remote memories can be connected to the electronic device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0186] The input device 730 can receive input digital or character information, and generate signals related to the user settings and function control of the electronic device. The output device 740 can include a display device such as a display screen.

[0187] The one or more modules are stored in the memory 720, and when executed by the one or more processors 710, execute the method for predicting the risk of coronary heart disease based on multimodal fundus images in any of the above method embodiments.

[0188] The above product can execute the method provided by the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiments of the present invention.

[0189] The electronic device in the embodiments of the present invention exists in various forms, including but not limited to:

[0190] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aiming to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0191] (2) Ultra-mobile personal computer devices: Such devices fall within the category of personal computers, have computing and processing capabilities, and generally also have the feature of mobile Internet access. Such terminals include: PDA, MID, and UMPC devices, etc.

[0192] (3) Portable entertainment devices: Such devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, as well as smart toys and portable in-vehicle navigation devices.

[0193] (4) Other airborne electronic devices with data interaction functions, such as in-vehicle device installed on vehicles.

[0194] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0195] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solutions, or rather the part that contributes to the related technologies, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0196] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A method for predicting the risk of coronary heart disease based on multimodal fundus images, characterized in that, The method includes: Obtaining a multi-modal fundus image group corresponding to a target user; the multi-modal fundus image group includes a first fundus image and a second fundus image; the first fundus image is obtained by color fundus photography, and the second fundus image is obtained by optical coherence tomography angiography; Based on a multi-scale attention model, extracting a fundus plane feature group corresponding to the first fundus image; Based on a region-guided attention model, extracting a fundus three-dimensional feature group corresponding to the second fundus image; Inputting the fundus plane feature group and the fundus three-dimensional feature group into a risk prediction model to determine a corresponding coronary heart disease risk prediction result, and the risk prediction model uses a multi-modal fusion deep learning model; The risk prediction model includes a feature fusion layer, a deep residual network layer, a bidirectional long short-term memory network layer, a multi-head self-attention layer, and a fully connected output layer; The feature fusion layer is used to fuse the fundus plane feature group and the fundus three-dimensional feature group: ; In the formula, represents the fundus plane feature group, represents the fundus three-dimensional feature group, and represents the fundus fusion feature group; The deep residual network layer includes multiple residual blocks, and each residual block includes multiple convolutional layers and short connections between different convolutional layers; the deep residual network layer is used to extract deep convolutional features corresponding to the fundus fusion feature group: ; In the formula, represents the convolution operation in the residual block, represents the weight of the -th convolutional layer in the residual block, and represents the deep convolutional feature; The bidirectional long short-term memory network layer is used to determine bidirectional temporal dependence features corresponding to the deep convolutional features: ; In the formula, represents the two-way temporal dependence feature, represents the output feature of the forward LSTM of represents the output feature of the backward LSTM of The multi-head self-attention layer is used to process the bidirectional temporal dependence features to obtain corresponding global attention features: ; ; ; ; ; In the formula, , and respectively represent the query weight matrix, the key weight matrix, and the value weight matrix, , and respectively represent the query matrix, the key matrix, and the value matrix; represents the softmax function, represents the dimension of the key vector, represents the i -th output of the attention head; represents the multi-head output weight matrix, h represents the total number of attention heads, represents the global attention feature; The fully connected output layer is used to process the global attention features to obtain a corresponding coronary heart disease risk probability: ; ; Wherein, represents the final output probability of coronary heart disease risk, and respectively represent the weight matrix and bias term of the fully connected layer; represents a non-linear activation function, and respectively represent the weight matrix and bias term of the output layer; represents the Sigmoid activation function; The region-guided attention model includes multiple region attention modules and a fusion module; each region attention module includes a cascaded region convolutional layer and a region-guided attention layer to extract region-guided attention features of a corresponding region; the fusion module is used to fuse the region-guided attention features of each region to obtain a corresponding fundus three-dimensional feature group; Through the region convolutional layer in each region attention module, extracting a region convolutional feature map of the second fundus image in each specific region: ; In the formula, , U is the total number of regional attention modules, and each regional attention module is used to process the corresponding delimited pixel region; represents the second fundus image; represents the regional convolution feature map extracted by the regional convolution layer in the th regional attention module, represents the convolution operation of the regional convolution layer in the th regional attention module; Through the region-guided attention layer in each region attention module, using a membrane layer region image to weight the region convolutional feature map to obtain a corresponding region-guided feature map: ; In the formula, Represents the membrane area image, Indicates A regional guided feature map determined by a regional guided attention layer in a regional attention module; Divide the region guiding feature map into multiple sub-region guiding feature maps , where represents the th sub-region; For each sub-region guiding feature map , the corresponding local feature vector is obtained through local average pooling : Calculate the local feature vector through the fully connected layer The corresponding local attention weight : ; In the formula, and respectively represent the weight matrix and bias term for calculating the attention weight of the fully connected layer for the th region and the th sub-region, represents the attention weight of the th region and the th sub-region; Guide the feature map for each sub-region Perform weighting to obtain the sub-region guided attention feature: ; In the formula, represents the sub-region guiding attention feature of the th sub-region in the th region; Stitching the attention features of all sub-regions to obtain a region-guided attention feature map and fusing it through the fusion module: ; ; In the formula, represents the region-guided attention feature map of the th region, S represents the total number of sub-regions included in the th region, represents the fusion feature based on the region-guided attention features of each region.

2. The method according to claim 1, characterized in that The fundus plane feature group includes at least one of the following: cup-to-disc ratio of the left and right eyes, vertical cup-to-disc ratio, optic cup area, optic cup volume, optic disc area, average thickness of the retinal nerve fiber layer, area of the foveal avascular zone, perimeter of the foveal avascular zone, and near-circularity index.

3. The method according to claim 1 or 2, characterized in that, The fundus three-dimensional feature group includes at least one of the following: total vessel diameter, artery diameter, vein diameter, total vessel fractal dimension, artery fractal dimension, vein fractal dimension, total vessel tortuosity, artery tortuosity, vein tortuosity, total vessel tortuosity density, artery tortuosity density, vein tortuosity density, and total vessel branching angle.

4. The method according to claim 1, wherein The loss function of the risk prediction model is as follows: ; ; ; Wherein, represents the focal loss term, represents the Bayesian optimization loss term, and respectively represent the importance coefficients of the corresponding loss terms; N represents the total number of samples in the data sample set; represents the -th sample's class weight, is the predicted probability that the model assigns the -th sample to the true class, represents the adjustment factor; is the prediction uncertainty of the -th sample, is the true label of the -th sample, is the predicted value of the -th sample.

5. The method according to claim 4, wherein The multi-scale attention model includes a plurality of cascaded attention modules and combination modules; each of the attention modules includes a convolutional layer and an attention layer for extracting attention features at corresponding scales; the combination module is used to determine corresponding multi-scale attention feature maps by combining the attention features of each scale, thereby obtaining a fundus plane feature group; Extract the multi-scale image features corresponding to the first fundus image through the convolutional layers in each of the attention modules: ; In the formula, , K represents the total number of attention modules, and the convolutional layers in each attention module have corresponding convolutional kernel sizes respectively; represents the first fundus image; represents the convolution operation of the convolutional layer in the th attention module, represents the feature map extracted by the th convolutional layer; Based on the first-scale attention module ranked first in each of the attention modules, perform self-attention feature operations to obtain corresponding self-attention features: ; ; ; ; In the formula, represents the global feature vector extracted by the first-scale attention module, and represent the weight matrix and bias term of the first-scale attention module respectively; represents the attention weight corresponding to the first-scale attention module, represents the self-attention feature output by the first-scale attention module; represents the feature map at the position the value of the th channel, represents the value of the th channel after global average pooling, , and represent respectively the height, width and number of channels; Based on each of the attention modules ranked non-first, respectively perform fusion attention calculations on the calculated attention features and the attention features output by the previous attention module, thereby establishing connections between features of different scales: ; In the formula, represents the query matrix output by the th attention module, and respectively represent the key matrix and value matrix output by the th attention module; , and respectively represent the query weight matrix, key weight matrix and value weight matrix of the th attention module; represents the attention weight of the th attention module to the th position, e > 1; Determine the attention weights by calculating the cosine similarity: ; Calculate the attention features according to the attention weights: ; In the formula, represents the attention feature of the th scale; Perform pooling operations on the attention features of all scales and then splice them to obtain a complete multi-scale attention feature map: ; ; In the formula, represents the pooling result of the attention features of the th scale, represents the multi-scale attention feature map.

6. The method according to claim 1, wherein The membrane layer region image is determined by the following method: For each pixel in the second fundus image, calculate the local mean within the window centered on this pixel: ; In the formula, represents the position of a pixel in the image, represents the local average pixel value at represents the second fundus image at the position the pixel value at, represents the window size; Calculate the local standard deviation: ; In the formula, represents the local standard deviation at Calculate the adaptive threshold using the local mean and standard deviation: ; In the formula, represents an adjustment parameter, represents the adaptive threshold at According to the adaptive threshold Perform binarization processing on the second fundus image to obtain the membrane layer region image: ; In the formula, represents the image of the film layer area The pixel value at And represents the image The pixel value at And 7. A coronary heart disease risk prediction system based on multimodal fundus images, characterized in that, The system includes: A multi-modal data acquisition unit for acquiring a multi-modal fundus image group corresponding to a target user; the multi-modal fundus image group includes a first fundus image and a second fundus image; the first fundus image is obtained by color fundus photography, and the second fundus image is obtained by optical coherence tomography angiography; A plane feature extraction unit for extracting the fundus plane feature group corresponding to the first fundus image based on the multi-scale attention model; A three-dimensional feature extraction unit for extracting the fundus three-dimensional feature group corresponding to the second fundus image based on the region-guided attention model; A risk prediction unit for inputting the fundus plane feature group and the fundus three-dimensional feature group into a risk prediction model to determine the corresponding coronary heart disease risk prediction result, and the risk prediction model uses a multi-modal fusion deep learning model; The risk prediction model includes a feature fusion layer, a deep residual network layer, a bidirectional long short-term memory network layer, a multi-head self-attention layer, and a fully connected output layer; The feature fusion layer is used to fuse the fundus plane feature group and the fundus three-dimensional feature group: ; In the formula, represents the fundus plane feature group, represents the fundus three-dimensional feature group, and represents the fundus fusion feature group; The deep residual network layer includes a plurality of residual blocks, and each residual block includes a plurality of convolutional layers and short connections between different convolutional layers; the deep residual network layer is used to extract deep convolutional features corresponding to the fundus fusion feature group; ; In the formula, represents the convolution operation in the residual block, represents the weight of the -th convolutional layer in the residual block, and represents the deep convolutional feature; The bidirectional long short-term memory network layer is used to determine the bidirectional temporal dependence features corresponding to the deep convolutional features: ; In the formula, represents the two-way temporal dependence feature, represents the output feature of the forward LSTM of represents the output feature of the backward LSTM of The multi-head self-attention layer is used to process the bidirectional temporal dependence features to obtain corresponding global attention features: ; ; ; ; ; Wherein, , and respectively represent a query weight matrix, a key weight matrix, and a value weight matrix, , and respectively represent a query matrix, a key matrix, and a value matrix; represents the softmax function, represents the dimension of the key vector, represents the i -th output of the attention head; represents the multi-head output weight matrix, h represents the total number of attention heads, represents the global attention feature; The fully connected output layer is used to process the global attention features to obtain the corresponding coronary heart disease risk probability: ; ; Wherein, represents the final output probability of coronary heart disease risk, and respectively represent the weight matrix and bias term of the fully connected layer; represents the non-linear activation function, and respectively represent the weight matrix and bias term of the output layer; represents the Sigmoid activation function; The described region-guided attention model includes multiple region attention modules and a fusion module; each of the region attention modules includes a cascaded region convolutional layer and a region-guided attention layer for extracting region-guided attention features of corresponding regions; the fusion module is used to fuse the region-guided attention features of each region to obtain a corresponding fundus stereo feature group; Through the region convolutional layers in each region attention module, region convolutional feature maps of the second fundus image in each specific region are extracted: ; In the formula, , U is the total number of regional attention modules, and each regional attention module is used to process the corresponding defined pixel region; represents the second fundus image; represents the regional convolutional feature map extracted by the regional convolutional layer in the th regional attention module, represents the convolution operation of the regional convolutional layer in the th regional attention module; Through the region-guided attention layers in each region attention module, the region convolutional feature maps are weighted using the membrane layer region images to obtain corresponding region-guided feature maps: ; In the formula, Represents the membrane area image, Indicates A regional guided feature map determined by a regional guided attention layer in a regional attention module; Divide the region guiding feature map into multiple sub-region guiding feature maps , where represents the th sub-region; For each sub-region guiding feature map , the corresponding local feature vector is obtained through local average pooling : Calculate the local feature vector through the fully connected layer The corresponding local attention weight : ; In the formula, , respectively represent the weight matrix and bias term for calculating the attention weight of the fully connected layer for the th region and the th sub-region, represents the attention weight of the th region and the th sub-region; Guide the feature map for each sub-region Perform weighting to obtain the sub-region guided attention feature: ; In the formula, represents the sub-region guiding attention feature of the th region and the th sub-region; The attention features of all sub-regions are concatenated to obtain a region-guided attention feature map, which is fused through the fusion module: ; ; In the formula, represents the region-guided attention feature map of the th region, S represents the total number of sub-regions included in the th region, represents the fused feature based on the region-guided attention features of each region.

Citation Information

Patent Citations

  • Hypertension risk prediction method, device, and equipment and medium

    CN113689954A

  • Method and system for screening coronary heart disease of type 2 diabetic patient based on retina morphology

    CN118098557A