System and method for screening multi-system diseases of the whole body based on a fundus color photo U-Transformer model

Through the U-Transformer model based on fundus color imaging, the limitations of fundus images in the prior art in screening of systemic diseases are solved, and the precise screening and classification of seven major systemic diseases in the whole body are achieved, which improves detection efficiency and real-time performance.

CN118781406BActive Publication Date: 2025-05-30GUANGDONG GENERAL HOSPITAL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410834324.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-26
Publication Date
2025-05-30
Estimated Expiration
2044-06-26

AI Technical Summary

Technical Problem

The prior art has problems such as limited disease type, insufficient feature extraction, and lack of secondary classification and refinement screening when using fundus images to screen systemic diseases, resulting in high cost, low efficiency and difficulty in precise positioning.

Method used

U-Transformer model based on fundus color illumination is adopted to establish and train models through deep learning methods to achieve efficient screening of fundus images. The model is constructed from an improved U-Net hybrid Transformer network, combined with gradient integral algorithm for visual analysis and feature positioning, improving the performance and recognition accuracy of the model.

Benefits of technology

It has achieved accurate screening and classification of seven major systemic diseases throughout the body, reduced testing costs, improved testing efficiency, simplified operating procedures, and significantly optimized real-time, suitable for the use of clinical emergencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118781406B_ABST
    Figure CN118781406B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of disease screening. More specifically, it relates to a method and system for screening multi-system diseases of the whole body based on a fundus color photo U-Transformer model, including: S1: Collecting fundus images of patients and data on the diagnosis of their systemic diseases in a public database to make a data set; S2: Performing data preprocessing on the data set obtained in the above step; S3: Inputting the preprocessed data into the U-Transformer model for training; S4: Inputting the fundus color photo to be tested into the trained U-Transformer model for detection to obtain the screening results of systemic diseases; The present invention realizes that only by inputting the fundus color photo into the model for detection can the diagnosis results of systemic diseases of the whole body be obtained, greatly reducing the detection cost and improving the detection efficiency. Moreover, this detection method is non-invasive, making the operation more convenient. It can not only achieve binary classification of health or disease, but also accurately screen and classify seven major categories of systemic diseases of the whole body.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of disease screening, and more particularly, to a method and system for screening multi-system diseases of the whole body based on a fundus color photo U-Transformer model. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, especially deep learning, the research of AI in the field of ophthalmology shows a trend of diversified disease types, extensive scenarios and in-depth research. AI has shown good performance in the research of ophthalmic diseases such as diabetic retinopathy, age-related macular degeneration, glaucoma, etc., demonstrating the great potential of ophthalmic AI. Fundus photos have been widely used for diagnosis and monitoring in the field of ophthalmology. Traditionally, doctors mainly use fundus color photos to detect eye diseases. However, these traditional methods mainly rely on the experience and skills of professional doctors and only focus on the diagnosis of eye diseases. But in clinical practice, many patients with systemic diseases have varying degrees of eye symptoms and signs, such as jaundice, corneal pigmentation ring, dry eye, etc., and the abnormal rate of fundus blood vessels is also significantly higher than that of the general population. Research shows that fundus images contain important information about the overall health status of the body, including but not limited to cardiovascular diseases, diabetes, hypertension, etc. And currently, the screening of multi-systemic diseases mainly relies on biochemical markers and imaging examinations, which are often invasive, costly and not easily popularized. Therefore, fundus examination, as a non-invasive, low-cost and easy-to-operate diagnostic method, provides a new way for the early screening of diseases.

[0003] However, the existing technology still has limitations in using fundus images for screening systemic diseases of the whole body. It is mostly limited to the auxiliary diagnosis application of single diseases such as cardiovascular diseases, and has disadvantages such as limited and single disease types for screening, insufficient feature extraction, and lack of secondary classification and refined screening. These problems may lead to adverse factors such as high cost, low efficiency and long time consumption, restricting the wide application of systemic disease screening of the whole body, and there are great deficiencies in the comprehensive screening of systemic diseases. In addition, the insufficient feature extraction will also make it impossible to accurately capture the tiny features related to multi-systemic diseases in fundus images, and the lack of methods for secondary classification and refined screening leads to difficulties in accurately positioning the systemic diseases to which the patients belong.

[0004] Therefore, the present invention establishes a method and system for screening multi-system diseases of the whole body based on a fundus color photo U-Transformer model to solve the above problems, and realizes the efficient screening of seven systemic diseases of the whole body, including cardiovascular and cerebrovascular diseases, nervous system diseases, mental diseases, urinary system diseases, respiratory system diseases, digestive system diseases and tumors. Summary of the Invention

[0005] The present invention aims to overcome at least one defect (shortcoming) of the above-mentioned prior art, and provides a method and system for screening multi-system diseases of the whole body based on a fundus color photo U-Transformer model, which is used to solve the problems in the prior art that when using fundus images for screening multi-system diseases of the whole body, the types of diseases that can be screened are limited and single, the feature extraction is insufficient, and there is a lack of secondary classification and refined screening, etc.

[0006] In the first aspect, the technical solution adopted by the present invention is a method for screening multi-system diseases of the whole body based on a fundus color photo U-Transformer model, including:

[0007] S1: Collect the fundus images of patients and the data of their systemic disease diagnoses in a public database to make a data set;

[0008] S2: Perform data preprocessing on the data set obtained in the above step;

[0009] S3: Input the preprocessed data into the U-Transformer model for training;

[0010] S4: Input the fundus color photo to be tested into the trained U-Transformer model for detection to obtain the screening results of systemic diseases;

[0011] This technical solution uses deep learning methods to establish and train a model, realizing that only by inputting the fundus color photo into the model for detection can the diagnosis results of systemic diseases of the whole body be obtained, greatly reducing the detection cost and improving the detection efficiency. And this detection method is non-invasive, making the operation more convenient. It can not only realize the binary classification of health or disease, but also accurately screen and classify seven major types of systemic diseases of the whole body.

[0012] Preferably, in step S3, it also includes using the gradient integral algorithm to perform visual analysis to accurately locate the imaging markers of different systemic diseases in the fundus image. Thus, according to the visual analysis results, the training range is targeted to be narrowed to the specific fundus regions of different systemic diseases to train the model and improve the performance of the network model.

[0013] Preferably, the U-Transformer model is constructed by an improved U-Net hybrid Transformer network, including:

[0014] Use U-Net to extract several feature information of the fundus color photos in the data set;

[0015] The feature map containing the above-mentioned several feature information extracted is used as a key value and input into the Transformer module. Then, through the decoder, the feature map is upsampled to the size of the original image, and a bridging layer is generated as the input of the TransformerHead.

[0016] The Transformer Head merges the features of the bridging layer through a convolutional layer, flattens the merged features into a sequence and passes it to the multi-head attention mechanism MHA, and passes the output of the multi-head attention mechanism MHA to the multi-layer perceptron MLP for processing.

[0017] The output of the multi-layer perceptron MLP is linearly upsampled and processed through the convolutional layer in the CBR block, thereby outputting the final prediction, and a U-Transformer model based on fundus color photos is constructed.

[0018] The architecture of the model includes multiple convolutional layers, pooling layers and fully connected layers. The specific settings of these layers are optimized for fundus image feature extraction, and an appropriate filter size, stride and activation function are selected for each layer to maximize the efficient extraction and representation of fundus features, and the inductive bias of convolutional images is used to avoid large-scale pre-training, effectively utilizing the ability of the Transformer to capture global feature relationships, realizing better localization of distinguishable features of diseases, and improving the accuracy of model recognition.

[0019] Preferably, the use of the gradient integration algorithm for visual analysis to accurately locate the imaging markers of different systemic diseases in fundus images includes:

[0020] S31: Set the fundus image as a completely black fundus image with pixel intensity of 0 as the reference input, apply the integrated gradient algorithm to each of the RGB channels separately for summation calculation to obtain the integration result, and the integration result is the attribution, thereby obtaining the pixel points that provide important contributions to predicting different systemic diseases in the fundus image.

[0021] S32: Grayscale the fundus image, and then add different attribution results to different color channels, thereby highlighting the distinguishable features of different systemic diseases in the input fundus image.

[0022] The gradient integration algorithm is used to attribute the category with the highest score in the output prediction categories, so as to study which pixel points provide the most important contributions when the network makes predictions. Then, different attribution results are added to different color channels, thereby locating the pixel points similar to the lesion nodes, highlighting the distinguishable features of the input pattern, and further enabling the model to learn more useful information during training, improving the recognition accuracy of the model.

[0023] Preferably, in the step S2, the data preprocessing includes:

[0024] S21: Crop the invalid area of the full dataset images to remove the invalid information in the images;

[0025] S22: Perform normalization on the cropped images;

[0026] S23: Perform image enhancement on the normalized images;

[0027] S24: Perform image stitching to fuse the binocular fundus color photos into one image to obtain binocular fundus information to the greatest extent.

[0028] When training the model, preprocessing the data first can effectively remove the invalid information in the images, enhance the detectability of relevant information, and simplify the data to the greatest extent, thereby improving the accuracy and reliability of feature extraction, image segmentation, matching, and recognition.

[0029] Further preferably, in step S23, the image enhancement at least includes:

[0030] Perform random horizontal translation and rotation on the normalized images;

[0031] Perform local average color removal on the normalized images to weaken image noise.

[0032] Through image enhancement operations, the visual effect of the images can be improved, the specific areas or features of the recognizable diseases in the images can be emphasized to make them more prominent and easy to identify. At the same time, irrelevant features can be suppressed, the noise, irrelevant details, or interference information in the images can be reduced, the main features can be made more prominent, the information content of the images can be increased, thereby improving the recognition rate and accuracy of the images.

[0033] Preferably, in step S4, it also includes first performing binary classification detection to divide into two major categories of healthy and diseased. If the binary classification detection result is healthy, the detection ends and it is prompted that the fundus color photo to be tested is healthy; if the binary classification detection result is diseased, perform systematic disease refinement detection and output the corresponding systematic disease diagnosis result.

[0034] Inputting the fundus images into the model for detection can obtain the health status of the body system of the input fundus images, effectively improving the efficiency of systematic disease detection, shortening the average analysis time to half of the original technology, and significantly optimizing the real-time performance, making it not only applicable to normal detections but also more suitable for use in clinical emergency situations, simplifying the operation process of doctors, and significantly improving the clinical work efficiency.

[0035] Preferably, supervised learning is applied during the model training process, and cross-validation is used to optimize the model parameters.

[0036] Using supervised learning to train a model by analyzing labeled training data, enabling the model to learn the relationship between inputs and outputs, so as to predict or classify new, unseen data, and realizing the ability to learn potential patterns and regularities from a large amount of fundus image data and its systemic disease diagnosis data, improving the prediction ability of the model. And cross-validation is used to reduce the risk of overfitting during model training, improving the reliability and generalization performance of the model.

[0037] In a second aspect, the present invention also provides a multi-system disease screening system for the whole body based on a fundus color photo U-Transformer model, including:

[0038] A data acquisition module: used to collect fundus images of patients and their systemic disease diagnosis data in a public database to make a data set;

[0039] A data preprocessing module: used to preprocess the data set obtained in the data acquisition module;

[0040] A model construction and training module: used to input the preprocessed data into the constructed U-Transformer model for training;

[0041] A detection module: used to input the fundus color photo to be tested into the trained U-Transformer model for detection to obtain the system disease screening result;

[0042] In this system, it is realized that only by inputting the fundus color photo can the diagnosis result of systemic diseases be obtained, greatly reducing the detection cost and improving the detection efficiency, and this detection method is non-invasive, making the operation more convenient.

[0043] The model construction and training module also includes a model construction unit, which is used to construct a U-Transformer model by mixing an improved U-Net and a Transformer network, including:

[0044] Using U-Net to extract several feature information of the fundus color photo in the data set;

[0045] Taking the feature map containing the above-mentioned several feature information obtained by extraction as the key value and inputting it into the Transformer module, and upsampling the feature map to the original image size through a decoder, and generating a bridging layer as the input of the TransformerHead;

[0046] The Transformer Head combines the features of the bridging layer through a convolutional layer, flattens the combined features into a sequence and passes it to the multi-head attention mechanism MHA, and passes the output of the multi-head attention mechanism MHA to the multi-layer perceptron MLP for processing;

[0047] Linearly upsample the output of the multi-layer perceptron (MLP) and process it through the convolutional layer in the CBR block to output the final prediction, thereby constructing a U-Transformer model based on fundus color photographs.

[0048] In the model construction unit, an improved U-Net hybrid Transformer network is adopted to maximize the efficient extraction and representation of fundus features, and the inductive bias of convolutional images is used to avoid large-scale pre-training. The ability of the Transformer to capture global feature relationships is effectively utilized to better localize the identifiable features of diseases and improve the accuracy of model recognition.

[0049] The model construction and training module also includes a marker localization unit. In the marker localization unit, the gradient integration algorithm is used for visual analysis to accurately locate the imaging markers of different systemic diseases in the fundus image. Then, according to the visual analysis results, the training range is specifically narrowed down to the fundus regions specific to different systemic diseases to train the model and improve the performance of the network model.

[0050] Preferably, the data preprocessing module includes a cropping unit, an image normalization and enhancement unit, and an image stitching unit, where:

[0051] Cropping unit: used to crop the invalid areas of the full dataset images and remove the invalid information in the images.

[0052] Image normalization and enhancement unit: used to perform normalization processing on the cropped images and perform image enhancement on the normalized images.

[0053] Image stitching unit: used to perform image stitching, fusing the fundus color photographs of both eyes into one image to obtain the fundus information of both eyes to the greatest extent.

[0054] Preprocessing the data through multiple preprocessing units in the data preprocessing module can effectively remove the invalid information in the images, enhance the detectability of relevant information, and simplify the data to the greatest extent, thereby improving the accuracy and reliability of feature extraction, image segmentation, matching, and recognition.

[0055] Preferably, the detection module includes a first detection output module and a second detection output module, where:

[0056] The first detection output module is used to output whether the fundus image to be tested is healthy or diseased. If the first detection result is healthy, the detection ends and it is prompted that the fundus color photograph to be tested is healthy; if the first detection result is diseased, the model is continued to be used for refined screening detection.

[0057] The second detection output module is used to output the corresponding systemic disease prevalence results after the model performs a refined detection of systemic diseases.

[0058] Through the above two detection output modules, the detection results of the input fundus images can be effectively and quickly viewed, improving the clinical work efficiency.

[0059] In a third aspect, the solution of the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the whole-body multi-system disease screening method based on the fundus color photo U-Transformer model as described in any one of the above.

[0060] The solution of the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the whole-body multi-system disease screening method based on the fundus color photo U-Transformer model as described in any one of the above.

[0061] The solution of the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the whole-body multi-system disease screening method based on the fundus color photo U-Transformer model as described in any one of the above.

[0062] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0063] (1) Through an advanced deep learning method, an improved U-Net hybrid Transformer network is used to construct a U-Transformer model, providing a hierarchical screening that surpasses traditional medical image analysis, realizing that the use of fundus color photos is not limited to detecting a single healthy or diseased binary classification, but can also accurately screen and classify seven major categories of systemic diseases in the whole body.

[0064] (2) Different from previous methods based on convolutional neural networks, the Transformer model is not only good at capturing global context information, but also shows strong adaptability to downstream tasks during large-scale pre-training. The optimized U-Transformer feature extraction algorithm significantly improves the model's ability to recognize subtle changes in fundus images, and uses the visualization analysis results of the gradient integration algorithm to specifically narrow the training range to improve network performance.

[0065] (3) The U-Transformer model has been significantly optimized in terms of real-time performance, and the average analysis time has been shortened to half of the original technology, making it more suitable for use in clinical emergencies. At the same time, the operation process is more convenient, significantly improving the clinical work efficiency.

[0066] (4) The whole-body multi-system disease screening method adopted in the present invention is a non-invasive diagnostic means, which is simple to operate and has high detection accuracy, effectively reducing the medical costs caused by delayed diagnosis and reducing the overall treatment costs of systemic diseases through early screening. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] The drawings are only for illustrative purposes and should not be construed as limitations on the present solution; for better illustration of the present solution, some components in the drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0068] Figure 1 It is a schematic flowchart of the whole-body multi-system disease screening method provided by the present solution.

[0069] Figure 2 It is a schematic flowchart of the data preprocessing of the whole-body multi-system disease screening method provided by the present solution.

[0070] Figure 3 It is a schematic diagram of the U-Transformer model of the whole-body multi-system disease screening method provided by the present solution.

[0071] Figure 4 It is a schematic diagram of the U-Net structure of the whole-body multi-system disease screening method provided by the present solution.

[0072] Figure 5 It is a schematic diagram of the Transformer module structure of the whole-body multi-system disease screening method provided by the present solution.

[0073] Figure 6 It is a schematic diagram of the detection accuracy of screening healthy and diseased individuals of the whole-body multi-system disease screening method provided by the present solution.

[0074] Figure 7 It is a schematic diagram of the detection accuracy of vascular diseases of the whole-body multi-system disease screening method provided by the present solution.

[0075] Figure 8 It is a schematic diagram of the detection accuracy of neurological diseases of the whole-body multi-system disease screening method provided by the present solution.

[0076] Figure 9 It is a schematic diagram of the detection accuracy of digestive system diseases of the whole-body multi-system disease screening method provided by the present solution.

[0077] Figure 10 It is a schematic diagram of the detection accuracy of mental system diseases of the whole-body multi-system disease screening method provided by the present solution.

[0078] Figure 11Schematic diagram of the detection accuracy of urinary system diseases in the whole-body multi-system disease screening method provided by this solution.

[0079] Figure 12 Schematic diagram of the detection accuracy of respiratory system diseases in the whole-body multi-system disease screening method provided by this solution.

[0080] Figure 13 Schematic diagram of the detection accuracy of cancer in the whole-body multi-system disease screening method provided by this solution.

[0081] Figure 14 Schematic diagram of the structure of the whole-body multi-system disease screening system provided by this solution.

[0082] Figure 15 Schematic diagram of the structure of the electronic device provided by this solution. Detailed implementation manners

[0083] The attached drawings of the present invention are only for illustrative purposes and should not be construed as a limitation to the present invention. To better illustrate the following embodiments, some components in the drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0084] Figure 1 Schematic diagram of the flow of the whole-body multi-system disease screening method provided by this solution. The following combines Figure 1 Describe a whole-body multi-system disease screening method based on the fundus color photo U-Transformer model adopted by the solution of this embodiment. As Figure 1 shown, the method includes:

[0085] Step S1: Collect the fundus images of patients and the data of their systemic disease diagnoses in a public database to make a dataset;

[0086] Specifically, more than 50,000 patients (personal identity information is hidden) in a large-scale multi-center public database are used for dataset collection, collecting 66,481 fundus color photo image data and systemic disease diagnosis data, mainly including whether they are completely healthy, whether there is an eye disease diagnosis and systemic diseases (cardiovascular and cerebrovascular diseases, neurodegenerative diseases, mental diseases, urinary system diseases, respiratory system diseases, digestive system diseases, tumors, etc.) diagnosis, as well as specific diagnoses. In addition, fundus image data and systemic disease diagnosis data of 10,000 patients (personal identity information is hidden) from 5 different centers are also collected, providing sufficient external verification samples for the training of the model to ensure the robustness and generalization ability of the model.

[0087] Step S2: Perform data preprocessing on the fundus photo data and the systemic disease diagnosis dataset obtained in the above steps;

[0088] Preferably, Figure 2 is a schematic diagram of the data preprocessing process for the multi-system disease screening method provided by this solution. As Figure 2 shown, in this embodiment, the data preprocessing in step S2 includes:

[0089] Step S21: Crop invalid regions: Use image processing software to crop the invalid regions in the full dataset images and remove the invalid information at the edges of the images;

[0090] Optionally, software such as matlab can be used to crop a large number of images and remove the invalid information in the images. However, it is not limited to using this method. For example, deep learning frameworks such as TensorFlow, PyTorch, and Keras can also perform preprocessing operations such as cropping on image data. In the specific implementation process, a suitable tool can be selected according to needs to process the images.

[0091] Step S22: Standardization processing: Perform standardization processing on the cropped images; in this embodiment, the resolution of the images is uniformly adjusted to 224×224 pixels. The specific image standard can be selected according to actual needs. Only one case is provided in this embodiment.

[0092] Step S23: Image enhancement: Perform image enhancement on the standardized images;

[0093] Preferably, in this embodiment, image enhancement includes at least the following two strategies:

[0094] Strategy 1: Randomly move the standardized images horizontally by 0 to 3 pixels, and randomly rotate them by 90 degrees, 180 degrees, and 270 degrees to expand the dataset;

[0095] Strategy 2: Perform an image preprocessing strategy based on local average color removal on the standardized images to weaken image noise.

[0096] Optionally, in this embodiment, the operations for image enhancement are not limited to the above strategies. Other strategies such as color adjustment, filter processing, and reinforcement learning can also be used to enhance the images. For example:

[0097] In color adjustment, the enhancement operations that can be used are:

[0098] Contrast enhancement: increasing the contrast of an image by adjusting the gray levels between pixels; Brightness adjustment: increasing or decreasing the brightness level of an image; Color balance: adjusting the color distribution of an image to make it more natural or have a specific hue, etc.

[0099] In filter processing, the enhancement operations that can be used are:

[0100] Gaussian blur: smoothing the image by applying a Gaussian filter; Sharpening: enhancing the clarity of the image by highlighting the edges and details in the image; Noise addition: adding Gaussian noise or other types of noise to the image to make the model more robust, etc.

[0101] In reinforcement learning, the enhancement operations that can be used are:

[0102] Contrast enhancement: using a contrast enhancement algorithm to enhance the contrast of an image; Histogram equalization: adjusting the histogram of the image to increase the dynamic range of the image, etc.

[0103] The above-listed are only common image enhancement operations. In the actual application process, the inventor can select one or more image enhancement methods for operation and processing according to actual needs. Specifically, in this embodiment, through image enhancement operations, the visual effect of the image can be significantly improved, the specific regions or features of the recognizable diseases in the image can be emphasized to make them more prominent and easy to identify. At the same time, irrelevant features can be suppressed, the noise, irrelevant details or interference information in the image can be reduced, the main features can be made more prominent, the information content of the image can be increased, thereby improving the recognition rate and accuracy of the image.

[0104] Step S24: Image stitching: fusing the color fundus photos of both eyes into one image to obtain the fundus information of both eyes to the greatest extent.

[0105] Before training the model, preprocessing the data can effectively eliminate the invalid information in the image, improve the image quality, enhance the detectability of relevant information, and simplify the data to the greatest extent, accelerate the training process, and make the model easier to process; increase the diversity of the data, which helps the model learn more data distribution situations and improve the generalization ability of the model; through random transformation and enhancement of the image, the model observes more data samples during the training process, reduces the dependence of the model on a specific data distribution, and thus reduces the risk of overfitting; improves the robustness of the model, making the model have better adaptability to small changes in the input image.

[0106] Preferably, in step S3: inputting the preprocessed data into the U-Transformer model for training;

[0107] Specifically, during the training process, the model extracts key features for detecting systemic diseases, such as vascular morphology, optic disc tilt, whether the optic disc boundary is clear, the total number of retinal hemorrhage points, etc., and further classifies the images into two major categories: healthy or systemic diseases, and further subdivides systemic diseases into seven systemic diseases.

[0108] Further preferably, as Figure 3 shown, Figure 3 is a schematic diagram of the U-Transformer model for the whole-body multi-system disease screening method provided by this solution. Combining Figure 3 it can be seen that in this embodiment, the U-Transformer model is constructed by an improved U-Net hybrid Transformer network, including:

[0109] First, use the U-Net network as Figure 4 shown to extract the vascular feature information and optic disc feature information of the binocular fundus color photos in the dataset;

[0110] Then, take the feature map containing the vascular feature information and optic disc feature information extracted as the key value and input it into the Transformer module as Figure 5 shown, and upsample the feature map to the original image size through the decoder, and generate a bridging layer as the input of the Transformer Head;

[0111] The Transformer Head merges the features of the bridging layer through the convolutional layer, flattens the merged features into a sequence and passes it to the multi-head attention mechanism MHA, and passes the output of the multi-head attention mechanism MHA to the multi-layer perceptron MLP for processing, mainly used to map the input features to the output features;

[0112] Finally, linearly upsample the output of the multi-layer perceptron MLP and process it through the convolutional layer in the CBR block, so as to output the final prediction, construct the U-Transformer model based on the fundus color photo, and finally realize the segmentation of the fundus blood vessels and optic disc by using the multi-head transformer structure, and realize the simultaneous extraction of disease features from multiple targets.

[0113] The hybrid architecture of the U-Transformer model includes multiple convolutional layers, pooling layers, and fully connected layers. The specific settings of these layers are optimized for fundus image feature extraction, and appropriate filter sizes, strides, and activation functions are selected for each layer to maximize the efficient extraction and representation of fundus features. The inductive bias of convolutional images is used to avoid large-scale pre-training, and the ability of the Transformer to capture global feature relationships is effectively utilized, which helps to better localize disease-specific vascular and optic disc features. Since missegmented regions are usually located at the boundaries of the regions of interest, high-resolution context information can play a crucial role in segmentation. Therefore, U-Transformer focuses on the self-attention module, which makes it possible to effectively process large-size feature maps. Instead of simply integrating the self-attention module on top of the feature maps from the UNet backbone, the Transformer module is applied to each level of the encoder and decoder to collect long-term dependencies from multiple scales, thereby achieving better localization of distinguishable features of diseases and improving the accuracy of model recognition.

[0114] Preferably, in step S3 of this embodiment, a gradient integration algorithm is further included to perform visual analysis for accurately locating imaging markers of different systemic diseases in fundus images. According to the visual analysis results, the training range can be targeted to narrow down to fundus regions specific to different systemic diseases to train the model and improve the performance of the network model. The formula of the gradient integration algorithm is as follows:

[0115]

[0116] Among them, assuming that the input is an n-dimensional vector x, a baseline input x' is defined, and the neural network f(x) maps the input to a probability value between 0 and 1. Then in the n-dimensional space, there is a straight-line path from x' to x, and it can be considered that there are countless samples between x and x'. For the predicted value, now it is necessary to calculate the attribution of the i-th feature x i of the sample x. It can be understood that this attribution result is the contribution degree of this feature to the final prediction result, and the derivative of f(x) with respect to the component x i can be obtained, and integrated between x i ' and xi. The integration result is the attribution.

[0117] More preferably, the use of the gradient integration algorithm to perform visual analysis for accurately locating imaging markers of different systemic diseases in fundus images includes:

[0118] Step S31: Set the input fundus image as a completely black fundus image with pixel intensity of 0 as the reference input. Apply the integrated gradient algorithm to each of the multiple RGB channels separately and sum the results. The integrated result is the attribution, thereby obtaining the pixel points in the fundus image that provide important contributions for predicting different systemic diseases.

[0119] Step S32: Grayscale the original fundus image, and then add different attribution results to different color channels. In this embodiment, the positive attribution results are added to the green channel, and the negative attribution results are added to the red channel. Among them, the outside of the lesion is positive attribution, and the inside of the lesion is negative attribution.

[0120] Use the gradient integration algorithm to attribute the category with the highest score in the output prediction categories, thereby studying which pixel points provide the most important contributions when the network makes predictions. Then add different attribution results to different color channels, thereby locating the pixel points similar to the lesion nodes, highlighting the distinguishable features of the input pattern, and further enabling the model to learn more useful information during training, improving the recognition accuracy of the model.

[0121] Preferably, in step S4: Input the fundus color photo to be measured into the trained U-Transformer model for detection to obtain the systemic disease screening result.

[0122] This technical solution uses deep learning methods to establish and train a model, realizing that only by inputting the fundus color photo into the model for detection can the diagnosis results of systemic diseases throughout the body be obtained, greatly reducing the detection cost and improving the detection efficiency. And this detection method is non-invasive, making the operation more convenient. It can not only achieve binary classification of health or disease, but also accurately screen and classify seven major categories of systemic diseases throughout the body.

[0123] Further preferably, as Figure 5 shown, step S4 also includes first performing binary classification detection to divide into two major categories of healthy and diseased. If the binary classification detection result is healthy, the detection is ended and it is prompted that the fundus color photo to be measured is healthy; if the binary classification detection result is diseased, then perform refined detection of systemic diseases and output the corresponding systemic disease diseased result.

[0124] Figures 6 - 13 They are respectively the schematic diagrams of the detection accuracy of screening healthy and diseased individuals and the detection accuracy of multi-system diseases provided by this solution. As can be seen from Figures 6 - 13 it, in this embodiment, the AUC of the disclosed U-Transformer model in distinguishing healthy and diseased individuals is 0.96, and the average AUC for further distinguishing multi-system diseases reaches above 0.90.

[0125] It can be seen that the technical solution provided in this embodiment has high detection accuracy for diseases and strong feasibility. In actual operation, the fundus image can be directly input into the model for detection to obtain the health status of the body system of the input fundus image, effectively improving the efficiency of system disease detection, shortening the average analysis time to half of the original technology, significantly optimizing the real-time performance, making it not only applicable to normal detection, but also more suitable for use in clinical emergencies, simplifying the operation process of doctors, and significantly improving the clinical work efficiency.

[0126] Preferably, supervised learning is applied during the model training process, and cross-validation is used to optimize the model parameters.

[0127] In this embodiment, five-fold cross-validation is selected to optimize the model parameters. When collecting data, fundus image data and the diagnosis of systemic diseases of a total of 10,000 patients (personal identity information is hidden) from 5 different central sources are collected to provide sufficient external validation samples as the test set for the training of the model, ensuring the robustness and generalization ability of the model.

[0128] Using supervised learning to train the model by analyzing the labeled training data, enabling the model to learn the relationship between the input and output, so as to predict or classify new and unseen data, and realizing that potential patterns and rules can be learned from a large amount of fundus image data and the data of its systemic disease diagnosis, improving the prediction ability of the model; and using cross-validation to reduce the risk of overfitting during the model training process and improve the reliability and generalization performance of the model.

[0129] In the second aspect, the present invention also provides a whole-body multi-system disease screening system based on the fundus color photo U-Transformer model, as Figure 14 shown, Figure 14 is the structural schematic diagram of the whole-body multi-system disease screening system provided by this solution. It can be seen from Figure 11 that this system includes:

[0130] Data acquisition module: used to collect the fundus images of patients and the data of their systemic disease diagnosis in a public database to make a data set;

[0131] Data preprocessing module: used to preprocess the data set obtained in the data acquisition module;

[0132] Model construction and training module: used to input the preprocessed data into the constructed U-Transformer model for training;

[0133] Detection module: used to input the fundus color photo to be detected into the trained U-Transformer model for detection to obtain the system disease screening result;

[0134] In this system, it is realized that only by inputting fundus color photos can the diagnosis results of systemic diseases be obtained, which greatly reduces the detection cost and improves the detection efficiency. Moreover, this detection method is non-invasive, making the operation more convenient.

[0135] The model construction and training module also includes a model construction unit for constructing a U-Transformer model by using an improved U-Net hybrid Transformer network, including:

[0136] Use U-Net to extract several feature information of fundus color photos in the dataset;

[0137] Take the extracted several feature information as key values and input them into the Transformer module. Then, upsample the feature map to the size of the original image through the decoder and generate a bridging layer as the input of the Transformer Head;

[0138] The Transformer Head combines the features of the bridging layer through a convolutional layer, flattens the combined features into a sequence and passes it to the multi-head attention mechanism MHA, and passes the output of the multi-head attention mechanism MHA to the multi-layer perceptron MLP for processing;

[0139] Linearly upsample the output of the multi-layer perceptron MLP and process it through the convolutional layer in the CBR block to output the final prediction, thus constructing a U-Transformer model based on fundus color photos;

[0140] In the model construction unit, using the improved U-Net hybrid Transformer network can maximize the efficient extraction and representation of fundus features, and use the inductive bias of convolutional images to avoid large-scale pre-training, effectively utilize the ability of Transformer to capture global feature relationships, realize better localization of distinguishable features of diseases, and improve the accuracy of model recognition.

[0141] The model construction and training module also includes a marker localization unit. In the marker localization unit, the gradient integration algorithm is used for visual analysis to accurately locate the imaging markers of different system diseases in the fundus image. Then, according to the visual analysis results, the training range is targeted to be narrowed down to the fundus regions specific to different system diseases to train the model and improve the performance of the network model.

[0142] Preferably, the data preprocessing module includes a cropping unit, an image normalization and enhancement unit, and an image stitching unit, where:

[0143] Cropping unit: used to crop the invalid area of the full dataset image and eliminate the invalid information in the image;

[0144] Image normalization and enhancement unit: used to perform normalization processing on the cropped image and perform image enhancement on the normalized image;

[0145] Image stitching unit: used to perform image stitching, fuse the binocular fundus color photos into one image, and obtain the binocular fundus information to the greatest extent.

[0146] In the data preprocessing module, preprocessing the data through multiple preprocessing units can effectively eliminate the invalid information in the image, enhance the detectability of relevant information, and simplify the data to the greatest extent, thereby improving the accuracy and reliability of feature extraction, image segmentation, matching, and recognition.

[0147] Preferably, the detection module includes a first detection output module and a second detection output module, where:

[0148] The first detection output module is used to output whether the fundus image to be measured is healthy or diseased. If the first detection result is healthy, the detection ends and it is prompted that the fundus color photo to be measured is healthy; if the first detection result is diseased, the model is continued to be used for refined screening detection;

[0149] The second detection output module is used to output the corresponding systemic disease prevalence result after the model performs refined detection of systemic diseases.

[0150] Through the above two detection output modules, the detection results of the input fundus image can be effectively and quickly viewed, the efficiency of systemic disease detection is improved, the average analysis time is shortened to half of the original technology, and the real-time performance is significantly optimized, making it not only applicable to normal detection, but also more applicable to clinical emergency situations, simplifying the doctor's operation process, and significantly improving the clinical work efficiency.

[0151] Figure 15 It is a schematic structural diagram of the electronic device provided by this solution. As Figure 15As shown in the figure, the electronic device may include: a processor 110, a communications interface 120, a memory 130, and a communication bus 140. Among them, the processor 110, the communications interface 120, and the memory 130 complete communication with each other through the communication bus 140. The processor 110 may call the logical instructions in the memory 130 to execute a method for screening systemic multi-system diseases based on the fundus color photo U-Transformer model. The method includes: S1: Collecting fundus images of patients and data on the diagnosis of their systemic diseases in a public database to produce a data set; S2: Performing data preprocessing on the data set obtained in the above steps;

[0152] S3: Inputting the preprocessed data into the U-Transformer model for training; S4: Inputting the fundus color photo to be tested into the trained U-Transformer model for detection to obtain the screening results of systemic diseases; Among them, in step S3, it also includes using the gradient integration algorithm to perform visual analysis to accurately locate different systemic disease imaging markers in the fundus image, so as to, according to the visual analysis results, targetedly narrow the training range to different systemic disease-specific fundus regions to train the model.

[0153] In addition, when the logical instructions in the above-mentioned memory 130 can be implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this solution, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this solution. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0154] On the other hand, this solution also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the whole-body multi-system disease screening method based on the fundus color photo U-Transformer model provided by the above-mentioned various methods. The method includes: S1: Collect the fundus images of patients and the data of their systemic disease diagnoses in a public database to make a data set; S2: Perform data preprocessing on the data set obtained in the above step; S3: Input the preprocessed data into the U-Transformer model for training; S4: Input the fundus color photo to be tested into the trained U-Transformer model for detection to obtain the system disease screening result; wherein, in step S3, it also includes using the gradient integration algorithm to perform visual analysis to accurately locate the imaging markers of different system diseases in the fundus image, so as to, according to the visual analysis result, specifically narrow the training range to the fundus regions specific to different system diseases to train the model.

[0155] On the other hand, this solution also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the whole-body multi-system disease screening method based on the fundus color photo U-Transformer model provided by the above-mentioned various methods. The method includes: S1: Collect the fundus images of patients and the data of their systemic disease diagnoses in a public database to make a data set; S2: Perform data preprocessing on the data set obtained in the above step; S3: Input the preprocessed data into the U-Transformer model for training; S4: Input the fundus color photo to be tested into the trained U-Transformer model for detection to obtain the system disease screening result; wherein, in step S3, it also includes using the gradient integration algorithm to perform visual analysis to accurately locate the imaging markers of different system diseases in the fundus image, so as to, according to the visual analysis result, specifically narrow the training range to the fundus regions specific to different system diseases to train the model.

[0156] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0157] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0158] Obviously, the above embodiments of the present solution are only examples for clearly explaining the present solution, rather than limitations on the implementation modes of the present solution. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation modes here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present solution shall be included in the protection scope of the claims of the present solution.

Claims

1. A whole-body multi-system disease screening method based on the U-Transformer model of fundus color photography, comprising: S1: Collect fundus images of patients and their systemic disease diagnosis data from public databases to create a dataset; S2: perform data preprocessing on the data set obtained in the above steps; S3: Input the preprocessed data into the U-Transformer model for training; S4: Input the fundus color photo to be tested into the trained U-Transformer model for detection to obtain the system disease screening results; The method is characterized in that step S3 also includes using a gradient integration algorithm to perform visual analysis to accurately locate imaging markers of different system diseases in fundus images, so as to specifically narrow the training scope to fundus areas specific to different system diseases to train the model based on the visual analysis results; The use of the gradient integration algorithm to perform visual analysis and accurately locate imaging markers of different system diseases in fundus images includes: S31: setting the fundus image to a completely black fundus image with a pixel intensity of 0 as a reference input, applying the integral gradient algorithm to the RGB multiple channels separately to perform summation calculation to obtain an integral result, which is the attribution, so as to obtain the pixel points in the fundus image that provide important contributions to the prediction of different system diseases; S32: grayscale the fundus image, and then add different attribution results to different color channels, so as to highlight the identifiable features of different system diseases in the input fundus image; The gradient integration algorithm formula is: Where x is the latitude vector input, x' represents the baseline input, and i is the number of features.

2. A method for whole body multi-system disease screening based on fundus color photography U-Transformer model according to claim 1, characterized in that: The U-Transformer model is constructed by an improved U-Net hybrid Transformer network, including: U-Net is used to extract some feature information of fundus color photos in the dataset; The extracted feature map containing the above-mentioned feature information is input into the Transformer module as the key value, and the feature map is upsampled to the original image size through the decoder, and a bridge layer is generated as the input of the Transformer Head; Transformer Head merges the features of the bridge layer through the convolution layer, flattens the merged features into a sequence and passes them to the multi-head attention mechanism MHA, and passes the output of the multi-head attention mechanism MHA to the multi-layer perceptron MLP for processing; The output of the multi-layer perceptron MLP is linearly upsampled and processed through the convolutional layer in the CBR block to output the final prediction and construct a U-Transformer model based on fundus color photos.

3. The whole body multi-system disease screening method based on fundus color photography U-Transformer model according to claim 1 is characterized in that: In step S2, the data preprocessing includes: S21: Crop the invalid area of ​​the image of the entire data set to remove invalid information in the image; S22: performing standardization processing on the cropped image; S23: performing image enhancement on the standardized image; S24: Perform image stitching to merge the fundus color photos of both eyes into one image, so as to obtain fundus information of both eyes to the greatest extent.

4. The method for whole body multi-system disease screening based on fundus color photography U-Transformer model according to claim 3, characterized in that: In step S23, the image enhancement at least includes: Perform random horizontal shift and rotation on the standardized image; The local average color removal is performed on the normalized image to weaken the image noise.

5. The method for whole body multi-system disease screening based on fundus color photography U-Transformer model according to claim 1, characterized in that: Step S4 also includes first performing a binary classification test to classify the patient into two categories: healthy and diseased. If the binary classification test result is healthy, then the test is terminated and a prompt is given that the fundus color photograph to be tested is healthy. If the binary classification test result is diseased, a systemic disease refinement test is performed and the corresponding systemic disease result is output.

6. A method for whole body multi-system disease screening based on fundus color photography U-Transformer model according to any one of claims 1 to 5, characterized in that: Supervised learning is applied during model training, and cross-validation is used to optimize model parameters.

7. A whole body multi-system disease screening system based on the fundus color photography U-Transformer model according to any one of claims 1 to 6, comprising: Data collection module: used to collect fundus images of patients and data on the diagnosis of their systemic diseases from public databases to create data sets; Data preprocessing module: used to preprocess the data set obtained in the data acquisition module; Model building and training module: used to input preprocessed data into the built U-Transformer model for training; Detection module: used to input the fundus color photo to be tested into the trained U-Transformer model for detection to obtain the system disease screening results; It is characterized in that The model building and training module also includes a model building unit for building a U-Transformer model using an improved U-Net hybrid Transformer network, including: U-Net is used to extract some feature information of fundus color photos in the dataset; The extracted feature map containing the above-mentioned feature information is input into the Transformer module as the key value, and the feature map is upsampled to the original image size through the decoder, and a bridge layer is generated as the input of the Transformer Head; Transformer Head merges the features of the bridge layer through the convolution layer, flattens the merged features into a sequence and passes them to the multi-head attention mechanism MHA, and passes the output of the multi-head attention mechanism MHA to the multi-layer perceptron MLP for processing; The output of the multi-layer perceptron MLP is linearly upsampled and processed through the convolutional layer in the CBR block to output the final prediction and construct a U-Transformer model based on the fundus color photo; The model building and training module also includes a marker positioning unit, in which a gradient integration algorithm is used to perform visual analysis to accurately locate image markers of different system diseases in fundus images, so that according to the visual analysis results, the training scope is narrowed down to fundus areas specific to different system diseases to train the model.

8. The whole body multi-system disease screening system based on fundus color photography U-Transformer model according to claim 7 is characterized in that: The data preprocessing module includes a cropping unit, an image standardization and enhancement unit, and an image stitching unit, wherein: Cropping unit: used to crop the invalid area of ​​the whole data set image and remove the invalid information in the image; Image standardization and enhancement unit: used to perform standardization on the cropped image and to enhance the image after the standardization; Image stitching unit: used for image stitching, fusing the fundus color photos of both eyes into one image, so as to obtain the fundus information of both eyes to the greatest extent.

9. According to the whole body multi-system disease screening system based on fundus color photography U-Transformer model in claim 8, the detection module comprises a first detection output module and a second output detection module, wherein: The first detection output module is used to output whether the fundus image to be tested is healthy or diseased. If the first detection result is healthy, the detection is terminated and a prompt is given that the fundus color photo to be tested is healthy; if the first detection result is diseased, the model is continued to be used for refined screening detection; The second detection output module is used to output the corresponding systemic disease results after the model performs detailed detection of systemic diseases.

Citation Information

Patent Citations

  • Detection method for eyeground multi-disease classification based on deep learning

    CN111938569A

  • Deep learning model performance verification method and device

    CN112434807A

  • Sugar net analysis method and system and electronic equipment

    CN113576399A