Retina image unsupervised anomaly detection method for early screening of diabetes mellitus

By constructing an unsupervised method of time-series vascular topology maps and hemodynamic simulation, the problems of neglecting time-series continuity and labeling dependence in early screening in existing technologies are solved, and efficient detection and risk assessment of early diabetic lesions are achieved.

CN120747019APending Publication Date: 2025-10-03TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510922215.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing technologies rely on static image analysis in early diabetes screening, ignoring the temporal continuity and dynamic accumulation of retinal pathology, making it difficult to capture early lesions. They also rely on massive amounts of manually labeled data and cannot be generalized to new or early fuzzy pathological morphologies. They lack hemodynamic analysis, making early diagnosis difficult.

Method used

By acquiring time-series retinal fundus images for longitudinal registration, constructing multi-scale attention-guided vascular segmentation and hemodynamic simulation, generating vascular skeleton maps and time-series vascular topology maps, performing unsupervised learning, identifying evolutionary events and abnormal patterns, and generating early diabetes risk scores.

Benefits of technology

Effectively capture early subtle changes in vascular structure, get rid of dependence on manual labeling, enhance the discriminability of pathological information, and improve the accuracy of early risk detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747019A_ABST
    Figure CN120747019A_ABST
Patent Text Reader

Abstract

The invention discloses a retina image unsupervised anomaly detection method for early screening of diabetes mellitus. The method comprises the following steps: carrying out registration and multi-scale attention-guided blood vessel segmentation on a longitudinal time sequence retina image of a patient; extracting a vascular skeleton and constructing a time sequence vascular topological graph, calculating geometric morphology and hemodynamic attributes of each vascular segment, identifying vascular morphology evolution characteristics by comparing topological graphs of adjacent time points, and calculating hemodynamic characteristics such as wall shear stress through simulation; the evolution and hemodynamic characteristics are jointly input into a time sequence encoder for unsupervised learning, and an early diabetes risk score is comprehensively generated by analyzing a reconstruction error, an abnormal score based on density estimation and a time sequence trajectory deviation degree of a potential space; the scheme of the invention does not depend on lesion labels, and can sensitively detect the tiny anomalies at the early stage of pathology from multi-dimensional dynamic changes, thereby providing an objective and quantitative new way for early screening and intervention of diabetes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to an unsupervised abnormality detection method for retinal images for early screening of diabetes. Background Art

[0002] Diabetes, a chronic metabolic disease affecting hundreds of millions of people worldwide, requires early screening and intervention to slow disease progression and prevent serious complications. The retina is the only part of the human body where vascular and neural tissue can be noninvasively and intuitively observed. Changes in its microvasculature are often considered early biomarkers for systemic vascular diseases, including diabetes. In recent years, computer-assisted diagnosis (CAD) based on retinal fundus images has become a research hotspot for early diabetes screening. Early technologies relied primarily on traditional image processing algorithms and machine learning models, manually designing feature extractors to quantify morphological metrics of retinal vessels. These methods then employed classifiers such as support vector machines (SVMs) and random forests for diagnosis. However, these approaches are highly dependent on the quality of feature engineering, have limited generalization capabilities, and struggle to capture complex, nonlinear pathological patterns. With scientific innovation and advancements, deep learning models, such as convolutional neural networks (CNNs), have become widely used. Through an end-to-end learning approach, these deep learning models can automatically learn high-level, abstract pathological features from images. Their expertise in detecting and grading diabetic retinopathy, particularly in identifying typical lesions such as microaneurysms, hemorrhages, and exudates, has reached levels comparable to those of experts in the field.

[0003] While existing technologies have been successful in identifying established diabetic retinopathy, they still face several limitations when targeting early-stage diabetes risk screening, where clinical symptoms are less apparent. First, existing methods generally rely on cross-sectional analysis of static images, essentially providing a snapshot assessment of a single, instantaneous disease state. This approach ignores the temporal continuity and dynamic cumulative nature of diabetic retinopathy pathology. In reality, the vascular network may undergo long-term, subtle structural remodeling and hemodynamic disturbances before a clear lesion appears. Therefore, relying solely on a single snapshot assessment can easily miss the golden window for disease detection. Second, most high-performance models rely on supervised learning, which requires massive amounts of pixel-level or lesion-level labels meticulously annotated by experts. This labeling effort is not only costly and time-consuming, but also subjectivity among experts in defining early, ambiguous pathological changes, making label quality difficult to guarantee. Furthermore, in early screening scenarios, abnormal patterns are unknown and diverse, making it nearly impossible to predefine and annotate all possible abnormalities. Third, existing analytical dimensions are overly limited to the morphological aspects of blood vessels, neglecting their functional properties as fluid transport networks. Furthermore, hemodynamic parameters such as wall shear stress and blood flow velocity are key physical and biological signals regulating endothelial function and vascular remodeling. However, existing technologies fail to couple the evolution of vascular topology with hemodynamic simulations, thus lacking the ability to deeply infer and predict pathology from the relationship between functional disorders and structural changes. Summary of the Invention

[0004] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of this application to avoid obscuring the purpose of this section, the abstract and the title of the invention, and such simplifications or omissions should not be used to limit the scope of the present invention.

[0005] In view of the above existing problems, the present invention is proposed. Therefore, the present invention provides an unsupervised abnormality detection method for retinal images for early screening of diabetes, which is used to solve the problems raised in the background technology.

[0006] To solve the above technical problems, the present invention provides the following technical solution: an unsupervised abnormality detection method for retinal images for early screening of diabetes, comprising: Acquiring time-series retinal fundus images of diabetic subjects and performing longitudinal registration to obtain a registered image sequence, performing multi-scale attention-guided vascular segmentation on the registered image sequence to generate a vascular segmentation map; Refining the blood vessel segmentation map to generate a blood vessel skeleton map, and constructing a time-series blood vessel topology map based on the blood vessel skeleton map; By performing probability matching on the time-series vascular topology map, evolving events are identified to obtain evolving features, and hemodynamic simulation is performed based on the time-series vascular topology map to obtain hemodynamic features; performing unsupervised learning on the evolution features and hemodynamic features, and learning normal evolution patterns through joint reconstruction loss; Based on the reconstruction error, abnormal evolution patterns are detected to generate an early diabetes risk score.

[0007] As a preferred embodiment of the unsupervised abnormality detection method for retinal images for early screening of diabetes described in the present invention, the step of obtaining time-series retinal fundus images of diabetic subjects and performing longitudinal registration includes: Preprocessing of multiple retinal fundus images collected from the same diabetic subject at different time points; By detecting the blood vessel bifurcation points in each pre-processed retinal fundus image, the translation, rotation and scaling transformations of the retinal fundus image are determined to obtain a registration image sequence.

[0008] As a preferred solution of the unsupervised abnormality detection method for retinal images for early diabetes screening described in the present invention, wherein: multi-scale attention-guided blood vessel segmentation is performed on the registered image sequence to generate a blood vessel segmentation map, including: Applying multi-scale convolution kernels to the registered image sequence to extract vascular features at different scales, and then fusing the vascular features and using the channel attention mechanism to obtain vascular feature weights; The weighted fused vascular features are fed into the visual transformer to process the full-scale context information and obtain an enhanced registered image sequence. Restoring the resolution of the registered image through upsampling and skip connections, and generating attention weights in combination with the enhanced registered image sequence; The segmentation results of small blood vessels and pathological new blood vessels generated by applying the attention weights are compared with the real blood vessels. If they are consistent, a blood vessel segmentation map is generated. Otherwise, the channel attention mechanism is reused to adjust the weights of blood vessel features at each scale.

[0009] As a preferred solution of the unsupervised abnormality detection method for retinal images for early diabetes screening according to the present invention, the vascular segmentation map is refined to generate a vascular skeleton map, including: Perform pixel neighborhood analysis on the vessel segmentation map, iteratively delete edge vessel pixels, and generate a preliminary skeleton map with a single pixel width; The analysis range of the pixel neighborhood is adjusted according to the blood vessel diameter, and curve fitting is performed with the preliminary skeleton image to generate a smooth blood vessel skeleton image.

[0010] As a preferred solution of the unsupervised abnormality detection method for retinal images for early screening of diabetes described in the present invention, constructing a time-series vascular topology map based on the vascular skeleton map includes: Perform pixel neighborhood analysis on the vascular skeleton image at each time point, use the intersections and endpoints obtained from the pixel neighborhood analysis as nodes, and the pixels between nodes as edges; Calculate the vascular segment attributes between nodes, including length, diameter, curvature, tortuosity, branching angle, and blood flow resistance, and generate a single-node vascular topology map with attributes; The vascular topology maps at each time point are stored in chronological order to form a time-series vascular topology map.

[0011] As a preferred embodiment of the unsupervised anomaly detection method for retinal images for early diabetes screening described in the present invention, the method includes: performing probability matching on the time-series vascular topology map, identifying evolution events, and obtaining evolution features, including: Based on the node positions at adjacent time points in the temporal vascular topology graph, the initial similarity between nodes is calculated to form a similarity matrix; According to the similarity matrix, the changes in the time series vascular topology map are analyzed, the evolution events are identified, and the evolution characteristics are obtained.

[0012] As a preferred solution of the unsupervised abnormality detection method for retinal images for early screening of diabetes described in the present invention, hemodynamic simulation is performed based on the time-series vascular topology map to obtain hemodynamic characteristics, including: Based on the diameter of each vascular segment in the time-series vascular topology, the inlet flow rate is allocated and the outlet pressure is set; By using the principles of flow conservation and pressure balance, the flow rate in each blood vessel segment is calculated, and the flow velocity and wall shear stress are calculated; Statistical features are extracted from the flow rate, flow velocity, and wall shear stress to generate hemodynamic features.

[0013] As a preferred embodiment of the unsupervised abnormality detection method for retinal images for early screening of diabetes described in the present invention, unsupervised learning is performed on the evolutionary features and hemodynamic features, and a normal evolutionary pattern is learned by joint reconstruction loss, including: A graph attention network is applied to the evolutionary features and hemodynamic features to extract the spatial dependencies between nodes and edges in the temporal vascular topology graph and the regional relationships of hemodynamic features. A time series encoder is used to process the temporal changes of evolutionary features and hemodynamic features, generating potential representations of vascular topology evolution and hemodynamics respectively. fusing the vascular topology evolution and hemodynamic potential representations through a cross-attention mechanism to generate a joint potential representation; The decoder is used to reconstruct the evolutionary and hemodynamic characteristics of the time series encoding input, reconstruct the error, and learn the normal evolution pattern.

[0014] As a preferred embodiment of the unsupervised anomaly detection method for retinal images for early diabetes screening described in the present invention, the method includes: detecting anomaly evolution patterns based on reconstruction errors and generating an early diabetes risk score, including: Analyze the reconstruction error distribution of samples in unsupervised learning and set anomaly detection thresholds; When the anomaly detection threshold is exceeded, the anomaly score of the sample is calculated by applying a density estimation method in the latent representation space, and the time series trajectory of the joint latent representation is analyzed to evaluate the degree of deviation of the evolution pattern; The reconstruction error, anomaly score, and trajectory deviation degree are integrated to generate an early diabetes risk score.

[0015] Compared with the prior art, the invention has the following beneficial effects: 1. This invention longitudinally analyzes time-series retinal images of subjects, constructs a time-series vascular topology map, and identifies evolutionary events. This allows for the capture of subtle, cumulative changes in vascular structure before the onset of clinically significant lesions. This effectively addresses the problem of existing technologies that rely on cross-sectional analysis and fail to perceive the dynamic development of the disease, thus missing the early diagnostic window. It enables risk assessment before irreversible diabetic lesions develop. 2. By learning normal evolution patterns and detecting abnormal deviations, it not only eliminates the reliance on massive, detailed, and costly manually labeled data, but also can identify abnormal patterns that deviate from the normal evolution trajectory, resolving the bottleneck problem that unsupervised learning methods can only identify predefined lesions and cannot generalize to new or early-stage ambiguous pathological forms. 3. The structural topological evolution characteristics of blood vessels and their hemodynamic characteristics are jointly analyzed. By expanding the analysis dimension from morphology to physical inducements, a dynamic relationship between structure and function is established, providing richer and more discriminative pathological information to enhance the accuracy of early risk detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort. Among them: Figure 1 This is an overall flow chart of an unsupervised anomaly detection method for retinal images for early diabetes screening according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.

[0018] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0019] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.

[0020] The present invention is described in detail with reference to schematic diagrams. For ease of illustration, cross-sectional views of device structures may be partially enlarged and not to scale when describing embodiments of the present invention. Furthermore, the schematic diagrams are merely illustrative and should not limit the scope of the present invention. Furthermore, in actual production, the three-dimensional dimensions of length, width, and depth should be included.

[0021] In the description of the present invention, it should be noted that the terms "upper, lower, inner, and outer" and other references to orientations or positional relationships are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first, second, or third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0022] In this disclosure, unless otherwise specified or limited, the terms "mounted," "connected," and "connected" should be interpreted broadly. For example, they may refer to fixed, removable, or integral connections. They may also refer to mechanical, electrical, or direct connections, indirect connections through an intermediary, or internal communication between two components. Those skilled in the art will understand the specific meanings of these terms in this disclosure.

[0023] Example 1 Reference Figure 1, which is the first embodiment of the present invention, provides an unsupervised abnormality detection method for retinal images for early diabetes screening, comprising: S1. Obtain time-series retinal fundus images of diabetic subjects and perform longitudinal registration to obtain a registered image sequence. Perform multi-scale attention-guided vascular segmentation on the registered image sequence to generate a vascular segmentation map. Furthermore, multiple retinal fundus images collected from the same diabetic subject at different time points were preprocessed; Specifically, the preprocessing process involves inputting the original retinal fundus images (RGB format, with a resolution of approximately 2048 pixels × 2048 pixels), calculating the grayscale mean and standard deviation of each retinal fundus image, normalizing the pixel values ​​to the range [0, 1], and applying a linear transformation (i.e., the new pixel value of the linear transformation = (original pixel value - mean) / standard deviation). This outputs a grayscale image with consistent brightness. Adaptive histogram equalization (limiting the contrast enhancement factor to 0.01 and the tile size to 8 × 8) is then applied to enhance the contrast between blood vessels and the background. Finally, a block matching three-dimensional filter (with a standard deviation set to 25) is used to preserve the details of the blood vessel edges and remove Gaussian and salt-and-pepper noise. The resulting set of preprocessed retinal fundus images is then output. Furthermore, the pre-processed retinal fundus image set is spatially aligned to generate a registered image sequence to eliminate eye movement or imaging angle differences; Specifically, by detecting the vascular bifurcation points in each preprocessed retinal fundus image, the translation, rotation and scaling transformations of the retinal fundus image are determined to obtain a registered image sequence; It should be explained that, since the vascular bifurcation point is a key structure in the retinal vascular network and has high local feature stability, it is suitable as a feature point for registration; Specifically, retinal fundus images at a first time point and a second time point are obtained from a set of preprocessed retinal fundus images, the retinal fundus image at the first time point is used as a reference image, and the retinal fundus image at the second time point is used as a floating image, vascular bifurcation points are extracted from the reference image and the floating image respectively, a corresponding feature vector is generated for each vascular bifurcation point, and the feature vector corresponding to the vascular bifurcation point of the reference image is matched with the feature vector corresponding to the floating image. The matching is performed using a nearest neighbor matching (NNM) algorithm to calculate the Euclidean distance between the feature vectors to determine a matched pair of vascular bifurcation points; It should be noted that the scale-invariant feature transform (SIFT) algorithm can be used to detect feature points. The SIFT algorithm can effectively extract key points that remain invariant to the brightness, scale, and rotation of the fundus image. It is very suitable for retinal images taken at different time points that may have slight posture changes. Specifically, using matched pairs of vascular bifurcation points, the random sampling consensus algorithm (RANSAC) is applied to estimate the homography matrix and remove incorrect matching pairs (set as a trigger when the Euclidean distance error is greater than 5 pixels). By calculating the homography matrix, the translation (x, y offset), rotation (angle), and scaling (scale) parameters of the registered retinal fundus images are decomposed and applied to each retinal fundus image to obtain a registered image sequence. It should be noted that the U-Net architecture can also be used to predict dense deformation fields and adjust the retinal fundus image pixel by pixel (i.e., the similarity between the registered image and the reference image is calculated, while constraining the smoothness of the deformation field and limiting excessive local deformation by calculating the gradient of the displacement vector). This ensures the local alignment of the vascular structure after the image is registered, thereby optimizing the similarity between the registered image and the reference image and the smoothness of the local deformation. In addition, the present invention designs a multi-scale feature fusion module to address the problems of class imbalance (where the number of blood vessel pixels is far less than that of background pixels) and hard-to-classify samples. Based on this multi-scale feature fusion module, an attention mechanism specifically targeting small blood vessels and pathological neovascularization is introduced to dynamically adjust the focus on vessels of different scales. Specifically, the architecture of the multi-scale feature fusion module is designed as follows, where the architecture adopts a sequential connection method: Encoder: Receives different scale features of a single retinal fundus image in the registered image sequence and creates a convolution block. The convolution block contains three parallel convolution paths, each using a convolution kernel of different sizes: convolution path 1: 3×3 convolution kernel; convolution path 2: 5×5 convolution kernel; convolution path 3: 7×7 convolution kernel; each convolution path contains a convolution layer (set stride to 1), batch normalization (set momentum to 0.1), followed by a ReLU activation function. Each convolution path outputs 64 feature channels, and the three convolution paths are combined. The outputs of the two paths (each with 64 channels) are spliced ​​to generate a 192-channel feature map; by applying 1×1 convolution, it is compressed to 128 channels to reduce the computational complexity of the encoder, and then the maximum pooling operation (2×2, stride 2) is used to halve the resolution of the input single retinal fundus image (e.g. 2048×2048 to 1024×1024), repeating the convolution block and downsampling 4 times to generate 4 multi-scale feature maps, where each layer represents a multi-scale feature map, namely: Layer 1: 1024×1024 , 128 channels; layer 2: 512×512, 256 channels; layer 3: 256×256, 512 channels; layer 4: 128×128, 1024 channels. Each multi-scale feature map contains different scale features of large and small blood vessels. Then, the channel attention mechanism is used to input each multi-scale feature map, perform global average pooling on each multi-scale feature map, compress its spatial dimension (such as 1024×1024 to 1×1), generate the average value of each channel, and obtain the feature vector of the output channel dimension. A two-layer fully connected network is applied to the feature vector. The neurons in the first layer compress the feature vector to 1 / 16 of the number of channels (for example, 128 channels are compressed to 8 dimensions), followed by a ReLU activation function. The neurons in the second layer expand the feature vector back to the original number of channels, use Sigmoid activation, and generate a vascular feature weight value between 0 and 1. This weight is used to represent the importance of each channel to vascular segmentation. The obtained weight value is multiplied element-by-element by the original multi-scale feature map to obtain the weight value of the vascular feature of the weighted multi-scale feature map. Bottleneck layer: Input the weighted multi-scale feature map of the last layer of the encoder (i.e., 128×128, 1024 channels), split the multi-scale feature map into non-overlapping blocks (patches) of 16×16 pixels, a total of 64×64=4096 blocks; flatten each block into a 256-dimensional vector (16×16×1), map it to a 512-dimensional embedding vector through a linear layer, and then add a position embedding (512-dimensional learnable vector) to preserve the spatial position information of the block; introduce the vision transformer (Vision Transformer) Transformer), a visual transformer that uses a 12-layer multi-head self-attention mechanism: each layer contains 8 attention heads, each of which is used to process 512 / 8=64-dimensional features and simultaneously calculate the query, key, and value vectors. It also fuses global context information based on the weight values ​​of the vascular features obtained by the encoder. It is followed by a feed-forward network (2048-dimensional intermediate layer, using ReLU activation) and layer normalization (setting eps=1e-6). Each layer outputs 4096 512-dimensional vectors, which are reshaped back into a 128×128×512 multi-scale feature map to enhance the weighted multi-scale feature map of the input. Decoder: Input the multi-scale feature map output by the bottleneck layer and the remaining weighted multi-scale feature maps, and double the resolution of the multi-scale feature map (e.g. 128×128 to 256×256) by upsampling operation using transposed convolution (2×2, set stride to 2). After each layer of upsampling operation, the number of channels of the multi-scale feature map is halved (e.g. 512→256), and two 3×3 convolutions are applied (set padding to 1, set stride to 1), followed by batch normalization and ReLU activation function; the upsampled features are compared with the weighted scale feature map of the corresponding encoder layer (same resolution) ) for splicing: Layer 1: 256×256, splicing the last multi-scale feature map of the input (if the remaining weighted multi-scale feature maps are F1, F2, and F3, then F3 is spliced ​​here), generating 768 channels (256+512); Layer 2: 512×512, as mentioned above, splicing F2 to generate 384 channels (128+256); Layer 3: 1024×1024, splicing F1 to generate 192 channels (64+128). After splicing, the number of channels is compressed by applying 1×1 convolution (e.g. 768→256, 384→128); Vascular Enhancement Attention Module: The decoder concatenates the weighted multi-scale feature map and the enhanced weighted multi-scale feature map, applies a 1×1 convolution to the enhanced weighted multi-scale feature map, and generates a single-channel attention map. The single-channel attention value ranges from 0 to 1 (values ​​close to 1 indicate vascular areas). The single-channel attention map is then upsampled to 1024×1024 through bilinear interpolation and element-wise multiplied with the decoder concatenates the weighted multi-scale feature map to generate attention weights for fine segmentation of small blood vessels and pathological neovascularization. The attention weights are applied to the enhanced weighted multi-scale feature map, mapped into a binary segmentation map, and generate segmentation results for small blood vessels and pathological neovascularization. Output layer: Input the segmentation results of small blood vessels and pathological new blood vessels, compare the segmentation results with the real blood vessels, and generate a blood vessel segmentation map if they are consistent. Otherwise, the channel attention mechanism is reused to adjust the weight values ​​of blood vessel features at each scale. It should be noted that the comparison method uses a weighted combination of binary cross entropy loss and Dice loss (Dice Loss); Specifically, the binary cross entropy loss is expressed as follows: Where N is the total number of pixels. Denote as the true label of pixel i, It is expressed as the probability of predicting pixel i as a blood vessel; represents the cross entropy between the predicted probability and the true label; Specifically, Dice loss is expressed as follows: Among them, Y represents the true blood vessel label set (binary image), P represents the predicted blood vessel probability map (0~1), Represented as the overlapping area of ​​Y and P, and denotes the sum of pixels in the true and predicted regions, is a smoothing constant to prevent the denominator from being 0; Specifically, the total loss value generated by the comparison is: in, and It is a hyperparameter with a value of 0.5, which is used to balance the two losses; Represents the total loss value, which is used to balance the pixel-level and region-level segmentation performance; S2, refining the vascular segmentation map to generate a vascular skeleton map, and constructing a time-series vascular topology map based on the vascular skeleton map; Furthermore, pixel neighborhood analysis is performed on the generated vessel segmentation map to check the connectivity of the pixel neighborhood in the vessel segmentation map. Edge vessel pixels are iteratively deleted, while retaining the single-pixel wide path in the vessel center to generate a preliminary skeleton map. This can enhance skeleton connectivity while reducing pseudo-branches. Specifically, a binary working image identical to the generated vascular segmentation map is created, and the generated vascular segmentation map is copied. For each vascular pixel in the binary working image, the 8 neighborhoods surrounding the vascular pixel are checked, and the number of vascular pixels (ranging from 0 to 8) and the transition number (i.e., the number of boundary changes from 0 to 1) in the neighborhood are calculated. The vascular pixels that meet the following two conditions are marked as edge vascular pixels: Condition 1: the number of vascular pixels is between 2 and 6 (to avoid endpoints), Condition 2: the transition number is 1 (the pixel is ensured to be exactly at the boundary), and then the edge vascular pixels are iteratively deleted. The iterative deletion is divided into two sub-iterations based on the above conditions: Sub-iteration 1: edge pixels that meet condition 1 are marked and deleted uniformly; Sub-iteration 2: edge pixels that meet condition 2 are marked and deleted uniformly; the sub-iterations are repeated until no edge vascular pixels can be deleted, then the iteration is stopped, and a preliminary skeleton map with a single pixel width is output. Furthermore, the analysis range of the pixel neighborhood is adjusted according to the vessel diameter, and curve fitting is performed with the preliminary skeleton map to generate a smooth vessel skeleton map; Specifically, a distance transform is applied to the copied vascular segmentation map. The distance from each vascular pixel to the nearest background pixel (denoted as 0) is calculated as the local radius. For the skeleton pixel in the preliminary skeleton map (denoted as 1), the distance value at the corresponding position of the pixel is extracted and multiplied by the distance multiple as the vascular diameter. Specifically, the classification of the blood vessel diameter is as follows: Small blood vessel diameter: <5 pixels; Medium blood vessel diameter: 5-10 pixels; Thick blood vessel diameter: >10 pixels; Specifically, the initial neighborhood size of the inspection is dynamically adjusted according to the classification of blood vessel diameter. For small blood vessel diameters, an 8-neighborhood (3×3) is used; for medium blood vessel diameters, the 8-neighborhood is expanded to a 12-neighborhood (5×5, including the outer two circles of edge pixels); for large blood vessel diameters, the 8-neighborhood or 12-neighborhood is expanded to a 24-neighborhood (7×7). It should be noted that when the neighborhood is adjusted, the number of blood vessel pixels calculated in the neighborhood is also adjusted accordingly (e.g., the original 0–8 becomes 0–12); Specifically, curve fitting was applied to segment the preliminary skeleton image into independent skeleton segments (from an endpoint to an intersection or another endpoint). For each skeleton segment (approximately 50–500 pixels long), B-spline curve fitting was applied. Ten control points were selected from each segment (using uniform sampling) each time. The control point positions were optimized using the least squares method to generate a smooth segment. This smooth segment was then resampled to ensure a pixel spacing of 1 and a single-pixel width. The curve fitting process was repeated three times, and all smooth segments were merged to generate a smoothed vascular skeleton image. Furthermore, pixel neighborhood analysis was performed on the smoothed vascular skeleton image at each time point. The intersection points (connecting multiple vascular segments) and endpoints (vascular segment termination) obtained from the pixel neighborhood analysis of the smoothed vascular skeleton image were used as nodes, and the pixels between the nodes were used as edges. It should be explained that an edge is also called a skeleton path, which means starting from a node, it moves along the continuous pixels (single-pixel width path) of the skeleton graph until it encounters another node, and records the pixel sequence on the movement path and the coordinate position of each pixel; Specifically, each skeleton pixel in the skeleton image is classified based on the skeleton path to improve the readability of the vascular topology map construction. The types of skeleton pixels are divided into: intersection point: connection number ≥ 3 (connecting 3 or more skeleton paths), endpoint: connection number = 1 (end of the skeleton path), and ordinary point: connection number = 2 (in the middle of the skeleton path); Furthermore, the vessel segment attributes between nodes are calculated, including length, diameter, curvature, tortuosity, branching angle, and blood flow resistance, to generate a single-node vessel topology map with attributes; Specifically, the number of pixels along the skeleton path is calculated (accumulated along the skeleton path) to obtain the length of the vessel segment; Specifically, since the blood vessel diameter is obtained through skeleton pixels in the solution of the present invention and therefore does not change, the blood vessel diameter is directly used as the diameter of the blood vessel segment; Specifically, a quadratic curve fitting operation is performed on the pixel sequence in the skeleton path, and the average curvature of the curve (i.e., the rate of change of the arc length of each pixel) is calculated to obtain the curvature of the blood vessel segment: in, Represents a point on the skeleton of a blood vessel segment The local curvature is based on The front and back neighbors and The circumcircle of the triangle formed; Represents a triangle The area of ​​​​(calculated by vector cross product), Expressed as the length of the three sides of the triangle, in pixels; Specifically, the ratio of the actual arc length of the vascular segment skeleton to the straight-line distance between the two end points is calculated to obtain the torsion T of the vascular segment: in, represents the pixel length of the skeleton path (accumulated pixels), Indicates the Euclidean distance between the start and end nodes, in pixels; Specifically, for the edges connected by the intersection, the angle between the adjacent edges at the intersection is calculated (based on the path tangent, ranging from 0 to 180°) as the branching angle of the vessel segment; Specifically, according to Poiseuille's law, without considering blood viscosity and other diseases (such as hypertension, arteriosclerosis, etc.), the relative blood flow resistance of the blood vessel segment is calculated, and the obtained blood flow resistance is normalized to 0-1 to reflect the blood flow resistance of the blood vessel segment; Specifically, blood flow resistance The formula is expressed as: Wherein, L represents the length of the vascular segment, and D represents the average diameter of the vascular segment; Furthermore, the single-node vascular topology graphs at each time point are stored in chronological order to form a time-series vascular topology graph; It should be noted that during the storage process, it is necessary to ensure that the node and edge formats of each single-node vascular topology map are consistent, and check the continuity of each time point (whether there is any missing) to ensure the data integrity of the time-series vascular topology map; S3. By performing probability matching on the time-series vascular topology map, the evolution events are identified and the evolution characteristics are obtained. At the same time, hemodynamic simulation is performed based on the time-series vascular topology map to obtain the hemodynamic characteristics; Furthermore, based on the node positions at adjacent time points in the temporal vascular topology graph, the initial similarities between nodes are calculated to form a similarity matrix; Specifically, the initial similarity between nodes is calculated by the Euclidean distance of the coordinates of adjacent time nodes. The maximum distance between nodes is set to the diagonal length of the binary working image. The smaller the distance, the higher the similarity (ranging from 0 to 100%). This similarity is then used as an element in the similarity matrix. The dimension of the similarity matrix is ​​determined by the size of the two adjacent time nodes themselves. Furthermore, based on the similarity matrix, the changes in the temporal vascular topology are analyzed to identify evolutionary events and obtain evolutionary features; Specifically, the nodes in the column with the highest cumulative sum of similarity elements in the similarity matrix are selected as potential matching nodes, and the other columns are marked as unmatched nodes; Specifically, the definition of evolution events is as follows: The unmatched nodes at the later time point in the temporal vascular topology map were identified as neovascular events, indicating the emergence of new vascular structures; The nodes that were not matched at the previous time point in the temporal vascular topology map were identified as vascular disappearance events, indicating the occlusion or degradation of the original vascular structure; Detecting whether there are attribute changes in the matching nodes and identifying changes exceeding a preset attribute threshold as morphological change events, indicating significant changes in vessel diameter, curvature, or tortuosity; It should be noted that the preset attribute thresholds are the length, diameter, curvature, tortuosity, branching angle, and blood flow resistance of the vascular segment attributes between nodes. The setting of the attribute thresholds needs to be determined based on the actual patient clinical personal information database. During this data acquisition process, the hospital's patient privacy regulations need to be complied with. The newly connected edge at the later time point in the temporal vascular topology diagram is regarded as an anastomotic branch formation event, indicating the connection of the original disconnected vascular segments; Furthermore, the above-defined neovascularization events, vessel disappearance events, morphological change events, and anastomotic branch formation events were summarized to generate evolutionary features; Furthermore, based on the diameter of each vascular segment in the time-series vascular topology map, the inlet flow rate is allocated and the outlet pressure is set; Specifically, we identify the entry nodes of the time-series vascular topology map and select the intersection points close to the retinal center (optic disc area). For the edges connected to the entry nodes, we obtain the blood vessel diameters, calculate the total entry flow, and set the flow of each edge at the entry node as the ratio of the total entry flow. For example, if the diameters of the blood vessels at the two entry edges are 10 pixels and 8 pixels, the flow ratio is 10. 3 :8 3 =1000:512, and then set the pressure of all outlet nodes to 0Pa to simulate the low-pressure state of blood flowing to the vein end; Furthermore, the flow rate in each blood vessel segment is calculated by the principles of flow conservation and pressure balance, and then the flow velocity and wall shear stress are calculated; Specifically, the entire vascular topology is considered as a circuit network, where each vascular segment is a resistor (i.e., blood flow resistance) and each node is a circuit node. According to the law of conservation of flow (Kirchhoff's first law), the total flow into any node is equal to the total flow out of that node. Also, according to the principle of pressure balance, the pressure drop ΔP between any two nodes is equal to the product of the blood flow and the resistance of the vascular segment (ΔP=Q×R). In addition, by setting the boundary conditions of the optic disc aortic vein (for example, the aortic inlet pressure is set to 1 and the venous outlet pressure is set to 0), a large sparse linear equation system A×B=I can be constructed, where A is the admittance matrix (composed of the inverse of the resistance of each vascular segment), B is the pressure vector of each node to be solved, and I is the source / sink flow vector. By solving this equation system, the pressure distribution of the entire network and the flow rate of each vascular segment can be obtained; Specifically, for each edge, calculate the blood vessel segment flow velocity (mm / s), where flow velocity = flow rate ÷ blood vessel segment cross-sectional area, blood vessel segment cross-sectional area = π × (blood vessel segment diameter ÷ 2) 2 , convert the obtained blood vessel cross section into the actual size (1 pixel ≈ 5 microns); Specifically, for each edge, the wall shear stress was calculated considering the blood viscosity, and the obtained shear stress unit was converted into Pa (representing the actual size); Specifically, the wall shear stress is expressed as: in, Expressed as wall shear stress, it reflects the friction of blood flow on the vessel wall; It is expressed as blood viscosity (physiological value is about 0.0035 Pa·s); Q is the blood flow rate of the vascular segment (mL / s); r is the average radius of the vascular segment; In addition, the wall shear stress (in terms of flow velocity) is expressed as: in, Expressed as the average flow velocity (mm / s, calculated as ); Furthermore, statistical features of flow velocity and wall shear stress are extracted to generate hemodynamic characteristics; Specifically, the statistical characteristics are the average value, maximum value and coefficient of variation of flow rate, flow velocity and wall shear stress; Specifically, coefficient of variation = standard deviation ÷ mean; It should be noted that statistical features can capture the dynamic changes of blood flow, while the coefficient of variation can reflect pathology (such as flow velocity fluctuations in the neovascular area); S4. Perform unsupervised learning on the evolution features and hemodynamic features, and learn the normal evolution pattern through joint reconstruction loss; It should be explained that unsupervised learning is a machine learning method that does not rely on any labeled data (i.e., it does not require pre-labeling of samples as normal or abnormal). This method learns by discovering inherent patterns, structures, or distributions in the data. In the present invention, unsupervised learning reconstructs input features through a time series encoder to learn the evolutionary patterns of normal blood vessels, while abnormal patterns are detected due to large reconstruction errors. Furthermore, a graph attention network is applied to the evolutionary features and hemodynamic features to extract the spatial dependencies between nodes and edges in the temporal vascular topology graph and the regional relationships of hemodynamic features, respectively. It should be explained that the spatial dependency relationship between nodes and edges refers to the mutual influence and association between nodes and edges in the spatial structure of the temporal vascular topology graph. This relationship reflects the local and global topological characteristics of the blood vessels, such as the connection pattern of branch points, vascular segment properties, and the contribution of nodes to graph connectivity. In addition, spatial dependency can also help unsupervised models identify key structures (such as abnormal branch points). Specifically, the graph attention network quantifies the spatial dependencies between nodes and edges in the temporal vascular topology graph by assigning different attention weights to each node's neighbors (connected edges and nodes). For example, a high-flow vascular segment contributes more to the importance of the intersection; It should be explained that the regional relationship of hemodynamic characteristics refers to the distribution and mutual influence of hemodynamic parameters such as flow velocity and wall shear stress between different regions (such as small, medium, and large blood vessels, or central and peripheral regions) in the time-series vascular topology map. This relationship reflects the dynamic characteristics of blood flow in space and its association with vascular structure. At the same time, regional relationships can help identify areas of abnormal blood flow (such as areas of high shear stress concentration, which are related to diabetes-related microvascular lesions). Specifically, the graph attention network models the relationship between regions by embedding hemodynamic features (such as flow velocity) into nodes and edges. For example, small blood vessel regions with high shear stress have a greater impact on the blood flow distribution of neighboring nodes. Furthermore, a time series encoder is used to process the temporal changes of evolutionary features and hemodynamic features, generating the vascular topology evolution potential representation and the hemodynamic potential representation respectively; Specifically, the Transformer-based time series encoder uses a self-attention mechanism to capture the long-term dependencies between evolutionary features and hemodynamic features at multiple time points (such as gradual changes in vascular diameter or fluctuations in flow), generating a potential representation of vascular topology evolution and hemodynamics. Specifically, independent time series encoders are applied to the vascular topology evolution and hemodynamic feature streams. Each encoder contains a three-layer architecture, eight attention heads, and a forward neural network with 768 hidden units. The input is the feature vector of n time points (usually 5-10). A learnable positional encoding is added (64 dimensions for the vascular topology evolution stream and 12 dimensions for the hemodynamic stream). The temporal order is preserved, and the encoder's built-in self-attention mechanism is used to model the relationship between time points (such as the association between flow changes and vascular diameter reduction). The potential representations of all time points are then averaged to generate a single 64-dimensional vascular topology evolution latent vector and a 12-dimensional hemodynamic latent vector. Furthermore, the vascular topology evolution and hemodynamic potential representation are fused through a cross-attention mechanism to generate a joint potential representation; It is worth explaining that the cross-attention mechanism is used to fuse the potential representation of vascular topology evolution with the potential representation of hemodynamics, calculate the correlation weight between the two, prioritize key features (such as the association between high shear stress and vascular remodeling), and generate a unified joint potential representation to capture the comprehensive information of vascular structure and its dynamic characteristics; Specifically, the cross-attention fusion operation uses the vascular topology evolution latent vector as the query vector and the hemodynamics latent vector as the key and value vectors. A multi-head attention module is applied, which uses four attention heads to calculate the similarity between the vascular topology evolution latent vector and the hemodynamics feature vector. The output values ​​of each attention head are concatenated to form a 32-dimensional intermediate vector. A feed-forward neural network (128 hidden units, ReLU activation) is then applied to optimize the fused representation, outputting a 32-dimensional joint latent representation. Furthermore, the decoder is used to reconstruct the evolutionary features and hemodynamic characteristics of the time series encoding input, reconstruct the error, and learn the normal evolution pattern; It should be explained that the decoder reconstructs the evolutionary features and hemodynamic characteristics of the input from the joint latent representation and learns the evolutionary pattern of normal blood vessels by minimizing the reconstruction loss; Specifically, the decoder is divided into an evolutionary feature decoder and a hemodynamic decoder. The evolutionary feature decoder reconstructs the input evolutionary features through a three-layer perceptron (128, 256, and 64 units, using ReLU activation); while the hemodynamic decoder reconstructs the input hemodynamic features through a two-layer perceptron (64 and 12 units, using ReLU activation). In addition, during the perceptron processing, dropout with a dropout rate of 0.3 can be applied to prevent overfitting. Specifically, the mean squared error (MSE) between the input features (evolution features, hemodynamic features) and the reconstructed features (evolution features, hemodynamic features) was calculated, and the KL divergence loss was added to constrain the potential representation (potential representation of vascular topology evolution and hemodynamics) to be close to the standard normal distribution (i.e., mean 0, variance 1). The total joint reconstruction loss value was obtained by weighted summation of the mean squared errors (initial evolution feature loss weight was 0.6, hemodynamic loss weight was 0.3, and KL loss weight was 0.1). The Adam optimizer (learning rate 0.0001, batch size 16, training 200 rounds) was used to train on healthy samples of the DRIVE dataset to learn the evolution pattern of normal blood vessels. Specifically, the total loss value of the joint reconstruction is Use the following formula to get: in, Expressed as the MSE of the evolutionary features, Expressed as the MSE of hemodynamic characteristics, Expressed as the KL divergence between the posterior distribution and the standard Gaussian distribution; 、 、 Represent the initial evolution feature loss weight, hemodynamic loss weight and KL loss weight respectively; is the evolutionary feature of the input, For the evolutionary characteristics of reconstruction; is the input hemodynamic characteristics, is the reconstructed hemodynamic feature; F represents the combined input feature, which is composed of the evolution feature and the hemodynamic feature; z represents the latent variable, which is the representation of F after being compressed by the encoder in the high-dimensional space; It is expressed as the prior distribution of the latent variable z, that is, a simple probability distribution pre-set for the latent variable z, usually the standard normal distribution; Expressed as the approximate posterior distribution of the latent variable; S5. Based on the reconstruction error, detect abnormal evolution patterns and generate early diabetes risk scores; Furthermore, we analyze the reconstruction error distribution of samples in unsupervised learning and set the anomaly detection threshold; Specifically, because normal samples have low reconstruction errors and a concentrated distribution, abnormal samples (such as neovascularization) have high errors due to their deviation from the normal evolution pattern. Here, the anomaly detection threshold is set to the mean plus three standard deviations (this setting has the advantage of covering 99.7% of normal samples, but assumes that the loss error is based on a Gaussian distribution). Furthermore, for samples that exceed the anomaly detection threshold, we calculate the anomaly score of the sample by applying density estimation methods in the latent representation space, and analyze the temporal trajectory of the joint latent representation to evaluate the degree of deviation from the evolution pattern. Specifically, the density estimation method is based on the Gaussian mixture model. For each potential anomaly sample's joint potential representation, its log-likelihood in the Gaussian mixture model is calculated, and the maximum likelihood value (normal sample mean) is mapped to 1, and the minimum likelihood value (outlier) is mapped to 0. The anomaly score = 1 - normalized likelihood value. It needs to be explained that considering the temporal trajectory of the joint latent representation can reflect the dynamic changes of vascular evolution and hemodynamics; Specifically, a normal trajectory template is created, several normal sample sequences are randomly selected, and the time series trajectory of their joint potential representation is calculated. K-means clustering (setting K=3) is used to divide the normal trajectories into three categories representing different but normal evolution patterns (such as slight expansion). The average value of each category of trajectory is taken to generate three normal trajectory templates. Then, for each potential abnormal sample trajectory, the dynamic time warping algorithm is used to calculate its minimum distance from the three normal trajectory templates. This minimum distance is used to measure the degree of trajectory deviation. For example, the deviation degree of the normal sample trajectory < the minimum distance, and the deviation degree of the abnormal sample trajectory > the minimum distance. Furthermore, the reconstruction error, abnormal score, and trajectory deviation degree are weighted to generate an early diabetes risk score to reflect the early risk of diabetes. A diabetic patient with a high score has a higher pathological risk, which can be used as a clinical reference. Specifically, the early diabetes risk score formula is expressed as: in, is the normalized total reconstruction error, and The first two terms of are combined to get, is the normalized anomaly score (based on the log-likelihood of the Gaussian mixture model), is the normalized trajectory deviation; 、 、 Respectively expressed as The weights of , take the classic values ​​as 0.3, 0.4, 0.3; For example, the risk score (0-100) is divided into the following levels: Risk score <30: indicates low risk; Risk score 30~60: indicates medium risk; Risk score>60: indicates high risk; It should be noted that this risk score can be adjusted based on the number of actual normal sample sequences to provide a more accurate risk assessment level; In addition, by inputting an early diabetes risk score, the risk area to which the early diabetes risk score belongs can be superimposed on the latest retinal image. For example, color coding can be superimposed on the latest retinal image to highlight potential pathological areas, making it easier for doctors to quickly identify them. At the same time, the evolution trend of the time-series vascular topology map can be displayed. For example, evolutionary events can be extracted from evolutionary features to generate multi-frame images to display the changes in vascular topology maps at 5 to 10 time points. For each time point, a vascular topology map is drawn, nodes are set as dots, edges are set as line segments, and evolutionary events are highlighted. For example, new blood vessels are flashing in red, blood vessel disappearance is faded in blue, morphological changes are highlighted in yellow, and anastomotic branches are bolded in purple to generate a visual clinical report.

[0024] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application may be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0025] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0026] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1The function specified in one or more boxes.

[0027] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0028] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0029] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. An unsupervised anomaly detection method for retinal images for early diabetes screening, characterized by: include: Acquiring time-series retinal fundus images of diabetic subjects and performing longitudinal registration to obtain a registered image sequence, performing multi-scale attention-guided vascular segmentation on the registered image sequence to generate a vascular segmentation map; Refining the blood vessel segmentation map to generate a blood vessel skeleton map, and constructing a time-series blood vessel topology map based on the blood vessel skeleton map; By performing probability matching on the time-series vascular topology map, evolving events are identified to obtain evolving features, and hemodynamic simulation is performed based on the time-series vascular topology map to obtain hemodynamic features; performing unsupervised learning on the evolution features and hemodynamic features, and learning normal evolution patterns through joint reconstruction loss; Based on the reconstruction error, abnormal evolution patterns are detected to generate an early diabetes risk score.

2. The unsupervised anomaly detection method for retinal images for early diabetes screening according to claim 1, characterized in that: The step of acquiring time-series retinal fundus images of a diabetic subject and performing longitudinal registration includes: Preprocessing of multiple retinal fundus images collected from the same diabetic subject at different time points; By detecting the blood vessel bifurcation points in each pre-processed retinal fundus image, the translation, rotation and scaling transformations of the retinal fundus image are determined to obtain a registration image sequence.

3. The unsupervised anomaly detection method for retinal images for early diabetes screening according to claim 2, characterized in that: Performing multi-scale attention-guided blood vessel segmentation on the registered image sequence to generate a blood vessel segmentation map, including: Applying multi-scale convolution kernels to the registered image sequence to extract vascular features at different scales, and then fusing the vascular features and using the channel attention mechanism to obtain vascular feature weights; The weighted fused vascular features are fed into the visual transformer to process the full-scale context information and obtain an enhanced registered image sequence. Restoring the resolution of the registered image through upsampling and skip connections, and generating attention weights in combination with the enhanced registered image sequence; The segmentation results of small blood vessels and pathological new blood vessels generated by applying the attention weights are compared with the real blood vessels. If they are consistent, a blood vessel segmentation map is generated. Otherwise, the channel attention mechanism is reused to adjust the weights of blood vessel features at each scale.

4. The unsupervised anomaly detection method for retinal images for early diabetes screening according to claim 3, characterized in that: Refining the blood vessel segmentation map to generate a blood vessel skeleton map includes: Perform pixel neighborhood analysis on the vessel segmentation map, iteratively delete edge vessel pixels, and generate a preliminary skeleton map with a single pixel width; The analysis range of the pixel neighborhood is adjusted according to the blood vessel diameter, and curve fitting is performed with the preliminary skeleton image to generate a smooth blood vessel skeleton image.

5. The unsupervised anomaly detection method for retinal images for early screening of diabetes according to claim 4, characterized in that: Constructing a time-series vascular topology map according to the vascular skeleton map includes: Perform pixel neighborhood analysis on the vascular skeleton image at each time point, use the intersections and endpoints obtained from the pixel neighborhood analysis as nodes, and the pixels between nodes as edges; Calculate the vascular segment attributes between nodes, including length, diameter, curvature, tortuosity, branching angle, and blood flow resistance, and generate a single-node vascular topology map with attributes; The vascular topology maps at each time point are stored in chronological order to form a time-series vascular topology map.

6. The unsupervised anomaly detection method for retinal images for early diabetes screening according to claim 5, characterized in that: By performing probability matching on the time-series vascular topology map, evolution events are identified and evolution features are obtained, including: Based on the node positions at adjacent time points in the temporal vascular topology graph, the initial similarity between nodes is calculated to form a similarity matrix; According to the similarity matrix, the changes in the time series vascular topology map are analyzed, the evolution events are identified, and the evolution characteristics are obtained.

7. The unsupervised anomaly detection method for retinal images for early diabetes screening according to claim 5, characterized in that: A hemodynamic simulation is performed based on the time-series vascular topology map to obtain hemodynamic characteristics, including: Based on the diameter of each vascular segment in the time-series vascular topology, the inlet flow rate is allocated and the outlet pressure is set; By using the principles of flow conservation and pressure balance, the flow rate in each blood vessel segment is calculated, and the flow velocity and wall shear stress are calculated; Statistical features are extracted from the flow rate, flow velocity, and wall shear stress to generate hemodynamic features.

8. The unsupervised anomaly detection method for retinal images for early screening of diabetes according to claim 6 or 7, characterized in that: The evolution characteristics and hemodynamic characteristics are unsupervisedly learned, and the normal evolution pattern is learned by joint reconstruction loss, including: A graph attention network is applied to the evolutionary features and hemodynamic features to extract the spatial dependencies between nodes and edges in the temporal vascular topology graph and the regional relationships of hemodynamic features. A time series encoder is used to process the temporal changes of evolutionary features and hemodynamic features, generating potential representations of vascular topology evolution and hemodynamics respectively. fusing the vascular topology evolution and hemodynamic potential representations through a cross-attention mechanism to generate a joint potential representation; The decoder is used to reconstruct the evolutionary and hemodynamic characteristics of the time series encoding input, reconstruct the error, and learn the normal evolution pattern.

9. The unsupervised anomaly detection method for retinal images for early diabetes screening according to claim 8, characterized in that: Based on the reconstruction error, abnormal evolution patterns are detected and an early diabetes risk score is generated, including: Analyze the reconstruction error distribution of samples in unsupervised learning and set anomaly detection thresholds; When the anomaly detection threshold is exceeded, the anomaly score of the sample is calculated by applying a density estimation method in the latent representation space, and the time series trajectory of the joint latent representation is analyzed to evaluate the degree of deviation of the evolution pattern; The reconstruction error, anomaly score, and trajectory deviation degree are integrated to generate an early diabetes risk score.

Citation Information

Cited By

  • Urine sugar concentration continuous monitoring method based on deep learning

    CN121231754A

  • Cardiovascular disease detection method based on image analysis

    CN121414757A