A granulocyte directed dual space semantic alignment clustering method for diabetic retinopathy
By employing a particle-sphere-guided dual-space semantic calibration clustering method, the in-modality and out-of-modality issues in the clustering of multimodal images of diabetic retinopathy (DR) were resolved. This method enables efficient extraction of multimodal features and adaptive clustering optimization, thereby improving the accuracy and clinical interpretability of pathological phenotype clustering.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANTONG UNIV
- Filing Date
- 2026-03-25
- Publication Date
- 2026-06-05
Smart Images

Figure CN122156695A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medical image processing and computer vision technology, and in particular to a particle-guided dual-space semantic calibration clustering method for diabetic retinopathy. Background Technology
[0002] Diabetic retinopathy (DR) is a common and serious microvascular complication of diabetes and a leading cause of blindness in working-age individuals. With the continued rise in global diabetes prevalence, early screening, accurate staging, and individualized treatment of DR have become major challenges in ophthalmology. DR exhibits significant pathological heterogeneity; different patients show vast differences in the manifestations of microaneurysms, hemorrhages, exudates, and neovascularization, and the rate of disease progression also varies within the same patient. This makes traditional staging diagnoses, which rely heavily on physician experience, not only time-consuming and laborious but also highly subjective, making it difficult to meet the needs of large-scale clinical screening.
[0003] Currently, deep learning-based methods for analyzing diabetic retinopathy (DR) have made significant progress, especially the fusion analysis of multimodal images (such as fundus photography, optical coherence tomography (OCT), and fluorescence angiography) with clinical indicators, which provides new possibilities for improving diagnostic accuracy. However, existing multimodal clustering and staging methods for DR still have many key bottlenecks, making it difficult to adapt to the complexity of DR pathological features and the heterogeneity of multimodal data. This problem was fully confirmed in the study "Prediction of diabeticretinopathy using machine learning techniques" published by T Jemima Jebaseeli et al. in the Journal of Engineering Research (Vol. 11, No. 2, 2023, pp. 27–37). This study, through a systematic analysis of existing retinal vessel and lesion segmentation algorithms, pointed out many inherent defects in current intelligent analysis technology for DR, which are highly consistent with the technical shortcomings of existing multimodal analysis methods. Specifically, the problems are as follows: most methods only encode single-modal features in a simple way, failing to fully explore the local neighborhood structure and biological topological relationships of lesion areas in images, and lacking effective means of dividing local semantic units. This makes the features susceptible to noise interference, with insufficient discriminative power, and unable to accurately characterize the complex lesion features of diabetic retinopathy (DR). The imaging principles and representation spaces of different modal data have natural differences. Existing methods lack effective semantic calibration mechanisms, and direct fusion can easily lead to semantic conflicts between modalities. They cannot effectively distinguish between consistent pathological information and unique pathological features between modalities, making it difficult to achieve cross-view semantic unification. Existing clustering methods mostly use predefined clustering prototypes or fixed update strategies, lacking guidance mechanisms based on lesion semantic units. They are unable to adapt to the complex distribution of DR pathology, resulting in loose intra-class samples and blurred inter-class boundaries, making it impossible to accurately capture the differences between different DR subtypes. Feature learning and clustering processes are disconnected from each other, failing to achieve collaborative optimization. They also lack a dynamic adjustment "exploration-utilization" strategy and cluster center update mechanism, making it difficult for the feature space and clustering structure to be adapted synchronously. Ultimately, this affects the reliability and clinical interpretability of staging results, and cannot meet the actual needs of accurate clinical diagnosis. Summary of the Invention
[0004] Purpose of the Invention: To address the problems of insufficient utilization of local structures within modalities, limited semantic alignment accuracy between modalities, rigid clustering strategies, and the disconnect between feature learning and the clustering process in existing multimodal image clustering methods for diabetic retinopathy (DR), this invention provides a sphere-guided dual-space semantic calibration clustering method for DR. The purpose of this invention is to overcome the technical bottlenecks of existing methods, achieve efficient extraction of multimodal features, accurate semantic alignment, and adaptive clustering optimization, thereby improving the accuracy and efficiency of clustering and stratification of DR pathological phenotypes and meeting the practical needs of large-scale clinical screening and accurate diagnosis.
[0005] This method includes the following steps:
[0006] Step 1: Design an independent encoder for multimodal diabetic retinopathy data, including fundus color images, optical coherence tomography (OCT), and clinical indicators, and extract view-specific features; define the multimodal dataset. ,in For the first The feature set of a modality For the first The sample at the th Feature vectors under various modes For the first Feature dimensions of a modality Represents the space of real numbers. The total number of samples; all modalities satisfy the sample-level alignment constraint, and the feature dimensions of different modalities are... They are all different; an independent encoder is designed for each mode;
[0007] Step 2: For the view-specific embedded features output by each independent encoder in Step 1, the view quality is evaluated from two dimensions: clustering separability and feature reliability. The confidence weight of each view is dynamically calculated. Based on the confidence ranking, the top two high-confidence views are selected as information sources, and information is supplemented for low-confidence views through a self-attention mechanism. At the same time, the feature variance distribution of low-confidence views is used to identify and filter redundant interference introduced by device noise and imaging artifacts in high-confidence views. The features of each view after bidirectional mutual enhancement are weighted and aggregated to generate a global fusion feature that has both completeness and robustness.
[0008] Step 3: Construct a semantic assignment matrix based on optimal transport, and achieve dual-space collaborative semantic calibration in the feature space and clustering space; perform K-means pre-clustering on the globally fused features to obtain initial cluster centers and construct a global common semantic space; in the feature space, strengthen the semantic consistency of multimodal features of the same patient through instance-level and feature-level comparative learning, shorten the feature distance of similar lesion samples, and widen the feature distance of different lesion samples; in the clustering space, calculate the optimal transport probability from each view feature to the cluster center based on optimal transport (OT) to generate a semantic assignment matrix; optimize the consistency of cluster pseudo-labels through highly reliable pseudo-label propagation, and use KL divergence to constrain the distribution alignment of the semantic assignment matrices of each view to achieve cross-view semantic unification;
[0009] Step 4: Divide the semantic units of diabetic retinopathy lesions through pre-modeling of spheres, and select initial cluster centers from high-confidence spheres; divide the features after dual-space semantic calibration into spheres, and based on the local semantic features of diabetic retinopathy (DR) lesions, aggregate samples that are close in distance and semantically similar in the feature space into two or more sphere units, with each sphere representing a set of samples with homogeneous lesion features; after completing the sphere construction, calculate the confidence of each sphere. The confidence is comprehensively judged based on the compactness of samples within the sphere, the separation between spheres, and the matching degree with the semantics of the clinical grading of diabetic retinopathy (DR), and select high-confidence spheres with confidence scores higher than a preset threshold, and use the center of the high-confidence sphere as the initial cluster center for two-stage clustering;
[0010] Step 5: Based on the confidence level of each particle, the exploration and utilization strategies and the cluster center update step size are dynamically adjusted, and combined with the semantic calibration signal, adaptive clustering optimization is achieved. During the iteration process, the exploration and utilization weights are dynamically allocated according to the confidence level of each particle. For particles with high confidence, the focus is on utilizing existing clustering information to stabilize the cluster centers. For particles with low confidence, the focus is on exploring potential cluster structures. At the same time, the cluster center update step size is dynamically adjusted according to the dual-space semantic calibration signal, and iterative updates are performed until the change in cluster centers is less than a preset threshold. Finally, the clustering results of diabetic retinopathy (DR) lesions are output.
[0011] In step 1, an independent encoder is designed for each mode. :
[0012] (1),
[0013] in, This represents the independent encoder function corresponding to the m-th mode. Indicates the first The sample at the th Original input data for each modality For the first The learnable parameters of an encoder, For the first The sample at the th View-specific embedding features extracted under various modalities For a unified embedding dimension.
[0014] Step 2 includes the following steps:
[0015] Step 2.1: Analyze the view-specific embedding features output by each independent encoder in Step 1. The quality of the view is evaluated from two dimensions: cluster separability and feature reliability, and the confidence weight of each view is dynamically calculated. First, calculate the view. Cluster separability score :
[0016] (2),
[0017] in, For the sample The average distance to other samples in the same cluster, For the sample The average distance to the nearest heterogeneous sample;
[0018] Secondly, calculate the view Feature reliability score :
[0019] (3),
[0020] in, For view Variance of features It is the first The variance of each view feature Indicated in all views The function that takes the maximum value above is synthesized to obtain the first... Confidence weight of each view :
[0021] (4),
[0022] in, It is a balance coefficient, and satisfies ,in Total number of views;
[0023] Step 2.2: Based on confidence level, select the top-2 high-confidence views as information sources, and use a self-attention mechanism to complete the information for low-confidence views. Features after completion The calculation formula is:
[0024] (5),
[0025] in, For low-confidence view indexes, For the first One sample in the low confidence view The embedded features after padding For the first One sample in the low confidence view The original embedding features below, For high-confidence view indexes, This is a set of Top-2 high-confidence views. For the first One sample in the high confidence view Embedded features, For self-attention mechanism functions, attention mechanism Defined as:
[0026] (6),
[0027] in, , , Represents the query vector, key vector, and value vector; softmax represents the activation function; Indicates transpose;
[0028] Step 2.3: Utilize the characteristic variance distribution of the low-confidence view to identify and filter redundant interference introduced by equipment noise and imaging artifacts in the high-confidence view, thus improving the high-confidence view. Features after denoising The calculation formula is:
[0029] (7),
[0030] in, For the first One sample in the high confidence view Denoising-reduced embedding features For element-wise product, For the first One sample in the high confidence view The noise mask below, noise mask Defined as:
[0031] (8),
[0032] in, Low confidence views The characteristic mean and standard deviation, This is the noise threshold coefficient. For indicator functions;
[0033] Step 2.4: Weighted aggregation of the view features after bidirectional mutual enhancement to generate global fusion features. .
[0034] In step 2.4, global fusion features Represented as:
[0035] (9),
[0036] in, For view The feature matrix after mutual enhancement.
[0037] Step 3 includes the following steps:
[0038] Step 3.1, global fusion features Perform K-means pre-clustering to obtain the initial cluster centers. Construct a global public semantic space and compute samples To the cluster center initial allocation probability :
[0039] (10)
[0040] in, For the first The global fusion feature vector corresponding to each sample For the first One initial cluster center, For the first One initial cluster center, For temperature coefficient, For the target number of clusters, It is a natural exponential function;
[0041] Step 3.2: In the feature space, semantic consistency of multimodal features of the same patient is enhanced through instance-level and feature-level contrastive learning. Instance-level contrastive loss is used. The calculation formula is:
[0042] (11),
[0043] in, For the sample Enhanced features of the same view For the sample The corresponding global fusion feature vector, similarity function Where a and b are intermediate parameters. Temperature coefficient;
[0044] Feature-level contrast loss The calculation formula is:
[0045] (12)
[0046] in, For the first The sample at the th Original embedded features under each view Total number of views;
[0047] Total feature space loss Represented as:
[0048] (13);
[0049] Step 3.3: In the clustering space, calculate the optimal transmission probability from each view feature to the cluster center based on optimal transmission, generate a semantic assignment matrix, and define the view. Transmission cost matrix :
[0050] (14)
[0051] in, For the first The sample at the th View down to the first The transmission cost of each cluster center;
[0052] The optimal transport plan matrix is solved by Sinkhorn iteration using the Sinkhorn algorithm. :
[0053] (15)
[0054] in, For the transmission probability matrix, To satisfy the sample distribution Distribution with cluster centers The set of feasible transfer matrices under constraints. For Frobenius inner product operations, For entropy regularization, for The sample was assigned to the first... The transmission probability of each cluster center This is the entropy regularization coefficient;
[0055] Step 3.4: Optimize the consistency of clustered pseudo-labels through highly reliable pseudo-label propagation, and use KL divergence constraints to align the distribution of semantic assignment matrices for each view. Calculated by the following formula:
[0056] (16)
[0057] in, For the first Pseudo-labels for each sample For the first The sample at the th The view is assigned to the first The transmission probability of each cluster center;
[0058] Distribution Alignment Loss Calculated by the following formula:
[0059] (17)
[0060] in, For KL divergence calculation, For the first Optimal transmission plan under each view;
[0061] Total clustering space loss Calculated by the following formula:
[0062] (18)
[0063] in, For the first The sample corresponds to the first Classification of pseudo-label one-hot encoded vectors in clusters. This is the balance coefficient;
[0064] Step 3.5: Jointly optimize the loss in the feature space and the clustering space to achieve collaborative semantic calibration in both spaces.
[0065] (19)
[0066] in, This is the loss for dual-space collaborative semantic calibration.
[0067] Step 4 includes the following steps:
[0068] Step 4.1, perform semantic calibration on the features in the dual-space model. Granulosphere segmentation is performed based on the local semantic features of diabetic retinopathy (DR) lesions. Samples that are close in distance and semantically similar in the feature space are aggregated into multiple granulosphere units. The construction of is represented by the following formula:
[0069] (20)
[0070] in, For the first One sample, For the first The calibrated global fusion feature vector corresponding to each sample For granules The initial center, Distance threshold The similarity threshold;
[0071] Step 4.2: Calculate the confidence score for each granule. The confidence score is determined comprehensively based on the compactness of the sample within the granule, the separation between granules, and the semantic match with the clinical grading of diabetic retinopathy (DR). Intragranular compactness... Calculated by the following formula:
[0072] (twenty one),
[0073] in, For granules Number of samples within;
[0074] intergranular separation Calculated by the following formula:
[0075] (twenty two),
[0076] in, For granules The initial center, the first Confidence level of individual balls Calculated by the following formula:
[0077] (twenty three),
[0078] in This is the balance coefficient;
[0079] Step 4.3: Filter out those with a confidence level higher than the preset threshold. The high-confidence spheres are used, and their centers are used as the initial cluster centers for two-stage clustering. .
[0080] In step 4.3, the initial cluster centers Represented as:
[0081] (twenty four).
[0082] Step 5 includes the following steps:
[0083] Step 5.1: Dynamically allocate exploration and utilization weights based on particle confidence levels to adaptively adjust exploration and utilization strategies during clustering iterations. For high-confidence particles, assign higher utilization weights to stabilize the clustering structure; for low-confidence particles, assign higher exploration weights to uncover potential cluster distributions. The allocation of exploration and utilization weights is calculated using the following formula:
[0084] (25)
[0085] in, For the first Confidence level of individual balls, For granules The use of weights, For granules The exploration weight;
[0086] Step 5.2: Combining the dual-space semantic calibration signal, dynamically adjust the update step size of the cluster centers to achieve precise adaptation for cluster optimization; based on the semantic consistency score between the feature space and the cluster space, determine the global update step size for each iteration, and the semantic consistency score... Calculated by the following formula:
[0087] (26)
[0088] in, For the first The optimal transmission probability in round iteration. For the first Feature assignment probabilities in round iterations, based on scores Calculate dynamic step size :
[0089] (27)
[0090] in, Based on step size, This is the step size adjustment factor;
[0091] Step 5.3: Utilize weights and dynamic step size to iteratively update cluster centers until convergence conditions are met. The first round Cluster centers The update formula is:
[0092] (28)
[0093] in, For the first The first iteration Cluster centers, For granules Belongs to the first A sample set of cluster centers For set Number of samples within, Gaussian noise for the exploration term;
[0094] Step 5.4: Determine whether the iterative change in cluster centers is less than the preset convergence threshold. The convergence criterion is expressed by the following formula:
[0095] (29)
[0096] If the convergence criterion is met, the iteration terminates, and the final cluster centers are determined. The corresponding clustering results are used as outputs to obtain the lesion clustering and lesion grading results of diabetic retinopathy.
[0097] The present invention also provides an electronic device, including a processor and a memory, the memory storing program code that, when executed by the processor, causes the processor to perform the steps of the method.
[0098] The present invention also provides a storage medium storing a computer program or instructions that, when the computer program or instructions are run on a computer, execute the steps of the method described.
[0099] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0100] (1) A multi-granularity feature discriminative enhancement mechanism was designed to address the heterogeneity of diabetic retinopathy (DR) pathology, thereby improving the accuracy and clinical interpretability of DR pathological phenotype clustering and staging. This invention extracts view-specific features through an independent encoder, and combines dual-space semantic calibration with particle-sphere guided clustering optimization to enable the feature space and clustering structure to synchronously adapt to the complex distribution of DR pathology, effectively distinguishing the lesion differences of different DR types, while improving the consistency between clustering results and clinical grading.
[0101] (2) A bidirectional mutual enhancement and dual-space semantic calibration mechanism for views is proposed to effectively solve the problems of multimodal heterogeneous fusion and semantic alignment. This invention realizes the information completion of low-confidence views by high-confidence views and the noise filtering of high-confidence views by low-confidence views through view confidence quantification, generating global fusion features with both integrity and robustness; at the same time, a semantic allocation matrix is constructed based on optimal transmission to realize dual-space collaborative semantic calibration in feature space and clustering space, effectively narrowing the semantic distance of multimodal features of the same patient and widening the semantic distance of different lesion samples, thereby improving the accuracy of cross-view semantic unification.
[0102] (3) A clustering initialization strategy guided by pre-modeling of granules is proposed to effectively avoid the clustering bias problem caused by random initialization. This invention divides lesion semantic units by pre-modeling of granules, and comprehensively judges the confidence of granules based on the compactness of samples within granules, the separation between granules, and the matching degree with the clinical grading of diabetic retinopathy (DR). High-confidence granule centers are selected as initial cluster centers, so that the clustering process starts from the semantic units of lesions, thereby improving the rationality and stability of cluster centers.
[0103] (4) An adaptive two-stage clustering optimization framework based on particle confidence is proposed, which effectively realizes the synergistic optimization of feature learning and clustering process. In the iterative process, the invention dynamically allocates the "exploration-utilization" weight according to the particle confidence, focusing on stabilizing the clustering structure for high-confidence particles and focusing on mining potential cluster distributions for low-confidence particles; at the same time, it dynamically adjusts the update step size of cluster centers by combining dual-space semantic calibration signals, so as to achieve accurate adaptation of clustering optimization and effectively improve the adaptability of clustering structure to the complex pathological distribution of diabetic retinopathy (DR). Attached Figure Description
[0104] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.
[0105] Figure 1 This is a data processing flowchart of a sphere-guided dual-space semantic calibration clustering method for diabetic retinopathy according to the present invention.
[0106] Figure 2 This is a diagram illustrating the overall framework of a granulosphere-guided dual-space semantic calibration clustering method for diabetic retinopathy, as described in this invention.
[0107] Figure 3 This is a detailed flowchart of a granulosphere-guided dual-space semantic calibration clustering method for diabetic retinopathy according to the present invention. Detailed Implementation
[0108] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. Of course, the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0109] like Figure 1 , Figure 2 and Figure 3 As shown, this embodiment provides a granuloma-guided dual-space semantic calibration clustering method for diabetic retinopathy, including the following steps:
[0110] Step 1: Design an independent encoder for multimodal diabetic retinopathy data, including fundus color images, optical coherence tomography (OCT), and clinical indicators, and extract view-specific features; define the multimodal dataset. ,in For the first The feature set of a modality For the first The sample at the th Feature vectors under various modes For the first Feature dimensions of a modality Represents the space of real numbers. The total number of samples; all modalities satisfy the sample-level alignment constraint, and the feature dimensions of different modalities are... They are all different; an independent encoder is designed for each mode;
[0111] Step 2: For the view-specific embedded features output by each independent encoder in Step 1, the view quality is evaluated from two dimensions: clustering separability and feature reliability. The confidence weight of each view is dynamically calculated. Based on the confidence ranking, the top two high-confidence views are selected as information sources, and information is supplemented for low-confidence views through a self-attention mechanism. At the same time, the feature variance distribution of low-confidence views is used to identify and filter redundant interference introduced by device noise and imaging artifacts in high-confidence views. The features of each view after bidirectional mutual enhancement are weighted and aggregated to generate a global fusion feature that has both completeness and robustness.
[0112] Step 3: Construct a semantic assignment matrix based on optimal transport, and achieve dual-space collaborative semantic calibration in the feature space and clustering space; perform K-means pre-clustering on the globally fused features to obtain initial cluster centers and construct a global common semantic space; in the feature space, strengthen the semantic consistency of multimodal features of the same patient through instance-level and feature-level comparative learning, shorten the feature distance of similar lesion samples, and widen the feature distance of different lesion samples; in the clustering space, calculate the optimal transport probability from each view feature to the cluster center based on optimal transport (OT) to generate a semantic assignment matrix; optimize the consistency of cluster pseudo-labels through highly reliable pseudo-label propagation, and use KL divergence to constrain the distribution alignment of the semantic assignment matrices of each view to achieve cross-view semantic unification;
[0113] Step 4: Semantic units of diabetic retinopathy (DR) lesions are pre-modeled using spheres. Initial cluster centers are selected from high-confidence spheres. After dual-space semantic calibration, features are sphere-divided. Based on the local semantic features of DR lesions, samples that are close in distance and semantically similar in the feature space are aggregated into two or more sphere units. Each sphere represents a set of samples with homogeneous lesion features. After sphere construction, the confidence level of each sphere is calculated. The confidence level is comprehensively determined based on the compactness of samples within the sphere, the separation between spheres, and the matching degree with the semantics of the clinical grading of DR. High-confidence spheres with confidence levels higher than a preset threshold are selected, and the centers of these high-confidence spheres are used as the initial cluster centers for two-stage clustering.
[0114] Step 5: Based on the confidence level of each particle, the exploration and utilization strategies and the cluster center update step size are dynamically adjusted, and combined with the semantic calibration signal, adaptive clustering optimization is achieved. During the iteration process, the exploration and utilization weights are dynamically allocated according to the confidence level of each particle. For particles with high confidence, the focus is on utilizing existing clustering information to stabilize the cluster centers. For particles with low confidence, the focus is on exploring potential cluster structures. At the same time, the cluster center update step size is dynamically adjusted according to the dual-space semantic calibration signal, and iterative updates are performed until the change in cluster centers is less than a preset threshold. Finally, the clustering results of diabetic retinopathy (DR) lesions are output.
[0115] In step 1, an independent encoder is designed for each mode. : (1), in, This represents the independent encoder function corresponding to the m-th mode. Indicates the first The sample at the th Original input data for each modality For the first The learnable parameters of an encoder, For the first The sample at the th View-specific embedding features extracted under various modalities For a unified embedding dimension.
[0116] Step 2 includes the following steps:
[0117] Step 2.1: Analyze the view-specific embedding features output by each independent encoder in Step 1. The quality of the view is evaluated from two dimensions: cluster separability and feature reliability, and the confidence weight of each view is dynamically calculated. First, calculate the view. Cluster separability score :
[0118] (2),
[0119] in, The total number of samples, For the sample The average distance to other samples in the same cluster, For the sample The average distance to the nearest heterogeneous sample;
[0120] Secondly, calculate the view Feature reliability score :
[0121] (3),
[0122] in, For view Variance of features For the first The variance of each view feature For all views The function that takes the maximum value above is synthesized to obtain the first... Confidence weight of each view :
[0123] (4),
[0124] in, It is a balance coefficient, and satisfies ,in Total number of views;
[0125] This example is validated based on a multimodal sample of 2500 cases of diabetic retinopathy (fundus photography + OCT imaging + clinical blood test indicators). Total sample size. Total number of views (referred to as visual) Figure 1 ,See Figure 2 ,See Figure 3 Balance coefficient The cluster separability score of each view is calculated according to formula (2). ,See Figure 1 ,See Figure 2 ,See Figure 3 of They are respectively , , A higher score indicates stronger view clustering separability. The feature reliability score is calculated according to formula (3). The variances of the view features are respectively , , The feature reliability scores are respectively , , The confidence weights are obtained by combining the results and then normalized. , , ,satisfy Constraints;
[0126] Step 2.2: Based on confidence level, select the top-2 high-confidence views as information sources, and use a self-attention mechanism to complete the information for low-confidence views. Features after completion The calculation formula is:
[0127] (5),
[0128] in, For low-confidence view indexes, For the first One sample in the low confidence view The embedded features after padding For the first One sample in the low confidence view The original embedding features below, For high-confidence view indexes, This is a set of Top-2 high-confidence views. For the first One sample in the high confidence view Embedded features, For self-attention mechanism functions, attention mechanism Defined as:
[0129] (6),
[0130] in, , , Represents the query vector, key vector, and value vector; softmax represents the activation function; Indicates transpose;
[0131] In this example, the Top-2 high-confidence views are selected as the view. Figure 1 ,See Figure 2 ( ={1,2}), low confidence view =3, feature dimension =64.
[0132] Taking the first 5 samples as an example, Figure 3Original feature matrix (64×2500 dimensions, select the first 5 dimensions × the first 5 samples):
[0133] ,
[0134] After combining high-confidence view features and completing the image using a self-attention mechanism, the view is obtained. Figure 3 Enhanced feature matrix (selecting the first 5 dimensions × the first 5 samples):
[0135] ,
[0136] After completing the completion of 2500 samples, the visual... Figure 3 Enhanced feature matrix The matrix elements take values ranging from 0.02 to 0.85.
[0137] Step 2.3: Utilize the characteristic variance distribution of the low-confidence view to identify and filter redundant interference introduced by equipment noise and imaging artifacts in the high-confidence view, thus improving the high-confidence view. Features after denoising The calculation formula is:
[0138] (7),
[0139] in, For the first One sample in the high confidence view Denoising-reduced embedding features For element-wise product, For the first One sample in the high confidence view The noise mask below, noise mask Defined as:
[0140] (8),
[0141] in, Low confidence views The characteristic mean and standard deviation, This is the noise threshold coefficient. For indicator functions;
[0142] In this example, a noise threshold coefficient is set. With low confidence Figure 3 Based on the statistic, the mean Standard deviation .
[0143] With view Figure 1 For example, the original feature matrix of the first 5 dimensions × the first 5 samples is:
[0144] ,
[0145] After noise masking (more than) After processing for noise, the denoised feature matrix (selecting the first 5 dimensions × the first 5 samples) is as follows:
[0146] ,
[0147] See Figure 2 Denoising feature matrix (selecting the first 5 dimensions × the first 5 samples):
[0148] ,
[0149] Denoising feature matrix The noise dimension (value > 0.62 or < 0.02) was set to 0, and the noise ratio was approximately 18%.
[0150] Step 2.4: Weighted aggregation of the view features after bidirectional mutual enhancement to generate global fusion features. Global fusion features Represented as:
[0151] (9),
[0152] in, For view The feature matrix after mutual enhancement;
[0153] In this example, the enhanced feature matrices of each view are weighted and aggregated with confidence weights to generate a global fused feature. .
[0154] The final global fusion feature matrix is obtained (selecting the first 5 dimensions × the first 5 samples):
[0155] ;
[0156] Step 3 includes the following steps:
[0157] Step 3.1, global fusion features Perform K-means pre-clustering to obtain the initial cluster centers. Construct a global public semantic space and compute samples To the cluster center initial allocation probability :
[0158] (10)
[0159] in, For the first The global fusion feature vector corresponding to each sample For the first One initial cluster center, For the first One initial cluster center, For temperature coefficient, For the target number of clusters, It is a natural exponential function;
[0160] In this example, the target number of clusters is set. Temperature coefficient .
[0161] For the global fusion feature matrix Perform K-means pre-clustering to obtain the initial cluster centers. All dimensions are 64. Substituting into formula (10) to calculate the sample... To the cluster center initial allocation probability This yields the initial probability matrix, where the sum of the elements in each row is 1, reflecting the initial membership degree of the sample to each cluster.
[0162] Step 3.2: In the feature space, semantic consistency of multimodal features of the same patient is enhanced through instance-level and feature-level contrastive learning. Instance-level contrastive loss is used. The calculation formula is:
[0163] (11),
[0164] in, For the sample Enhanced features of the same view For the sample The corresponding global fusion feature vector, similarity function Where a and b are intermediate parameters. Temperature coefficient;
[0165] Feature-level contrast loss The calculation formula is:
[0166] (12)
[0167] in, For the first The sample at the th Original embedded features under each view Total number of views;
[0168] Total feature space loss Represented as:
[0169] (13);
[0170] In this example, the temperature coefficient is set. Based on samples Same view enhancement features With global fusion features Cosine similarity, calculate instance-level contrast loss Based on the original embedding features of each view With global fusion features The cosine similarity is used to calculate the feature-level contrast loss;
[0171] Step 3.3: In the clustering space, calculate the optimal transmission probability from each view feature to the cluster center based on optimal transmission, generate a semantic assignment matrix, and define the view. Transmission cost matrix :
[0172] (14)
[0173] in, For the first The sample at the th View down to the first The transmission cost of each cluster center;
[0174] The optimal transport plan matrix is solved by Sinkhorn iteration using the Sinkhorn algorithm. :
[0175] (15)
[0176] in, For the transmission probability matrix, To satisfy the sample distribution Distribution with cluster centers The set of feasible transfer matrices under constraints. For Frobenius inner product operations, For entropy regularization, for The sample was assigned to the first... The transmission probability of each cluster center This is the entropy regularization coefficient;
[0177] In this example, the entropy regularization coefficient is set. ;
[0178] Step 3.4: Optimize the consistency of clustered pseudo-labels through highly reliable pseudo-label propagation, and use KL divergence constraints to align the distribution of semantic assignment matrices for each view. Calculated by the following formula:
[0179] (16)
[0180] in, For the first Pseudo-labels for each sample For the first The sample at the th The view is assigned to the first The transmission probability of each cluster center;
[0181] Distribution Alignment Loss Calculated by the following formula:
[0182] (17)
[0183] in, For KL divergence calculation, For the first Optimal transmission plan under each view;
[0184] Total clustering space loss Calculated by the following formula:
[0185] (18)
[0186] in, For the first The sample corresponds to the first Classification of pseudo-label one-hot encoded vectors in clusters. This is the balance coefficient;
[0187] In this example, the balance coefficient is set. ;
[0188] Step 3.5: Jointly optimize the loss in the feature space and the clustering space to achieve collaborative semantic calibration in both spaces.
[0189] (19)
[0190] in, For dual-space collaborative semantic calibration loss;
[0191] Step 4 includes the following steps:
[0192] Step 4.1, perform semantic calibration on the features in the dual-space model. Granulosphere segmentation is performed based on the local semantic features of diabetic retinopathy (DR) lesions. Samples that are close in distance and semantically similar in the feature space are aggregated into multiple granulosphere units. The construction of is represented by the following formula:
[0193] (20)
[0194] in, For the first One sample, For the first The calibrated global fusion feature vector corresponding to each sample For granules The initial center, Distance threshold The similarity threshold;
[0195] In this example, a distance threshold is set. Similarity threshold Particle-sphere segmentation is performed based on the semantic features of diabetic retinopathy lesions. Features after dual-space semantic calibration are then analyzed. Based on formula (20), samples that are close in distance and semantically similar are aggregated into granular units. The final division was obtained Each sphere contains approximately 70 to 80 samples;
[0196] Step 4.2: Calculate the confidence score for each granule. The confidence score is determined comprehensively based on the compactness of the sample within the granule, the separation between granules, and the semantic match with the clinical grading of diabetic retinopathy (DR). Intragranular compactness... Calculated by the following formula:
[0197] (twenty one),
[0198] in, For granules Number of samples within;
[0199] intergranular separation Calculated by the following formula:
[0200] (twenty two),
[0201] in, For granules The initial center, the first Confidence level of individual balls Calculated by the following formula:
[0202] (twenty three),
[0203] in As a balance coefficient, in this example, a balance coefficient is set. ;
[0204] Step 4.3: Filter out those with a confidence level higher than the preset threshold. The high-confidence spheres are used, and their centers are used as the initial cluster centers for two-stage clustering. :
[0205] (twenty four);
[0206] In this example, a confidence threshold is set. Nineteen high-confidence particles with confidence levels above the threshold were selected. Based on formula (24), the centers of these 19 high-confidence particles were... As the initial cluster centers of two-stage clustering ,Right now ;
[0207] Step 5 includes the following steps:
[0208] Step 5.1: Dynamically allocate exploration and utilization weights based on particle confidence levels to adaptively adjust exploration and utilization strategies during clustering iterations. For high-confidence particles, assign higher utilization weights to stabilize the clustering structure; for low-confidence particles, assign higher exploration weights to uncover potential cluster distributions. The allocation of exploration and utilization weights is calculated using the following formula:
[0209] (25)
[0210] in, For the first Confidence level of individual balls, For granules The use of weights, For granules The exploration weight;
[0211] Quantification explanation: In formula (25), weights are used. confidence level of particles There is a positive correlation. When the confidence level of the particle is... At higher levels, molecules The larger, The closer the value is to 1, the higher the utilization weight, thus stabilizing the clustering structure; when the particle confidence level... The lower the value, the more molecules... The smaller, The smaller the value, the greater the exploration weight. The larger the value, the higher the exploration weight, thereby uncovering potential cluster distributions.
[0212] For example, if the global maximum confidence is 0.9, the high-confidence particles ( Utilizing weights (=0.8) =0.8 / 0.9≈0.89, significantly higher than the exploration weight. =0.11; Low confidence granules ( Exploration weight (=0.3) =1−0.3 / 0.9≈0.67, significantly higher than using weights ≈0.33.
[0213] Step 5.2: Combining the dual-space semantic calibration signal, dynamically adjust the update step size of the cluster centers to achieve precise adaptation for cluster optimization; based on the semantic consistency score between the feature space and the cluster space, determine the global update step size for each iteration, and the semantic consistency score... Calculated by the following formula:
[0214] (26)
[0215] in, For the first The optimal transmission probability in round iteration. For the first Feature assignment probabilities in round iterations, based on scores Calculate dynamic step size :
[0216] (27)
[0217] in, Based on step size, This is the step size adjustment factor;
[0218] In this example, a basic step size is set. Step size adjustment coefficient Calculate the first according to formula (26). Round semantic consistency score Substituting into formula (27) yields the dynamic step size. ;
[0219] Step 5.3: Utilize weights and dynamic step size to iteratively update cluster centers until convergence conditions are met. The first round Cluster centers The update formula is:
[0220] (28)
[0221] in, For the first The first iteration Cluster centers, For granules Belongs to the first A sample set of cluster centers For set Number of samples within, To explore the Gaussian noise, in this example, the Gaussian noise variance is set. ;
[0222] Step 5.4: Determine whether the iterative change in cluster centers is less than the preset convergence threshold. The convergence criterion is expressed by the following formula:
[0223] (29)
[0224] If the convergence criterion is met, the iteration terminates, and the final cluster centers are determined. The corresponding clustering results are used as outputs to obtain the lesion clustering and lesion grading results of diabetic retinopathy.
[0225] In the example, a convergence threshold is set. Based on formula (29), the iteration change is determined. After approximately 35 iterations, the convergence condition is met, the iteration is terminated, and the final cluster centers are output. The method presented in this invention exhibited significant advantages in convergence efficiency and clustering accuracy on 2500 multi-view samples of diabetic retinopathy lesions, along with corresponding clustering and grading results. Multi-dimensional comparative data are shown in Table 1.
[0226] Table 1
[0227]
[0228] In this example, all hyperparameters involved in each step of the method (including but not limited to view confidence balance coefficient, temperature coefficient, particle-sphere splitting threshold, clustering convergence threshold, etc.) are set as preset initial values. The above hyperparameters can be adaptively and iteratively optimized in combination with the sample size, multi-view modality characteristics, convergence of dual-space collaborative semantic calibration loss during model training, diabetic retinopathy staging accuracy, clustering contour coefficient, and other indicators. Through periodic evaluation and dynamic update mechanisms, the values of each hyperparameter are adjusted. For example, when the clustering contour coefficient is lower than the preset threshold, the particle-sphere distance and similarity threshold are automatically fine-tuned to enhance the local semantic aggregation effect. When the staging accuracy fluctuates greatly, the balance coefficient and step size coefficient are optimized to stabilize the model convergence direction. Thus, while taking into account the computational complexity, the staging accuracy of diabetic retinopathy lesions and the generalization stability of the model are continuously improved.
[0229] This example also provides a system for accurate staging of multimodal diabetic retinopathy using multi-view data. The system encompasses modules for data acquisition, view feature enhancement, dual-space semantic calibration, particle-sphere segmentation and confidence assessment, particle-sphere-guided clustering, and result output. Addressing existing multi-view clustering methods' problems such as unbalanced utilization of view information, insufficient feature semantic alignment, and clustering's tendency to fall into local optima and its difficulty in adapting to the semantic distribution of medical lesions, this system employs an "exploration-utilization" clustering strategy—dynamic allocation of view confidence weights, dual-space collaborative semantic calibration, and particle-sphere-guided clustering—to achieve efficient fusion and accurate lesion staging of multimodal data including fundus images and clinical indicators. Compared to existing technologies, this system not only solves the problem of low-quality views interfering with clustering results but also avoids local optima through particle-sphere confidence guidance. Furthermore, it enhances the model's stability under different sample sizes and modal distributions through a hyperparameter adaptive optimization mechanism.
[0230] This invention provides a granulocyte-sphere guided dual-space semantic calibration clustering method for diabetic retinopathy. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment of the invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.
Claims
1. A granulosphere-guided dual-space semantic calibration clustering method for diabetic retinopathy, characterized in that, Includes the following steps: Step 1: Design an independent encoder for multimodal diabetic retinopathy data, including fundus color images, optical coherence tomography (OCT), and clinical indicators, and extract view-specific features; define the multimodal dataset. ,in For the first The feature set of a modality For the first The sample at the th Feature vectors under various modes For the first Feature dimensions of a modality Represents the space of real numbers. The total number of samples; all modalities satisfy the sample-level alignment constraint, and the feature dimensions of different modalities are... They are all different; an independent encoder is designed for each mode; Step 2: For the view-specific embedded features output by each independent encoder in Step 1, the view quality is evaluated from two dimensions: clustering separability and feature reliability. The confidence weight of each view is dynamically calculated. Based on the confidence ranking, the top two high-confidence views are selected as information sources, and information is supplemented for low-confidence views through a self-attention mechanism. At the same time, the feature variance distribution of low-confidence views is used to identify and filter redundant interference introduced by device noise and imaging artifacts in high-confidence views. The features of each view after bidirectional mutual enhancement are weighted and aggregated to generate a global fusion feature that has both completeness and robustness. Step 3: Construct a semantic allocation matrix based on optimal transmission, and achieve dual-space collaborative semantic calibration in the feature space and clustering space; perform K-means pre-clustering on the globally fused features to obtain initial cluster centers and construct a global common semantic space; in the feature space, strengthen the semantic consistency of multimodal features of the same patient through instance-level and feature-level comparative learning, shorten the feature distance of similar lesion samples, and widen the feature distance of different lesion samples; in the clustering space, calculate the optimal transmission probability from each view feature to the cluster center based on optimal transmission OT, and generate a semantic allocation matrix. By using highly reliable pseudo-label propagation, the consistency of clustered pseudo-labels is optimized, and KL divergence is used to constrain the distribution alignment of semantic assignment matrices of each view, thereby achieving cross-view semantic unification. Step 4: Divide the semantic units of diabetic retinopathy lesions through pre-modeling of spheres, and select initial cluster centers from high-confidence spheres; divide the features after dual-space semantic calibration into spheres, and based on the local semantic features of diabetic retinopathy (DR) lesions, aggregate samples that are close in distance and semantically similar in the feature space into two or more sphere units, with each sphere representing a set of samples with homogeneous lesion features; after completing the sphere construction, calculate the confidence of each sphere. The confidence is comprehensively judged based on the compactness of samples within the sphere, the separation between spheres, and the matching degree with the semantics of the clinical grading of diabetic retinopathy (DR), and select high-confidence spheres with confidence scores higher than a preset threshold, and use the center of the high-confidence sphere as the initial cluster center for two-stage clustering; Step 5: Based on the confidence level of each particle, the exploration and utilization strategies and the cluster center update step size are dynamically adjusted, and combined with the semantic calibration signal, adaptive clustering optimization is achieved. During the iteration process, the exploration and utilization weights are dynamically allocated according to the confidence level of each particle. For particles with high confidence, the focus is on utilizing existing clustering information to stabilize the cluster centers. For particles with low confidence, the focus is on exploring potential cluster structures. At the same time, the cluster center update step size is dynamically adjusted according to the dual-space semantic calibration signal, and iterative updates are performed until the change in cluster centers is less than a preset threshold. Finally, the clustering results of diabetic retinopathy (DR) lesions are output.
2. The method according to claim 1, characterized in that, In step 1, an independent encoder is designed for each mode. : (1), in, This represents the independent encoder function corresponding to the m-th mode. Indicates the first The sample at the th Original input data for each modality For the first The learnable parameters of an encoder, For the first The sample at the th View-specific embedding features extracted under various modalities For a unified embedding dimension.
3. The method according to claim 2, characterized in that, Step 2 includes the following steps: Step 2.1: Analyze the view-specific embedding features output by each independent encoder in Step 1. The quality of the view is evaluated from two dimensions: cluster separability and feature reliability, and the confidence weight of each view is dynamically calculated. First, calculate the view. Cluster separability score : (2), in, For the sample The average distance to other samples in the same cluster, For the sample The average distance to the nearest heterogeneous sample; Secondly, calculate the view Feature reliability score : (3), in, For view Variance of features It is the first The variance of each view feature Indicated in all views The function that takes the maximum value above is synthesized to obtain the first... Confidence weight of each view : (4), in, It is a balance coefficient, and satisfies ,in Total number of views; Step 2.2: Based on confidence level, select the top-2 high-confidence views as information sources, and use a self-attention mechanism to complete the information for low-confidence views. Features after completion The calculation formula is: (5), in, For low-confidence view indexes, For the first One sample in the low confidence view The embedded features after padding For the first One sample in the low confidence view The original embedding features below, For high-confidence view indexes, This is a set of Top-2 high-confidence views. For the first One sample in the high confidence view Embedded features, For self-attention mechanism functions, attention mechanism Defined as: (6), in, , , Represents the query vector, key vector, and value vector; softmax represents the activation function; Indicates transpose; Step 2.3: Utilize the characteristic variance distribution of the low-confidence view to identify and filter redundant interference introduced by equipment noise and imaging artifacts in the high-confidence view, thus improving the high-confidence view. Features after denoising The calculation formula is: (7), in, For the first One sample in the high confidence view Denoising-reduced embedding features For element-wise product, For the first One sample in the high confidence view The noise mask below, noise mask Defined as: (8), in, Low confidence views The characteristic mean and standard deviation, This is the noise threshold coefficient. For indicator functions; Step 2.4: Weighted aggregation of the view features after bidirectional mutual enhancement to generate global fusion features. .
4. The method according to claim 3, characterized in that, In step 2.4, global fusion features Represented as: (9), in, For view The feature matrix after mutual enhancement.
5. The method according to claim 4, characterized in that, Step 3 includes the following steps: Step 3.1, global fusion features Perform K-means pre-clustering to obtain the initial cluster centers. Construct a global public semantic space and compute samples To the cluster center initial allocation probability : (10), in, For the first The global fusion feature vector corresponding to each sample For the first One initial cluster center, For the first One initial cluster center, For temperature coefficient, For the target number of clusters, It is a natural exponential function; Step 3.2: In the feature space, semantic consistency of multimodal features of the same patient is enhanced through instance-level and feature-level contrastive learning. Instance-level contrastive loss is used. The calculation formula is: (11), in, For the sample Enhanced features of the same view For the sample The corresponding global fusion feature vector, similarity function Where a and b are intermediate parameters. Temperature coefficient; Feature-level contrast loss The calculation formula is: (12), in, For the first The sample at the th Original embedded features under each view Total number of views; Total feature space loss Represented as: (13); Step 3.3: In the clustering space, calculate the optimal transmission probability from each view feature to the cluster center based on optimal transmission, generate a semantic assignment matrix, and define the view. Transmission cost matrix : (14), in, For the first The sample at the th View down to the first The transmission cost of each cluster center; The optimal transport plan matrix is solved by Sinkhorn iteration using the Sinkhorn algorithm. : (15), in, For the transmission probability matrix, To satisfy the sample distribution Distribution with cluster centers The set of feasible transfer matrices under constraints. For Frobenius inner product operations, For entropy regularization, for The sample was assigned to the first... The transmission probability of each cluster center This is the entropy regularization coefficient; Step 3.4: Optimize the consistency of clustered pseudo-labels through highly reliable pseudo-label propagation, and use KL divergence constraints to align the distribution of semantic assignment matrices for each view. Calculated by the following formula: (16), in, For the first Pseudo-labels for each sample For the first The sample at the th The view is assigned to the first The transmission probability of each cluster center; Distribution Alignment Loss Calculated by the following formula: (17), in, For KL divergence calculation, For the first Optimal transmission plan under each view; Total clustering space loss Calculated by the following formula: (18), in, For the first The sample corresponds to the first Classification of pseudo-label one-hot encoded vectors in clusters. This is the balance coefficient; Step 3.5: Jointly optimize the loss in the feature space and the clustering space to achieve collaborative semantic calibration in both spaces. (19), in, This is the loss for dual-space collaborative semantic calibration.
6. The method according to claim 5, characterized in that, Step 4 includes the following steps: Step 4.1, perform semantic calibration on the features in the dual-space model. Granulosphere segmentation is performed based on the local semantic features of diabetic retinopathy (DR) lesions. Samples that are close in distance and semantically similar in the feature space are aggregated into multiple granulosphere units. The construction of is represented by the following formula: (20), in, For the first One sample, For the first The calibrated global fusion feature vector corresponding to each sample For granules The initial center, Distance threshold The similarity threshold; Step 4.2: Calculate the confidence score for each granule. The confidence score is determined comprehensively based on the compactness of the sample within the granule, the separation between granules, and the semantic match with the clinical grading of diabetic retinopathy (DR). Intragranular compactness... Calculated by the following formula: (21), in, For granules Number of samples within; intergranular separation Calculated by the following formula: (22), in, For granules The initial center, the first Confidence level of individual balls Calculated by the following formula: (23), in This is the balance coefficient; Step 4.3: Filter out those with a confidence level higher than the preset threshold. The high-confidence spheres are used, and their centers are used as the initial cluster centers for two-stage clustering. .
7. The method according to claim 6, characterized in that, In step 4.3, the initial cluster centers Represented as: (24)。 8. The method according to claim 7, characterized in that, Step 5 includes the following steps: Step 5.1: Dynamically allocate exploration and utilization weights based on particle confidence levels to adaptively adjust exploration and utilization strategies during clustering iterations. For high-confidence particles, assign higher utilization weights to stabilize the clustering structure; for low-confidence particles, assign higher exploration weights to uncover potential cluster distributions. The allocation of exploration and utilization weights is calculated using the following formula: (25), in, For the first Confidence level of individual balls, For granules The use of weights, For granules The exploration weight; Step 5.2: Combining the dual-space semantic calibration signal, dynamically adjust the update step size of the cluster centers to achieve precise adaptation for cluster optimization; based on the semantic consistency score between the feature space and the cluster space, determine the global update step size for each iteration, and the semantic consistency score... Calculated by the following formula: (26), in, For the first The optimal transmission probability in round iteration. For the first Feature assignment probabilities in round iterations, based on scores Calculate dynamic step size : (27), in, Based on step size, This is the step size adjustment factor; Step 5.3: Utilize weights and dynamic step size to iteratively update cluster centers until convergence conditions are met. The first round Cluster centers The update formula is: (28), in, For the first The first iteration Cluster centers, For granules Belongs to the first A sample set of cluster centers For set Number of samples within, Gaussian noise for the exploration term; Step 5.4: Determine whether the iterative change in cluster centers is less than the preset convergence threshold. The convergence criterion is expressed by the following formula: (29), If the convergence criterion is met, the iteration terminates, and the final cluster centers are determined. The corresponding clustering results are used as outputs to obtain the lesion clustering and lesion grading results of diabetic retinopathy.
9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing program code that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, It stores a computer program or instructions that, when run on a computer, perform the steps of the method as described in any one of claims 1 to 8.