Rare vitreoretinopathy severity assessment system based on large model
By employing large-model-based visual recalibration, semantic location awareness, and cosine contrast regularization clustering, the problems of data scarcity and feature loss in RVD diagnosis are addressed, achieving high-accuracy evaluation without training data and enhancing the model's interpretability and applicability.
Patent Information
- Application Number
- CN202510932696.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-11-04
AI Technical Summary
Existing deep learning methods face challenges in diagnosing rare vitreoretinal diseases (RVD), including data scarcity, reliance on training data and prompts, lack of fine-grained features, and insufficient model generalization ability, resulting in poor diagnostic accuracy.
A rare vitreoretinal disease severity assessment system based on a large model is adopted. Through a visual recalibration module, a semantic location awareness module, a visual-semantic fusion module, and cosine contrast regularization clustering, severity assessment without training data and training prompts can be achieved.
It breaks through the dependence on large amounts of labeled data, significantly improves the accuracy of severity assessment for rare vitreoretinal diseases, and enhances the interpretability and applicability of the model, making it suitable for complex medical image analysis.
Smart Images

Figure CN120894288A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of medical information visual processing, and particularly relates to a rare vitreoretinal disease severity evaluation system based on a large model, in particular, a rare vitreoretinal disease severity evaluation algorithm (TPFree) based on general large model assistance. BACKGROUND
[0002] Rare vitreoretinal disease (RVD) is a rare ophthalmic disease that mainly affects the vitreous and retinal structures, and is mostly a genetic disease. Due to its low incidence and complex pathology, the diagnosis and treatment of RVD have always been a great challenge. The existing clinical diagnosis mainly relies on the experience of doctors and traditional medical image analysis methods such as fundus photography, retinal fluorescence angiography, etc., but these methods are limited by the lack of clinical data, the complexity of the lesion type and the difference in the experience of doctors in making diagnoses, resulting in difficulty in making accurate diagnoses on some special types of retinal lesions.
[0003] With the continuous development of deep learning and computer vision technology, computer-aided diagnosis (CA) methods based on image processing have gradually been applied in the field of ophthalmology. Existing deep learning models, especially those based on convolutional neural networks (CNN), have made some progress in the diagnosis of common ophthalmic diseases such as diabetic retinopathy. However, many existing methods require a large amount of labeled data for feature extraction and model training, often relying on a large amount of training data and explicit training cues, and usually relying on specific training cues to adjust the behavior of the model. This dependence limits the adaptability of the method, making it difficult for them to play an effective role in the case of data scarcity. Especially in the case of insufficient data or lack of clear training cues, the diagnostic performance is poor.
[0004] The patent document "Diabetic retinal image classification method and system based on deep learning" (CN108615051A) discloses a hemorrhagic lesion recognition model and an exudative lesion recognition model, which are trained by labeling the hemorrhagic lesion area and exudative lesion area in the fundus image, respectively, and then inputting the FCN model. Although this reduces the requirement for the description ability of the network model, making the model easy to train, it still needs a standard data pre-trained model, and the training model still needs classification recognition and labeled data.
[0005] Due to the scarcity of RVD image data and the diversity of lesion structures, existing deep learning methods still face some challenges. For example, the rarity of RVD makes clinical data scarce, and it is very difficult to obtain enough labeled data, so many traditional deep learning methods cannot obtain enough labeled data to train an efficient diagnostic model, and thus cannot achieve ideal performance in the computer-aided diagnosis of RVD.
[0006] For such a diversified and complex disease as RVD, traditional deep learning methods often fail to accurately capture the fine-grained features of the lesion area. Even some advanced visual models may have diagnostic errors due to lack of fine perception of pathological structures. Existing deep learning models, especially in the diagnosis of rare diseases, often face the problem of overfitting. This means that when the model performs well on the training set, its performance on real-world data or test sets of different lesion types is often poor.
[0007] The patent document "Retinal diabetic retinopathy deep network detection method based on genetic fuzzy tree" (CN114494196A) discloses that through enhancement processing, the image is segmented to construct an interpretable fuzzy decision tree, encode and construct a fitness function, and perform combination optimization based on a genetic algorithm. Although it can accurately identify retinal diabetic retinopathy blood vessel endings and improve detection classification accuracy, it still lacks fine-grained features in the image, especially lesion edges and important structures, and relies on real diagnosis results for training.
[0008] Therefore, a rare vitreoretinal disease severity evaluation algorithm based on general large model assistance (TPFree) is proposed to realize severity evaluation without training prompts and training data based on general models, and to improve the accuracy of RVD severity evaluation. SUMMARY
[0009] In view of the defects in the prior art, the purpose of the present application is to provide a rare vitreoretinal disease severity evaluation system based on a large model.
[0010] According to the rare vitreoretinal disease severity evaluation system based on a large model provided by the present application, the image data is input into the visual recalibration module to extract visual features, the segmentation mask is calibrated through segmentation, the semantic features are extracted by the semantic position perception module, the saliency map is generated and hierarchical pooling filtering is performed.
[0011] The visual recalibration module extracts visual features from the input image data, calibrates the feature map through segmentation mask, and the semantic position perception module extracts semantic features to generate a saliency map and perform hierarchical pooling filtering.
[0012] The visual-semantic fusion module hierarchically fuses the feature map and the saliency map to generate a semantic feature vector.
[0013] The cosine contrast regularization clustering calculates the cosine distance and density regularization according to the semantic feature vector, and clusters to evaluate the severity.
[0014] Preferably, the visual recalibration module generates a preliminary segmentation mask from the input image data through a large-scale pre-trained visual model, and extracts three feature maps of three different spatial scales from the preliminary segmentation mask and upsampling and element-wise addition operations are performed:
[0015]
[0016] Channel-wise mean calibration is performed on the feature map:
[0017]
[0018] wherein H represents the feature height;
[0019] W represents the feature weight;
[0020] C represents the feature channel;
[0021] represents the feature dimension;
[0022] i represents the layer sequence number;
[0023] I represents the number of feature map layers;
[0024] represents upsampling;
[0025] represents element-level addition;
[0026] represents the pixel set;
[0027] Mask (x) represents the xth pixel area of the segmentation mask obtained by reaching the model;
[0028] p, q each represent the value of the coordinate pole, ranging from -2 to 2;
[0029] i, j each represent the value of the coordinate pole, ranging from ;
[0030] represents the visual feature.
[0031] Preferably, the semantic position perception module extracts semantic features of text and images through a text-image model, and for input text (i) and img, three saliency maps
[0032]
[0033] m∈{1,2,3}
[0034] The hierarchical pooling layer is used to filter coarse, medium and fine-grained features to obtain lesion area positioning information:
[0035]
[0036] wherein, denotes a text-image large model;
[0037] denotes a text encoder;
[0038] text (i) denotes the i-th layer text feature;
[0039] img denotes an image semantic feature;
[0040] H denotes a feature height;
[0041] W denotes a feature weight;
[0042] C denotes a feature channel;
[0043] σ denotes an activation function;
[0044] denotes a feature dimension;
[0045] h, w respectively denote the length and width of the image;
[0046] P (x) denotes a pooling operation;
[0047] x denotes a pooling size, taking 16, 8 and 4;
[0048] denotes the m-th layer saliency map;
[0049] denotes the m-th layer saliency map at the (i, j) position;
[0050] M i,j , S i,j both denote whether the position (i, j) is located in the target region, if yes, 1, if not, 0;
[0051] ||M (m) ||0 denotes the number of non-zero elements.
[0052] Preferably, the visual-semantic fusion module comprises a CVR module and an SLP module;
[0053] The SLP module down-samples the saliency map to obtain a semantic feature:
[0054]
[0055] The CVR module processes the feature map, hierarchically fuses the visual feature and the semantic feature, generates a feature vector, and reduces the dimension:
[0056]
[0057] wherein, represents the I-th layer feature map;
[0058] represents down-sampling;
[0059] ⊙ represents Hadamard product;
[0060] ⊕ represents element-wise addition;
[0061] represents feature dimension;
[0062] represents the I-th layer saliency map;
[0063] represents the I-th layer semantic feature intermediate value;
[0064] represents the I-th layer visual feature
[0065] represents the I-th layer semantic feature.
[0066] Preferably, the cosine contrast regularization clustering comprises cosine distance calculation and regularization clustering;
[0067] The cosine distance calculation is to calculate the cosine distance between the feature map and the semantic feature
[0068]
[0069] The regularization clustering is a contrast cosine distance optimization clustering operation, and the severity evaluation result is obtained by a clustering method:
[0070]
[0071] wherein, C represents a feature channel;
[0072] c represents a feature channel ordinal number;
[0073] d h,w,h',w′ represents the distance between two pixels located at (h', w') and (h, w);
[0074] h ' , w' respectively represent another pixel position different from h, w position in the same mask category cluster;
[0075] d max represents the maximum distance between two pixels;
[0076] L cluster represents the target function between clusters;
[0077] L density denotes density regularization.
[0078] According to the application, a rare vitreoretinal pathology severity assessment method based on a large model is provided, comprising:
[0079] a visual recalibration step, extracting visual features of image data, and calibrating feature maps through segmentation mask segmentation;
[0080] a semantic position perception step, inferring semantic features, generating a saliency map and performing hierarchical pooling filtering on the saliency map;
[0081] a fusion step, hierarchically fusing the feature maps of the visual features and the saliency map of the semantic features to obtain a semantic feature vector;
[0082] an evaluation step, calculating cosine distance and density regularization according to the semantic feature vector, and clustering to evaluate the severity.
[0083] Preferably, in the visual recalibration step, the input image data is subjected to a large-scale pre-trained visual model to generate a preliminary segmentation mask, from which three feature maps of three different spatial scales are extracted and up-sampling and element-wise addition operations are performed:
[0084]
[0085] The feature maps are subjected to channel-wise mean calibration:
[0086]
[0087] where H represents the feature height;
[0088] W represents the feature weight;
[0089] C represents the feature channel;
[0090] denotes the feature dimension;
[0091] i represents the layer number;
[0092] I represents the number of feature map layers;
[0093] denotes up-sampling;
[0094] ⊕ denotes element-wise addition;
[0095] denotes pixel aggregation;
[0096] Mask (x)represents the xth segmentation mask pixel region obtained by the arrival model;
[0097] p, q represent the value of the coordinate pole, the range is between -2, 2;
[0098] i, j represent the value of the coordinate pole, the range is between ;
[0099] represents the visual feature.
[0100] Preferably, in the semantic position perception step, the semantic features of the text and the image are extracted by a text-image model, and for input text (i) and img, three saliency maps
[0101]
[0102] m belongs to {1, 2, 3}
[0103] The hierarchical pooling layer is used to filter the coarse, medium and fine-grained features to obtain the lesion area positioning information:
[0104]
[0105] wherein, represents a text-image large model;
[0106] represents a text encoder;
[0107] text (i) represents the i-th layer text feature;
[0108] img represents the image semantic feature;
[0109] H represents the feature height;
[0110] W represents the feature weight;
[0111] C represents the feature channel;
[0112] σ represents the activation function;
[0113] represents the feature dimension;
[0114] h, w represent the length and width of the image respectively;
[0115] P (x) represents the pooling operation;
[0116] x represents the pooling size, which is 16, 8 and 4;
[0117] represents the m-th layer saliency map;
[0118] represents the m-th layer saliency map at the (i,j) position;
[0119] M i,j , S i,j represents whether the position (i,j) is located in the target region, 1 if located, 0 if not located;
[0120] ||M (m) ||0 represents the number of non-zero elements.
[0121] Preferably, in the fusion step, the saliency map is down-sampled to obtain semantic features:
[0122]
[0123] The feature map is processed, the visual features and the semantic features are hierarchically fused to generate a feature vector, and the dimension is reduced:
[0124]
[0125] wherein, represents the I-th layer feature map;
[0126] represents down-sampling;
[0127] ⊙ represents Hadamard product;
[0128] ⊕ represents element-level addition;
[0129] represents the feature dimension;
[0130] represents the I-th layer saliency map;
[0131] represents the I-th layer semantic feature intermediate value;
[0132] represents the I-th layer visual feature
[0133] represents the I-th layer semantic feature.
[0134] Preferably, the cosine distance is calculated as the cosine distance between the feature map and the semantic feature :
[0135]
[0136] The clustering evaluation is a contrastive cosine distance optimization clustering operation, and the severity evaluation result is obtained by a clustering method:
[0137]
[0138] C represents a feature channel;
[0139] c represents a feature channel ordinal;
[0140] d h,w,h′,w′ represents the distance between the two pixels located at (h', w') and (h, w);
[0141] h', w' respectively represent another pixel position different from h, w within the same mask category cluster;
[0142] d max represents the maximum distance between the two pixels;
[0143] L cluster represents the target function between clusters;
[0144] L density represents density regularization.
[0145] Compared with the prior art, the present application has the following beneficial effects:
[0146] 1. The present application can diagnose without training data and training hints by adopting the TPFree method, breaking through the dependence of traditional methods on a large amount of labeled data.
[0147] 2. The present application can capture fine-grained features in images, especially lesion edges and important structures, through multi-module cooperation, significantly improving the accuracy of severity evaluation.
[0148] 3. The present application adopts visual-semantic fusion and hierarchical pooling technology, and the model can provide clear feature mapping and explanation when diagnosing, enhancing its explainability and credibility in clinical application. It is not only suitable for the diagnosis of RVD and other rare diseases, but also can be extended to other complex medical image analysis fields. BRIEF DESCRIPTION OF DRAWINGS
[0149] Other features, objects and advantages of the present application will become more apparent through reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings:
[0150] Figure 1 is a schematic diagram of a rare vitreoretinal disease severity evaluation system based on a large model;
[0151] Figure 2 is a schematic diagram of a visual recalibration process;
[0152] Figure 3 is a schematic diagram of a semantic position perception process;
[0153] Figure 4 The visual semantic fusion process schematic diagram is shown in the following figure;
[0154] Figure 5 The cosine distance contrast regularization process schematic diagram is shown in the following figure. DETAILED DESCRIPTION
[0155] The application will be described in detail below with specific embodiments. The following examples will help those skilled in the art to further understand the application, but do not limit the application in any form. It should be pointed out that for those skilled in the art, without departing from the concept of the application, a number of changes and improvements can be made. These all belong to the protection scope of the application.
[0156] The existing deep learning and computer vision technology has the problems of data scarcity, dependence on training data and prompts, lack of fine-grained features, poor model generalization ability, etc. In view of these deficiencies of the prior art, the present application proposes a rare vitreoretinal disease (RVD) severity assessment system based on a large model. The rare vitreoretinal disease severity assessment algorithm (TPFree) based on general large model assistance is adopted. By introducing a variety of technical modules, the training difficulty caused by data scarcity in the existing method is overcome, and fine-grained lesion features can be accurately captured. In the absence of training data and training prompts, the prior knowledge of general large-scale visual and semantic models is used for RVD severity assessment, improving the accuracy of RVD severity assessment. Taking Figure 1 For example, it includes:
[0157] Visual recalibration (CVR) module, semantic location perception (SLP) module, visual-semantic fusion (VSF) module, cosine distance contrast regularization clustering. Each module contains specific technical steps and formulas to ensure the realization of RVD image severity assessment without training prompts.
[0158] Specifically, the purpose of the visual recalibration (CVR) module is to generate a preliminary feature map through a visual model (such as SAM) and to enhance the edge information in the low-dimensional feature map in fine granularity. Taking Figure 2 For example, based on SAM segmentation, the recalibrated visual features are refined through upsampling, including:
[0159] Preliminary feature extraction: input image data is passed through a large-scale pre-trained visual model (such as SAM) to generate a preliminary segmentation mask and feature map.
[0160] Extracting multi-scale features: three feature maps of three different spatial scales are extracted from the segmentation mask
[0161] Wherein, represents The feature map of the last stage I. The feature dimension is represented by H, the feature height is represented by W, the feature weight is represented by C, and the layer sequence number is represented by i. Since the edge detail features are missing in the low-dimensional feature map, the up-sampling And the element-level addition ⊕ is introduced from the maximum spatial dimension feature map at the earliest stage.
[0162]
[0163] Channel-wise mean recalibration: The channel-wise mean calibration is performed on the feature map. The formula is as follows:
[0164]
[0165] Wherein, represents the pixel set, Mask (x) represents the xth segmentation mask pixel area obtained by reaching the model; p and q both represent the value of the coordinate pole, ranging from -2 to 2, i and j both represent the value of the coordinate pole, ranging from
[0166] Fine-grained edge enhancement: Through up-sampling and element-wise addition operation, the edge information in the low-dimensional feature map is enhanced to ensure that the details of the lesion area are retained.
[0167] The semantic location perception (SLP) module improves the semantic perception ability of visual features by introducing semantic information based on a text-image model (such as CLIP), especially in the perception of multi-scale lesion areas, to obtain a hierarchical down-sampled location perception semantic map based on an explicit saliency map. Taking Figure 3 For example, it specifically includes:
[0168] Semantic feature extraction: Through a text-image model (such as CLIP), the semantic features of text and image are extracted from the image, and the implicit medical concept is perceived as an explicit text representation through the semantic location perception (SLP) module. For the input text (i) and img, three saliency maps
[0169]
[0170] Wherein, represents a text-image large model, represents a text encoder. h and w represent the length and width of the image, respectively. σ represents the activation function Sotfmax function.
[0171] Multi-scale pooling: In order to perceive different lesion areas (retinal area, macular area, optic disc area, etc.), a hierarchical pooling layer P(16) , P (8) , P (4) Filtering the coarse, medium and fine granularity features of the macular region Retinal region And intervertebral disc region The formula is as follows:
[0172]
[0173] Where, P (x) Indicates the pooling operation, and the pooling size is x. The feature representation of each pixel is padded x and pooled. Indicates m th Layer filtering saliency map, i.e. the m-th layer saliency map, M i,j Indicates whether the position (i,j) is located in the target region (1 if located, otherwise 0). ||M (m) ||0 represents the number of non-zero elements. Indicates the m-th layer saliency map at the (i,j) position, i,j represents the corresponding coordinates, that is, the saliency map is a mask map.
[0174] The semantic prior knowledge of multi-scale filtering is introduced, and it has the effect of region independence, which enhances the dispersion of the feature vector in the parameter-free evaluation.
[0175] Hierarchical feature perception: through the semantic feature enhanced by multi-level pooling processing, the perception of the lesion region is enhanced, and finally more fine lesion region positioning information is obtained.
[0176] The visual-semantic fusion (VSF) module fuses visual features and semantic features to generate a unified feature vector to enhance the expression ability of the lesion region. Based on average pooling, the visual features and semantic positions are hierarchically fused to Figure 4 For example, specifically includes:
[0177] Fusion of visual and semantic features: since the visual recalibration feature representation cannot filter noise and the semantic perception feature representation cannot present details, after the CVR upsampling operation, a visual semantic fusion (VSF) module is proposed. The semantic feature is down-sampled to match the dimension. The input visual feature is obtained by processing through the CVR module, and the semantic feature is obtained by the SLP module. The formula is as follows:
[0178]
[0179] Where, Indicates the intermediate value of the visual semantic fusion module, which is the semantic feature.
[0180] Dimension reduction processing: in order to unify the fused feature vector to a lower dimensional space, a dimension reduction operation is performed for subsequent clustering analysis:
[0181]
[0182] wherein, represents down-sampling, and represents Hadamard product, represents semantic features. In this way, the gap between visual feature refinement (prior knowledge restriction) and semantic position perception (training phase requirement) is bridged.
[0183] In order to optimize the severity evaluation, the cosine contrast regularization clustering performs cosine distance contrast and density regularization, which is used to enhance the feature separation degree of different severity levels, so as to enhance the performance embedded in the clustering. Specifically, taking Figure 5 as an example, it includes:
[0184] Cosine distance calculation: calculate the cosine distance between visual -semantic features
[0185]
[0186] Regularization clustering: the clustering operation is optimized by cosine distance contrast to ensure that the features of different severity levels are effectively separated in the embedding space, and finally the severity evaluation result is obtained through the clustering method. The formula is as follows:
[0187]
[0188] wherein, d h,w,h′,w′ represents the distance between two pixels, and h', w' respectively represent another pixel position different from h, w within the same mask category cluster. L cluster represents the target function between clusters calculated by using a classical clustering algorithm, so as to optimize the classical clustering operation.
[0189] By fusing semantic position and visual position, a pre-trained model of a general image is used to obtain a relatively satisfactory result with high generalization ability in a high-difficulty medical scene through a feature extractor without training.
[0190] According to the rare vitreoretinal disease (RVD) severity evaluation method based on a large model provided by the application, the prior knowledge of a large-scale pre-trained visual and semantic model is used to realize the severity evaluation without training prompt and training data. Specifically includes:
[0191] The visual recalibration step, the multi-scale visual features of the image data are extracted from the SAM encoder, and by segmenting re-calibrate them
[0192] In more preferred examples, the input image data is passed through a large pre-trained visual model to generate a preliminary segmentation mask from which three feature maps of three different spatial scales are extracted and up-sampling and element-wise addition operations are performed:
[0193]
[0194] The feature maps are channel-wise mean calibrated:
[0195]
[0196] where H represents the feature height; W represents the feature weight; C represents the feature channel; represents the feature dimension; i represents the layer number; I represents the number of feature map layers; represents up-sampling; and represents element-wise addition. represents the pixel set; Mask (x) represents the xth pixel area of the segmentation mask obtained by reaching the model; p and q both represent the value of the coordinate pole, ranging between -2, 2; i and j both represent the value of the coordinate pole, ranging between ; represents the visual feature.
[0197] Semantic position perception step, inferring semantic saliency map and performing hierarchical average pooling semantic weight on it
[0198] In more preferred examples, semantic features of text and images are extracted by a text-image model, for input text (i) and img, three saliency maps are obtained
[0199]
[0200] m∈{1,2,3}
[0201] The hierarchical pooling layer is used to filter coarse, medium and fine-grained features to obtain lesion area positioning information:
[0202]
[0203] wherein, represents a text-image large model; represents a text encoder; text (i)represents the i-th layer text feature; img represents the image semantic feature; H represents the feature height; W represents the feature weight; C represents the feature channel; and σ represents the activation function. represents the feature dimension; h and w represent the length and width of the image, respectively; (x) represents the pooling operation; x represents the pooling size, which is 16, 8, and 4; represents the m-th layer saliency map; represents the m-th layer saliency map at the (i, j) position; M i,j , S i,j both represent whether the position (i, j) is located in the target region, and if so, 1, and if not, 0; ||M (m) ||0 represents the number of non-zero elements.
[0204] The fusion step is to efficiently perceive the structure, realize the robust feature representation, and fuse the visual feature and the semantic weight .
[0205] In more preferred examples, the saliency map is down-sampled to obtain the semantic feature:
[0206]
[0207] The processing feature map, the visual feature and the semantic feature are hierarchically fused to generate a feature vector, and are reduced in dimension:
[0208]
[0209] wherein, represents the I-th layer feature map; represents the down-sampling; ⊙ represents the Hadamard product; and ⊕ represents the element-level addition. represents the feature dimension; represents the I-th layer saliency map; represents the I-th layer semantic feature intermediate value; represents the I-th layer visual feature represents the I-th layer semantic feature.
[0210] The evaluation step is to use the cosine distance operation and the density regularization L density for clustering.
[0211] In more preferred examples, the calculation of the cosine distance is to calculate the cosine distance between the feature map and the semantic feature :
[0212]
[0213] The clustering evaluation is a contrast cosine distance optimization clustering operation, and the severity evaluation result is obtained through the clustering method:
[0214]
[0215] Wherein, C represents a feature channel; c represents a feature channel ordinal number; d h,w,h′,w′ represents a distance between two pixels located at (h', w') and (h, w); h', w' respectively represent another pixel position different from h, w in the same mask category cluster; d max represents the maximum distance between two pixels; L cluster represents an objective function between clusters; L density represents density regularization.
[0216] The specific embodiments of the present application are described above. It needs to be understood that the present application is not limited to the above specific embodiments, and various changes or modifications can be made by those skilled in the art within the scope of the claims, which does not affect the essential content of the present application. The embodiments of the present application and the features in the embodiments can be arbitrarily combined with each other without conflict.
Claims
1. A system for assessing the severity of rare vitreoretinal diseases based on a large model, characterized in that, include: Visual recalibration module, semantic location awareness module, visual-semantic fusion module, cosine contrast regularization clustering; Image data is input into the visual recalibration module to extract visual features. The feature map is calibrated by segmentation mask. The semantic location awareness module extracts semantic features, generates a saliency map, and performs hierarchical pooling filtering on it. The visual-semantic fusion module fuses feature maps and saliency maps in layers to generate semantic feature vectors; Cosine contrastive regularization clustering calculates cosine distance and density regularization based on semantic feature vectors, and assesses the severity of clustering.
2. The system for assessing the severity of rare vitreoretinal lesions based on a large model according to claim 1, characterized in that, The visual recalibration module processes the input image data through a large-scale pre-trained visual model to generate a preliminary segmentation mask, from which three feature maps at three different spatial scales are extracted. And perform upsampling and element-wise addition operations: Channel-wise mean calibration of the feature map: Where H represents the feature height; W represents the feature weight; C represents the feature channel; Indicates the feature dimension; i represents the level number; I represents the number of feature layers; Indicates upsampling; This represents element-wise addition. Represents a collection of pixels; Mask (x) This represents the x-th segmentation mask pixel region obtained from the arriving model; p and q both represent the extreme points of the coordinates, ranging from -2 to 2; i and j both represent the extreme points of the coordinates, ranging from... Inside; Indicates visual characteristics.
3. The system for assessing the severity of rare vitreoretinal lesions based on a large model according to claim 1, characterized in that, The semantic location awareness module extracts semantic features from text and images using a text-image model. For the input text... (i) And img, we get three saliency maps. m∈{1,2,3} By using a hierarchical pooling layer to filter coarse, medium, and fine-grained features, lesion region localization information is obtained. in, Represents a large text-image model; Indicates a text encoder; text (i) Represents the text features of the i-th layer; img represents the semantic features of an image; H represents the feature height; W represents the feature weight; C represents the feature channel; σ represents the activation function; Indicates the feature dimension; h and w represent the length and width of the image, respectively; P (x) Indicates pooling operation; x represents the pooling size, which can be 16, 8, or 4. This represents the saliency map of the m-th layer; This represents the m-th saliency map at position (i,j); M i,j S i,j Both indicate whether the position (i,j) is within the target area; if it is, the value is 1, and if it is not, the value is 0. ||M (m) ||0 represents the number of non-zero elements.
4. The system for assessing the severity of rare vitreoretinal lesions based on a large model according to claim 1, characterized in that, The visual-semantic fusion module includes a CVR module and an SLP module; The SLP module downsamples the saliency map to obtain semantic features: The CVR module processes the feature map, fusing visual and semantic features in layers to generate feature vectors and reduce their dimensionality. in, This represents the feature map of layer I; Indicates downsampling; ⊙ represents the Hadamard product; This represents element-wise addition. Indicates the feature dimension; This represents the saliency map of layer I; This represents the intermediate value of the semantic features at layer I; Represents the visual features of layer I This represents the semantic features of the I-th layer.
5. The system for assessing the severity of rare vitreoretinal lesions based on a large model according to claim 1, characterized in that, The cosine comparison regularized clustering includes cosine distance calculation and regularized clustering; The cosine distance is calculated as a feature map. and semantic features Cosine distance between them: The regularized clustering is a contrastive cosine distance optimized clustering operation, and the severity assessment results are obtained through the clustering method: Where C represents the feature channel; c represents the feature channel ordinal number; d h,w,h',w' This represents the distance between two pixels located at (h',w') and (h,w); h' and w' represent different pixel positions within the same mask category cluster than h and w, respectively; d max This represents the maximum distance between two pixels; L cluster Represent the objective function between clusters; L density This indicates density regularization.
6. A method for assessing the severity of rare vitreoretinal lesions based on a large model, characterized in that, include The visual recalibration step involves extracting visual features from image data and calibrating the feature map using a segmentation mask. The semantic location awareness step involves inferring semantic features, generating a saliency map, and then performing hierarchical pooling filtering on it. The fusion step involves layer-by-layer fusion of the feature maps of visual features and the saliency maps of semantic features to obtain semantic feature vectors. The evaluation steps include calculating cosine distance and density regularization based on semantic feature vectors, and clustering to assess severity.
7. The method for assessing the severity of rare vitreoretinal lesions based on a large model according to claim 6, characterized in that, In the visual recalibration step, the input image data is processed through a large-scale pre-trained visual model to generate a preliminary segmentation mask, from which three feature maps of different spatial scales are extracted. And perform upsampling and element-wise addition operations: Channel-wise mean calibration of the feature map: Where H represents the feature height; W represents the feature weight; C represents the feature channel; Indicates feature dimension; i represents the level number; I represents the number of feature layers; Indicates upsampling; This represents element-wise addition. Represents a collection of pixels; Mask (x) This represents the x-th segmentation mask pixel region obtained from the arriving model; p and q both represent the extreme points of the coordinates, ranging from -2 to 2; i and j both represent the extreme points of the coordinates, ranging from... Inside; Indicates visual characteristics.
8. The method for assessing the severity of rare vitreoretinal lesions based on a large model according to claim 6, characterized in that, In the semantic location awareness step, semantic features of text and image are extracted using a text-image model. For the input text... (i) And img, we get three saliency maps. m∈{1,2,3} By using a hierarchical pooling layer to filter coarse, medium, and fine-grained features, lesion region localization information is obtained. in, Represents a large text-image model; Indicates a text encoder; text (i) Represents the text features of the i-th layer; img represents the semantic features of an image; H represents the feature height; W represents the feature weight; C represents the feature channel; σ represents the activation function; Indicates the feature dimension; h and w represent the length and width of the image, respectively; P (x) Indicates pooling operation; x represents the pooling size, which can be 16, 8, or 4. This represents the saliency map of the m-th layer; This represents the m-th saliency map at position (i,j); M i,j S i,j Both indicate whether the position (i,j) is within the target area; if it is, the value is 1, and if it is not, the value is 0. ‖M (m) ||0 represents the number of non-zero elements.
9. The method for assessing the severity of rare vitreoretinal lesions based on a large model according to claim 6, characterized in that, In the fusion step, the saliency map is downsampled to obtain semantic features: The feature maps are processed by fusing visual and semantic features in layers to generate feature vectors and then reducing their dimensionality. in, This represents the feature map of layer I; Indicates downsampling; ⊙ represents the Hadamard product; This represents element-wise addition. Indicates feature dimension; This represents the saliency map of layer I; This represents the intermediate value of the semantic features at layer I; Represents the visual features of layer I This represents the semantic features of the I-th layer.
10. The method for assessing the severity of rare vitreoretinal lesions based on a large model according to claim 6, characterized in that, The calculated cosine distance is used to calculate the feature map. and semantic features Cosine distance between them: The clustering evaluation involves a comparative cosine distance optimization clustering operation, and the severity assessment result is obtained through the clustering method. Where C represents the feature channel; c represents the feature channel ordinal number; d h,w,h′,w′ This represents the distance between two pixels located at (h',w') and (h,w); h' and w' represent different pixel positions within the same mask category cluster than h and w, respectively; d max This represents the maximum distance between two pixels; L cluster Represent the objective function between clusters; L density This indicates density regularization.
Citation Information
Patent Citations
Deep learning-based diabetic retina image classification method and system
CN108615051A
Genetic fuzzy tree-based retinal diabetes mellitus variable depth network detection method
CN114494196A