Fundus image segmentation system based on multi-stage encoding alignment and lesion perception

CN122176000BActive Publication Date: 2026-08-21ZHEJIANG UNIV OF TECH +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610644487.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-08-21
Estimated Expiration
2046-05-12

AI Technical Summary

Technical Problem

[0007]为解决现有眼底图像分割方法在有限标注数据条件下,对视网膜层状拓扑结构、细节纹理及病灶区域刻画能力不足,以及单次特征对齐易导致特征失准与语义错位的问题,提供一种能够高效利用无标签数据、融合多尺度全局语义与局部结构-细节互补特征、并通过多阶段渐进对齐增强特征一致性的半监督分割系统,从而显著提升对视网膜层边界、微小病灶等复杂结构的泛化能力与分割精度

Benefits of technology

[0029]本发明的总损失函数强制多尺度编码器中三个不同尺度的输出特征两两语义一致,消除尺度间的表征鸿沟,提升多尺度融合质量。通过双网络多阶段对齐,使两个分支的编码融合特征向历史平均水平收敛,避免过拟合单一样本,增强泛化能力。联合优化,同时实现有标签数据监督、无标签数据互蒸馏、多尺度内部对齐、双网络渐进对齐,形成完整的半监督训练闭环。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122176000B_ABST
    Figure CN122176000B_ABST
Patent Text Reader

Abstract

The application discloses a fundus image segmentation system based on multi-stage coding alignment and lesion perception, constructs a double-branch encoder containing a multi-scale encoder and a wavelet decomposition encoder, extracts multi-scale global semantic features, separates low-frequency topological structures and high-frequency detailed textures through wavelet decomposition, and respectively models and fuses them, so as to form multi-level coding fusion features complementary in structure and details; after adaptive enhancement of the fusion features by a lesion perception expert module, the decoding encoder gradually recovers the spatial resolution to obtain a segmentation result. In the training stage, a double-network semi-supervised mutual learning framework is adopted, a multi-stage coding alignment module is introduced, the coding features of the historical stages are averaged as alignment targets by using a sliding window memory storage unit, the fusion features are gradually aligned in cosine similarity, and multi-scale feature alignment loss and cross pseudo-supervision loss are combined to optimize, so that the ability to depict the retinal layered topological structure and the lesion boundary is enhanced under limited labeled data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical image processing technology, specifically relating to a fundus image segmentation system based on multi-stage coding alignment and lesion perception. Background Technology

[0002] The incidence of fundus diseases, such as diabetic macular edema and age-related macular degeneration, is increasing year by year, becoming one of the major causes of visual impairment and even blindness. This places higher demands on the early diagnosis and precise intervention of retinal diseases. Commonly used fundus examination methods include fundus color photography, ultra-wide-angle imaging, and fluorescein angiography, etc. However, traditional fundus images can usually only observe the surface information of the retina. But some diseases can only show lesions in the deep layers of the retina in the early stages, so the examination of the retinal layer structure is particularly important. Optical coherence tomography (OCT) can observe the deep structure of the retina, providing more comprehensive structural information. It has the characteristics of being non-invasive, high-resolution, and real-time imaging, providing new possibilities for the early detection and diagnosis of lesions.

[0003] Accurate retinal layer segmentation and lesion identification are crucial prerequisites for quantitatively analyzing disease progression and developing treatment plans. In recent years, deep learning, especially convolutional neural networks and visual Transformers, has made significant progress in OCT image segmentation tasks, but it still faces the following technical challenges: high cost of acquiring high-quality pixel-level labeled data, limited information provided by single-modal imaging, and insufficient generalization ability of models to fine structures and lesion regions. To address these issues, researchers have proposed improvement schemes from different perspectives.

[0004] For example, patent CN112750141A discloses a 3D iris surface reconstruction and quantization method and segmentation network based on anterior segment OCT images. This scheme utilizes a wavelet transform-based U-shaped segmentation network to extract the upper edge of the iris and performs 3D iris surface reconstruction using a guided optimization algorithm with adaptive Poisson disk sampling. Finally, curvature-related features are extracted from the 3D surface for angle-closure glaucoma screening. By introducing wavelet transform into the encoder-decoder skip connections of the segmentation network, redundant information is reduced and edge details are preserved by decomposing high- and low-frequency subbands. However, this method targets the iris structure of anterior segment OCT and is not directly applicable to retinal layer and lesion segmentation; furthermore, its application is limited to 3D surface reconstruction and classification, without involving semi-supervised learning or multimodal information fusion.

[0005] Patent application CN121305066A discloses a semi-supervised feature fusion segmentation network based on OCT-OCTA multimodal imaging. This scheme uses a pre-trained large model with an embedded adapter as the encoder to extract geometric features from both OCT and OCT (Optical Coherence Tomography) bimodal imaging. Through an intramodal information enhancement module (convolutional attention) and an intermodal wavelet extraction and fusion module, high- and low-frequency components are separated using Haar wavelet transform, and then spatial and frequency domain information is fused via a cross-attention mechanism. Finally, a contrastive consistency learning module regularizes the multi-path outputs. This scheme achieves good results in semi-supervised multimodal vessel segmentation. However, the intermodal fusion of this network mainly relies on the concatenation of wavelet decomposition and cross-attention, lacking explicit separation and modeling of topological information and texture details. Furthermore, its encoder relies only on a single large model and is not specifically designed for multi-scale features of the retinal layer, limiting its ability to characterize the continuity and boundary fineness of layered structures.

[0006] In summary, existing technologies face the following challenges. Firstly, retinal layer structures and lesion regions in OCT images often exhibit significant scale differences, blurred interlayer boundaries, and low local contrast. This is particularly true in lesion regions, where complex tissue morphology and subtle grayscale variations make it difficult for models to achieve accurate boundary localization and region differentiation. Secondly, precise annotation of OCT images relies on experienced clinicians; the pixel-by-pixel annotation process for each retinal layer boundary and lesion region is time-consuming and labor-intensive, thus limiting the acquisition of high-quality annotated data. Furthermore, OCT images are susceptible to speckle noise, signal attenuation, and imaging artifacts during acquisition, further increasing the difficulty of stable segmentation under complex conditions. Therefore, there is an urgent need for a fundus image segmentation system capable of integrating multi-scale semantic and structural-detail complementary information, and possessing efficient semi-supervised learning capabilities. Summary of the Invention

[0007] To address the shortcomings of existing fundus image segmentation methods in depicting the layered topology, detailed textures, and lesion regions of the retina under limited labeled data conditions, as well as the problem that single feature alignment can easily lead to feature inaccuracies and semantic misalignments, this paper proposes a semi-supervised segmentation system that can efficiently utilize unlabeled data, fuse multi-scale global semantics with complementary features of local structure and detail, and enhance feature consistency through multi-stage progressive alignment. This significantly improves the generalization ability and segmentation accuracy for complex structures such as retinal layer boundaries and small lesions.

[0008] To achieve the above-mentioned objectives, this invention provides a fundus image segmentation system based on multi-stage coding alignment and lesion perception, comprising a first segmentation network and a second segmentation network. The first segmentation network and the second segmentation network have the same structure and each independently includes: The data acquisition and processing module is used to acquire OCT fundus images and perform preprocessing. A dual-branch encoder is used to receive preprocessed OCT fundus images. Multi-scale features are extracted by a multi-scale encoder and then fused at the pixel level to obtain multi-scale features. At the same time, a wavelet decomposition encoder is used to extract topological features and detail texture features from the preprocessed fundus image data and then fused at the pixel level to obtain wavelet encoded features. The feature fusion module is used to fuse multi-scale features and wavelet-coded features to obtain coded fusion features; The lesion perception expert module is used to receive encoded fusion features, generate routing weights through a pre-trained lesion recognition module, and then use the expert network to adaptively weight the encoded fusion features based on the routing weights to obtain enhanced encoded features. The decoder receives the enhanced encoded features, restores the spatial resolution step by step, and finally outputs a pixel-level segmentation prediction map to complete the segmentation of the fundus image.

[0009] Preferably, the fundus image segmentation system further includes a training module, which comprises a multi-stage encoding alignment module, a cross-supervision unit, and a total loss calculation unit, for jointly training the first segmentation network and the second segmentation network using labeled and unlabeled data.

[0010] Preferably, the multi-stage encoding alignment module is used to receive a first encoding fusion feature and a second encoding fusion feature, perform channel-level fusion of the first encoding fusion feature and the second encoding fusion feature, generate a mapping feature through a mapper, and store it in a sliding window-type memory storage unit; calculate the average value of all historical mapping features in the memory storage unit as the alignment target, and align the first encoding fusion feature and the second encoding fusion feature to the alignment target by cosine similarity, and output the aligned feature; The feature fusion module of the first segmentation network outputs the first encoded fusion feature, and the feature fusion module of the second segmentation network outputs the second encoded fusion feature.

[0011] This invention generates low-dimensional mapping features through channel-level fusion and mapping, reducing redundancy. A sliding window memory storage unit stores mapping features from multiple historical stages, constructing a temporal feature structure. The average value of all historical features is used as the alignment target, avoiding drastic perturbations in a single alignment and achieving gradual, smooth alignment. Cosine similarity loss (rather than mean squared error loss) is employed, focusing on feature direction rather than amplitude, making it more suitable for aligning high-dimensional semantic features and maintaining the consistency of the fused features encoded by the two networks, thereby improving the stability of semi-supervised learning.

[0012] Optionally, the mapper is a convolutional layer.

[0013] Optionally, the memory storage unit is a sliding window queue with a fixed capacity, and when the number of stored items exceeds a preset capacity threshold, the earliest stored feature at the head of the queue is removed.

[0014] Furthermore, the preset capacity threshold is 500.

[0015] Preferably, the multi-stage encoding alignment loss is , , , In the formula, The encoding loss of the first segmentation network, The encoding loss of the second segmentation network, For pixel index, For the number of pixels, For cosine similarity, This is the first encoding fusion feature. This is the second encoding fusion feature. It represents the average characteristic after aggregation of memory storage units.

[0016] Preferably, the preprocessing includes: normalizing the size, normalizing the pixel value, and performing data enhancement on the acquired OCT fundus image to obtain the preprocessed OCT fundus image.

[0017] Preferably, the multi-scale encoder scales the preprocessed OCT fundus image to three sizes, inputs them into three ViT encoders respectively, extracts multi-scale features, and performs pixel-level addition and fusion. Each ViT encoder divides the scaled OCT fundus image into fixed-size image blocks and maps them to a high-dimensional feature space. It performs context modeling through a multi-layer Transformer coding structure to obtain the serialized feature representation at the corresponding scale. The serialized feature representation at each scale is rearranged to restore the spatial structure and aligned to the same spatial size using upsampling or interpolation methods. The features are then fused element by element to obtain multi-scale features.

[0018] Preferably, the wavelet decomposition encoder includes: The wavelet decomposition unit is used to decompose the preprocessed OCT fundus image into low-frequency components, horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components. The topology extraction unit is used to process the low-frequency component and the horizontal high-frequency component sequentially through a linear layer, a deep convolutional layer and a two-dimensional state space layer, and then add them together to obtain the topology features. The detail texture extraction unit is used to map the vertical high-frequency components and diagonal high-frequency components through a normalization layer, a convolutional layer, and a GELU activation function to obtain vertical features and diagonal features respectively; the vertical features and diagonal features are added at the pixel level to generate fused features; the fused features are used as the query vector, the vertical features are used as the key, and the diagonal features are used as the value, and the detail texture features are calculated through an attention mechanism. The fusion output unit is used to add the topological features and detailed texture features, and after pooling downsampling, input them into the next level of wavelet decomposition. After repeating the process for multiple levels, wavelet encoded features are obtained.

[0019] To address the anisotropy of the retinal layered structure, low-frequency (overall contour) and horizontal high-frequency (along the layer direction) signals are combined and processed through a two-dimensional state-space layer (state-space sequence modeling) to capture layered continuity and consistency, thus enhancing topological features. Vertical high-frequency and diagonal high-frequency (edge ​​abrupt change direction) signals are combined; local responses are first extracted through convolution, and then the fused features generated by these two signals are used to adaptively enhance detail texture through an attention mechanism. Through multi-level cascaded downsampling, topological and detail information is abstracted and fused layer by layer, significantly improving the ability to characterize weak boundaries and small lesions.

[0020] Preferably, the pre-trained lesion recognition module consists of a ViT encoder and an adapter, used to extract lesion type features from the pre-processed OCT fundus images and output the routing weights of each expert network; wherein, the pre-training process of the lesion recognition module is independent of the main network training, and the contrastive learning loss function InfoNCE is used to optimize the routing during pre-training.

[0021] An independent pre-trained lesion recognition module (ViT and adapter) employs contrastive learning (InfoNCE) to optimize routing, enabling adaptive allocation of routing weights based on the lesion type in the input image. Top-K expert selection reduces computational redundancy while preserving diversity; weighted summation enhances lesion-specific features. The filtering response can be dynamically adjusted for different lesion regions, improving the segmentation model's robustness to changes in lesion morphology.

[0022] More preferably, the route is optimized during pre-training using the contrastive learning loss function InfoNCE, including: Healthy fundus images, age-related macular degeneration fundus images, and diabetic macular edema fundus images were divided into three groups. In each training round, one group was randomly set as the positive sample group and the other two groups were set as the negative sample groups. Positive and negative samples were randomly selected and passed through an expert network and a lesion recognition module to obtain positive and negative sample feature vectors. The InfoNCE loss function was used for optimization.

[0023] More preferably, the InfoNCE loss function is calculated as follows: , In the formula, It is an exponential function. For cosine similarity, and These are the feature vectors of two positive samples. For the first The feature vector of each negative sample The total number of negative sample feature vectors. This is the temperature coefficient.

[0024] Optionally, the adapter employs a multilayer perceptron to assign corresponding weights to the expert module based on the features generated by the lesion recognition module; The expert network described above employs a feedforward neural network to adaptively extract discriminative features based on different lesion types.

[0025] More preferably, a Top-K filtering strategy is adopted to filter multiple expert networks based on the routing weights.

[0026] Preferably, the decoder includes multiple upsampling layers and corresponding residual layers. The upsampling layers expand the spatial resolution of the feature map by transposed convolution or interpolation. The residual layers include convolutional layers and batch normalization layers, and establish jump connections between the input and output. Each upsampling layer is followed by a residual layer for feature reconstruction and refinement. The output of the last upsampling layer is passed through a convolutional layer to obtain a pixel-level segmentation prediction map.

[0027] Preferably, the cross-supervision unit is used to supervise the second segmentation network by using the segmentation prediction map of the first segmentation network as a pseudo-label based on unlabeled data, and at the same time, to supervise the first segmentation network by using the segmentation prediction map of the second segmentation network as a pseudo-label, thus preventing the back propagation of pseudo-labels in both directions.

[0028] Preferably, the total loss calculation unit is used to perform a weighted summation of the supervised loss, the cross-pseudo-supervised loss, the multi-scale feature alignment loss, and the multi-stage encoding alignment loss to jointly train the fundus image segmentation system. The supervised loss is composed of the average of the Dice loss and cross-entropy loss of the two segmentation networks on the labeled data, and is expressed as: , , , In the formula, For the purpose of supervision loss, and These are the supervised losses for the first and second segmentation networks, respectively. For Dice's loss, For cross-entropy loss, and These are the segmentation results from the two segmentation networks, respectively. Indicates the actual value of the label; The cross-supervision loss is a combination of the Dice loss from mutual supervision of the two segmentation networks on unlabeled data and the cross-entropy loss, expressed as: , , , In the formula, For cross-monitoring loss, and The cross-pseudo-supervision losses of the first and second segmentation networks are respectively. and These are the segmentation results from the two segmentation networks, respectively. and These are pseudo-labels that prevent backpropagation; The multi-scale feature alignment loss is achieved by calculating the mean squared error loss pairwise for each multi-scale feature in the multi-scale encoder, and is expressed as follows: , In the formula, These are the multi-scale features in a multi-scale encoder. This is the mean squared error loss function.

[0029] The total loss function of this invention forces the output features of the three different scales in the multi-scale encoder to be semantically consistent pairwise, eliminating the representation gap between scales and improving the quality of multi-scale fusion. Through multi-stage alignment of the two networks, the encoded fusion features of the two branches converge to the historical average level, avoiding overfitting to a single sample and enhancing generalization ability. Joint optimization simultaneously achieves supervised training with labeled data, cross-distillation of unlabeled data, internal alignment within the multi-scale, and progressive alignment of the two networks, forming a complete semi-supervised training loop.

[0030] Among them, the joint optimization of supervised loss and cross pseudo-supervised loss clarifies the gradient blocking mechanism of pseudo-labels (preventing backpropagation), avoids the two networks from falling into self-consistency collapse, stabilizes the mutual supervision training process, and enables the two networks to learn from each other while maintaining independent decision boundaries.

[0031] Compared with the prior art, the beneficial effects of the present invention include at least the following: This invention addresses the challenges of segmenting the retinal layer and lesions in OCT retinal imaging by proposing a fundus image segmentation system based on multi-stage coding alignment and lesion perception, effectively segmenting the retinal layer and lesions. Based on preprocessed OCT fundus images, a multi-scale coding alignment module and a wavelet decomposition state-space attention module are used to extract multi-scale features, topological structure features, and texture detail features, respectively, which are then fused into a coded fusion feature. To ensure that the two segmentation networks extract effective segmentation features, the coded fusion feature obtained from the two encoders is input into the multi-stage coding alignment module. Step-by-step alignment of the multi-stage fusion feature avoids feature inaccuracies and semantic misalignments caused by one-time alignment. At the bottleneck layer, the coded fusion feature is input into the lesion perception expert module, thereby optimizing the extraction of features for different lesions. In the decoder, the spatial resolution of the feature map is gradually restored through multiple upsampling layers and a residual convolution module, ultimately outputting a pixel-level segmentation prediction map of the same size as the input image. Finally, cross-pseudo-supervision is used at the output end to utilize unlabeled data, improving the model's segmentation performance. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0033] Figure 1 This is an overall framework diagram of the fundus image segmentation system based on multi-stage coding alignment and lesion perception provided by the present invention.

[0034] Figure 2 This is a detailed structural diagram of the fundus image segmentation system based on multi-stage coding alignment and lesion perception provided by the present invention; wherein, Figure 2 Figure (A) shows a dual-branch encoder. Figure 2 Figure (B) shows the lesion perception expert module. Figure 2 Figure (C) in the diagram represents the multi-stage coding alignment module.

[0035] Figure 3 This is the visualization result of the retinal layer and lesion segmentation comparison experiment of the present invention. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and given in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0037] The inventive concept of this invention is as follows: Addressing the shortcomings of existing fundus image segmentation methods under limited labeled data conditions in depicting the layered topology, detailed textures, and lesion regions of the retina, and the problem that single feature alignment easily leads to feature inaccuracies and semantic misalignments, this invention develops a fundus image segmentation system based on multi-stage coding alignment and lesion perception. This system constructs a dual-branch encoder composed of a multi-scale ViT encoder and a wavelet decomposition encoder to extract multi-scale global semantic features. Wavelet decomposition separates low-frequency topological structures and high-frequency detailed textures, which are then modeled and fused separately, forming a multi-level coded fusion feature that complements both structure and detail. After adaptive weighting enhancement by a lesion perception expert module, the decoder progressively restores the spatial resolution to obtain the segmentation result. During the training phase, a dual-network semi-supervised mutual learning framework is adopted, introducing a multi-stage coding alignment module. A sliding window memory storage unit is used to average the coding features from historical stages as the alignment target. The fused features of the two networks are progressively aligned using cosine similarity, and multi-scale feature alignment loss and cross-pseudo-supervised loss are combined for joint optimization. This enhances the model's ability to depict the layered topology and lesion boundaries of the retina under limited labeled data.

[0038] like Figure 1 As shown in the embodiment, a fundus image segmentation system based on multi-stage encoding alignment and lesion perception is provided, including a first segmentation network and a second segmentation network. The first segmentation network and the second segmentation network have the same structure, and each independently includes: a data acquisition and processing module, a dual-branch encoder, a feature fusion module, a lesion perception expert module, a decoder, and a training module. Specifically, as follows: 1. Data acquisition and processing module, used to acquire OCT fundus images and perform preprocessing. Details are as follows: The preprocessing process includes: firstly, interpolating and scaling the acquired raw OCT fundus images to unify images of different resolutions to a preset size of 256×256, ensuring consistent input size for subsequent networks; then, normalizing the images to map pixel values ​​to a uniform numerical range [0,1], reducing numerical differences between different samples and improving the stability of model training; further, to enhance the model's adaptability to imaging differences, data augmentation processing is performed on the images, including random flipping and random adjustment of brightness and contrast. Specifically, brightness and contrast adjustment are achieved by linearly transforming the image pixel values ​​to simulate image distribution under different imaging conditions, thereby improving the model's generalization ability.

[0039] 2. A dual-branch encoder receives preprocessed OCT fundus images and extracts multi-scale features using a multi-scale encoder, then performs pixel-level fusion. Simultaneously, a wavelet decomposition encoder extracts topological and detail texture features from the preprocessed fundus image data and fuses them pixel-level. Details are as follows: like Figure 2 As shown, the multi-scale encoder is used to extract multi-scale feature information from an image. The preprocessed image is scaled to multiple sizes and input into different ViT encoders to obtain multi-scale encoded features. The specific process is as follows: like Figure 2 As shown in (A), the preprocessed image input is scaled to resolutions of 64, 128, and 256. Three ViT encoders are used to extract image coding features at three different scales. Each ViT encoder divides the scaled OCT fundus image into fixed-size image blocks and maps them to a high-dimensional feature space. Context modeling is performed through a multi-layer Transformer coding structure to obtain the serialized feature representation at the corresponding scale. The serialized feature representation at each scale is rearranged to restore the spatial structure and aligned to the same spatial size using upsampling or interpolation methods. Finally, the image coding features at different scales are pixel-level added and fused to obtain multi-scale features.

[0040] like Figure 2 As shown in (A), the wavelet decomposition method is used to decompose the preprocessed OCT fundus image into low-frequency components, horizontal high-frequency components, vertical high-frequency components and diagonal high-frequency components in the wavelet decomposition unit. Among them, the low-frequency components can effectively preserve the overall layered structure of the retina and its topological relationship information, while the horizontal high-frequency components have a significant response in the horizontal direction and are consistent with the extension direction of the retinal layered structure, which is conducive to enhancing the ability to extract the interlayer continuity and boundary orientation. The topology extraction unit sequentially passes the low-frequency components and horizontal high-frequency components through a linear layer, a deep convolutional layer, and a two-dimensional state space layer (SS2D layer). It performs feature combination processing on the low-frequency and horizontal high-frequency components, and unifies and re-encodes the channel dimensions through a linear mapping layer. A deep convolutional layer is introduced to enhance the modeling ability of topologically related features. Based on this, the processed features are input into the SS2D layer for modeling. The state space mechanism is used to model the long-range dependencies of the spatial sequence, capturing the continuity and consistency features of the retinal layer structure, thereby obtaining a topological feature representation. The vertical and diagonal high-frequency components effectively characterize the edge changes of retinal layer boundaries and lesion areas, exhibiting a strong response to fine-grained textures and structural abrupt changes, which is beneficial for improving the model's ability to extract detailed features. The detail texture extraction unit processes the vertical and diagonal high-frequency components through a normalization layer, a 3×3 convolutional layer, and a GELU activation function to obtain vertical and diagonal features, respectively. The normalization layer standardizes the feature distribution to improve training stability and accelerate convergence. Subsequently, a 3×3 convolutional kernel is used to model the local receptive fields of the two types of features, extracting vertical and diagonal features. The two features are then fused pixel-level to obtain a query vector. The vertical and diagonal features are used as keys and values, respectively, and input into the attention module, thereby achieving adaptive modeling and enhancement of detail information to obtain detail texture features. The fusion output unit is used to perform pixel-level addition and fusion of topological structure features and detail texture features, and then inputs them into the next level of wavelet decomposition after pooling downsampling. Through a multi-level cascaded wavelet decomposition and feature modeling process, the final encoded fused feature is output in the fourth level wavelet decomposition module. This provides richer and more discriminative feature representations for subsequent segmentation tasks.

[0041] This module addresses the anisotropy of the retinal layered structure by combining low-frequency (overall contour) and horizontal high-frequency (along the layer direction) signals. A two-dimensional state-space layer (state-space sequence modeling) captures the layered continuity and consistency, enhancing topological features. Vertical high-frequency signals are combined with diagonal high-frequency signals (edge ​​abrupt changes). Local responses are first extracted through convolution, and then the fused features generated by these two methods are used to adaptively enhance detailed textures through an attention mechanism. Multi-level cascaded downsampling allows for the layer-by-layer abstraction and fusion of topological and detailed information, significantly improving the ability to characterize weak boundaries and minute lesions.

[0042] 3. Feature fusion module, used to fuse the fused multi-scale features and the fused wavelet-coded features to obtain coded fused features.

[0043] 4. The lesion perception expert module receives the encoded fusion features, generates routing weights through a pre-trained lesion recognition module, and then uses the expert network to adaptively weight the encoded fusion features based on these routing weights to obtain enhanced encoded features. Details are as follows: like Figure 2 As shown in (B), the lesion perception expert module introduces a routing mechanism to adaptively allocate input features, guiding images of different lesion types to the corresponding expert network, thereby achieving targeted feature extraction and dynamic modeling to obtain encoded features.

[0044] Specifically, it involves encoding and fusing features. The pre-trained lesion recognition module is input, which obtains lesion features from the pre-trained ViT encoder and inputs them into the adapter to obtain corresponding expert weights, and then the encoded features are fused. Input the top-K expert networks and weight the expert features according to their own weights to obtain the encoded features.

[0045] The pre-training process of the lesion recognition module is independent of the training process of the main network. Pre-training is completed first, and the parameters are kept fixed or undergo limited fine-tuning during the overall model training phase. First, the pre-processed images are divided into three groups: healthy images, age-related macular degeneration (AMD) fundus images, and diabetic macular edema (DME) fundus images. Two images from each group are randomly selected as positive samples, and six images from the other two groups are selected as negative samples. These are then input into a feedforward neural network, which consists of a three-level feedforward module and a compression activation module. The feedforward neural module includes a batch normalization module and a module of size [missing information]. It consists of a convolutional kernel, a ReLU activation module, and an average pooling module.

[0046] Input two positive samples and six negative samples to obtain the positive samples. and negative samples .

[0047] Specifically, the adapter consists of two fully connected layers, with a ReLU activation function added after each fully connected layer. The encoded fused features are input into the lesion-sensing routing to obtain positive samples. and negative samples Each vector dimension is Then, the InfoNCE loss function is used to optimize the routing, as shown in the following formula: , In the formula, It is an exponential function. For cosine similarity, and These are the feature vectors of two positive samples. For the first The feature vector of each negative sample The total number of negative sample feature vectors. This is the temperature coefficient.

[0048] 4. Decoder: Receives the enhanced encoded features, restores the spatial resolution step by step, and finally outputs a pixel-level segmentation prediction map to complete the fundus image segmentation.

[0049] The encoded features are first upsampled to gradually restore spatial resolution, and then input into the residual module for feature reconstruction and refinement, resulting in the first-level decoded features. Subsequently, these decoded features are passed through subsequent upsampling layers and the corresponding residual modules, continuously optimizing feature representation while restoring spatial information layer by layer. Finally, the last level outputs decoded features with the same resolution as the original image, thus reconstructing both structural and detail information. Finally, a 1×1 convolutional layer is used to obtain the segmentation result.

[0050] 5. Training module, which includes a multi-stage encoding alignment module, a cross-supervision unit, and a total loss calculation unit, is used to jointly train the first segmentation network and the second segmentation network using labeled and unlabeled data.

[0051] The total loss calculation unit calculates the cross-supervised loss and supervised loss at the output, and simultaneously introduces multi-scale feature alignment loss into the multi-scale encoder to fuse the encoded features of the two segmentation networks. Multi-stage encoding alignment loss is used to jointly optimize various losses, thereby improving model performance.

[0052] Specifically, for tagged output and Using conventional supervised training, the supervised loss is a combination of the average Dice loss and cross-entropy loss of the two segmentation networks on the labeled data, expressed as: , , , In the formula, For the purpose of supervision loss, and These are the supervised losses for the first and second segmentation networks, respectively. For Dice's loss, For cross-entropy loss, and These are the segmentation results from the two segmentation networks, respectively. Indicates the actual value of the label; The cross-supervision unit employs a dual-segmentation network cross-supervision mechanism for unlabeled data. This allows the two networks to use each other's predictions as supervision signals for collaborative training, thereby improving the model's learning ability and prediction consistency on unlabeled samples. The cross-supervision loss is a combination of the Dice loss from mutual supervision between the two segmentation networks on unlabeled data and the cross-entropy loss, expressed as: , , , In the formula, For cross-monitoring loss, and The cross-pseudo-supervision losses of the first and second segmentation networks are respectively. and These are the segmentation results from the two segmentation networks, respectively. and These are pseudo-labels used to prevent backpropagation.

[0053] In multi-scale encoders, features extracted at different scales differ in spatial resolution and semantic representation. Therefore, this invention uses feature alignment to eliminate these differences by mapping features from different ViT encoders.

[0054] Specifically, the obtained feature map is first deformed to one dimension. Then, each scale feature is adjusted to a consistent channel dimension using its respective mapper. Feature alignment is then used to eliminate semantic differences. The mapper is a multilayer perceptron. The multi-scale feature alignment loss is achieved by calculating the mean squared error loss pairwise for each multi-scale feature in the multi-scale encoder, and is expressed as follows: , In the formula, These are the multi-scale features in a multi-scale encoder. This is the mean squared error loss function.

[0055] To optimize the encoder's extraction of multi-scale features of images, as well as the topological structure and detailed texture features of the retina, this study introduces a multi-stage coding feature alignment module.

[0056] like Figure 2 As shown in (C), specifically, the encoded fusion features obtained from the two segmentation networks are... Channel-level feature fusion is performed, and the resulting mapped feature, which is a convolutional layer of size 1, is mapped to the original feature and has the same dimension as the encoded fused feature. This mapped feature can represent the features that need to be aligned, and the feature differences are reduced by fusing the original features, thus reducing the difficulty of alignment. This mapped feature map is then placed at the end of a memory storage unit of size 500, constructing a temporal feature map structure containing current features and features from previous stages. When the storage capacity of the memory storage unit reaches its limit, the topmost layer, i.e., the earliest added mapped feature, is removed. This method enables the storage of feature maps across multiple stages. Finally, the average value of all stage feature maps in the entire memory storage unit is calculated and used as the alignment target for the two encoded fusion feature maps, achieving multi-stage feature alignment. Specifically, the update mechanism of the memory storage unit includes: storing the current mapping feature map at the end of the storage unit; when the number of feature maps in the storage unit exceeds a preset capacity threshold of 500, simultaneously removing the earliest stored feature map from the head of the queue, ensuring that the storage unit always contains feature information from the most recent 500 stages. Each storage location corresponds to the encoded fusion feature map of a historical stage, and the storage unit as a whole constitutes a sliding window-like multi-stage feature sequence. In the feature alignment process, the average value of all stage feature maps within the memory storage unit is used as the alignment target and fused with the current mapped feature map to obtain the aligned enhanced features, thus achieving multi-stage feature alignment. The formula is as follows: , , , In the formula, The encoding loss of the first segmentation network, The encoding loss of the second segmentation network, For pixel index, For the number of pixels, For cosine similarity, This is the first encoding fusion feature. This is the second encoding fusion feature. It represents the average characteristic after aggregation of memory storage units.

[0057] Adding the above loss functions together yields the final total loss function, as shown in the formula: .

[0058] To further verify the feasibility and effectiveness of the fundus image segmentation system based on multi-stage coding alignment and lesion perception proposed in this invention, comparison results between this invention (Ours) and other comparative models are presented. For example... Figure 3 As shown in the figure, this figure presents a visualization comparison of the OCT images of the retinal layer and lesion segmentation of the proposed model and the comparison model. According to the visualization results, the model of the present invention has better recognition results for the layer and lesion area, and can still achieve a relatively accurate segmentation result for small lesions.

[0059] Table 1 shows the quantitative results of the comparative experiment. The metrics used were the Dice coefficient and the HD95 (95% Hausdorff distance) value.

[0060] The Dice coefficient is used to predict the overall spatial overlap between the segment and the ground truth annotation, and its calculation formula is as follows: , In the formula, Represents the predicted value. Represents the actual value.

[0061] The maximum deviation between the predicted boundary and the actual boundary of the HD95 value is calculated using the following formula: , , , , In the formula, Represents the set of predicted values. Represents the set of true values. This represents the 95th percentile.

[0062] Table 1. Quantitative results of the comparative experiment on retinal layer and lesion segmentation.

[0063] According to the data in Table 1, our invention outperforms all the comparison methods in overall performance, with an average Dice coefficient of 0.902 and an HD95 reduced to 4.352, indicating that the model has significant advantages in both segmentation accuracy and boundary localization.

[0064] In terms of various metrics, Ours achieved high accuracy in most retinal layers, such as the ganglion cell layer and inner plexiform layer (GCL-IPL), outer nuclear layer (ONL), and pigment epithelium (RPE). Its performance was particularly outstanding in lesion areas (Drusen, 0.903) and fluid accumulation areas (Fluid, 0.938). Other layers, such as the retinal nerve fiber layer (RNFL), inner nuclear layer (INL), outer plexiform layer (OPL), outer peripheral membrane (ELM), inner photoreceptor segment (IS), and outer photoreceptor segment (OS), also showed excellent segmentation results, indicating that this method can effectively improve the segmentation ability of complex lesion areas. Furthermore, compared to the higher HD95 values ​​of other methods, the model proposed in this invention significantly reduced boundary errors, further validating its superiority in structural continuity and detail extraction.

[0065] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A fundus image segmentation system based on multi-stage coding alignment and lesion perception, comprising a first segmentation network and a second segmentation network, wherein the first segmentation network and the second segmentation network have the same structure, characterized in that, The first and second segmentation networks independently include: The data acquisition and processing module is used to acquire OCT fundus images and perform preprocessing. A dual-branch encoder is used to receive preprocessed OCT fundus images. Multi-scale features are extracted by a multi-scale encoder and then fused at the pixel level to obtain multi-scale features. At the same time, a wavelet decomposition encoder is used to extract topological features and detail texture features from the preprocessed fundus image data and then fused at the pixel level to obtain wavelet encoded features. The feature fusion module is used to fuse multi-scale features and wavelet-coded features to obtain coded fusion features; The lesion perception expert module is used to receive encoded fusion features, generate routing weights through a pre-trained lesion recognition module, and then use the expert network to adaptively weight the encoded fusion features based on the routing weights to obtain enhanced encoded features. The decoder receives the enhanced encoded features, restores the spatial resolution step by step, and finally outputs a pixel-level segmentation prediction map to complete the fundus image segmentation. The training module includes a multi-stage encoding alignment module, a cross-supervision unit, and a total loss calculation unit, used to jointly train the first segmentation network and the second segmentation network using labeled and unlabeled data. The multi-stage encoding alignment module receives the first and second encoding fusion features, performs channel-level fusion of the first and second encoding fusion features, generates mapping features through a mapper, and stores them in a sliding window-type memory storage unit. It calculates the average value of all historical mapping features in the memory storage unit as the alignment target, aligns the first and second encoding fusion features to the alignment target using cosine similarity, and outputs the aligned features. Specifically, the feature fusion module of the first segmentation network outputs the first encoding fusion feature, and the feature fusion module of the second segmentation network outputs the second encoding fusion feature.

2. The fundus image segmentation system based on multi-stage coding alignment and lesion perception according to claim 1, characterized in that, Multi-stage encoding alignment loss is , , In the formula, The encoding loss of the first segmentation network, The encoding loss of the second segmentation network, For pixel index, For the number of pixels, For cosine similarity, This is the first encoding fusion feature. This is the second encoding fusion feature. It represents the average characteristic after aggregation of memory storage units.

3. The fundus image segmentation system based on multi-stage coding alignment and lesion perception according to claim 1, characterized in that, The multi-scale encoder scales the preprocessed OCT fundus image to three sizes, inputs them into three ViT encoders respectively, extracts multi-scale features, and performs pixel-level addition and fusion. Each ViT encoder divides the scaled OCT fundus image into fixed-size image blocks and maps them to a high-dimensional feature space. It performs context modeling through a multi-layer Transformer coding structure to obtain the serialized feature representation at the corresponding scale. The serialized feature representation at each scale is rearranged to restore the spatial structure and aligned to the same spatial size using upsampling or interpolation methods. The features are then fused element by element to obtain multi-scale features.

4. The fundus image segmentation system based on multi-stage coding alignment and lesion perception according to claim 1, characterized in that, The wavelet decomposition encoder includes: The wavelet decomposition unit is used to decompose the preprocessed OCT fundus image into low-frequency components, horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components. The topology extraction unit is used to process the low-frequency component and the horizontal high-frequency component sequentially through a linear layer, a deep convolutional layer and a two-dimensional state space layer, and then add them together to obtain the topology features. The detail texture extraction unit is used to map the vertical high-frequency components and diagonal high-frequency components through a normalization layer, a convolutional layer, and a GELU activation function to obtain vertical features and diagonal features respectively; the vertical features and diagonal features are added at the pixel level to generate fused features; the fused features are used as the query vector, the vertical features are used as the key, and the diagonal features are used as the value, and the detail texture features are calculated through an attention mechanism. The fusion output unit is used to add the topological features and detailed texture features, and after pooling downsampling, input them into the next level of wavelet decomposition. After repeating the process for multiple levels, wavelet encoded features are obtained.

5. The fundus image segmentation system based on multi-stage coding alignment and lesion perception according to claim 1, characterized in that, The pre-trained lesion recognition module, consisting of a ViT encoder and an adapter, is used to extract lesion type features from preprocessed OCT fundus images and output the routing weights of each expert network. The pre-training process of the lesion recognition module is independent of the main network training. During pre-training, the contrastive learning loss function InfoNCE is used to optimize the routing.

6. The fundus image segmentation system based on multi-stage coding alignment and lesion perception according to claim 1, characterized in that, The decoder comprises multiple upsampling layers and corresponding residual layers. The upsampling layers expand the spatial resolution of the feature maps by using transposed convolution or interpolation. The residual layers include convolutional layers and batch normalization layers, and establish jump connections between the input and output. Each upsampling layer is followed by a residual layer for feature reconstruction and refinement. The output of the last upsampling layer is passed through a convolutional layer to obtain a pixel-level segmentation prediction map.

7. The fundus image segmentation system based on multi-stage coding alignment and lesion perception according to claim 1, characterized in that, The cross-supervision unit is used to supervise the second segmentation network by using the segmentation prediction map of the first segmentation network as a pseudo-label based on unlabeled data, and at the same time, to supervise the first segmentation network by using the segmentation prediction map of the second segmentation network as a pseudo-label. The back propagation of pseudo-labels is prevented in both directions.

8. The fundus image segmentation system based on multi-stage coding alignment and lesion perception according to claim 1, characterized in that, The total loss calculation unit is used to perform a weighted summation of supervised loss, cross-pseudo-supervised loss, multi-scale feature alignment loss, and multi-stage encoding alignment loss to jointly train the fundus image segmentation system. The supervised loss is a combination of the average of the Dice loss and cross-entropy loss of the two segmentation networks on the labeled data; The cross pseudo-supervision loss is composed of the Dice loss of mutual supervision between the two segmentation networks on unlabeled data and the cross-entropy loss. The multi-scale feature alignment loss is achieved by calculating the mean squared error loss of each pair of multi-scale features in the multi-scale encoder.

Citation Information

Patent Citations

  • 3D iris surface reconstruction and quantification method and segmentation network based on AS-OCT image

    CN112750141A

  • Semi-supervised feature fusion segmentation network based on OCT-OCTA multi-mode

    CN121305066A

  • Medical image segmentation method based on wavelet enhancement and multi-scale feature fusion

    CN120563825A

  • Pancreatic postoperative diabetes prediction system based on supervised deep subspace learning

    US20240395408A1