Method and apparatus for segmenting cerebral choroid plexus
By introducing operations such as feature alignment, adaptation processing, and feature fusion into the encoder structure, combined with the decoder structure, the problem of low segmentation accuracy of the choroid plexus in the brain in the prior art is solved, and accurate segmentation of the choroid plexus region in brain images is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANTOU UNIV
- Filing Date
- 2025-05-12
- Publication Date
- 2026-04-14
AI Technical Summary
Existing fully convolutional neural networks lack the ability to extract features in brain vascular bundle segmentation, resulting in low segmentation accuracy and difficulty in effectively capturing vascular bundle features of small targets and complex structures.
A segmentation model combining encoder and decoder structures is adopted. Through feature alignment, adaptation, feature extraction and feature fusion operations, combined with a large model and convolutional neural network, the feature extraction capability is improved.
It achieves precise segmentation of the choroid plexus region in brain images, improves segmentation accuracy, and can comprehensively capture complex and subtle choroid plexus features.
Smart Images

Figure CN120672659B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical image processing technology, and in particular to a method and apparatus for segmenting the choroid plexus of the brain. Background Technology
[0002] The choroid plexus is a structure within the ventricles of the brain, composed of the pia mater, vascular plexus, and ependymal epithelium. It is responsible for producing cerebrospinal fluid. Its morphological characteristic is the repeated branching of blood vessels into a plexus-like structure that protrudes into the ventricular cavity. In related technologies, existing fully convolutional neural networks are used to process brain images and segment them to obtain the choroid plexus outline. However, the choroid plexus is characterized by its small size and complex structure, while existing fully convolutional neural networks often use convolutional neural networks as their encoders, which are insufficient in capturing the features of the choroid plexus, resulting in low segmentation accuracy for the brain's choroid plexus. Summary of the Invention
[0003] This application provides a method and apparatus for segmenting the choroid plexus of the brain, which improves the segmentation accuracy of the choroid plexus of the brain.
[0004] On one hand, embodiments of this application provide a method for segmenting the choroid plexus of the brain, including the following steps:
[0005] Acquire brain images;
[0006] The brain image is input into a pre-trained segmentation model to obtain a choroid plexus contour map;
[0007] The segmentation model includes:
[0008] The encoder structure comprises multiple encoders arranged sequentially. The first encoder is used to perform feature alignment and feature extraction on the brain image to obtain the output of the first encoder. The other encoders besides the first encoder are used to perform adaptation processing, feature extraction, and feature fusion on the inputs of the other encoders to obtain the outputs of the other encoders.
[0009] A decoder structure is connected in a skip connection to the encoder structure. The decoder structure is provided with multiple decoders in sequence. Each decoder is used to decode the input of the decoder to obtain the output of the decoder.
[0010] The output layer is used to obtain the ventricle contour map based on the output of the last decoder.
[0011] On the other hand, embodiments of this application provide a brain choroid plexus segmentation device, including:
[0012] The acquisition module is used to acquire brain images;
[0013] The segmentation module is equipped with a pre-trained segmentation model. The segmentation module is used to input the brain image into the segmentation model to obtain a choroid plexus contour map.
[0014] The segmentation model includes:
[0015] The encoder structure comprises multiple encoders arranged sequentially. The first encoder is used to perform feature alignment and feature extraction on the brain image to obtain the output of the first encoder. The other encoders besides the first encoder are used to perform adaptation processing, feature extraction, and feature fusion on the inputs of the other encoders to obtain the outputs of the other encoders.
[0016] A decoder structure is connected in a skip connection to the encoder structure. The decoder structure is provided with multiple decoders in sequence. Each decoder is used to decode the input of the decoder to obtain the output of the decoder.
[0017] The output layer is used to obtain the ventricle contour map based on the output of the last decoder.
[0018] According to the brain choroid plexus segmentation method and apparatus provided in this application, a brain image is acquired and input into a pre-trained segmentation model to obtain a choroid plexus contour map. The segmentation model includes an encoder structure, a decoder structure, and an output layer. The encoder structure has multiple encoders arranged sequentially. The first encoder performs feature alignment and feature extraction on the brain image to obtain its output. Other encoders perform adaptation processing, feature extraction, and feature fusion on their inputs to obtain their outputs. The decoder structure is skip-connected to the encoder structure and has multiple decoders arranged sequentially. Each decoder decodes its input to obtain its output. The output layer obtains the choroid plexus contour map based on the output of the last decoder. According to the technical solution of this application, by introducing feature alignment, adaptation processing, feature extraction, and feature fusion operations into the encoder structure and combining them with the decoding operations of the decoder structure, segmentation processing of the choroid plexus region in the brain image is achieved. This can accurately capture small targets and complex choroid plexus features in the brain image, improve the feature extraction capability of the segmentation model, and thus effectively improve the segmentation accuracy of the brain choroid plexus.
[0019] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description
[0020] Figure 1This is a flowchart of a brain choroid plexus segmentation method provided in this application;
[0021] Figure 2 This is a structural diagram of the segmentation model provided in this application;
[0022] Figure 3 This is a schematic diagram of the feature alignment layer provided in this application;
[0023] Figure 4 This is a schematic diagram of the adaptive feature extraction module provided in this application;
[0024] Figure 5 This is a schematic diagram of the frequency convolution operation provided in this application;
[0025] Figure 6 This is a schematic diagram of the wavelet transform provided in this application;
[0026] Figure 7 This is a schematic diagram of the feature fusion layer provided in this application;
[0027] Figure 8 This is a structural diagram of the decoder provided in this application. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0029] The present application will be further described below with reference to the accompanying drawings and specific embodiments. The described embodiments should not be considered as limitations on the present application, and all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the present application.
[0030] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0032] It should be noted that in all specific embodiments of this disclosure, when processing based on data such as brain images is required, the permission or consent of the subject is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. For example, when an embodiment of this disclosure needs to obtain data such as brain images, separate permission or consent from the subject is obtained through pop-ups or redirection to a confirmation page. Only after obtaining the subject's separate permission or consent is the necessary data, such as brain images, required for the proper functioning of the embodiments of this disclosure be obtained.
[0033] The choroid plexus is a structure within the ventricles of the brain, composed of the pia mater, vascular plexus, and ependymal epithelium. It is responsible for producing cerebrospinal fluid. Its morphological characteristic is the repeated branching of blood vessels into a plexus-like structure that protrudes into the ventricular cavity. In related technologies, existing fully convolutional neural networks are used to process brain images and segment them to obtain the choroid plexus contour map. Taking U-Net as an example, U-Net employs an encoder-decoder structure. The encoding part extracts features and reduces resolution through convolutional and pooling layers, while the decoding part gradually recovers the spatial information of the image through upsampling or deconvolution, thereby segmenting the choroid plexus contour map from the brain image.
[0034] However, the choroid plexus is characterized by its small size and complex structure. Furthermore, brain images are typically obtained through magnetic resonance imaging (MRI) or computed tomography (CT), which are often noisy and have low contrast. This results in indistinct contrast between the choroid plexus boundaries and surrounding tissues, significantly increasing the difficulty of choroid plexus segmentation. Existing fully convolutional neural networks often use convolutional neural networks as their encoders, which have insufficient feature extraction capabilities. They struggle to effectively capture the features of small, complex choroid plexuses, leading to issues such as blurred boundaries, local omissions, and loss of high-frequency details and edge information. Consequently, related technologies achieve low segmentation accuracy for the brain's choroid plexus.
[0035] In view of this, embodiments of this application provide a method and apparatus for segmenting the choroid plexus of the brain, which aims to effectively improve the segmentation accuracy of the choroid plexus of the brain.
[0036] First, the following will describe in detail, with reference to the accompanying drawings, a brain choroid plexus segmentation method provided by the embodiments of this application.
[0037] This application provides a brain choroid plexus segmentation method, which can be applied to terminals, servers, or software running on either terminal or server. Terminals can be tablets, laptops, desktop computers, etc., but are not limited to these. Servers can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Furthermore, a server can be a node server in a blockchain network, but is not limited to these. Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms.
[0038] Reference Figure 1 and Figure 2 The brain choroid plexus segmentation method may include the following steps S101-S102:
[0039] S101, acquire brain images;
[0040] S102, the brain image is input into a pre-trained segmentation model to obtain a venous plexus contour map; wherein, the segmentation model includes: an encoder structure, a decoder structure, and an output layer; the encoder structure has multiple encoders arranged sequentially; the first encoder is used to perform feature alignment and feature extraction on the brain image to obtain the output of the first encoder; the other encoders besides the first encoder are used to perform adaptation processing, feature extraction, and feature fusion on the inputs of other encoders to obtain the outputs of other encoders; the decoder structure is skip-connected to the encoder structure, and the decoder structure has multiple decoders arranged sequentially, each decoder is used to decode the input of the decoder to obtain the output of the decoder; the output layer is used to obtain the venous plexus contour map based on the output of the last decoder.
[0041] In this embodiment, firstly, brain images are acquired through a preset database, which pre-stores multiple brain images. These brain images refer to magnetic resonance imaging (MRI) images or computed tomography (CT) images associated with the brain. Then, the brain images are input into a pre-trained segmentation model. This segmentation model is a neural network model trained using multiple preset brain image samples and the corresponding label information for each sample. The label information refers to the choroidal plexus contour map in the brain image sample. The segmentation model can segment the choroidal plexus contour map from the brain image; the choroidal plexus contour map is a binary mask image.
[0042] Specifically, the segmentation model can include an encoder structure, a decoder structure, and an output layer. The encoder and decoder structures are connected in a skip connection, with multiple encoders sequentially arranged in the encoder structure and multiple decoders sequentially arranged in the decoder structure. In the encoder structure, the input of the first encoder is a brain image. The first encoder performs feature alignment and feature extraction on its input to obtain the output of the first encoder. Here, the feature alignment operation can fully explore the details of local choroidal plexus and global choroidal plexus structure information at different directions and scales in the brain image, and ensure the feature format. The feature extraction operation can initially capture choroidal plexus regions with complex structures and low contrast. The inputs of the other encoders besides the first encoder are the outputs of the previous encoder. The other encoders perform adaptation processing, feature extraction, and feature fusion on their inputs to obtain the outputs of the other encoders. Here, the adaptation operation is introduced to adapt the feature dimensions of the input feature map, effectively reducing computation while comprehensively capturing and enhancing key features associated with choroidal plexuses. Feature extraction further locates structurally complex, low-contrast choroidal plexus regions, thus comprehensively capturing complex and subtle choroidal plexus features. Feature fusion further integrates multi-dimensional choroidal plexus features and passes them to the corresponding decoder, improving decoding performance. In the decoder structure, the input is decoded to obtain the decoder output. Here, the feature maps output by each encoder are high-level features with lower spatial resolution but higher semantic level, while the feature maps output by each decoder are low-level features with higher spatial resolution and more local choroidal plexus details. Decoding effectively integrates high-level and low-level features, extracting more valuable choroidal plexus features and reconstructing choroidal plexus regions in brain images. In the output layer, the output of the last decoder is used as a reference to generate the final choroidal plexus contour map, which is a binary mask image.
[0043] In summary, the embodiments of this application introduce operations such as feature alignment, adaptation processing, feature extraction, and feature fusion into the encoder structure, and combine them with the decoding operations of the decoder structure to achieve segmentation processing of the choroid plexus region in brain images. This can accurately capture small targets and complex choroid plexus features in brain images, improve the feature extraction capability of the segmentation model, and thus effectively improve the segmentation accuracy of the choroid plexus in the brain.
[0044] The specific implementation of the above segmentation model will be further explained below.
[0045] With the success of large models such as the Segment Anything Model (SAM) in multi-scene, multi-object segmentation tasks, their powerful generalization ability has attracted widespread attention. A few related techniques have attempted to introduce large models into the segmentation task of the brain's choroid plexus. However, large models often require pre-training on large-scale general datasets, while prior knowledge of specific medical scenarios (such as the brain's choroid plexus) is difficult to obtain, leading to poor performance of large models in choroid plexus segmentation tasks. Furthermore, few techniques integrate large models with convolutional neural networks into the same segmentation model.
[0046] To address the shortcomings of existing technologies, such as insufficient feature extraction capabilities and the difficulty in integrating large models and convolutional neural networks into a single segmentation model, this application proposes a dual-branch fusion segmentation model that combines a large model with a convolutional neural network, specifically addressing the challenges of small volume and complex structure in choroidal plexus regions of brain images. The encoder of this model consists of a large model branch and a convolutional network branch. The large model branch is configured with a pre-trained large model, which is adaptively modified through feature alignment and adaptation operations to fully utilize the extensive feature representation capabilities of the large model and enhance the feature extraction capabilities of the segmentation model. The type of large model can be flexibly set according to actual conditions; for example, it can be a SAM2 model, but is not limited to this. The convolutional network branch is implemented based on a convolutional neural network structure and is configured with traditional convolution operations and frequency convolution operations based on wavelet transform, which can further enhance the extraction of details of choroidal plexus regions and their edges in small targets in brain images.
[0047] Simply put, refer to Figure 2 ,exist Figure 2In this code, "feature alignment" refers to the feature alignment layer of the first encoder, "convolutional module" refers to the traditional convolutional layer of the first encoder, and "frequency convolution" connected to the "convolutional module" refers to the first frequency convolutional layer of the first encoder. "Feature alignment," "convolutional module," and "frequency convolution" together constitute the first encoder. "SAM2 adjustment module" refers to the adaptation feature extraction module in other encoders besides the first encoder. "Frequency convolution," running parallel to the "SAM2 adjustment module," is the second frequency convolutional layer in other encoders besides the first encoder. "Feature fusion," connecting the "SAM2 adjustment module" and the "frequency convolution," refers to the feature fusion layer in other encoders besides the first encoder. A single "SAM2 adjustment module," a single "frequency convolution," and a single "feature fusion" together constitute another encoder. A brain image with an original size of 1×256×256 is first simultaneously input into two parallel branches: the large model branch and the convolutional network branch. The convolutional network branch corresponds to the traditional convolutional layer and the first-frequency convolutional layer of the first encoder, and the second-frequency convolutional layer of the other encoders. The large model branch corresponds to the feature alignment layer of the first encoder, and the adaptation feature extraction module of the other encoders. Both the large model branch and the convolutional network branch output feature maps of different sizes. The output scales of the large model branch are 144×64×64, 288×32×32, 576×16×16, and 1152×8×8, respectively, while the output scales of the convolutional network branch are 36×256×256, 72×128×128, 144×64×64, 288×32×32, 576×16×16, and 1152×8×8, respectively. In the encoder, output layers with feature maps of the same scale have a fusion layer that merges the feature maps from the two branches into a 64-channel feature map, which is then used as the encoder output. Subsequently, starting with the smallest feature map, the system gradually upsamples the image using the decoder, eventually obtaining a choroid plexus contour map of the same size as the brain image.
[0048] (I) Encoder structure.
[0049] In some implementations, refer to Figure 2 The output of the first encoder may include the first output, second output, and third output of the first encoder; in the aforementioned first encoder, feature alignment and feature extraction are performed on the brain image to obtain the output of the first encoder, which may include:
[0050] Feature alignment is performed on the brain image to obtain the first output of the first encoder;
[0051] Performing a traditional convolution operation on the brain image yields the second output of the first encoder;
[0052] The third output of the first encoder is obtained by performing a frequency convolution operation on the second output of the first encoder.
[0053] In this embodiment, the first encoder may include a feature alignment layer, a conventional convolutional layer, and a first frequency convolutional layer. In the first encoder, firstly, the brain image is feature-aligned using the feature alignment layer to obtain the first output of the first encoder. This fully extracts local choroidal bundle details and global choroidal bundle structural information at different directions and scales in the brain image, while ensuring feature format. Simultaneously, the brain image is subjected to conventional convolution operations using the conventional convolutional layer to obtain the second output of the first encoder. This initially enhances the capture of choroidal bundle edges and textures, obtaining feature information associated with these edges and textures. The convolutional kernel of the conventional convolutional layer can be flexibly set according to actual conditions; for example, the kernel can be 3×3, but it is not limited to this. Subsequently, the second output of the first encoder is subjected to frequency convolution operations using the first frequency convolutional layer to obtain the third output of the first encoder. This initially captures choroidal bundle regions with complex structures and low contrast. Among them, the feature alignment layer belongs to the large model branch, that is, the first output of the first encoder belongs to the large model branch, while the traditional convolutional layer and the first frequency convolutional layer belong to the convolutional network branch, that is, the second and third outputs of the first encoder belong to the convolutional network branch.
[0054] Here, compared to conventional encoders that rely solely on convolutional neural networks in related technologies, this implementation introduces feature alignment and feature extraction operations in the first encoder. The feature extraction operation can be divided into traditional convolution and frequency convolution operations. Feature alignment fully extracts details of local choroidal plexus at different directions and scales, as well as global choroidal plexus structure information in brain images. Considering that the feature alignment layer directly connects to the subsequent construction of large model branches, feature alignment ensures that the feature format is adapted to the subsequent construction of large model branches. Traditional convolution initially enhances the capture of choroidal plexus edges and textures, obtaining feature information associated with these edges and textures. Subsequently, traditional convolution is followed by frequency convolution, which, based on the feature information associated with choroidal plexus edges and textures, locates structurally complex, low-contrast choroidal plexus regions. Thus, shallow feature extraction is achieved, enabling the initial extraction of valuable and diverse choroidal plexus feature information, thereby improving the feature extraction capability of the segmentation model.
[0055] In some implementations, refer to Figure 3 In the aforementioned feature alignment layer, feature alignment is performed on the brain image to obtain the first output of the first encoder, which may include:
[0056] Multi-scale convolution operations are performed on brain images to obtain multiple convolutional feature maps;
[0057] Multiple convolutional feature maps are fused together to obtain a composite feature map.
[0058] The composite feature map is compressed to obtain the first output of the first encoder.
[0059] In this embodiment, for the input 1×256×256 single-channel brain image, a multi-scale feature alignment layer is designed to fully extract local vascular bundle details and global vascular bundle structure information at different directions and scales in the brain image. Simultaneously, considering that the feature alignment layer belongs to a large model branch and needs to connect with subsequent constructions of the large model branch, the feature alignment layer provides a suitable feature format for subsequent connection with the pre-trained large model, further fully leveraging the capabilities of the large model.
[0060] Specifically, firstly, brain images are processed in parallel using convolutional layers of multiple sizes to capture diverse choroid plexus information. The convolutional kernels of these layers are 1×N, N×1, 3×3, 5×5, and 7×7, respectively. Using 1×N and N×1 kernels enhances the detection of lateral edges, vertical edges, and linear structures in the choroid plexus region, which helps to accurately extract choroid plexus edge features in different directions. Using 3×3, 5×5, and 7×7 kernels covers a larger receptive field, thus more sensitively capturing texture, structural changes, and global contextual information in the choroid plexus region. In implementation, this method fully utilizes the aforementioned convolutional kernels of various sizes, performing multi-scale convolution operations on the brain image using these kernels to obtain multiple convolutional feature maps. These convolutional feature maps encode brain image information from different angles and sizes, forming a rich multi-scale choroid plexus feature representation.
[0061] Then, these convolutional feature maps are fused into a composite feature map. The fusion process can be flexibly configured according to the actual situation; for example, concatenating the convolutional feature maps obtained from different kernels along the channel dimension can generate a composite feature map, but this is not a limitation. Considering the large number of channels in the fused composite feature map, to ensure that the output features are consistent with the input required by the subsequent pre-trained large model, this implementation further reconstructs the channel number of the composite feature map, so that the composite feature map is compressed into the first output of the first encoder. The compression process can be flexibly configured according to the actual situation; for example, using a 1×1 convolutional kernel to compress the composite feature map into the first output of the first encoder, but this is not a limitation. In this way, the multi-scale features of the convolutional cluster region are preserved while meeting the input format requirements of downstream tasks.
[0062] For example, in the feature alignment layer, N in the multi-scale convolution operation can take values of 3 and 5, resulting in 7 types of convolution kernels: 7×7, 3×3, 5×5, 1×3, 3×1, 1×5, and 5×1. Eight filters are set for each kernel, resulting in a total of 56 convolutional feature maps. Next, these convolutional feature maps obtained from different kernels are concatenated along the channel dimension to generate a composite feature map with 56 channels. Finally, the composite feature map is reconstructed using a 1×1 convolution kernel, compressing or expanding the 56 channels into 3 channels while maintaining a spatial size of 256×256, thus obtaining the first output of the first encoder.
[0063] Here, multi-scale and multi-shape convolutions are performed on the original brain images under different receptive fields. The feature maps obtained from the convolutions are then fused and compressed. This operation serves two purposes. First, considering that choroid plexuses are usually small targets and easily interfered with by complex backgrounds, the feature alignment layer can initially capture the low-level features of the choroid plexus, such as edges, gradients, and local saliency, and accurately locate these fine-grained information. This provides the necessary local contextual information for accurately locating small targets, thereby enabling the large model to accurately locate the choroid plexus region and amplify features in advance, which helps to better enable the large model to perform. Second, the training data of large models mostly comes from natural images, which are difficult to align with the features of medical brain images. The output of the feature alignment layer is connected with the subsequent construction of the large model branches. The feature alignment layer can reduce the distribution difference between the "large model pre-training data domain" and the "medical image data domain", so that the features of the large model can be more naturally "aligned" with the texture distribution of medical images, laying a good foundation for subsequent processing.
[0064] In some implementations, refer to Figure 2 The output of the first encoder mentioned above may include the first and third outputs of the first encoder, and the outputs of other encoders may include the first and second outputs of other encoders. In the other encoders mentioned above (excluding the first encoder), the inputs of the other encoders are adapted, feature-extracted, and feature-fused to obtain the outputs of the other encoders, which may include:
[0065] The first inputs of other encoders are adapted and features are extracted to obtain the first outputs of other encoders; where the first input of the second encoder is the first output of the first encoder, and the first input of the encoders after the second encoder is the first output of the previous encoder;
[0066] The second inputs of other encoders are subjected to frequency convolution to obtain the second outputs of other encoders; wherein the second input of the second encoder is the third output of the first encoder, and the second input of the encoders after the second encoder is the second output of the previous encoder.
[0067] In this embodiment, for a single other encoder, the input of the other encoder may include a first input and a second input, and its output may include a first output and a second output. The first input and the first output correspond to the large model branch, and the second input and the second output correspond to the convolutional network branch. Considering that the large model branches of each encoder are connected sequentially, and the convolutional network branches of each encoder are also connected sequentially, the first input of the second encoder is the first output of the first encoder, the first input of the encoder after the second encoder is the first output of the previous encoder, the second input of the second encoder is the third output of the first encoder, and the second input of the encoder after the second encoder is the second output of the previous encoder.
[0068] Other encoders may include an adaptive feature extraction module and a second frequency convolutional layer. In these other encoders, firstly, the adaptive feature extraction module adapts and extracts features from the first input, yielding the first output. This effectively reduces computation while comprehensively capturing and enhancing key features associated with the vein clusters. Simultaneously, the second frequency convolutional layer performs frequency convolution on the second input, yielding the second output. This further locates structurally complex, low-contrast vein cluster regions, thus comprehensively capturing complex and subtle vein cluster features. The adaptive feature extraction module belongs to the large model branch, while the second frequency convolutional layer belongs to the convolutional network branch. This means that the first output of the other encoder possesses richer global vein cluster semantic information, while the second output contains local details and edge textures of the vein clusters.
[0069] Here, compared to conventional encoders that rely solely on convolutional neural networks in related technologies, this implementation introduces adaptive feature extraction and frequency convolution operations in encoders other than the first encoder. Adaptive feature extraction effectively reduces computational cost while comprehensively capturing and enhancing key features associated with vein clusters, thereby extracting richer global vein cluster semantic information. Frequency convolution operations further locate structurally complex, low-contrast vein cluster regions, thus comprehensively capturing complex and subtle vein cluster features, such as local details and edge textures. This achieves deep feature extraction, further enhancing the feature extraction capabilities of the segmentation model.
[0070] In some implementations, refer to Figure 4 , Figure 4 The "adaptation block" refers to the adaptation module. In the aforementioned adaptation feature extraction module, adaptation processing and feature extraction are performed on the first input of other encoders to obtain the first output of those other encoders. This can include:
[0071] The first input of other encoders is subjected to adaptation processing based on downsampling and upsampling to obtain an adaptation feature map;
[0072] The adapted feature map is fused with the first input of other encoders to obtain the feature map to be processed;
[0073] The feature map to be processed is subjected to multiple feature extractions based on a large model to obtain the first output of other encoders.
[0074] In this embodiment, for a single other encoder, the aforementioned adaptation feature extraction module includes an adaptation module, an adaptation fusion layer, and multiple sequentially connected feature extraction modules. The feature extraction modules are implemented based on a large model, and their type can be flexibly set according to actual needs. For example, the feature extraction module can be the Hierarchical Vision Transformer in the SAM2 model, which retains the core of the Transformer (i.e., self-attention and forward propagation) while making specific modifications to the input and hierarchy for images (i.e., patch segmentation and fusion downsampling), thus enabling the architecture to perform excellently in image tasks.
[0075] To maintain the general visual features learned by the large model during pre-training while effectively fine-tuning brain images (especially fine structures such as choroidal plexuses), this implementation design includes a single-ended adaptation module. This module performs downsampling and upsampling adaptation processing on the first input of other encoders to obtain an adapted feature map. The core idea is to map the feature dimension from the original dimension (dim) to a lower dimension (e.g., 16 or 32) and then back to the original dimension, thereby reducing computation and enhancing key choroidal plexus features. This also enables the choroidal plexus features to be adapted to the subsequent feature extraction module based on the large model.
[0076] Specifically, in the adaptation module, firstly, considering that the feature sequences output by large models are usually high-dimensional, the first input of other encoders is downsampled to obtain a downsampled feature map. This downsampling operation, using components such as feedforward neural networks or linear mapping layers, compresses the original feature dimension (dim) to a low dimension (e.g., 16 or 32). This operation effectively reduces the computational resources required for subsequent processing while preserving key information about the choroidal plexus to the maximum extent. Then, to highlight the responses of subtle structures such as the choroidal plexus in brain images, the downsampled feature map is activated to obtain an activated feature map. Here, the dimensionality-reduced features are input into activation functions such as Sigmoid, ReLU, or LeakyReLU, using nonlinear transformations to enhance the expressive power of key choroidal plexus features. Next, to ensure the final output aligns with the original large model features in terms of dimension, the activation feature map is upsampled to obtain an upsampled feature map. This upsampling operation, using linear or fully connected layers, raises the low-dimensional (e.g., 16 or 32) space back to the original high-dimensional space (dim), ensuring smooth integration with subsequent modules (e.g., the decoder). Finally, the upsampled feature map is activated to obtain an adapted feature map. Here, the upsampled feature map is again processed by activation functions such as Sigmoid or ReLU to further stabilize and sparsify the distribution of vein cluster features, while suppressing noise that may be introduced during upsampling. This ensures that the output features maintain the global visual prior while also focusing more on the local details of the vein cluster region.
[0077] After the adaptation process is completed, the adapted feature map is fused with the first input of the other encoders in the adaptation fusion layer to obtain the feature map to be processed. Here, the adapted feature map represents the adjusted vein cluster feature information, while the first input of the other encoders represents the shallow original information. Fusing the adjusted feature information with the shallow original information can effectively increase the diversity of vein cluster features and improve the segmentation model's ability to express complex vein cluster features. The fusion operation can be flexibly set according to the actual situation; for example, fusion can be concatenated along the channel dimension, but it is not limited to this.
[0078] Subsequently, the feature map to be processed is input into the first feature extraction module, and then passes through multiple feature extraction modules in sequence to perform further feature extraction operations, ultimately obtaining the first output of other encoders, thereby further capturing key features associated with the ventricle. In simple terms, each feature extraction module can include two normalization layers, one attention layer, and one linear layer. The input to the feature extraction module is processed by the first normalization layer and then enters the attention layer. The feature map processed by the attention layer is concatenated with the input of the feature extraction module along the channel dimension. The concatenated result is input into the second normalization layer for normalization processing. The normalized concatenated result is then processed by the linear layer, and the output of the linear layer is concatenated with the concatenated result along the channel dimension to obtain the output of the feature extraction module.
[0079] Optionally, during segmentation model training, the parameters of the adaptation module can be trained, while the parameters of each feature extraction module cannot be trained.
[0080] Here, the adaptation feature extraction operation in this embodiment can be divided into adaptation operation and feature extraction operation, with the feature extraction operation relying on the large model. Through the adaptation operation, while maintaining the general visual features learned by the large model during pre-training, it is possible to effectively fine-tune brain images (especially fine structures such as choroid plexuses). This reduces computational load, enhances key choroid plexus features, and enables these features to be adapted to the subsequent feature extraction module based on the large model. The feature extraction operation further captures key features associated with choroid plexuses. Thus, deep feature extraction is achieved, effectively ensuring the segmentation model's ability to extract features from the semantic information of the global choroid plexus.
[0081] In some implementations, refer to Figure 5 and Figure 6 In the first frequency convolutional layer or the second frequency convolutional layer described above, the steps of the frequency convolution operation may include:
[0082] Wavelet transform is applied to the input of the frequency convolution operation to obtain multiple sub-band feature maps;
[0083] The feature maps of multiple sub-bands are reconstructed to obtain the reconstructed feature maps;
[0084] Perform multiple traditional convolution operations on the reconstructed feature map to obtain the output of the frequency convolution operation.
[0085] In this embodiment, for a single encoder, in the frequency convolution operation, firstly, the input of the frequency convolution operation is processed by Haar wavelet transform to obtain multiple sub-band feature maps. Haar wavelet transform is a simple and efficient discrete wavelet transform method. It captures the low-frequency and high-frequency features of the image by performing averaging and difference operations on the image data. For two-dimensional feature maps, Haar wavelet transform can divide the feature map into four frequency band sub-maps, that is, four sub-band feature maps. Each sub-band feature map represents different frequency band information, as shown in the following formula (1):
[0086]
[0087] In Equation (1), I(x, y) represents the pixel value of the feature map at coordinates (x, y); I(x+1, y) represents the pixel value of the feature map at coordinates (x+1, y); I(x, y+1) represents the pixel value of the feature map at coordinates (x, y+1); I(x+1, y+1) represents the pixel value of the feature map at coordinates (x+1, y+1); LL represents the low-frequency approximation subband; LH represents the vertical detail subband; HL represents the horizontal detail subband; and HH represents the diagonal detail subband.
[0088] A two-dimensional Haar wavelet transform is performed on the input of the current feature layer (i.e., the input of the frequency convolution operation). The core idea is to perform addition and subtraction operations on each 2×2 region of the input feature map, resulting in four subbands: a low-frequency approximation subband (LL), a vertical detail subband (LH), a horizontal detail subband (HL), and a high-frequency detail subband (HH) diagonal detail subband. Since each 2×2 pixel block is remapped into four subbands, the spatial resolution is halved (e.g., from 256×256 to 128×128), while the number of channels increases by four times, thus ensuring the complete preservation of low-frequency and high-frequency information. This transformation is also known as lossless downsampling because it reconstructs features through an inverse operation.
[0089] Next, to maintain consistency with the features of the subsequent large model branch, the four expanded sub-band feature maps need to be reconstructed. Specifically, each of the four sub-band feature maps represents complete image information. These four sub-band feature maps are concatenated along the channel dimension to obtain a feature map with 4 channels. Then, a 1×1 convolution kernel is used to reconstruct the channel count of this feature map, restoring a preset original channel count (e.g., maintaining the same channel count as the output of the large model branch) from a value four times the original. This results in the reconstructed feature map. A 1×1 convolution kernel can be viewed as a linear combination of the multi-channel vectors at each pixel location; it does not change the spatial size but effectively compresses or rearranges channel information.
[0090] After channel reconstruction, the system proceeds to the secondary convolution stage, where multiple consecutive conventional convolution operations are performed on the reconstructed feature map to obtain the output of the frequency convolution operation. This reduces the size of the feature map while further enhancing the extraction of edge and texture features of small structures such as vein clusters. The number of conventional convolution operations can be flexibly set according to the actual situation; for example, it can be two operations, but it is not limited to this. Furthermore, the conventional convolution operations can be flexibly set according to the actual situation. For example, a 3×3 convolution kernel can be used for convolution operations, followed by normalization (e.g., BatchNorm, LayerNorm) and activation function (e.g., ReLU). Here, the 3×3 convolution kernel offers a good balance in receptive field, smoothness, and detail capture. Combined with normalization and activation functions, it can suppress noise and highlight the response of key regions.
[0091] It should be understood that for the first frequency convolutional layer, the input of its frequency convolution operation is the second output of the first encoder, and the output of its frequency convolution operation is the third output of the first encoder; for the second frequency convolutional layer, the input of its frequency convolution operation is the second input of the other encoders, and the output of its frequency convolution operation is the second output of the other encoders.
[0092] Here, in order to enhance the recognition of small targets in medical images, especially to capture finer-grained textures in complex and low-contrast regions such as vascular bundles, this implementation introduces wavelet transform into the convolution operation, compared to the encoders in related technologies that rely solely on convolutional neural networks. This transform splits the original feature map into multiple frequency bands, highlighting the edge information and details of the vascular bundles while avoiding the loss of vascular bundle features. Then, through reconstruction processing, the feature maps of each frequency band are integrated into a single reconstructed feature map, maximizing the preservation of vascular bundle information. Finally, through multiple traditional convolution operations, complex and low-contrast vascular bundle regions are further captured, enhancing the vascular bundle feature representation.
[0093] In some implementations, refer to Figure 7 The outputs of other encoders may include the first output, second output, and third output of other encoders; among the other encoders mentioned above, in addition to the first encoder, the inputs of other encoders are adapted, feature extracted, and feature fused to obtain the outputs of other encoders, which may include:
[0094] The first and second outputs of other encoders are concatenated to obtain a concatenated feature map.
[0095] Channel compression is performed on the spliced feature map to obtain a compressed feature map;
[0096] An activation operation is performed on the compressed feature map to obtain the third output of the other encoders.
[0097] In this embodiment, the segmentation model is configured with two parallel branches: a large model branch and a convolutional network branch. In addition to the first encoder, other encoders may also include feature fusion layers. The feature fusion layer is designed for multi-layer and multi-stage fusion. When both large model feature output and convolutional network feature output are available, the features of the two branches are fused. Then, channel information is integrated and semantic discriminativeness is enhanced through 1×1 convolution and activation function.
[0098] Specifically, for a single other encoder, the first output of the other encoder can be obtained from the large model branch, and the second output of the other encoder can be obtained from the convolutional network branch. Both have the same spatial dimension (e.g., H×W) but differ in the number of channels (C): the large model branch often has richer global context cluster semantic information, while the convolutional network branch is better at extracting local details and edge textures of the context cluster. To fuse and complement these two different scales of context cluster feature information, firstly, the first and second outputs of the other encoder are concatenated along the channel dimension, resulting in a concatenated feature map with twice the original number of channels (2C). Next, to avoid the additional computational burden caused by redundant channels and to make the fused feature map more compact and efficient, the number of channels in the concatenated feature map is compressed to the same scale (C) as the single-branch output using a 1×1 convolutional kernel, resulting in a compressed feature map. In this way, a 1×1 convolutional kernel can linearly combine multi-channel vectors at each pixel location without changing the spatial size, effectively fusing global contextual semantics from large model branches and local details from convolutional network branches. This results in a more reasonable convolutional cluster feature representation that combines both macroscopic and microscopic information, while also having a more reasonable number of channels and computational cost. Finally, activation functions such as ReLU and Leaky ReLU are used to activate the compressed feature map, obtaining the third output of other encoders. This activation process further stabilizes and sparsifies the distribution of the convolutional cluster feature representation.
[0099] In this embodiment, the rich global context cluster semantic information is complementaryly fused with the local details and edge textures of the context cluster in the feature fusion layer of other encoders. This enables the third output of other encoders to inherit the general visual priors learned by the large model on large-scale data, and to utilize the fine context cluster features brought by the convolutional network branches. The third output of other encoders is then passed to the decoder structure along with the second and third outputs of the first encoder. This further improves the decoding effect, thereby providing richer and more balanced feature support for downstream tasks (such as segmentation, detection, or classification), ensuring the excellent performance of the segmentation model in context cluster segmentation tasks.
[0100] (ii) Decoder structure.
[0101] In some implementations, refer to Figure 2 The first decoder is connected in a skip connection to the last encoder and the penultimate encoder, respectively. The input of the first decoder may include the third output of the last encoder and the third output of the penultimate encoder. The last decoder and the penultimate decoder are both connected in a skip connection to the first encoder. The input of the last decoder may include the second output of the first encoder and the output of the penultimate decoder. The input of the penultimate decoder may include the third output of the first encoder and the output of the third-to-last decoder. The remaining decoders, except for the first decoder, the last decoder, and the penultimate decoder, correspond one-to-one with the remaining encoders, except for the first encoder, the last encoder, and the penultimate encoder, and are connected in a skip connection. The remaining decoders, except for the first decoder, the last decoder, and the penultimate decoder, include the third output of the encoder corresponding to the remaining decoder and the output of the previous decoder of the remaining decoder.
[0102] In this embodiment, such as Figure 2 As shown, with the brain image as the top, the encoder structure consists of the first encoder to the last encoder from top to bottom, and the decoder structure consists of the last decoder to the first decoder from top to bottom.
[0103] For the first decoder, it has skip connections to the last encoder and the penultimate encoder, respectively. That is, the input of the first decoder includes the third output of the last encoder and the third output of the penultimate encoder. The third output of each encoder other than the first encoder is the output of its feature fusion layer. For the last decoder, it has skip connections to the first encoder, meaning its input includes the second output of the first encoder and the output of the penultimate decoder. The second output of the first encoder is the output of its conventional convolutional layer. For the penultimate decoder, it has skip connections to the first encoder, meaning its input includes the third output of the first encoder and the output of the third-to-last decoder. The third output of the first encoder is the output of its first frequency convolutional layer. For the remaining decoders (excluding the first, last, and penultimate decoders), they have one-to-one skip connections with each of the other encoders (excluding the first, last, and penultimate encoders), meaning their input includes the third output of the encoder corresponding to each other and the output of the previous decoder. The third output of each encoder other than the first encoder is the output of its feature fusion layer. This forms a special type of skip connection.
[0104] For example, such as Figure 2As shown, there are five encoders and five decoders. Using the brain image as the top, the encoder structure consists of encoders one through five from the top down, and the decoder structure consists of decoders five through one from the top down. The first decoder is connected to the fourth and fifth encoders in a skip connection; the second decoder is connected to the third decoder in a skip connection; the third decoder is connected to the second encoder in a skip connection; the fourth decoder is connected to the first encoder in a skip connection; and the fifth decoder is connected to the first encoder in a skip connection.
[0105] In some implementations, refer to Figure 2 and Figure 8 In the decoder described above, the input to the decoder is decoded to obtain the output of the decoder, which may include:
[0106] The input to the decoder is fused to obtain a fused feature map;
[0107] The decoder output is obtained by performing multiple traditional convolution operations on the fused feature map.
[0108] In this embodiment, in the decoder structure, while receiving the fused high-dimensional feature map, the decoder also utilizes two types of scale information produced by different levels in the segmentation model. One type is high-level features produced by the corresponding encoder, which have a smaller spatial resolution but a higher semantic level (as the number of encoder layers increases, its resolution gradually decreases, and the semantic information gradually transforms into abstract high-level semantics). The other type is low-level features produced by the previous decoder, which have a larger spatial resolution and contain more local details. By comprehensively utilizing these two types of features, a segmentation result consistent with the resolution of the original image is recovered in the final output stage. Accordingly, each decoder can include a fusion input layer and a convolutional output layer. The input of the fusion input layer includes at least two types of feature maps. First, the fusion input layer fuses the decoder input to fuse feature maps of different scales, resulting in a fused feature map. Then, the convolutional input layer performs multiple traditional convolution operations on the fused feature map to reintegrate and extract more valuable ventricle features, ultimately obtaining the decoder output.
[0109] Specifically, the first decoder receives the third output of the last encoder and the third output of the penultimate encoder. In the fusion input layer of the first decoder, the third output of the last encoder is upsampled to restore its size, and then the upsampled third output of the last encoder is concatenated with the third output of the penultimate encoder on the same channel to obtain the fusion feature map of the first decoder. This aims to combine the ventricle cluster feature information of the two lowest-level encoders and pass it to the first decoder, thereby capturing more feature information associated with the ventricle cluster region. Subsequently, in the convolutional output layer of the first decoder, multiple consecutive traditional convolution operations are performed on the fusion feature map of the first decoder to obtain the output of the first decoder, which aims to initially restore the ventricle cluster region.
[0110] Here, the first decoder is the bottom-level decoder, which relies on the ventricle feature information from different encoders to restore the ventricle region. In this process, the information from different encoders needs to be re-integrated through convolution and the more valuable ventricle features are extracted. In multiple convolution operations, the number of channels is controlled while restoring the ventricle region, which helps to keep the number of feature layers matched throughout the decoding process. This enables the bottom-level decoder to generate the initial local details.
[0111] The second to third-to-last decoders are defined as the remaining decoders. Each remaining decoder receives the third output from its corresponding encoder and the output of its preceding decoder. In the fusion input layer of the remaining encoder, the output of the preceding decoder is upsampled to match its spatial dimensions with the third output of the encoder, ensuring smooth concatenation of the two feature maps along the channel dimension. Then, the upsampled output of the preceding decoder is concatenated with the third output of its corresponding encoder along the channel dimension to obtain the fusion feature map of the remaining decoders. This aims to transfer the network cluster feature information from the encoder to the decoder and fully combine the network cluster feature information from both the decoder and encoder, thereby capturing more feature information associated with the network cluster region. Subsequently, in the convolutional output layer of the remaining decoders, multiple consecutive conventional convolution operations are performed on the fusion feature map to obtain the output of the remaining decoders, aiming to gradually reconstruct the network cluster region.
[0112] Here, the second to third-to-last decoders are intermediate-level decoders, whose task is to refine the local details of the vein cluster region. These intermediate-level decoders receive both high-level semantics and local details from the previous layer decoder. Therefore, they can rely on vein cluster feature information from the corresponding encoder and the previous layer decoder to reconstruct the vein cluster region. In this process, the information from the decoder and encoder needs to be reintegrated through convolution to extract the more valuable vein cluster features. Controlling the number of channels while reconstructing the vein cluster region during multiple convolution operations helps maintain a matching number of feature layers throughout the decoding process. This allows the intermediate-level encoder to gradually refine the local details of the vein cluster region, thus better reconstructing it.
[0113] The penultimate decoder receives the third output from the first encoder and the output from the penultimate decoder. In the fusion input layer of the penultimate decoder, the output of the penultimate decoder is upsampled to match the spatial dimensions of the third output from the first encoder, ensuring smooth concatenation of the two feature maps along the channel dimension. Then, the upsampled output of the penultimate decoder is concatenated with the third output from the first encoder along the channel dimension to obtain the fusion feature map of the penultimate decoder. This aims to transfer the complex, low-contrast vesicle feature information from the encoder to the decoder and fully combine the vesicle feature information from the decoder with the aforementioned complex, low-contrast vesicle feature information, thereby capturing more detailed feature information associated with the vesicle region. Subsequently, in the convolutional output layer of the penultimate decoder, multiple consecutive traditional convolution operations are performed on the fusion feature map of the penultimate decoder to obtain the output of the penultimate decoder, aiming to further reconstruct the vesicle region.
[0114] Here, the penultimate decoder is the top-level decoder. Unlike the middle-level decoders, the top-level decoder is closer to the end of the reconstruction. Its task is no longer to refine the local details of the vein cluster region, but to fuse high-level semantics and local details (i.e., multi-scale feature fusion) to balance the relationship between the two and promote the generation of accurate vein cluster boundaries. To achieve multi-scale feature fusion and improve the fusion effect, the penultimate decoder receives information from the previous decoder and information from the frequency convolution operation. The frequency convolution operation can accurately locate the vein cluster region with complex structure and low contrast, and can initially compensate for spatial details. By fusing these two types of information through convolution, more detailed vein cluster features can be extracted, especially for vein cluster regions with complex structure, low contrast, and small targets. In multiple convolution operations, the number of channels is controlled while reconstructing the vein cluster region, which helps to keep the number of feature layers matched throughout the decoding process. This allows the top-level decoder to fuse high-level semantics and local details, so that semantic information and spatial details are balanced, and more accurate vein cluster boundaries are generated.
[0115] The last decoder receives the second output from the first encoder and the output from the penultimate decoder. In the fusion input layer of the last decoder, the output of the penultimate decoder is upsampled to match its spatial dimensions with the second output of the first encoder, ensuring smooth concatenation of the two feature maps along the channel dimension. Then, the upsampled output of the penultimate decoder and the second output of the first encoder are concatenated along the channel dimension to obtain the fusion feature map of the last decoder. This aims to transfer the feature information from the encoder that is moderately or highly correlated with the ventricle region to the decoder, and to fully combine the ventricle feature information from the decoder with the aforementioned feature information, thereby ultimately capturing the most detailed feature information associated with the ventricle region. Subsequently, in the convolutional output layer of the last decoder, multiple consecutive traditional convolution operations are performed on the fusion feature map of the last decoder to obtain the output of the last decoder, thus finally reconstructing the ventricle contour map.
[0116] Here, the last decoder is the top-level decoder, whose task is to fuse high-level semantics and local details (i.e., multi-scale feature fusion) to balance the relationship between the two and promote the generation of accurate vein cluster boundaries. To achieve multi-scale feature fusion and improve the fusion effect, the last decoder receives information from the previous decoder and information from traditional convolutional operations. Traditional convolutional operations can initially capture shallow features associated with vein cluster texture and edges, and can further compensate for spatial details. By fusing these two types of information through convolution, the most detailed vein cluster features can be extracted. In multiple convolutional operations, the number of channels is controlled while restoring the vein cluster region, which helps to keep the number of feature layers matched throughout the decoding process. This allows the top-level decoder to fully fuse high-level semantics and local details to generate the final vein cluster boundary.
[0117] Optionally, the upsampling operation in the fusion input layer of the decoder can be flexibly set according to the actual situation. For example, the upsampling operation can be linear interpolation upsampling, but it is not limited to this.
[0118] Optionally, the number of traditional convolution operations can be flexibly set according to the actual situation. For example, a traditional convolution operation can be performed twice, but it is not limited to this. In addition, the traditional convolution operation can be flexibly set according to the actual situation. For example, a 3×3 convolution kernel can be used for convolution operation, and normalization processing (such as BatchNorm, LayerNorm, etc.) and activation function processing (such as ReLU, etc.) can be performed in sequence after convolution.
[0119] In summary, considering that most related technologies' fully convolutional neural networks only have simple cascaded and dense connections, it is difficult to capture local details of the ventricle bundle, resulting in low segmentation accuracy. To address this, in this embodiment, the feature fusion operation of the encoder structure can extract more abstract ventricle bundle semantics, while the decoder structure can extract ventricle bundle details such as color, texture, and edges. Through special skip connections between the decoder and encoder structures, and the decoding processing of the decoder structure, the ventricle bundle semantics can be passed to the decoder structure, and the ventricle bundle semantics and details can be combined to accurately capture small targets and complex ventricle bundle features in brain images, improving the feature extraction capability of the segmentation model and effectively improving the segmentation accuracy of brain ventricles.
[0120] (III) Output layer.
[0121] In some implementations, refer to Figure 2 , Figure 2 The "output" in this context refers to the output layer. In this output layer, based on the output of the last decoder, a network contour map is obtained, which may include:
[0122] The output of the last decoder is activated to obtain the activation feature map.
[0123] For the intensity value of each pixel in the activation feature map, if the intensity value of the current pixel is greater than the preset intensity value, the value of the current pixel is determined as the first value; otherwise, the value of the current pixel is determined as the second value, thus obtaining the ventricle contour map.
[0124] In this embodiment, an output layer is connected after the last decoder. In the output layer, functions such as Sigmoid and Softmax are used to activate the output of the last decoder, aiming to compress the discrete intensity values of each pixel in the output of the last decoder to [0, 1], thus obtaining an activation feature map. Subsequently, for the intensity value of each pixel in the activation feature map, it is determined whether the intensity value of the current pixel is greater than a preset intensity value. If so, it indicates that the pixel belongs to a vein cluster region, and the value of the current pixel is determined as the first value; otherwise, it indicates that the pixel belongs to the background region, and the value of the current pixel is determined as the second value. By traversing all pixels in the activation feature map, the values of all pixels can be determined, thus obtaining a vein cluster contour map represented by a binary mask. This is because the essence of the vein cluster segmentation task is the detection of vein cluster regions; it is only necessary to know whether a certain pixel position belongs to a vein cluster. Therefore, setting an output layer after decoding is to convert the discrete output feature map into a binary label map.
[0125] Optionally, the preset intensity value can be flexibly set according to the actual situation. For example, the preset intensity value can be 0.5, but it is not limited to this. In addition, the first value and the second value can also be flexibly set according to the actual situation. For example, the first value can be 1 and the second value can be 0, but it is not limited to this.
[0126] (iv) Training of the segmentation model.
[0127] In some embodiments, the above-described choroid plexus segmentation method may further include:
[0128] Multiple brain image samples were acquired, each with a corresponding choroid plexus contour map label;
[0129] Multiple brain image samples were preprocessed to obtain multiple preprocessed brain image samples;
[0130] The segmentation model was trained using multiple preprocessed brain image samples.
[0131] In this implementation, firstly, brain image samples of different sizes are collected, and the choroid plexus contour maps of each brain image sample are manually drawn. For example, clinicians use professional software such as 3D Slicer to manually draw the choroid plexus contour maps. Doctors rely on their extensive clinical experience to establish annotation standards to ensure accurate delineation of the choroid plexus regions, and cross-validation is used to further reduce subjective bias. Simultaneously, all brain image samples undergo quality checks, and low-quality samples due to noise, motion artifacts, or annotation errors are removed. Then, to ensure the quality of the brain image samples, each original brain image sample needs to be preprocessed. Specifically, through methods such as padding or resizing, all brain image samples are adjusted to a fixed size of 256×256. Simultaneously, all brain image samples are normalized, and data augmentation techniques such as rotation, translation, and mirroring are combined to achieve preprocessing operations. This results in multiple preprocessed brain image samples, providing high-quality input for the model and improving its robustness and generalization ability. Finally, multiple preprocessed brain image samples are input into the segmentation model to train it.
[0132] In some embodiments, the above-described choroid plexus segmentation method may further include:
[0133] The Dice loss function is used to backpropagate the segmentation model.
[0134] In this embodiment, the Dice loss function is used as the loss function throughout the model training process. When the prediction result matches the real label perfectly, the Dice coefficient is 1, and the corresponding Dice loss function is 0; otherwise, the larger the Dice loss function, the lower the overlap between the prediction result and the label. The trainable parameters in the segmentation model (such as feature alignment layer, adaptation module and convolution weights in convolution branches) are continuously adjusted through the backpropagation algorithm, while keeping some of the preprocessing weights of the large model frozen, so as to achieve a balance between generalization ability and targeted fine-tuning. During the training process, strategies such as learning rate decay and early stopping can also be used, and cross-validation and independent test sets are used to evaluate the model performance to ensure stability and robustness on real clinical data. Among them, the Dice loss function satisfies the following formula (2):
[0135]
[0136] In equation (2), DSC(p, g) represents the Dice loss function; DSC(p, g) represents the Dice coefficients; p i G represents the probability of the i-th pixel (or voxel) in the prediction result; iRepresents the value of the i-th pixel (or voxel) in the actual result; ∈ is a very small positive number used to prevent the denominator from being zero.
[0137] (V) Explanation of the principle of the segmentation model.
[0138] Reference Figure 2 A single-channel MRI brain image of 1×256×256 is simultaneously input into two parallel branches of the segmentation model: the large model branch and the convolutional branch. Then, the features output by these branches are fused at multiple scale levels. Finally, the decoder gradually upsamples and restores the choroid plexus segmentation result of the same size as the original image.
[0139] The following application scenario illustrates the overall data flow of the segmentation model.
[0140] In the segmentation model, there are five encoders and five decoders. An output layer is placed after the last decoder. Using the brain image as the top layer, the encoder structure consists of five encoders from top to bottom, and the decoder structure consists of five decoders from top to bottom, and so on. The first decoder has skip connections to the fourth and fifth encoders; the second decoder has skip connections to the third decoder; the third decoder has skip connections to the second encoder; the fourth decoder has skip connections to the first encoder; and the fifth decoder has skip connections to the first encoder.
[0141] Each encoder has a main model branch and a convolutional network branch, with the main model branches and convolutional network branches of each encoder connected sequentially. Using the brain image as the top, the main model branch, from top to bottom, consists of the feature alignment layer of the first encoder and the adaptation feature extraction modules of the second to fifth encoders. The convolutional network branches, from top to bottom, consist of the traditional convolutional layer and the first frequency convolutional layer of the first encoder, and the second frequency convolutional layer of the second to fifth encoders. In addition, each encoder except the first encoder has a feature fusion layer, and the output of the feature fusion layer of each encoder is passed to the corresponding decoder. Furthermore, the outputs of the traditional convolutional layer and the first frequency convolutional layer of the first encoder are also passed to the corresponding decoder.
[0142] In the large model branch, before the 1×256×256 brain image formally enters the adaptation feature extraction module of the encoders other than the first encoder, such as Figure 3As shown, in the feature alignment layer of the first encoder, the brain image is first subjected to parallel convolution using multi-scale, multi-shape convolutional kernels (e.g., 1×n, n×1, 3×3, 5×5, 7×7) to capture edge and texture information under different orientations and receptive fields. After concatenating these features along the channel dimension, a 1×1 convolutional kernel is used to compress the number of channels to 3 to match the input format of the adaptive feature extraction module. The first output of the first encoder is obtained through the feature alignment layer, which is then passed to the adaptive feature extraction module of the second encoder and serves as the first input of the second encoder.
[0143] The convolutional network branch focuses more on extracting local details from the ventricle region. It directly performs pure convolution operations on the 1×256×256 brain image through the traditional convolutional layer of the first encoder, such as several 3×3 convolutions. After each convolution, normalization and activation functions are applied, resulting in a feature map that maintains a spatial size of 256×256. This output serves as the second output of the first encoder, which is then passed to the fifth decoder and used as one of its inputs. This allows for initial enhancement of the capture of ventricle edges and textures without downsampling. Subsequently, in the first frequency convolutional layer of the first encoder, a frequency convolution operation based on Haar wavelet transform is performed on the second output of the first encoder, such as... Figure 5 and Figure 6 As shown, the specific operation is as follows: Each 2×2 region is decomposed into four sub-bands (LL, LH, HL, HH) using Haar wavelet transform, thus halving the spatial resolution but expanding the number of channels to four times the original. Then, a 1×1 convolutional kernel is used to restore the number of channels to the desired target value. Several 3×3 convolutions (also combined with normalization and activation functions) are then applied to further enhance the convolutional feature representation, resulting in the third output of the first encoder. This output is then passed to the second frequency convolutional layer of the second encoder and serves as the second input to the second encoder. This allows for the acquisition of multi-band convolutional edge and texture information at deeper levels, outputting multi-scale feature maps with the same resolution as the large model branches.
[0144] Next, the first output of the first encoder is used as the first input of the second encoder, and it is fed into the adaptation feature extraction module of the second encoder. For example... Figure 4As shown, in this adaptive feature extraction module, to balance the global semantics of the large model with the small target vascular bundle features of medical images, an adaptation operation is introduced. This involves downsampling and upsampling the first input from other encoders to obtain an adaptive feature map. Specifically, higher-dimensional features are first compressed to 16 or 32, then processed by an activation function to restore the original dimensions, followed by a second activation. Subsequently, the adapted feature map is input into the first feature extraction module, and then sequentially passed through multiple feature extraction modules to perform further feature extraction operations, resulting in the output of the adaptive feature extraction module, which is the first output of the second encoder. This output is then passed to the adaptive feature extraction module of the third encoder and serves as the first input to the third encoder. The feature extraction module here is a hierarchical visual Transformer in the SAM2 model, which captures key features associated with vascular bundles level by level through its hierarchical visual Transformer structure. This preserves the macroscopic priors of the large model while making the features more suitable for the segmentation needs of fine structures such as vascular bundles.
[0145] The third output of the first encoder is then used as the second input of the second encoder, fed into the second frequency convolutional layer of the second encoder. The second frequency convolutional layer outputs the second output of the second encoder, which is then passed to the second frequency convolutional layer of the third encoder and used as the second input of the third encoder. The implementation of the second frequency convolutional layer is the same as that of the first frequency convolutional layer, and will not be described in detail here.
[0146] After processing by the second encoder, the corresponding first output (corresponding to the large model branch) and second output (corresponding to the convolutional network branch) are obtained. The first output of the second encoder serves as the first input of the third encoder, and the second output of the second encoder serves as the second input of the third encoder. The third encoder then undergoes the same data flow operation as the second encoder. This process continues, with the third, fourth, and fifth encoders sequentially performing the encoding operations.
[0147] Since the large model branch and the convolutional network branch each output feature maps with the same resolution, feature fusion is performed at each corresponding scale to combine their advantages. This fusion is executed by the feature fusion layers of the encoders other than the first encoder, such as... Figure 7As shown, the specific method is as follows: First, feature maps from the large model branch and the convolutional network branch (i.e., the first and second outputs of other encoders) are concatenated along the channel dimension to obtain a concatenated feature map with double the number of channels. Then, a 1×1 convolution is used to compress the number of channels back to the scale of a single branch output, and an activation function is used for feature recombination and filtering to obtain the third output of other encoders, which will be passed to the corresponding decoder and used as one of the inputs of the decoder. In this way, the global context of the large model branch and the local ventricle details of the convolutional network branch can be used simultaneously to obtain a comprehensive feature representation that preserves semantic information and highlights the edges of small targets in the ventricle.
[0148] In summary, after processing using the encoder structure described above, the output directions of each encoder are:
[0149] First encoder: First output, passed to the adaptation feature extraction module of the second encoder; Second output, passed to the first frequency convolutional layer and the fifth decoder of the first encoder; Second output, passed to the second frequency convolutional layer and the fourth decoder of the second encoder.
[0150] The second encoder has the following outputs: the first output is passed to the adaptive feature extraction module of the third encoder and the feature fusion layer of the second encoder; the second output is passed to the second frequency convolutional layer of the third encoder and the feature fusion layer of the second encoder; and the third output is passed to the third decoder.
[0151] The third encoder has the following outputs: the first output is passed to the adaptation feature extraction module of the fourth encoder and the feature fusion layer of the third encoder; the second output is passed to the second frequency convolutional layer of the fourth encoder and the feature fusion layer of the third encoder; and the third output is passed to the second decoder.
[0152] The fourth encoder has the following outputs: First output, which is passed to the adaptation feature extraction module of the fifth encoder and the feature fusion layer of the fourth encoder; Second output, which is passed to the second frequency convolutional layer of the fifth encoder and the feature fusion layer of the fourth encoder; Third output, which is passed to the first decoder.
[0153] The fifth encoder has three outputs: the first output is passed to the feature fusion layer of the fifth encoder; the second output is passed to the feature fusion layer of the fifth encoder; and the third output is passed to the first decoder.
[0154] Accordingly, the output of the encoder structure is input into the decoder structure. Each decoder processes high-level features (small resolution, strong semantics) and low-level features (large resolution, rich details) sequentially, such as... Figure 8As shown, this is specifically achieved by first upsampling high-level features to the same spatial size as low-level features, then concatenating them across channels, and continuously performing 3×3 convolution, normalization, and activation function operations to repair artifacts caused by upsampling and refine the edges of the vein clusters. Through upsampling and concatenation by multiple decoders, the decoder structure ultimately outputs a vein cluster segmentation map of the same size as the original image (1×256×256), and uses the Sigmoid or Softmax function in the output layer to classify each pixel, generating an accurate segmentation mask for the vein cluster region.
[0155] Finally, the generated segmentation mask may require further optimization using post-processing techniques (such as morphological operations and region connectivity analysis) to remove noise and ensure the continuity and smoothness of the segmentation boundaries. The segmentation results can be used not only for clinical diagnosis but also for scientific research analysis and further model optimization and validation.
[0156] This complete process fully utilizes the macroscopic priors of the large model and the fine capture of details of the convolutional network branches. Through multi-level fusion and progressive refinement by the decoder, it ultimately achieves higher accuracy and robustness in the convolutional network segmentation task.
[0157] Furthermore, embodiments of this application provide a brain choroid plexus segmentation device, which may include:
[0158] The acquisition module is used to acquire brain images;
[0159] The segmentation module is equipped with a pre-trained segmentation model. The segmentation module is used to input brain images into the segmentation model to obtain venous plexus contour maps.
[0160] The segmentation model includes an encoder structure, a decoder structure, and an output layer. The encoder structure has multiple encoders arranged sequentially. The first encoder is used to perform feature alignment and feature extraction on the brain image, resulting in the output of the first encoder. The other encoders besides the first encoder are used to adapt, extract, and fuse the inputs of other encoders, resulting in the outputs of other encoders. The decoder structure is connected to the encoder structure in a skip connection. The decoder structure has multiple decoders arranged sequentially. Each decoder is used to decode the input of the decoder, resulting in the output of the decoder. The output layer is used to obtain the ventricle contour map based on the output of the last decoder.
[0161] The content of the above method embodiments is applicable to the device embodiments. The specific functions implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0162] In summary, the embodiments of this application have at least the following technical effects:
[0163] (1) This application proposes a dual-branch adaptive fusion of the large model branch and the convolutional network branch in the segmentation model. In related technologies, most use at least one convolutional neural network structure, U-Net, or fine-tune the large model separately, and rarely implement parallel branches and adaptive fusion of the two. Accordingly, this application sets up corresponding large model branches and convolutional network branches for each encoder. The large model branches of each encoder are connected sequentially, and the convolutional network branches of each encoder are also connected sequentially. In this way, through the dual-branch design in the segmentation model, the convolutional neural network can be guaranteed to depict the fine-grained network boundaries while ensuring the strong prior ability of the large model. At the same time, this application designs feature fusion operations for encoders other than the first encoder. Through multiple fusions, the general visual features of the large model and the specific details of the network cluster are combined layer by layer and passed to the corresponding decoder. This allows the segmentation model to retain the high-level semantics of the large model and preserve the details of the network cluster by combining the convolutional network branches when processing small targets, thereby reducing the probability of omission or misclassification of the network cluster region of small targets.
[0164] (2) This application introduces Haar wavelet transform into the convolution operation. In related technologies, the convolution operation often involves max pooling or average pooling, which can easily lose the edges of the vein cluster. In response, this application introduces Haar wavelet transform into the convolution operation of each encoder, which can well preserve the edge features of the vein cluster in different frequency bands and is more sensitive to the detection of vein cluster edges of low contrast and small targets. Thus, compared with simple pooling, this application avoids the edge blurring caused by multiple pooling in the convolution layer, and preserves more high-frequency details associated with the vein cluster region without significantly increasing the computational cost.
[0165] (3) This application proposes a feature alignment operation and an adaptation operation. In related technologies, when applying large models to medical imaging scenarios, it is often necessary to fine-tune the large model. However, directly fine-tuning the large model requires unfreezing a large number of weights, which leads to a significant increase in computational cost and memory consumption. In addition, simply freezing the large model cannot make it finely applicable to choroid plexus segmentation tasks. To address this, this application designs a feature alignment operation for the first encoder. Through the feature alignment operation, the difference between the brain image domain and the natural image domain is reduced. Furthermore, a large model is configured for the encoders other than the first encoder, and an adaptation operation is inserted into the large model. This large model is trainable, so only key parameters are fine-tuned. In this way, without completely destroying the existing generalization ability of the large model, customized adaptation is performed for small target scenarios such as choroid plexus. Compared with the traditional direct fine-tuning of large models, the training cost is lower and the ability to characterize segmentation boundaries is stronger. Moreover, compared with the approach of simultaneously unfreezing large models or increasing the depth of convolutional networks, this application only inserts adaptation operations in key parts for fine-tuning, reducing computational and memory overhead, and maintaining good compatibility with images of various resolutions.
[0166] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.
[0167] The above is a detailed description of the preferred embodiments of this application, but this application is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application, and these equivalent modifications or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A method of segmenting a cerebral choroid plexus, characterized by, Includes the following steps: Acquire brain images; The brain image is input into a pre-trained segmentation model to obtain a choroid plexus contour map; The segmentation model includes: The encoder structure comprises multiple encoders arranged sequentially. The first encoder is used to perform feature alignment and feature extraction on the brain image to obtain the output of the first encoder. The other encoders besides the first encoder are used to perform adaptation processing, feature extraction, and feature fusion on the inputs of the other encoders to obtain the outputs of the other encoders. A decoder structure is connected in a skip connection to the encoder structure. The decoder structure is provided with multiple decoders in sequence. Each decoder is used to decode the input of the decoder to obtain the output of the decoder. An output layer is used to obtain the ventricle contour map based on the output of the last decoder. The output of the first encoder includes a first output, a second output, and a third output; the step of performing feature alignment and feature extraction on the brain image to obtain the output of the first encoder includes: The brain image is feature-aligned to obtain the first output of the first encoder; Perform a conventional convolution operation on the brain image to obtain the second output of the first encoder; A frequency convolution operation is performed on the second output of the first encoder to obtain the third output of the first encoder; The outputs of the other encoders include the first and second outputs of the other encoders; the process of adapting, extracting features, and fusing features from the inputs of the other encoders to obtain their outputs includes: The first inputs of the other encoders are adapted and feature extracted to obtain the first outputs of the other encoders; wherein the first input of the second encoder is the first output of the first encoder, and the first input of the encoders after the second encoder is the first output of the previous encoder; A frequency convolution operation is performed on the second input of the other encoders to obtain the second output of the other encoders; wherein the second input of the second encoder is the third output of the first encoder, and the second input of the encoders after the second encoder is the second output of the previous encoder.
2. The method according to claim 1, characterized in that, The step of aligning features in the brain image to obtain the first output of the encoder includes: The brain image is subjected to multi-scale convolution operations to obtain multiple convolutional feature maps; The multiple convolutional feature maps are fused to obtain a composite feature map; The composite feature map is compressed to obtain the first output of the encoder.
3. The method according to claim 1, characterized in that, The process of adapting and extracting features from the first input of the other encoders to obtain the first output of the other encoders includes: The first input of the other encoders is subjected to adaptation processing based on downsampling and upsampling to obtain an adaptation feature map; The adapted feature map is fused with the first input of the other encoders to obtain the feature map to be processed. The feature map to be processed is subjected to multiple feature extractions based on a large model to obtain the first output of the other encoders.
4. The method according to claim 1, characterized in that, The steps of the frequency convolution operation include: Wavelet transform is applied to the input of the frequency convolution operation to obtain multiple sub-band feature maps; The multiple sub-band feature maps are reconstructed to obtain reconstructed feature maps; The reconstructed feature map is subjected to multiple conventional convolution operations to obtain the output of the frequency convolution operation.
5. The method according to claim 1, characterized in that, The outputs of the other encoders include the first output, second output, and third output of the other encoders; the process of adapting, extracting features, and fusing features from the inputs of the other encoders to obtain their outputs includes: The first and second outputs of the other encoders are spliced together to obtain a spliced feature map. Channel compression is performed on the spliced feature map to obtain a compressed feature map; An activation operation is performed on the compressed feature map to obtain the third output of the other encoders.
6. The method according to claim 1, characterized in that, The first decoder is connected in a skip connection to the last encoder and the penultimate encoder, respectively, and the input of the first decoder includes the third output of the last encoder and the third output of the penultimate encoder; The last decoder and the penultimate decoder are both connected to the first encoder in a skip connection. The input of the last decoder includes the second output of the first encoder and the output of the penultimate decoder. The input of the penultimate decoder includes the third output of the first encoder and the output of the third-to-last decoder. The remaining decoders, except for the first, last, and penultimate decoders, and the remaining encoders, except for the first, last, and penultimate encoders, correspond one-to-one and are connected in a skip connection; The remaining decoders, excluding the first, last, and penultimate decoders, include the third output of the encoder corresponding to the remaining decoders and the output of the previous decoder of the remaining decoders.
7. The method according to claim 6, characterized in that, The process of decoding the input of the decoder to obtain the output of the decoder includes: The input to the decoder is fused to obtain a fused feature map; The decoder output is obtained by performing multiple conventional convolution operations on the fused feature map.
8. A brain choroid plexus segmentation device, characterized in that, include: The acquisition module is used to acquire brain images; The segmentation module is equipped with a pre-trained segmentation model. The segmentation module is used to input the brain image into the segmentation model to obtain a choroid plexus contour map. The segmentation model includes: The encoder structure comprises multiple encoders arranged sequentially. The first encoder is used to perform feature alignment and feature extraction on the brain image to obtain the output of the first encoder. The other encoders besides the first encoder are used to perform adaptation processing, feature extraction, and feature fusion on the inputs of the other encoders to obtain the outputs of the other encoders. A decoder structure is connected in a skip connection to the encoder structure. The decoder structure is provided with multiple decoders in sequence. Each decoder is used to decode the input of the decoder to obtain the output of the decoder. An output layer is used to obtain the ventricle contour map based on the output of the last decoder. The output of the first encoder includes a first output, a second output, and a third output; the step of performing feature alignment and feature extraction on the brain image to obtain the output of the first encoder includes: The brain image is feature-aligned to obtain the first output of the first encoder; Perform a conventional convolution operation on the brain image to obtain the second output of the first encoder; A frequency convolution operation is performed on the second output of the first encoder to obtain the third output of the first encoder; The outputs of the other encoders include the first and second outputs of the other encoders; the process of adapting, extracting features, and fusing features from the inputs of the other encoders to obtain their outputs includes: The first inputs of the other encoders are adapted and feature extracted to obtain the first outputs of the other encoders; wherein the first input of the second encoder is the first output of the first encoder, and the first input of the encoders after the second encoder is the first output of the previous encoder; A frequency convolution operation is performed on the second input of the other encoders to obtain the second output of the other encoders; wherein the second input of the second encoder is the third output of the first encoder, and the second input of the encoders after the second encoder is the second output of the previous encoder.
Citation Information
Patent Citations
Cell nucleus segmentation method of cascade coding segmentation network based on large model guidance
CN118366153A
Depth learning-based vein clump image segmentation method and system
CN118485675A