Brain venation plexus segmentation method and device

By introducing operations such as feature alignment, adaptation processing, and feature fusion into the encoder structure and combining it with the decoder structure, the problem of insufficient feature extraction of fully convolutional neural networks in brain choroid plexus segmentation is solved, and accurate segmentation of the choroid plexus area in brain images is achieved, thereby improving the segmentation accuracy.

CN120672659AActive Publication Date: 2025-09-19SHANTOU UNIV +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510605463.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-09-19
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

Existing fully convolutional neural networks have insufficient feature extraction capabilities in brain choroid plexus segmentation, resulting in low segmentation accuracy and difficulty in effectively capturing the choroid plexus features of small targets and complex structures.

Method used

A segmentation model that combines an encoder structure with a decoder structure is used. Through feature alignment, adaptation processing, feature extraction, and feature fusion operations, the large model is combined with a convolutional neural network to improve the feature extraction capability.

Benefits of technology

It achieves accurate segmentation of the choroid plexus area in brain images, improves segmentation accuracy, and can fully capture complex and subtle choroid plexus features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672659A_ABST
    Figure CN120672659A_ABST
Patent Text Reader

Abstract

The invention discloses a brain venation plexus segmentation method and device, and is applied to the technical field of medical image processing, and the method comprises the steps: inputting a brain image into a pre-trained segmentation model, and obtaining a venation plexus contour map; wherein the segmentation model comprises an encoder structure, a decoder structure and an output layer, the encoder structure is sequentially provided with a plurality of encoders, and the first encoder performs feature alignment and feature extraction on the brain image to obtain the output of the first encoder; the other encoders carry out adaptation processing, feature extraction and feature fusion on the input of the other encoders to obtain the output of the other encoders; the decoder structure is in jump connection with the encoder structure, the decoder structure is sequentially provided with a plurality of decoders, and each decoder decodes the input of the decoder to obtain the output of the decoder; and an output layer obtains a venation cluster contour diagram according to the output of the last decoder. The method effectively improves the segmentation precision of the venation plexus region in the brain image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of medical image processing technology, and in particular to a method and device for segmenting brain choroid plexus. Background Art

[0002] The choroid plexus is a structure within the brain's ventricles, composed of the pia mater, vascular plexus, and ependymal epithelium. It is responsible for producing cerebrospinal fluid. Its morphology is characterized by repeated branching of blood vessels into a plexus-like structure that protrudes into the ventricular cavity. In related techniques, brain images are processed using existing fully convolutional neural networks to segment and obtain the contours of the choroid plexus in brain images. However, the choroid plexus has characteristics such as small size and complex structure, and existing fully convolutional neural networks often use convolutional neural networks as their encoders. Their ability to capture choroid plexus features is insufficient, resulting in low segmentation accuracy of the choroid plexus in the brain. Summary of the Invention

[0003] The embodiments of the present application provide a method and apparatus for segmenting the choroid plexus of the brain, which are used to improve the segmentation accuracy of the choroid plexus of the brain.

[0004] In one aspect, an embodiment of the present application provides a method for segmenting a brain choroid plexus, comprising the following steps:

[0005] Obtain brain images;

[0006] inputting the brain image into a pre-trained segmentation model to obtain a choroid plexus contour map;

[0007] Wherein, the segmentation model includes:

[0008] The encoder structure is sequentially provided with a plurality of encoders; the first encoder is used to perform feature alignment and feature extraction on the brain image to obtain the output of the first encoder; the encoders other than the first encoder are used to perform adaptation processing, feature extraction and feature fusion on the inputs of the other encoders to obtain the outputs of the other encoders;

[0009] A decoder structure is jump-connected to the encoder structure, wherein the decoder structure is sequentially provided with a plurality of decoders, each of which is used to decode the input of the decoder to obtain the output of the decoder;

[0010] The output layer is used to obtain the choroid plexus contour map according to the output of the last decoder.

[0011] In another aspect, an embodiment of the present application provides a brain choroid plexus segmentation device, comprising:

[0012] an acquisition module, for acquiring brain images;

[0013] a segmentation module configured with a pre-trained segmentation model, the segmentation module being used to input the brain image into the segmentation model to obtain a choroid plexus contour map;

[0014] Wherein, the segmentation model includes:

[0015] The encoder structure is sequentially provided with a plurality of encoders; the first encoder is used to perform feature alignment and feature extraction on the brain image to obtain the output of the first encoder; the encoders other than the first encoder are used to perform adaptation processing, feature extraction and feature fusion on the inputs of the other encoders to obtain the outputs of the other encoders;

[0016] A decoder structure is jump-connected to the encoder structure, wherein the decoder structure is sequentially provided with a plurality of decoders, each of which is used to decode the input of the decoder to obtain the output of the decoder;

[0017] The output layer is used to obtain the choroid plexus contour map according to the output of the last decoder.

[0018] According to a brain choroid plexus segmentation method and device provided by the present application, a brain image is obtained, the brain image is input into a pre-trained segmentation model, and a choroid plexus contour map is obtained. The segmentation model includes an encoder structure, a decoder structure, and an output layer. The encoder structure is sequentially provided with multiple encoders. The first encoder performs feature alignment and feature extraction on the brain image to obtain the output of the first encoder. The other encoders perform adaptation processing, feature extraction, and feature fusion on the input of other encoders to obtain the output of other encoders. The decoder structure is jump-connected to the encoder structure. The decoder structure is sequentially provided with multiple decoders. Each decoder decodes the decoder input to obtain the decoder output. The output layer obtains the choroid plexus contour map based on the output of the last decoder. According to the technical solution of the present application, by introducing operations such as feature alignment, adaptation processing, feature extraction, and feature fusion into the encoder structure and combining them with the decoding operation of the decoder structure, the choroid plexus region in the brain image is segmented. The segmentation processing can accurately capture small targets and complex choroid plexus features in the brain image, improve the feature extraction capability of the segmentation model, and thus effectively improve the segmentation accuracy of the brain choroid plexus.

[0019] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be achieved and obtained through the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1This is a flow chart of a brain choroid plexus segmentation method provided by the present application;

[0021] Figure 2 It is a structural diagram of the segmentation model provided by this application;

[0022] Figure 3 This is a schematic diagram of the feature alignment layer provided by this application;

[0023] Figure 4 This is a schematic diagram of the adaptive feature extraction module provided by this application;

[0024] Figure 5 This is a schematic diagram of the frequency convolution operation provided by this application;

[0025] Figure 6 It is a schematic diagram of the wavelet transform provided by this application;

[0026] Figure 7 This is a schematic diagram of the feature fusion layer provided by this application;

[0027] Figure 8 This is a structural diagram of the decoder provided by this application. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0029] The present application is further described below in conjunction with the accompanying drawings and specific embodiments. The described embodiments should not be considered as limiting the present application. All other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0030] In the following description, reference is made to “some embodiments,” which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0032] It should be noted that in various specific embodiments of the present disclosure, when it comes to the need to perform relevant processing based on relevant data such as brain images, the subject's permission or consent will be obtained first, and the collection, use and processing of such data will comply with relevant laws, regulations and standards. For example, when the present disclosure embodiment needs to obtain relevant data such as brain images, the subject's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the subject's separate permission or consent, the relevant data such as brain images necessary for the normal operation of the present disclosure embodiment will be obtained.

[0033] The choroid plexus is a structure within the brain's ventricles, composed of the pia mater, vascular plexus, and ependymal epithelium. It is responsible for producing cerebrospinal fluid. Its morphology is characterized by repeated branching of blood vessels into a plexus-like structure that protrudes into the ventricular cavity. In related technologies, brain images are processed using existing fully convolutional neural networks to segment and obtain the outline of the choroid plexus in the brain image. Taking U-Net as an example, U-Net uses an encoder-decoder structure, in which the encoder extracts features and reduces the resolution through convolutional and pooling layers, while the decoder gradually restores the spatial information of the image through upsampling or deconvolution, thereby segmenting and obtaining the outline of the choroid plexus in the brain image.

[0034] However, the choroid plexus is small and complex in structure, and brain images are typically obtained from magnetic resonance imaging (MRI) or computed tomography (CT), which are often high in noise and low in contrast. This results in a lack of contrast between the choroid plexus's boundaries and surrounding tissues, significantly increasing the difficulty of choroid plexus segmentation. Existing fully convolutional neural networks often use convolutional neural networks as their encoders, which have insufficient feature extraction capabilities and are unable to effectively capture small, complex choroid plexus features. This makes it prone to blurred boundaries, localized missed segments, and loss of high-frequency details and edge information, resulting in low segmentation accuracy for the brain's choroid plexus.

[0035] In view of this, embodiments of the present application provide a method and apparatus for brain choroid plexus segmentation, aiming to effectively improve the segmentation accuracy of the brain choroid plexus.

[0036] First, a brain choroid plexus segmentation method provided by an embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0037] The brain choroid plexus segmentation method provided in the embodiments of the present application can be applied to a terminal, a server, or software running in a terminal or a server. The terminal can be a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. In addition, the server can also be a node server in a blockchain network, but is not limited thereto. Blockchain is a new application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms.

[0038] Reference Figure 1 and Figure 2 The brain choroid plexus segmentation method may include the following steps S101-S102:

[0039] S101, acquiring brain images;

[0040] S102, inputting the brain image into a pre-trained segmentation model to obtain a choroid plexus contour map; wherein the segmentation model includes: an encoder structure, a decoder structure, and an output layer; the encoder structure is sequentially provided with multiple encoders; the first encoder is used to perform feature alignment and feature extraction on the brain image to obtain the output of the first encoder; the encoders other than the first encoder are used to perform adaptation processing, feature extraction, and feature fusion on the inputs of the other encoders to obtain the outputs of the other encoders; the decoder structure is jump-connected to the encoder structure, and the decoder structure is sequentially provided with multiple decoders, each decoder is used to decode the decoder input to obtain the decoder output; the output layer is used to obtain the choroid plexus contour map based on the output of the last decoder.

[0041] In an embodiment of the present application, a brain image is first acquired through a preset database, which pre-stores a plurality of brain images. The brain image refers to a magnetic resonance imaging image or computed tomography image to be measured and associated with the brain. The brain image is then input into a pre-trained segmentation model. The segmentation model is a neural network model trained using a plurality of preset brain image samples and label information corresponding to each brain image sample. The label information refers to a choroid plexus contour map within the brain image sample. The segmentation model can be used to segment the brain image to obtain a choroid plexus contour map, which is a binary mask map.

[0042] Specifically, the segmentation model can include an encoder structure, a decoder structure, and an output layer. The encoder structure and the decoder structure are jump-connected. The encoder structure is sequentially provided with multiple encoders, and the decoder structure is sequentially provided with multiple decoders. In the encoder structure, the input of the first encoder is a brain image. The first encoder performs feature alignment and feature extraction on its input to obtain the output of the first encoder. Here, the introduction of the feature alignment operation can fully explore the local choroid plexus details and global choroid plexus structural information in different directions and scales in the brain image and ensure the feature format. The feature extraction operation can initially capture the complex structure and low-contrast choroid plexus areas. The input of each encoder other than the first encoder is the output of the previous encoder. The input of each encoder is adapted, feature extracted, and fused by the other encoders to obtain the output of the other encoders. Here, the adaptation operation adapts the feature dimensions of the input feature maps, effectively reducing computational effort while fully capturing and enhancing key features associated with the choroid plexus. Feature extraction further locates structurally complex, low-contrast choroid plexus regions, thereby fully capturing complex and subtle choroid plexus features. Feature fusion further fuses multi-dimensional choroid plexus features and passes them to the corresponding decoder, improving decoding performance. In the decoder architecture, the decoder decodes its input to produce a decoder output. The feature maps output by each encoder are high-level features with lower spatial resolution but a higher semantic level, while the feature maps output by each decoder are low-level features with higher spatial resolution and more local choroid plexus details. The decoding operation effectively integrates these high-level and low-level features, extracting more valuable choroid plexus features and thereby restoring the choroid plexus region in brain images. In the output layer, the output of the last decoder is used as a reference to generate the final choroid plexus contour map, which is a binary mask.

[0043] In summary, the embodiments of the present application implement segmentation of the choroid plexus region in brain images by introducing operations such as feature alignment, adaptation processing, feature extraction, and feature fusion into the encoder structure, and combining them with the decoding operations of the decoder structure. This can accurately capture small targets and complex choroid plexus features in brain images, enhance the feature extraction capability of the segmentation model, and thus effectively improve the segmentation accuracy of the choroid plexus in the brain.

[0044] The specific implementation of the above segmentation model will be further explained below.

[0045] With the success of large models such as the Segment Anything Model (SAM) in multi-scenario and multi-target segmentation tasks, their powerful generalization capabilities have attracted widespread attention. A few related technologies have attempted to introduce large models to the task of segmenting the choroid plexus. However, large models often require pre-training on large-scale general datasets, and prior knowledge of specific medical scenarios (such as the choroid plexus) is difficult to obtain, resulting in poor performance of large models in the choroid plexus segmentation task. Moreover, few related technologies have integrated large models with convolutional neural networks into the same segmentation model.

[0046] In response to the shortcomings of related technologies such as insufficient feature extraction capabilities and the difficulty in integrating large models and convolutional neural networks into the same segmentation model, the embodiments of the present application address the difficulties of the small size and complex structure of the choroid plexus region in brain images and propose a dual-branch fusion segmentation model that combines a large model with a convolutional neural network. The encoder of this model is divided into a large model branch and a convolutional network branch. The large model branch is configured with a pre-trained large model, which is adaptively modified through feature alignment and adaptation operations to fully utilize the extensive feature representation capabilities of the large model to enhance the feature extraction capabilities of the segmentation model. The type of large model can be flexibly set according to actual conditions. For example, the large model can be a SAM2 model, but is not limited to this. The convolutional network branch is implemented based on a convolutional neural network structure and is configured with traditional convolution operations and frequency convolution operations based on wavelet transforms. It can further enhance the extraction of details of the choroid plexus region and the edges of the choroid plexus region of small targets in brain images.

[0047] In simple terms, refer to Figure 2 ,exist Figure 2In the figure, "feature alignment" refers to the feature alignment layer of the first encoder, "convolution module" refers to the traditional convolution layer of the first encoder, and "frequency convolution" connected to the "convolution module" refers to the first frequency convolution layer of the first encoder. Together, "feature alignment," "convolution module," and "frequency convolution" constitute the first encoder. "SAM2 adjustment module" refers to the adaptive feature extraction module in encoders other than the first encoder. "Frequency convolution" running in parallel with the "SAM2 adjustment module" refers to the second frequency convolution layer of encoders other than the first encoder. "Feature fusion" connecting the "SAM2 adjustment module" and "frequency convolution" refers to the feature fusion layer of encoders other than the first encoder. A single "SAM2 adjustment module," a single "frequency convolution," and a single "feature fusion" together constitute an additional encoder. A brain image of original size 1×256×256 is first fed simultaneously into two parallel branches: the large model branch and the convolutional network branch. The convolutional network branch corresponds to the traditional convolutional layers and first-frequency convolutional layers of the first encoder, as well as the second-frequency convolutional layers of all encoders except the first. The large model branch corresponds to the feature alignment layer of the first encoder and the adapted feature extraction modules of all encoders except the first. Both the large model branch and the convolutional network branch output feature maps of different sizes: the large model branch outputs 144×64×64, 288×32×32, 576×16×16, and 1152×8×8, respectively; the convolutional network branch outputs 36×256×256, 72×128×128, 144×64×64, 288×32×32, 576×16×16, and 1152×8×8, respectively. In the encoder output layer, feature maps of the same scale are fused together by a fusion layer into a 64-channel feature map, which serves as the encoder output. Then, starting from the smallest feature map, the decoder is gradually upsampled to obtain a choroid plexus contour map of the same size as the brain image.

[0048] (1) Encoder structure.

[0049] In some embodiments, reference Figure 2 , the output of the first encoder may include the first output, the second output, and the third output of the first encoder; in the first encoder, feature alignment and feature extraction are performed on the brain image to obtain the output of the first encoder, which may include:

[0050] Perform feature alignment on the brain image to obtain the first output of the first encoder;

[0051] Perform a traditional convolution operation on the brain image to obtain the second output of the first encoder;

[0052] A frequency convolution operation is performed on the second output of the first encoder to obtain a third output of the first encoder.

[0053] In this embodiment, the first encoder may include a feature alignment layer, a traditional convolution layer, and a first frequency convolution layer. In the first encoder, first, the feature alignment layer is used to perform feature alignment on the brain image to obtain the first output of the first encoder. This can fully explore the local choroid plexus details and global choroid plexus structural information in different directions and scales in the brain image, and ensure the feature format. At the same time, the traditional convolution layer is used to perform a traditional convolution operation on the brain image to obtain the second output of the first encoder. This can preliminarily enhance the capture of the choroid plexus edges and textures and obtain feature information associated with the choroid plexus edges and textures. The convolution kernel of the traditional convolution layer can be flexibly set according to actual conditions. For example, the convolution kernel of the traditional convolution layer can be 3×3, but is not limited to this. Subsequently, the first frequency convolution layer is used to perform a frequency convolution operation on the second output of the first encoder to obtain the third output of the first encoder. This can preliminarily capture the choroid plexus area with complex structure and low contrast. Among them, the feature alignment layer belongs to the large model branch, that is, the first output of the first encoder belongs to the large model branch, while the traditional convolution layer and the first frequency convolution layer belong to the convolutional network branch, that is, the second output and third output of the first encoder belong to the convolutional network branch.

[0054] Compared to conventional encoders based solely on convolutional neural networks in related art, this embodiment introduces feature alignment and feature extraction operations within the first encoder. The feature extraction operations can be divided into traditional convolution and frequency convolution. The feature alignment operation fully exploits local choroid plexus details at different orientations and scales in brain images, as well as global choroid plexus structural information. Furthermore, given that the feature alignment layer directly connects with the subsequent construction of the large model branch, the feature alignment operation ensures that the feature format is compatible with the subsequent construction of the large model branch. Traditional convolution operations initially enhance the capture of choroid plexus edges and textures, obtaining feature information associated with these edges and textures. Subsequently, the traditional convolution operation is followed by a frequency convolution operation. This frequency convolution operation, based on these feature information, locates structurally complex, low-contrast choroid plexus regions. This achieves shallow feature extraction, enabling the initial extraction of valuable and diverse choroid plexus feature information, thereby enhancing the feature extraction capabilities of the segmentation model.

[0055] In some embodiments, reference Figure 3 In the feature alignment layer, the brain image is feature aligned to obtain the first output of the first encoder, which may include:

[0056] Perform multi-scale convolution operations on brain images to obtain multiple convolution feature maps;

[0057] Multiple convolutional feature maps are fused to obtain a composite feature map;

[0058] The composite feature map is compressed to obtain the first output of the first encoder.

[0059] In this implementation, a multi-scale feature alignment layer is designed for the input 1×256×256 single-channel brain image to fully exploit local choroid plexus details at different orientations and scales within the brain image, as well as global choroid plexus structural information. Furthermore, considering that the feature alignment layer belongs to the large model branch and needs to be integrated with the subsequent construction of the large model branch, the feature alignment layer provides a suitable feature format for subsequent integration with the pre-trained large model, further leveraging the capabilities of the large model.

[0060] Specifically, first, brain images are processed in parallel through convolutional layers of multiple sizes to capture diverse choroid plexus information. The convolution kernels of these convolutional layers are 1×N, N×1, 3×3, 5×5, and 7×7, respectively. The use of 1×N and N×1 convolution kernels respectively strengthens the detection of horizontal edges, vertical edges, and linear structures of the choroid plexus region, which helps to accurately extract choroid plexus edge features in different directions; the use of 3×3, 5×5, and 7×7 convolution kernels can cover a larger receptive field, thereby more sensitively capturing the texture, structural changes, and global contextual information of the choroid plexus region. In the implementation process, this embodiment makes full use of the above-mentioned convolution kernels of various sizes, and uses these convolution kernels to perform multi-scale convolution operations on brain images to obtain multiple convolution feature maps. These convolution feature maps encode brain image information from different angles and sizes, forming a rich multi-scale choroid plexus feature expression.

[0061] Then, these convolution feature maps are fused into a composite feature map. The fusion process can be flexibly set according to the actual situation, for example, the convolution feature maps obtained by different convolution kernels are spliced ​​in the channel dimension to generate a composite feature map, but it is not limited to this. Taking into account that the number of channels of the composite feature map obtained by fusion is large, in order to make the output features consistent with the input required by the subsequent pre-trained large model, this embodiment further reconstructs the number of channels of the composite feature map so that the composite feature map is compressed to the first output of the first encoder. The compression process can be flexibly set according to the actual situation, for example, a 1×1 convolution kernel is used to compress the composite feature map to the first output of the first encoder, but it is not limited to this. In this way, the multi-scale features of the choroid plexus region are retained, and the input format requirements of downstream tasks are met.

[0062] For example, in the feature alignment layer, the N value in its multi-scale convolution operation can be 3 and 5, respectively. This results in seven convolution kernels: 7×7, 3×3, 5×5, 1×3, 3×1, 1×5, and 5×1. Each convolution kernel is configured with 8 filters, resulting in a total of 56 convolution feature maps. Next, these convolution feature maps obtained by different convolution kernels are concatenated along the channel dimension to generate a composite feature map with 56 channels. Finally, the composite feature map is reconstructed using a 1×1 convolution kernel to compress or expand the 56 channels into 3 channels while maintaining the spatial size of 256×256, thus obtaining the first output of the first encoder.

[0063] Here, multi-scale and multi-shape convolutions are used to convolve the original brain image with different receptive fields, and the resulting feature maps are then fused and compressed. This operation, on the one hand, takes into account that the choroid plexus is typically small and easily obscured by complex backgrounds. The feature alignment layer can initially capture the underlying features of the choroid plexus, such as edges, gradients, and local saliency, and accurately locate this fine-grained information. This provides the necessary local context for precise localization of small targets, allowing the large model to accurately locate the choroid plexus region and amplify features early on, helping to better utilize the large model's capabilities. On the other hand, the training data for the large model mostly comes from natural images, making it difficult to align features with medical brain images. The output of the feature alignment layer is connected to the subsequent construction of the large model branches. The feature alignment layer can reduce the distribution differences between the "large model pre-training data domain" and the "medical imaging data domain," allowing the large model features to more naturally "align" with the texture distribution of medical images, laying a solid foundation for subsequent processing.

[0064] In some embodiments, reference Figure 2 The output of the first encoder may include the first output and the third output of the first encoder, and the output of the other encoders may include the first output and the second output of the other encoders; in the encoders other than the first encoder, the inputs of the other encoders are adapted, feature extracted, and fused to obtain the outputs of the other encoders, which may include:

[0065] Performing adaptation processing and feature extraction on the first inputs of the other encoders to obtain the first outputs of the other encoders; wherein the first input of the second encoder is the first output of the first encoder, and the first input of the encoder after the second encoder is the first output of the previous encoder;

[0066] Performing a frequency convolution operation on the second input of the other encoders to obtain the second output of the other encoders; wherein the second input of the second encoder is the third output of the first encoder, and the second input of the encoder after the second encoder is the second output of the previous encoder.

[0067] In this embodiment, for a single other encoder, the input of the other encoder may include a first input and a second input, and the output of the other encoder may include a first output and a second output, wherein the first input and the first output correspond to the large model branch, and the second input and the second output correspond to the convolutional network branch. Considering that the large model branch of each encoder is connected sequentially, and the convolutional network branch of each encoder is also connected sequentially, the first input of the second encoder is the first output of the first encoder, the first input of the encoder after the second encoder is the first output of the previous encoder, the second input of the second encoder is the third output of the first encoder, and the second input of the encoder after the second encoder is the second output of the previous encoder.

[0068] The other encoder may include an adaptive feature extraction module and a second frequency convolution layer. In the other encoder, the adaptive feature extraction module first performs adaptive processing and feature extraction on the first input of the other encoder to obtain the first output of the other encoder. This effectively reduces the amount of computation while comprehensively capturing and enhancing key features associated with the choroid plexus. Simultaneously, the second frequency convolution layer performs a frequency convolution operation on the second input of the other encoder to obtain the second output of the other encoder. This further locates choroid plexus regions with complex structures and low contrast, thereby comprehensively capturing complex and subtle choroid plexus features. The adaptive feature extraction module belongs to the large model branch, while the second frequency convolution layer belongs to the convolutional network branch. This means that the first output of the other encoder contains richer global choroid plexus semantic information, while the second output of the other encoder contains local details and edge textures of the choroid plexus.

[0069] Compared to conventional encoders based solely on convolutional neural networks (CNNs) in related art, this embodiment introduces adaptive feature extraction and frequency convolution operations in all encoders except the first one. The adaptive feature extraction operation effectively reduces computational complexity while comprehensively capturing and enhancing key features associated with the choroid plexus, thereby extracting richer global semantic information about the choroid plexus. The frequency convolution operation further locates structurally complex, low-contrast choroid plexus regions, thereby comprehensively capturing complex and subtle choroid plexus features, such as local details and edge textures. This achieves deep feature extraction, further enhancing the segmentation model's feature extraction capabilities.

[0070] In some embodiments, reference Figure 4 , Figure 4 The "adaptation block" in the above is an adaptation module. In the adaptation feature extraction module, the first input of the other encoder is adapted and feature extracted to obtain the first output of the other encoder, which may include:

[0071] Performing adaptation processing based on downsampling and upsampling on the first input of the other encoders to obtain an adapted feature map;

[0072] Fusing the adapted feature map with the first input of other encoders to obtain a feature map to be processed;

[0073] Multiple large-model-based feature extractions are performed on the feature map to be processed to obtain the first output of other encoders.

[0074] In this embodiment, for a single other encoder, the above-mentioned adaptive feature extraction module includes an adaptation module, an adaptive fusion layer, and multiple sequentially connected feature extraction modules. Among them, the feature extraction module is implemented based on a large model, and its type can be flexibly set according to actual conditions. For example, the feature extraction module can be a hierarchical visual transformer in the SAM2 model. It retains the core of the transformer (i.e., self-attention and forward propagation), while making special image-specific modifications in the input and hierarchy (i.e., tile segmentation and fusion downsampling), so that the architecture performs well in image tasks.

[0075] To maintain the general visual features learned by the large model during pre-training while effectively fine-tuning brain images (especially subtle structures such as the choroid plexus), this implementation design incorporates a single-ended adaptation module. This module performs downsampling and upsampling adaptation processing on the first input of the other encoders to obtain an adapted feature map. The core idea is to map the feature dimension from the original dimension (dim) to a lower dimension (e.g., 16 or 32) and then back to the original dimension, thereby reducing computational effort and enhancing key choroid plexus features. This allows these features to be adapted to subsequent feature extraction modules based on the large model.

[0076] Specifically, in the adaptation module, first, considering that the feature sequences output by large models are typically of higher dimensionality, the first input of the other encoders is downsampled to obtain a downsampled feature map. This downsampling operation compresses the original feature dimension (dim) to a lower dimension (e.g., 16 or 32) through components such as feedforward neural networks or linear mapping layers. This operation maximizes the preservation of key choroid plexus information while effectively reducing the computational resources required for subsequent processing. Then, to highlight the responses of subtle structures such as the choroid plexus in brain images, the downsampled feature map is activated to obtain an activated feature map. Here, the reduced-dimensional features are input into activation functions such as Sigmoid, ReLU, or LeakyReLU, using nonlinear transformations to enhance the expressive power of key choroid plexus features. Afterwards, to ensure that the final output is dimensionally aligned with the original large model features, the activated feature map is upsampled to obtain an upsampled feature map. Here, the upsampling operation is introduced to increase the low-dimensionality (such as 16 or 32) to the original high-dimensional space (dim) through components such as linear layers or fully connected layers. This ensures smooth docking with subsequent modules (such as the decoder). Finally, the upsampled feature map is activated to obtain an adapted feature map. Here, the upscaled feature map is again processed by activation functions such as Sigmoid or ReLU to further stabilize and sparse the distribution of choroid plexus features, while suppressing noise that may be introduced during the upsampling process. This ensures that the output features maintain the global visual prior while focusing more on the local details of the choroid plexus region.

[0077] After the adaptation process is complete, the adapted feature map is fused with the first input of the other encoders in the adaptation fusion layer to produce the processed feature map. Here, the adapted feature map represents the adjusted choroid plexus feature information, while the first input of the other encoders represents the shallow raw information. Fusing the adjusted feature information with the shallow raw information effectively increases the diversity of choroid plexus features and improves the segmentation model's ability to represent complex choroid plexus features. The fusion operation can be flexibly configured based on actual conditions. For example, fusion can be performed by splicing along the channel dimension, but this is not limited to this.

[0078] Subsequently, the feature map to be processed is input into the first feature extraction module and sequentially passes through multiple feature extraction modules to perform further feature extraction operations, ultimately obtaining the first output of other encoders, thereby further capturing key features associated with the choroid plexus. Simply put, each feature extraction module can include two normalization layers, an attention layer, and a linear layer. The input of the feature extraction module is processed by the first normalization layer and then enters the attention layer. The feature map processed by the attention layer is spliced ​​with the input of the feature extraction module in the channel dimension. The spliced ​​result is input into the second normalization layer for normalization processing. The spliced ​​result after normalization is then processed by the linear layer. The output of the linear layer and the spliced ​​result are spliced ​​in the channel dimension to obtain the output of the feature extraction module.

[0079] Optionally, when training the segmentation model, the parameters of the adaptation module are trainable, but the parameters of each feature extraction module are not trainable.

[0080] The adaptive feature extraction operation in this embodiment can be divided into an adaptation operation and a feature extraction operation. The feature extraction operation is implemented based on the large model. The adaptation operation maintains the general visual features learned by the large model during pre-training while effectively fine-tuning the brain image (especially subtle structures such as the choroid plexus). This reduces computational effort, enhances key choroid plexus features, and enables the adaptation of choroid plexus features to subsequent feature extraction modules based on the large model. The feature extraction operation further captures key features associated with the choroid plexus. This achieves deep feature extraction, effectively ensuring the segmentation model's ability to extract global choroid plexus semantic information.

[0081] In some embodiments, reference Figure 5 and Figure 6 In the first frequency convolution layer or the second frequency convolution layer, the frequency convolution operation may include:

[0082] Perform wavelet transform on the input of the frequency convolution operation to obtain multiple sub-band feature maps;

[0083] Reconstructing multiple sub-band feature maps to obtain a reconstructed feature map;

[0084] Perform multiple traditional convolution operations on the reconstructed feature map to obtain the output of the frequency convolution operation.

[0085] In this embodiment, for a single encoder, in the frequency convolution operation, first, the input of the frequency convolution operation is subjected to Haar wavelet transform processing to obtain multiple sub-band feature maps. Haar wavelet transform is a simple and efficient discrete wavelet transform method that captures the low-frequency and high-frequency features of an image by performing averaging and differential operations on image data. For a two-dimensional feature map, Haar wavelet transform can separate the feature map into four frequency band sub-maps, i.e., four sub-band feature maps, each of which represents different frequency band information, as shown in the following formula (1):

[0086]

[0087] In formula (1), I(x, y) represents the pixel value of the feature map at the coordinate (x, y); I(x+1, y) represents the pixel value of the feature map at the coordinate (x+1, y); I(x, y+1) represents the pixel value of the feature map at the coordinate (x, y+1); I(x+1, y+1) represents the pixel value of the feature map at the coordinate (x+1, y+1); LL represents the low-frequency approximation subband; LH represents the vertical detail subband; HL represents the horizontal detail subband; HH represents the diagonal detail subband.

[0088] A two-dimensional Haar wavelet transform is performed on the input of the current feature layer (i.e., the input of the frequency convolution operation). The core idea is to perform addition and subtraction operations on each 2×2 region in the input feature map to obtain four subbands: a low-frequency approximation subband (LL), a vertical detail subband (LH), a horizontal detail subband (HL), and a high-frequency detail subband (HH). Since each 2×2 pixel block is remapped into four subbands, the spatial resolution is halved (for example, from 256×256 to 128×128), while the number of channels is increased to 4 times the original value, thus ensuring that both low-frequency and high-frequency information are fully preserved. This transform is also called lossless downsampling because it reconstructs features through the inverse operation.

[0089] Next, in order to keep consistent with the features of the subsequent large model branch, the four expanded sub-band feature maps need to be reconstructed. Specifically, the four sub-band feature maps all represent complete image information. The four sub-band feature maps are spliced ​​in the channel dimension to obtain a feature map with 4 channels. The number of channels of the feature map is then reconstructed through a 1×1 convolution kernel to restore 4 times the original number of channels to the preset original number of channels (for example, keep the same number of channels as the output of the large model branch), thereby obtaining a reconstructed feature map. The 1×1 convolution kernel can be regarded as a linear combination of the multi-channel vectors at each pixel position. It does not change the spatial size, but can effectively compress or rearrange the channel information.

[0090] After completing the channel reconstruction, the secondary convolution part is entered, and multiple continuous traditional convolution operations are performed on the reconstructed feature map to obtain the output of the frequency convolution operation. In this way, the size of the feature map can be reduced while further enhancing the extraction of edge and texture features of tiny structures such as the choroid plexus. Among them, the number of traditional convolution operations can be flexibly set according to actual conditions. For example, the traditional convolution operation can be 2 times, but it is not limited to this. In addition, the traditional convolution operation can be flexibly set according to actual conditions. For example, a 3×3 convolution kernel is used for convolution operation, and normalization processing (such as BatchNorm, LayerNorm, etc.) and activation function (such as ReLU, etc.) are performed in sequence after convolution. Here, the 3×3 convolution kernel has a good balance in receptive field, smoothness and detail capture. After combining normalization and activation function, it can suppress noise and highlight the response of key areas.

[0091] It should be understood that, for the first frequency convolution layer, the input of its frequency convolution operation is the second output of the first encoder, and the output of its frequency convolution operation is the third output of the first encoder; for the second frequency convolution layer, the input of its frequency convolution operation is the second input of the other encoders, and the output of its frequency convolution operation is the second output of the other encoders.

[0092] Here, in order to enhance the recognition of small targets in medical images, especially to capture finer-grained textures of regions with complex structures and low contrast, such as the choroid plexus, compared with the encoder in the related art that only relies on convolutional neural networks, this embodiment introduces wavelet transform in the convolution operation to split the original feature map into multiple frequency bands, highlighting the edge information and details of the choroid plexus while avoiding the loss of choroid plexus features. The feature maps of each frequency band are then integrated into a reconstructed feature map through reconstruction processing, so that the choroid plexus information is maximized. Finally, after multiple traditional convolution operations, the choroid plexus regions with complex structures and low contrast are further captured, thereby enhancing the choroid plexus feature representation.

[0093] In some embodiments, reference Figure 7 , the output of the other encoders may include the first output, the second output, and the third output of the other encoders; in the above-mentioned other encoders other than the first encoder, the input of the other encoders is adapted, the feature is extracted, and the feature is fused to obtain the output of the other encoders, which may include;

[0094] Concatenate the first and second outputs of other encoders to obtain a concatenated feature map;

[0095] Perform channel compression on the spliced ​​feature map to obtain a compressed feature map;

[0096] The compressed feature map is activated to obtain the third output of the other encoders.

[0097] In this embodiment, the segmentation model is configured with two parallel branches: a large model branch and a convolutional network branch. In addition to the first encoder, other encoders can also include a feature fusion layer. The feature fusion layer is designed for multi-layer and multi-stage fusion. When there are both large model feature outputs and convolutional network feature outputs, the features of the two branches are fused, and then the channel information is integrated and the semantic distinction is enhanced through 1×1 convolution and activation functions.

[0098] Specifically, for a single other encoder, the first output of the other encoder is obtained from the large model branch, and the second output is obtained from the convolutional network branch. These two have the same spatial dimensions (e.g., H×W) but different characteristics in terms of the number of channels (C): the large model branch often contains richer global semantic information about the choroid plexus, while the convolutional network branch is better at extracting local details and edge textures of the choroid plexus. To fuse and complement these two different-scale choroid plexus feature information, the first and second outputs of the other encoder are first concatenated along the channel dimension, resulting in a concatenated feature map with twice the number of channels (2C). Next, to avoid the additional computational burden of redundant channels and make the fused feature map more compact and efficient, a 1×1 convolution kernel is used to compress the number of channels of the concatenated feature map to the same size (C) as the single-branch output, resulting in a compressed feature map. In this way, the 1×1 convolution kernel can linearly combine the multi-channel vectors at each pixel position without changing the spatial scale, effectively fusing the global contextual semantics from the large model branch with the local details of the convolutional network branch. This results in a choroid plexus feature representation that combines both macroscopic and microscopic information while being more reasonable in terms of channel number and computational complexity. Finally, activation functions such as ReLU and Leaky ReLU are used to activate the compressed feature map to obtain the third output of the other encoders. This activation process further stabilizes and sparsifies the distribution of the choroid plexus feature representation.

[0099] In this way, this embodiment complementarily fuses the rich global choroid plexus semantic information with the local details and edge texture of the choroid plexus in the feature fusion layer of other encoders, so that the third output of other encoders inherits the general visual prior learned by the large model on large-scale data, and can utilize the fine choroid plexus features brought by the convolutional network branch. The third output of other encoders will be passed to the decoder structure together with the second and third outputs of the first encoder, which can further improve the effect of the decoding operation, thereby providing richer and more balanced feature support for downstream tasks (such as segmentation, detection or classification), ensuring the excellent performance of the segmentation model in the choroid plexus segmentation task.

[0100] (2) Decoder structure.

[0101] In some embodiments, reference Figure 2 , the above-mentioned first decoder is jump-connected with the last encoder and the penultimate encoder respectively, and the input of the first decoder may include the third output of the last encoder and the third output of the penultimate encoder; the last decoder and the penultimate decoder are both jump-connected with the first encoder, and the input of the last decoder may include the second output of the first encoder and the output of the penultimate decoder, and the input of the penultimate decoder may include the third output of the first encoder and the output of the penultimate decoder; the remaining decoders except the first decoder, the last decoder and the penultimate decoder, and the remaining encoders except the first encoder, the last encoder and the penultimate encoder have one-to-one correspondence and jump connection; the remaining decoders except the first decoder, the last decoder and the penultimate decoder include the third output of the encoder corresponding to the remaining decoders and the output of the previous decoder of the remaining decoders.

[0102] In this embodiment, if Figure 2 As shown, with the brain image as the top, in the encoder structure, from the top to the bottom, there are the first encoder to the last encoder, and in the decoder structure, from the top to the bottom, there are the last decoder to the first decoder.

[0103] The first decoder has a skip connection to the last and penultimate encoders, respectively. This means its inputs include the third output of the last and penultimate encoders, and the third outputs of the encoders other than the first encoder are the outputs of the feature fusion layers of the encoders other than the first. The last decoder has a skip connection to the first encoder, meaning its inputs include the second output of the first encoder and the output of the penultimate decoder, and the second output of the first encoder is the output of the traditional convolutional layer of the first encoder. The penultimate decoder has a skip connection to the first encoder, meaning its inputs include the third output of the first encoder and the output of the third-to-last decoder, and the third output of the first encoder is the output of the first frequency convolutional layer of the first encoder. The remaining decoders, excluding the first, last, and penultimate decoders, have a one-to-one correspondence and skip connection with the remaining encoders, excluding the first, last, and penultimate encoders. This means their inputs include the third output of the encoder corresponding to the remaining decoder and the output of the previous decoder of the remaining decoder. The third output of the encoders other than the first encoder is the output of the feature fusion layer of the encoders other than the first encoder. This forms a special skip connection.

[0104] For example, if Figure 2As shown, there are five encoders and five decoders. With the brain image at the top, the encoder structure is organized from the first to the fifth encoder, while the decoder structure is organized from the fifth to the first decoder. The first decoder has a skip connection to the fourth and fifth encoders, the second decoder has a skip connection to the third decoder, the third decoder has a skip connection to the second encoder, the fourth decoder has a skip connection to the first encoder, and the fifth decoder has a skip connection to the first encoder.

[0105] In some embodiments, reference Figure 2 and Figure 8 In the above decoder, decoding the decoder input to obtain the decoder output may include:

[0106] Fuse the decoder input to obtain a fused feature map;

[0107] Perform multiple traditional convolution operations on the fused feature map to obtain the output of the decoder.

[0108] In this embodiment, in the decoder structure, while receiving the fused high-dimensional feature maps, the decoder also utilizes two types of scale information produced by different layers in the segmentation model: high-level features with lower spatial resolution but higher semantic levels, produced by the corresponding encoder (as the number of encoder layers increases, the resolution gradually decreases, and the semantic information gradually transforms into abstract high-level semantics); and low-level features produced by the previous decoder, with higher spatial resolution and more local details. By combining these two types of features, a segmentation result consistent with the original image resolution is restored at the final output stage. Accordingly, each decoder can include a fusion input layer and a convolutional output layer, where the input of the fusion input layer includes at least two feature maps. First, the decoder input is fused by the fusion input layer to fuse the feature maps of different scales to obtain a fused feature map. Then, the convolutional input layer performs multiple traditional convolution operations on the fused feature map to reintegrate and extract more valuable choroid plexus features, ultimately obtaining the decoder output.

[0109] Specifically, the first decoder receives the third output of the last encoder and the third output of the penultimate encoder. In the fusion input layer of the first decoder, the third output of the last encoder is upsampled to restore its third output size. The upsampled third output of the last encoder is then concatenated with the third output of the penultimate encoder on the channel level to produce the fused feature map of the first decoder. This combines the choroid plexus feature information from the two lowest-level encoders and passes it to the first decoder, thereby capturing more feature information associated with the choroid plexus region. Subsequently, in the convolutional output layer of the first decoder, multiple consecutive traditional convolution operations are performed on the fused feature map of the first decoder to produce the output of the first decoder, aiming to initially restore the choroid plexus region.

[0110] Here, the first decoder is the bottom-level decoder, which relies on the choroid plexus feature information from different encoders to restore the choroid plexus area. In this process, the information from different encoders needs to be reintegrated through convolution to extract the more valuable choroid plexus features. In multiple convolution operations, the number of channels is controlled while restoring the choroid plexus area, which helps to keep the number of feature layers matching throughout the decoding process, so that the bottom-level decoder can generate the initial local details.

[0111] The second to third-to-last decoders are defined as the rest decoders. The rest decoders receive the third output of the encoder corresponding to the rest decoders and the output of the previous decoder. In the fusion input layer of the rest encoder, the output of the previous decoder is upsampled to align its spatial size with the third output of the encoder, ensuring smooth concatenation of the two feature maps along the channel dimension. The upsampled output of the previous decoder is then concatenated with the third output of the encoder corresponding to the rest decoder along the channel dimension to produce the fused feature map of the rest decoder. This aims to convey the encoder's choroid plexus feature information to the decoder and fully combine the decoder's choroid plexus feature information with the encoder's choroid plexus feature information to capture more feature information associated with the choroid plexus region. Subsequently, in the convolutional output layer of the rest decoder, multiple consecutive traditional convolution operations are performed on the fused feature map of the rest decoder to produce the output of the rest decoder, aiming to gradually restore the choroid plexus region.

[0112] Here, the second to third-to-last decoders are mid-level decoders, tasked with refining local details of the choroid plexus region. The mid-level decoders receive both high-level semantics and local details from the previous decoder layer. Therefore, they can rely on choroid plexus feature information from the corresponding encoder and the previous decoder layer to restore the choroid plexus region. During this process, information from the decoder and encoder is reintegrated through convolution to extract the more valuable choroid plexus features. Multiple convolution operations simultaneously restore the choroid plexus region while controlling the number of channels, helping to maintain a consistent number of feature layers throughout the decoding process. This allows the mid-level encoder to gradually refine local details of the choroid plexus region, leading to better restoration of the choroid plexus region.

[0113] The penultimate decoder receives the third output from the first encoder and the output from the third-to-last decoder. In the penultimate decoder's fusion input layer, the penultimate decoder's output is upsampled to align its spatial dimensions with the first encoder's third output, ensuring smooth channel-wise concatenation of the two feature maps. The upsampled penultimate decoder output is then concatenated with the first encoder's third output channel-wise to produce the penultimate decoder's fused feature map. This aims to convey the encoder's complex, low-contrast choroid plexus features to the decoder, effectively combining the decoder's choroid plexus features with the previously described complex, low-contrast choroid plexus features to capture more detailed feature information associated with the choroid plexus region. Subsequently, in the penultimate decoder's convolutional output layer, multiple consecutive traditional convolution operations are performed on the penultimate decoder's fused feature map to produce the penultimate decoder's output, further restoring the choroid plexus region.

[0114] Here, the penultimate decoder is the top decoder. Unlike the decoders in the middle layers, the top decoder is near the end of restoration. Its task is no longer to refine local details of the choroid plexus region, but to fuse high-level semantics and local details (i.e., multi-scale feature fusion), thereby balancing the two and generating precise choroid plexus boundaries. To achieve multi-scale feature fusion and improve its effectiveness, the penultimate decoder receives information from the previous decoder and from frequency convolution operations. Frequency convolution operations accurately locate complex, low-contrast choroid plexus regions, providing a preliminary complement to spatial detail. By fusing these two types of information through convolution, more detailed choroid plexus features can be extracted, especially in regions with complex structures, low contrast, and small objects. Multiple convolution operations simultaneously restore the choroid plexus region while controlling the number of channels, helping to maintain a consistent number of feature layers throughout the decoding process. This allows the top decoder to fuse high-level semantics and local details, achieving a balance between semantic information and spatial detail, and generating more precise choroid plexus boundaries.

[0115] The final decoder receives the second output from the first encoder and the output of the penultimate decoder. In the final decoder's fusion input layer, the penultimate decoder's output is upsampled to align its spatial dimensions with the first encoder's second output, ensuring smooth concatenation of the two feature maps along the channel dimension. The upsampled penultimate decoder output and the first encoder's second output are then concatenated along the channels to produce the final decoder's fused feature map. This aims to convey feature information from the encoder that is moderately or highly correlated with the choroid plexus region to the decoder, and to fully combine the decoder's choroid plexus feature information with the aforementioned feature information to ultimately capture the most detailed feature information associated with the choroid plexus region. Subsequently, in the final decoder's convolutional output layer, multiple consecutive traditional convolution operations are performed on the final decoder's fused feature map to produce the final decoder's output, ultimately restoring the choroid plexus contour map.

[0116] Here, the last decoder is the top-level decoder, whose task is to fuse high-level semantics and local details (i.e., multi-scale feature fusion), thereby balancing the relationship between the two and promoting the generation of accurate choroid plexus boundaries. To achieve multi-scale feature fusion and improve the fusion effect, the last decoder receives information from the previous decoder and from traditional convolution operations. Traditional convolution operations can initially capture shallow features associated with choroid plexus texture and edges, which can further compensate for spatial details. By fusing these two types of information through convolution, the most detailed choroid plexus features can be extracted. Multiple convolution operations restore the choroid plexus area while controlling the number of channels, helping to ensure that the number of feature layers matches throughout the decoding process. This allows the top-level decoder to fully fuse high-level semantics and local details to generate the final choroid plexus boundary.

[0117] Optionally, the upsampling operation in the fusion input layer of the decoder can be flexibly set according to actual conditions. For example, the upsampling operation can be linear interpolation upsampling, but is not limited thereto.

[0118] Optionally, the number of traditional convolution operations can be flexibly set according to actual conditions, for example, the traditional convolution operation can be performed twice, but is not limited thereto. In addition, the traditional convolution operation can be flexibly set according to actual conditions, for example, using a 3×3 convolution kernel for convolution operation, and performing normalization processing (such as BatchNorm, LayerNorm, etc.) and activation function (such as ReLU, etc.) after convolution.

[0119] In summary, considering that most related fully convolutional neural networks only have simple cascades and dense connections, which makes it difficult to capture local details of the choroid plexus, resulting in low choroid plexus segmentation accuracy, in this embodiment, the feature fusion operation of the encoder structure can extract more abstract choroid plexus semantics, while the decoder structure can extract choroid plexus details such as color, texture, and edges. Through special skip connections between the decoder and encoder structures and the decoding processing of the decoder structure, the choroid plexus semantics can be transmitted to the decoder structure, and the choroid plexus semantics and choroid plexus details can be combined to accurately capture small objects and complex choroid plexus features in brain images, enhance the feature extraction capability of the segmentation model, and effectively improve the segmentation accuracy of the brain choroid plexus.

[0120] (3) Output layer.

[0121] In some embodiments, reference Figure 2 , Figure 2 The "output" in the above is the output layer. In the above output layer, the choroid plexus contour map is obtained according to the output of the last decoder, which can include:

[0122] Perform activation processing on the output of the last decoder to obtain an activation feature map;

[0123] For the intensity value of each pixel in the activation feature map, if the intensity value of the current pixel is greater than the preset intensity value, the value of the current pixel is determined as the first value; otherwise, the value of the current pixel is determined as the second value, thereby obtaining the choroid plexus contour map.

[0124] In this embodiment, an output layer is connected after the last decoder. In this output layer, activation processing is performed on the output of the last decoder using functions such as Sigmoid and Softmax. The purpose is to compress the discrete intensity values ​​of each pixel in the output of the last decoder to [0, 1], thereby obtaining an activation feature map. Subsequently, for each pixel intensity value in the activation feature map, a determination is made as to whether the intensity value of the current pixel is greater than a preset intensity value. If so, it indicates that the pixel belongs to the choroid plexus region, and the value of the current pixel is determined as a first value. Otherwise, it indicates that the pixel belongs to the background region, and the value of the current pixel is determined as a second value. By traversing all pixels in the activation feature map, the values ​​of all pixels can be determined, and then a choroid plexus outline image represented by a binary mask image can be obtained. This is because the essence of the choroid plexus segmentation task is to detect the choroid plexus region. It is only necessary to determine whether a certain pixel location is a choroid plexus. Therefore, the output layer is provided after decoding to convert the discrete output feature map into a binary label image.

[0125] Optionally, the preset intensity value can be flexibly set according to actual conditions, for example, the preset intensity value can be 0.5, but is not limited thereto. In addition, the first value and the second value can also be flexibly set according to actual conditions, for example, the first value can be 1, and the second value can be 0, but is not limited thereto.

[0126] (4) Training of segmentation model.

[0127] In some embodiments, the above-mentioned brain choroid plexus segmentation method may further include:

[0128] Acquire a plurality of brain image samples, each brain image sample having a corresponding choroid plexus contour map label;

[0129] Preprocessing the multiple brain image samples to obtain multiple preprocessed brain image samples;

[0130] The segmentation model is trained using multiple preprocessed brain image samples.

[0131] In this embodiment, first, brain image samples of different sizes are collected, and the choroid plexus contour map labels of each brain image sample are manually outlined. For example, a clinician manually outlines the choroid plexus contour map using professional software such as 3D Slicer. The doctor relies on rich clinical experience to develop annotation standards to ensure that the boundaries of the choroid plexus area are accurately outlined, and cross-validation is used to further reduce subjective bias. At the same time, all brain image samples are quality checked, and low-quality samples caused by noise, motion artifacts, or annotation errors are eliminated. Then, in order to ensure the quality of the brain image samples, each original brain image sample needs to be preprocessed. Specifically, all brain image samples are adjusted to a fixed size of 256×256 by padding or resampling, and all brain image samples are normalized. In addition, data augmentation techniques such as rotation, translation, and mirroring are combined to implement preprocessing operations. In this way, multiple preprocessed brain image samples are obtained, thereby providing high-quality input for the model and improving the robustness and generalization ability of the model. Finally, multiple preprocessed brain image samples are input into the segmentation model to train the segmentation model.

[0132] In some embodiments, the above-mentioned brain choroid plexus segmentation method may further include:

[0133] The Dice loss function is used to backpropagate the segmentation model.

[0134] In this embodiment, the Dice loss function is used as the loss function during the entire model training process. When the prediction result completely matches the true annotation, the Dice coefficient is 1, and the corresponding Dice loss function is 0; otherwise, the larger the Dice loss function, the lower the overlap between the prediction result and the annotation. The trainable parameters in the segmentation model (such as the feature alignment layer, the adaptation module, and the convolution weights in the convolution branch) are continuously adjusted through the backpropagation algorithm, while keeping the preprocessing weights of the large model partially frozen, to achieve a balance between generalization ability and targeted fine-tuning. Strategies such as learning rate decay and early stopping can also be used during training, and cross-validation and independent test sets are used to evaluate the model performance to ensure stability and robustness on real clinical data. Among them, the Dice loss function satisfies the following formula (2):

[0135]

[0136] In formula (2), represents the Dice loss function; DSC(p, g) represents the Dice coefficient; p i represents the probability of the i-th pixel (or voxel) in the prediction result; g iRepresents the value of the i-th pixel (or voxel) in the true result; ∈ is a small positive number used to prevent the denominator from being zero.

[0137] (5) Description of the principle of the segmentation model.

[0138] Reference Figure 2 , a 1×256×256 single-channel MRI brain image is simultaneously input into the large model branch and the convolution branch of the segmentation model, two parallel branches. Then, the features output by them are fused at multiple scale levels, and finally the decoder gradually upsamples and restores the choroid plexus segmentation result of the same size as the original image.

[0139] The following uses an application scenario to illustrate the overall data flow of the segmentation model.

[0140] In the segmentation model, there are five encoders and five decoders. The output layer is located after the last decoder, with the brain image at the top. The encoder structure is organized from the top to the fifth encoder, and the decoder structure is organized from the fifth decoder to the first decoder. The first decoder has a skip connection to the fourth and fifth encoders, the second decoder has a skip connection to the third decoder, the third decoder has a skip connection to the second encoder, the fourth decoder has a skip connection to the first encoder, and the fifth decoder has a skip connection to the first encoder.

[0141] Each encoder is equipped with a large model branch and a convolutional network branch, and the large model branches of each encoder are connected in sequence, and the convolutional network branches of each encoder are connected in sequence. With the brain image as the top, the structure of the large model branch from the top to the bottom is the feature alignment layer of the first encoder, the adaptive feature extraction module of the second encoder to the fifth encoder; the structure of the convolutional network branch from the top to the bottom is the traditional convolution layer and the first frequency convolution layer of the first encoder, and the second frequency convolution layer of the second encoder to the fifth encoder. In addition, the encoders other than the first encoder are also equipped with a feature fusion layer, and the output of the feature fusion layer of each encoder is passed to the corresponding decoder. Moreover, the output of the traditional convolution layer and the first frequency convolution of the first encoder is also passed to the corresponding decoder.

[0142] In the large model branch, before the 1×256×256 brain image formally enters the adaptation feature extraction module of other encoders except the first encoder, such as Figure 3As shown in the figure, in the feature alignment layer of the first encoder, the brain image is first convolved in parallel using multi-scale and multi-shape convolution kernels (e.g., 1×n, n×1, 3×3, 5×5, 7×7) to capture edge and texture information in different directions and receptive fields. After these features are concatenated in the channel dimension, a 1×1 convolution kernel is used to compress the number of channels to 3 to match the input format of the adaptive feature extraction module. The first output of the first encoder can be obtained through the feature alignment layer, which will be passed to the adaptive feature extraction module of the second encoder and serve as the first input of the second encoder.

[0143] The convolutional network branch focuses more on extracting local details of the choroid plexus area. It directly performs pure convolution operations on the 1×256×256 brain image through the traditional convolution layer of the first encoder, such as several 3×3 convolutions. Each convolution is followed by normalization and activation functions to output a feature map that still maintains a spatial size of 256×256 as the second output of the first encoder, which will be passed to the fifth decoder and used as one of the inputs of the fifth decoder. In this way, the capture of the edges and textures of the choroid plexus can be preliminarily enhanced without downsampling. Subsequently, in the first frequency convolution layer of the first encoder, a frequency convolution operation based on the Haar wavelet transform is performed on the second output of the first encoder, such as Figure 5 and Figure 6 As shown in the figure, the specific operation is as follows: each 2×2 region is decomposed into four sub-bands (LL, LH, HL, HH) through Haar wavelet transform, thereby halving the spatial resolution but expanding the number of channels to 4 times the original. Then, a 1×1 convolution kernel is used to restore the number of channels to the required target value. Several 3×3 convolutions (also with normalization and activation functions) are then performed to further enhance the choroid plexus feature representation, thereby obtaining the third output of the first encoder, which will be passed to the second frequency convolution layer of the second encoder and serve as the second input of the second encoder. In this way, multi-band choroid plexus edge and texture information can be obtained at a deeper level, and a multi-scale feature map with the same resolution as the large model branch can be output.

[0144] Next, the first output of the first encoder is used as the first input of the second encoder, which is fed into the adaptation feature extraction module of the second encoder. Figure 4As shown, in this adaptive feature extraction module, an adaptation operation is introduced to balance the global semantics of the large model with the features of the small choroid plexus in medical images. This involves downsampling and upsampling the first input of the other encoders to produce an adapted feature map. Specifically, higher-dimensional features are first compressed to 16 or 32, then processed by an activation function and reactivated back to their original dimensions. The adapted feature map is then input into the first feature extraction module and passes through multiple feature extraction modules for further feature extraction. The output of the adapted feature extraction module, i.e., the first output of the second encoder, is passed to the adapted feature extraction module of the third encoder and serves as the first input of the third encoder. The feature extraction module here is the hierarchical visual transformer in the SAM2 model. Through this hierarchical visual transformer structure, key features associated with the choroid plexus are captured step by step. This approach preserves the macroscopic priors of the large model while making the features more suitable for the segmentation of fine structures such as the choroid plexus.

[0145] The third output of the first encoder serves as the second input of the second encoder and is fed into the second frequency convolution layer of the second encoder. The second frequency convolution layer then outputs the second output of the second encoder, which is then passed to the second frequency convolution layer of the third encoder and serves as the second input of the third encoder. The implementation of the second frequency convolution layer here is similar to that of the first frequency convolution layer and will not be repeated here.

[0146] After being processed by the second encoder, the corresponding first output (corresponding to the large model branch) and second output (corresponding to the convolutional network branch) are obtained. The first output of the second encoder serves as the first input of the third encoder, and the second output of the second encoder serves as the second input of the third encoder. The third encoder performs the same operation as the data stream of the second encoder. Similarly, the third encoder, the fourth encoder, and the fifth encoder perform the encoding processing operation in sequence.

[0147] Since the large model branch and the convolutional network branch output feature maps with the same resolution respectively, in order to combine the advantages of the two, feature fusion is performed once at each corresponding scale, which is performed by the feature fusion layer of other encoders except the first encoder, such as Figure 7As shown in the figure, the specific method is: first, the feature maps from the large model branch and the convolutional network branch (i.e., the first and second outputs of the other encoders) are concatenated in the channel dimension to obtain a concatenated feature map with twice the number of channels; then, a 1×1 convolution is used to compress the number of channels back to the scale of a single output, and an activation function is used to reorganize and filter the features to obtain the third output of the other encoder, which is passed to the corresponding decoder and serves as one of the decoder's inputs. This allows the global context of the large model branch and the local context details of the convolutional network branch to be simultaneously utilized, resulting in a comprehensive feature representation that preserves semantic information while highlighting the edges of small objects in the context.

[0148] In summary, after the processing of the above encoder structure, the output direction of each encoder is:

[0149] First encoder: The first output is passed to the adaptation feature extraction module of the second encoder; the second output is passed to the first frequency convolution layer and the fifth decoder of the first encoder; the second output is passed to the second frequency convolution layer and the fourth decoder of the second encoder.

[0150] Second encoder: The first output is passed to the adaptation feature extraction module of the third encoder and the feature fusion layer of the second encoder; the second output is passed to the second frequency convolution layer of the third encoder and the feature fusion layer of the second encoder; the third output is passed to the third decoder;

[0151] The third encoder: The first output is passed to the adaptive feature extraction module of the fourth encoder and the feature fusion layer of the third encoder; the second output is passed to the second frequency convolution layer of the fourth encoder and the feature fusion layer of the third encoder; the third output is passed to the second decoder;

[0152] The fourth encoder: The first output is passed to the adaptive feature extraction module of the fifth encoder and the feature fusion layer of the fourth encoder; the second output is passed to the second frequency convolution layer of the fifth encoder and the feature fusion layer of the fourth encoder; the third output is passed to the first decoder;

[0153] Fifth encoder: The first output is passed to the feature fusion layer of the fifth encoder; the second output is passed to the feature fusion layer of the fifth encoder; the third output is passed to the first decoder.

[0154] Based on this, the output of the encoder structure is input into the decoder structure. Each decoder will process high-level features (small resolution, strong semantics) and low-level features (large resolution, rich details) step by step, such as Figure 8As shown in the figure, the high-level features are first upsampled to the same spatial size as the low-level features, and then the two are spliced ​​across the channels. 3×3 convolution, normalization, and activation functions are then performed continuously to repair upsampling artifacts and refine the edges of the choroid plexus. Through upsampling and splicing of multiple decoders, the decoder structure ultimately outputs a choroid plexus segmentation map of the same size as the original image (1×256×256). The output layer uses a Sigmoid or Softmax function to classify each pixel, generating an accurate segmentation mask for the choroid plexus region.

[0155] Finally, the generated segmentation mask may require further optimization through post-processing techniques (such as morphological operations and regional connectivity analysis) to remove noise and ensure the continuity and smoothness of the segmentation boundaries. The segmentation results can be used not only for clinical diagnosis but also for scientific analysis and further model optimization and validation.

[0156] Such a complete process fully utilizes the macro-prior of the large model and the fine capture of choroid plexus details by the convolutional network branch. Through multi-level fusion and step-by-step refinement of the decoder, it ultimately achieves higher accuracy and robustness in the choroid plexus segmentation task.

[0157] In addition, an embodiment of the present application provides a brain choroid plexus segmentation device, which may include:

[0158] an acquisition module, for acquiring brain images;

[0159] a segmentation module configured with a pre-trained segmentation model, the segmentation module being used to input a brain image into the segmentation model to obtain a choroid plexus contour map;

[0160] Among them, the segmentation model includes: an encoder structure, a decoder structure and an output layer; the encoder structure is sequentially provided with multiple encoders; the first encoder is used to perform feature alignment and feature extraction on the brain image to obtain the output of the first encoder; other encoders except the first encoder are used to perform adaptation processing, feature extraction and feature fusion on the input of other encoders to obtain the output of other encoders; the decoder structure is jump-connected to the encoder structure, and the decoder structure is sequentially provided with multiple decoders, each decoder is used to decode the input of the decoder to obtain the output of the decoder; the output layer is used to obtain the choroid plexus contour map based on the output of the last decoder.

[0161] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0162] In summary, the embodiments of the present application have at least the following technical effects:

[0163] (1) This application proposes a dual-branch adaptive fusion of a large model branch and a convolutional network branch in the segmentation model. In the related art, most of them will use at least one convolutional neural network structure, U-Net, or fine-tune the large model alone, and rarely achieve parallel branches and adaptive fusion of the two. Accordingly, this application sets up corresponding large model branches and convolutional network branches for each encoder. The large model branches of each encoder are connected in sequence, and the convolutional network branches of each encoder are also connected in sequence. In this way, through the dual-branch design in the segmentation model, while ensuring the convolutional neural network's depiction of fine-grained choroid plexus boundaries, the powerful prior ability of the large model is ensured. At the same time, this application designs a feature fusion operation for encoders other than the first encoder. Through multiple fusions in the middle, the general visual features of the large model and the specific details of the choroid plexus are combined layer by layer and passed to the corresponding decoder. This allows the segmentation model to not lose the high-level semantics of the large model when processing small targets, and can also retain the choroid plexus details in combination with the convolutional network branches, thereby reducing the probability of missing or misclassifying the choroid plexus area of ​​the small target.

[0164] (2) The present application introduces Haar wavelet transform in the convolution operation. In the related art, the convolution operation often has maximum pooling or average pooling, which easily loses the edge of the choroid plexus. In response to this, the present application introduces Haar wavelet transform in the convolution operation for each encoder, which can well preserve the edge features of the choroid plexus in different frequency bands and is more sensitive to the choroid plexus edge detection of low-contrast and small targets. Compared with simple pooling, the present application avoids the edge blurring caused by multiple pooling of the convolution layer, and retains more high-frequency details associated with the choroid plexus area without significantly increasing the amount of calculation.

[0165] (3) This application proposes a feature alignment operation and an adaptation operation. In the related art, when applying a large model to a medical imaging scene, it is often necessary to fine-tune the large model. Directly fine-tuning the large model requires unfreezing a large number of weights, which leads to a significant increase in computational complexity and video memory consumption. In addition, simply freezing the large model cannot make it finely applicable to the choroid plexus segmentation task. In response to this, this application designs a feature alignment operation for the first encoder. Through the feature alignment operation, the difference between the brain image domain and the natural image domain is first reduced. In addition, a large model is configured for other encoders except the first encoder, and an adaptation operation is inserted into the large model. It is trainable, so only key parameters are fine-tuned. In this way, customized adaptation is performed for small target scenes such as the choroid plexus without completely destroying the existing generalization ability of the large model. Compared with the direct fine-tuning of the traditional large model, the training cost is lower and the ability to depict the segmentation boundary is stronger. Moreover, compared with the practice of significantly unfreezing the large model or expanding the depth of the convolutional network at the same time, this application only inserts adaptation operations in key parts for fine-tuning, reducing computational and video memory overhead, and maintaining good compatibility with images of multiple resolutions.

[0166] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and intent of the present application, and that the scope of the present application is defined by the claims and their equivalents.

[0167] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.

Claims

1. A brain choroid plexus segmentation method, characterized in that: The following steps are involved: Obtain brain images; inputting the brain image into a pre-trained segmentation model to obtain a choroid plexus contour map; Wherein, the segmentation model includes: The encoder structure is sequentially provided with a plurality of encoders; the first encoder is used to perform feature alignment and feature extraction on the brain image to obtain the output of the first encoder; the encoders other than the first encoder are used to perform adaptation processing, feature extraction and feature fusion on the inputs of the other encoders to obtain the outputs of the other encoders; A decoder structure is jump-connected to the encoder structure, wherein the decoder structure is sequentially provided with a plurality of decoders, each of which is used to decode the input of the decoder to obtain the output of the decoder; The output layer is used to obtain the choroid plexus contour map according to the output of the last decoder.

2. The method according to claim 1, characterized in that The output of the first encoder includes a first output, a second output, and a third output of the first encoder; and performing feature alignment and feature extraction on the brain image to obtain the output of the first encoder includes: Performing feature alignment on the brain image to obtain a first output of the first encoder; performing a conventional convolution operation on the brain image to obtain a second output of the first encoder; A frequency convolution operation is performed on the second output of the first encoder to obtain a third output of the first encoder.

3. The method according to claim 2, characterized in that The performing feature alignment on the brain image to obtain a first output of the first encoder includes: Performing a multi-scale convolution operation on the brain image to obtain multiple convolution feature maps; Fusing the multiple convolutional feature maps to obtain a composite feature map; The composite feature map is compressed to obtain a first output of the first encoder.

4. The method according to claim 1, wherein The output of the first encoder includes the first output and the third output of the first encoder, and the outputs of the other encoders include the first output and the second output of the other encoders; The step of performing adaptation processing, feature extraction, and feature fusion on the input of the other encoders to obtain the output of the other encoders includes: Performing adaptation processing and feature extraction on the first inputs of the other encoders to obtain first outputs of the other encoders; wherein the first input of the second encoder is the first output of the first encoder, and the first input of the encoder after the second encoder is the first output of the previous encoder; Performing a frequency convolution operation on the second input of the other encoders to obtain the second output of the other encoders; wherein the second input of the second encoder is the third output of the first encoder, and the second input of the encoder after the second encoder is the second output of the previous encoder.

5. The method according to claim 4, characterized in that The performing adaptation processing and feature extraction on the first input of the other encoder to obtain the first output of the other encoder includes: Performing adaptation processing based on downsampling and upsampling on the first input of the other encoders to obtain an adapted feature map; Fusing the adapted feature map with the first input of the other encoders to obtain a feature map to be processed; Perform multiple large-model-based feature extractions on the feature graph to be processed to obtain the first output of the other encoders.

6. The method according to any one of claims 2 or 4, characterized in that The steps of the frequency convolution operation include: Performing wavelet transform processing on the input of the frequency convolution operation to obtain multiple sub-band feature maps; Reconstructing the plurality of sub-band feature maps to obtain a reconstructed feature map; Perform multiple traditional convolution operations on the reconstructed feature map to obtain the output of the frequency convolution operation.

7. The method according to claim 1, characterized in that The outputs of the other encoders include the first output, the second output, and the third output of the other encoders; and the adaption processing, feature extraction, and feature fusion of the inputs of the other encoders to obtain the outputs of the other encoders include: Splicing the first output and the second output of the other encoders to obtain a spliced ​​feature map; Performing channel compression on the spliced ​​feature map to obtain a compressed feature map; An activation operation is performed on the compressed feature map to obtain a third output of the other encoder.

8. The method according to claim 1, characterized in that The first decoder is jump-connected to the last encoder and the penultimate encoder respectively, and the input of the first decoder includes the third output of the last encoder and the third output of the penultimate encoder; The last decoder and the penultimate decoder are both jump-connected to the first encoder, the input of the last decoder includes the second output of the first encoder and the output of the penultimate decoder, and the input of the penultimate decoder includes the third output of the first encoder and the output of the third-to-last decoder; The remaining decoders except the first decoder, the last decoder and the penultimate decoder, and the remaining encoders except the first encoder, the last encoder and the penultimate encoder have a one-to-one correspondence and are jump-connected; The remaining decoders except the first decoder, the last decoder and the penultimate decoder include the third output of the encoder corresponding to the remaining decoders and the output of the previous decoder of the remaining decoders.

9. The method according to claim 8, characterized in that The decoding process of the decoder input to obtain the decoder output includes: Fusing the decoder input to obtain a fused feature map; Perform multiple traditional convolution operations on the fused feature map to obtain the output of the decoder.

10. A brain choroid plexus segmentation device, characterized in that: include: an acquisition module, for acquiring brain images; a segmentation module configured with a pre-trained segmentation model, the segmentation module being used to input the brain image into the segmentation model to obtain a choroid plexus contour map; Wherein, the segmentation model includes: The encoder structure is sequentially provided with a plurality of encoders; the first encoder is used to perform feature alignment and feature extraction on the brain image to obtain the output of the first encoder; the encoders other than the first encoder are used to perform adaptation processing, feature extraction and feature fusion on the inputs of the other encoders to obtain the outputs of the other encoders; A decoder structure is jump-connected to the encoder structure, wherein the decoder structure is sequentially provided with a plurality of decoders, each of which is used to decode the input of the decoder to obtain the output of the decoder; The output layer is used to obtain the choroid plexus contour map according to the output of the last decoder.

Citation Information

Patent Citations

  • Brain glioma segmentation model and segmentation method based on deep learning

    CN112419267A

  • Cell nucleus segmentation method of cascade coding segmentation network based on large model guidance

    CN118366153A

  • Depth learning-based vein clump image segmentation method and system

    CN118485675A

  • Medical image segmentation method, system and equipment based on space-spectrum cross-domain coding and entropy perception double decoding

    CN118552565A

  • Method, device, and storage medium for lesion segmentation and recist diameter prediction via click-driven attention and dual-path connection

    US20220335600A1