Ovarian junctional tumor focus segmentation method based on deep learning
By constructing a multi-center standardized dataset and designing multi-strategy data augmentation, combined with SegNet+ network and attention mechanism, the problem of insufficient segmentation performance of BOTs lesions was solved, achieving high-precision pixel-level segmentation, adapting to imaging protocols of different hospitals, and reducing labor costs.
Patent Information
- Application Number
- CN202511860060.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-01-13
AI Technical Summary
In existing technologies, the segmentation performance of BOTs lesions is poor, lacking large-scale multi-center datasets and being insufficiently targeted, resulting in low efficiency and poor consistency of manual segmentation.
We constructed a multi-center standardized BOTs MRI dataset, designed multi-strategy data augmentation, and combined it with the SegNet+ image segmentation network, incorporating location and channel attention mechanisms to optimize lesion feature representation and achieve pixel-level accurate segmentation.
It improves the accuracy of BOT lesion segmentation (mIoU=0.855, DICE=0.845), reduces labor costs, adapts to the differences in MRI imaging protocols of different hospitals, and ensures the stability of segmentation performance.
Smart Images

Figure CN121329985A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image segmentation technology, and in particular to a deep learning-based method for segmenting borderline ovarian tumor lesions. Background Technology
[0002] Borderline ovarian tumors (BOTs) are a special type of lesion between benign ovarian cysts and malignant tumors, accounting for about 10% of all epithelial ovarian tumors. In their clinical diagnosis and treatment, the location, size and margin information of the lesion are crucial to the formulation of surgical plans for women of childbearing age, and directly affect the choice between unilateral cyst resection or more extensive resection. It is a key assessment content in the clinical diagnosis and treatment process.
[0003] Currently, the segmentation of BOTs (Body Occult Organs) lesions in clinical practice mainly relies on manual annotation by radiologists based on MRI images. This method is the fundamental means for obtaining lesion information, assisting in diagnosis, and surgical planning, providing direct image analysis basis for the clinical diagnosis and treatment of BOTs. In recent years, deep learning technology has made breakthrough progress in the field of medical image segmentation. Mainstream models such as the encoder-decoder architecture represented by U-Net, Deeplabv3+, and FCN have been successfully applied to various medical image segmentation tasks, including brain tumors and multiple abdominal organs, improving segmentation accuracy through mechanisms such as multi-scale feature fusion and semantic information extraction. However, due to the lack of large-scale multi-center datasets and insufficient focus on BOT lesion segmentation, the segmentation performance for BOTs lesions is unsatisfactory, and an automated segmentation scheme that can be directly applied in clinical practice has not yet been formed, failing to solve the efficiency and consistency problems caused by manual segmentation.
[0004] Therefore, there is an urgent need to provide a deep learning-based method for segmenting borderline ovarian tumor lesions. Summary of the Invention
[0005] To address the current issues of insufficient large-scale multicenter datasets and inadequate focus on lesion segmentation for borderline ovarian tumors (BOTs), resulting in poor segmentation performance for BOT lesions, this invention provides a deep learning-based method for segmenting ovarian borderline tumor lesions.
[0006] On the one hand, a deep learning-based method for segmenting borderline ovarian tumor lesions is provided, the method comprising: A dataset consisting of MRI images of borderline ovarian tumor lesions from multiple hospitals was obtained; The samples in the dataset are preprocessed and subjected to random transformation enhancement to obtain a dataset containing several target samples; Several target samples are input into the SegNet+ image segmentation network to train a tumor segmentation model; wherein the SegNet+ image segmentation network contains an encoder, an attention module and a decoder.
[0007] On the other hand, a deep learning-based segmentation device for ovarian borderline tumor lesions is provided, based on the steps described in any embodiment of the method in the specification. The device includes: The acquisition unit is used to acquire a dataset consisting of MRI images of ovarian borderline tumor lesions from multiple hospitals; An enhancement unit is used to preprocess and randomly transform the samples in the dataset to obtain a dataset containing several target samples. The segmentation unit is used to input several target samples into the SegNet+ image segmentation network to train a tumor segmentation model; wherein the SegNet+ image segmentation network contains an encoder, an attention module, and a decoder.
[0008] On the other hand, a computer device is provided, the computer device including a memory and a processor, the memory for storing a computer program, and the processor for executing the computer program stored in the memory to implement the steps of the method described above.
[0009] On the other hand, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements the steps of the method described above.
[0010] On the other hand, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method described above.
[0011] The technical solution provided by this invention can bring at least the following beneficial effects: By constructing a multi-center standardized BOTs MRI dataset and designing multi-strategy data augmentation, the model can adapt to the differences in MRI imaging protocols of different hospitals and maintain stable segmentation performance on multi-center test data. The deep learning method integrates traditional convolutional neural networks and attention mechanisms, and adds position and channel attention mechanisms on the basis of the classic SegNet to enhance the feature representation of BOTs and achieve pixel-level accurate segmentation of BOTs lesions. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1This is a flowchart of a deep learning-based method for segmenting ovarian borderline tumor lesions, provided by an embodiment of the present invention. Figure 2 This is a schematic diagram of a tumor segmentation model architecture provided in an embodiment of the present invention. Figure 3 This is a structural diagram of a deep learning-based ovarian borderline tumor lesion segmentation device provided in an embodiment of the present invention; Figure 4 This is a hardware architecture diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0015] The following describes the specific implementation of the above concept.
[0016] Please refer to Figure 1 This invention provides a deep learning-based method for segmenting borderline ovarian tumor lesions, the method comprising: Step 100: Obtain a dataset consisting of MRI images of borderline ovarian tumor lesions from multiple hospitals; Step 102: Preprocess and randomly transform the samples in the dataset to obtain a dataset containing several target samples; Step 104: Input several target samples into the SegNet+ image segmentation network to train a tumor segmentation model; the SegNet+ image segmentation network contains an encoder, an attention module and a decoder.
[0017] In this embodiment of the invention, by constructing a multi-center standardized BOTs MRI dataset and designing multi-strategy data augmentation, the model can adapt to the differences in MRI imaging protocols of different hospitals and maintain stable segmentation performance on multi-center test data. The deep learning method integrates traditional convolutional neural networks and attention mechanisms, and adds position and channel attention mechanisms on the basis of the classic SegNet to enhance the feature representation of BOTs and achieve pixel-level accurate segmentation of BOTs lesions.
[0018] The following description Figure 1 The execution method for each step is shown.
[0019] For step 100: The construction of the multi-center dataset is as follows: four tertiary hospitals (Peking University Third Hospital, Peking University People's Hospital, Qilu Hospital of Shandong University, and the Third Affiliated Hospital of Kunming Medical University (Yunnan Cancer Hospital)) jointly collected preoperative MRI images (including T2WI and DWI sequences) of patients with pathologically confirmed BOTs. The lesion areas were labeled by radiologists to form pixel-level labels.
[0020] Regarding step 102: In this embodiment, MRI images of borderline ovarian tumors (BOTs) often include surrounding tissues such as the uterus, and multicenter data exhibit differences in imaging equipment models and scanning parameters, leading to inconsistencies in image signal intensity and resolution. Furthermore, BOT lesions themselves exhibit morphological heterogeneity, such as mixed cystic and solid lesions and blurred boundaries. Directly inputting these into the model can easily lead to interference from redundant information, reducing segmentation accuracy. To enable the model to focus on learning BOT lesion features, standardization and enhancement strategies are employed to improve data quality and diversity.
[0021] In this step, the DICOM format images from each center are uniformly converted to NIfTI format; the Z-score algorithm is used to normalize the intensity of the sequence images, thereby eliminating signal differences between devices.
[0022] The normalization formula is as follows: Where I represents the original pixel value, μ represents the image mean, and σ represents the standard deviation. Simultaneously, all images are standardized to a uniform resolution of 512×512 pixels to ensure consistency between data augmentation and model input size.
[0023] In some implementations, the random transformation enhancement methods include at least: probability horizontal flipping, random rotation, random scaling, random cropping, and Gaussian filtering.
[0024] Multi-strategy augmentation is performed on the training set to increase sample diversity, including: horizontal flipping with a 50% probability, random rotation within the range of [-10°, 10°], random scaling of 0.8-1.2 times, random cropping of 448×448 pixels, and Gaussian filtering with a value of 0-1.5 sigma. The augmented images are transformed synchronously with the corresponding labels to avoid label misalignment. The processed target samples are used for training the tumor segmentation module.
[0025] Regarding step 104: refer to Figure 2 In some implementations, the model architecture involves inputting several target samples into a SegNet+ image segmentation network to train a tumor segmentation model, including: For each target sample, execute: The current target sample is input into the encoder, which performs hierarchical feature extraction on the current target sample from local details to global context, and the pooling index is input into the corresponding decoder; The target feature map output from the last layer of the encoder is input into the attention module to perform positional attention enhancement and channel attention enhancement on the target feature map, and outputs an enhanced feature map; The decoder uses pooling indexes to upsample the enhanced feature maps layer by layer to obtain the segmented image; The difference between the label of the segmented image and the label of the current target sample is calculated using the category-balanced cross-entropy loss function. This difference is then used to train the network parameters of the encoder, attention module, and decoder until a tumor segmentation model that meets the expectations is obtained.
[0026] This model addresses the segmentation challenges of borderline ovarian tumors (BOTs) on MRI, characterized by blurred lesion boundaries, heterogeneous signal intensity, and low image coverage (≤25%). It optimizes the SegNet encoder-decoder architecture to create SegNet+, with the overall model architecture as follows: Figure 2 As shown, by adding an attention mechanism to enhance the focus on lesion features and combining it with a targeted loss function to solve the class imbalance problem, accurate pixel-level lesion segmentation is achieved.
[0027] In this embodiment of the invention, the encoder contains five first convolutional blocks, each containing a convolutional layer, a normalization layer, and a pooling layer; the decoder contains second convolutional blocks that correspond one-to-one with the first convolutional blocks of the encoder, each second convolutional block containing an upsampling layer and a convolutional layer, each second convolutional block receiving a pooling index corresponding to the first convolutional block, and the upsampling layer being used to upsample the enhanced feature map output by the attention module based on the pooling index, replacing the learnable upsampling kernel.
[0028] In this embodiment, the encoder contains five 3×3 first convolutional blocks, each containing convolution, ReLU activation, batch normalization, and pooling layers, realizing hierarchical feature extraction from local details to global context; each pooling layer is a 2×2 max pooling (stride=2), and the pooling index is retained to preserve the position of the maximum value after pooling, providing spatial location basis for accurate upsampling by the decoder, reducing computational complexity while retaining core features.
[0029] The decoder takes the pooled index retained by the encoder and the attention-optimized feature map as input, and restores the feature map resolution to the original MRI image size through progressive upsampling and convolutional refinement.
[0030] Each decoder's second convolutional block uses the pooling index retained from the corresponding encoder's first convolutional block to perform precise upsampling on the attention-optimized low-resolution feature map, without relying on a learnable upsampling kernel. After upsampling, the feature map is refined through a 3×3 convolutional layer to increase the density of the sparse upsampled output. The decoder consists of five second convolutional blocks, corresponding one-to-one with the encoder's first convolutional blocks. The resolution of the output feature map in each second convolutional block is progressively increased, ultimately outputting a feature map with the same size as the original MRI image.
[0031] In this embodiment of the invention, the attention module is located between the normalization layer and the pooling layer of the last first convolutional block of the encoder; The attention module includes a positional attention submodule and a channel attention submodule, which are used to perform positional attention enhancement and channel attention enhancement on the target feature map output by the last layer of the encoder, respectively, and then weighted and superimposed to output an enhanced feature map.
[0032] In this embodiment, the attention mechanism is based on the dual-dimensional collaborative optimization of "position + channel". The position attention submodule (PAM) and channel attention submodule (CAM) are applied in parallel between the normalization layer and the pooling layer of the last first convolutional block of the encoder. The dynamic weight allocation strengthens the lesion features and suppresses background interference. Both modules contain learnable parameters and residual connections to ensure the optimization effect while avoiding feature loss.
[0033] The core computational logic of the attention mechanism can be uniformly expressed by the following formula: In this context, Q (Query), K (Key), and V (Value) represent different forms of input features, dk is a scaling factor used to mitigate numerical bias in inner product calculations, and Softmax is used to generate attention weights and weight the output V. The main difference between PAM and CAM lies in the input sources of Q, K, and V.
[0034] Position Attention Module (PAM): The input feature map has a dimension of 1. (B represents batch size, C represents the number of channels, and H / W represents the feature map size). Q, K, and V are all generated from X through a 1×1 convolution, where the number of channels in Q and K is compressed to C / 8, while V remains unchanged.
[0035] Channel Attention Module (CAM): The input feature map has the same dimensions as PAM. Q, K, and V are directly reshaped from X. All three are reshaped into (B, C, H×W).
[0036] Both modules scale the feature maps using a learnable parameter gamma, and then add them to the original feature maps to achieve residual connections.
[0037] In some implementations, the class-balanced cross-entropy loss function is: In the formula, For category weights, For pixel-based true labels, represents the pixel class probability predicted by the model, where i=1 represents the lesion class and i=2 represents the background class.
[0038] In this embodiment, differentiated weights (w1=4, w2=1) are assigned to the background class and the lesion class to alleviate the problem that the model is biased towards predicting the background due to the small proportion of lesions.
[0039] In summary, the specific process of this invention is as follows: First, the MRI data preprocessing and enhancement module provides high-quality and diverse input data for model training through multi-center case collection, standardized format conversion, and multi-strategy data enhancement; second, the tumor segmentation module is based on the SegNet framework, which enhances lesion spatial and channel features through encoder hierarchical feature extraction, PAM+CAM dual attention mechanism, and decoder accurate restoration of boundary details. It also solves the class imbalance problem by combining class-balanced cross-entropy loss, and finally achieves pixel-level accurate segmentation of BOTs lesions.
[0040] A standardized multi-center BOTs MRI dataset was constructed, and multi-strategy data augmentation was designed to enable the model to adapt to the differences in MRI imaging protocols across different hospitals. It maintained stable segmentation performance on multi-center test data, laying the foundation for widespread clinical application. Secondly, traditional manual segmentation methods relying on radiologists are not only time-consuming and labor-intensive but also susceptible to individual differences among radiologists, resulting in poor consistency in lesion annotation. This invention, based on deep learning methods, significantly reduces manual costs and improves the accuracy of BOTs segmentation (mIoU=0.855, DICE=0.845).
[0041] Please refer to Figure 3 This invention provides a deep learning-based device for segmenting ovarian borderline tumor lesions, used to implement the steps of any method embodiment in the specification. The device includes: The acquisition unit 301 is used to acquire a dataset consisting of MRI images of ovarian borderline tumor lesions from multiple hospitals; Enhancement unit 302 is used to preprocess and randomly transform the samples in the dataset to obtain a dataset containing several target samples; The segmentation unit 303 is used to input several target samples into the SegNet+ image segmentation network to train a tumor segmentation model; the SegNet+ image segmentation network contains an encoder, an attention module and a decoder.
[0042] It should be noted that the above device embodiments and method embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0043] Embodiments of this application also provide a computer device, please refer to... Figure 4 The computer device includes a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, at least one program, code set or instruction set being loaded and executed by the processor to implement the deep learning-based ovarian borderline tumor lesion segmentation method provided in the above-described method embodiments.
[0044] The embodiments of this application also provide a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the deep learning-based ovarian borderline tumor lesion segmentation method provided in the above-described method embodiments.
[0045] Embodiments of this application also provide a computer program product, which includes a computer program. A processor of a computer device reads the computer program from a computer-readable storage medium and executes the computer program, causing the computer device to perform any of the deep learning-based ovarian borderline tumor lesion segmentation methods described in the above embodiments.
[0046] For ease of description, the above devices or apparatuses are described separately according to their functions, divided into various modules or units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0047] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of the embodiments of this application.
[0048] Finally, it should be noted that in this document, relational terms such as first, second, third, and fourth are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0049] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A deep learning-based method for segmenting borderline ovarian tumor lesions, characterized in that, include: A dataset consisting of MRI images of borderline ovarian tumor lesions from multiple hospitals was obtained; The samples in the dataset are preprocessed and subjected to random transformation enhancement to obtain a dataset containing several target samples; Several target samples are input into the SegNet+ image segmentation network to train a tumor segmentation model; wherein the SegNet+ image segmentation network contains an encoder, an attention module and a decoder.
2. The method as described in claim 1, characterized in that, Random transformation enhancement methods include at least: probabilistic horizontal flipping, random rotation, random scaling, random cropping, and Gaussian filtering.
3. The method as described in claim 1, characterized in that, The step of inputting several target samples into the SegNet+ image segmentation network to train a tumor segmentation model includes: For each target sample, execute: The current target sample is input into the encoder, which performs hierarchical feature extraction on the current target sample from local details to global context, and the pooling index is input into the corresponding decoder; The target feature map output from the last layer of the encoder is input into the attention module to perform positional attention enhancement and channel attention enhancement on the target feature map, and output an enhanced feature map. The decoder upsamples the enhanced feature map layer by layer based on the pooling index to obtain a segmented image; The difference between the label of the segmented image and the label of the current target sample is calculated using the category-balanced cross-entropy loss function. This difference is then used to train the network parameters of the encoder, the attention module, and the decoder until a tumor segmentation model that meets the expectations is obtained.
4. The method as described in claim 1, characterized in that, The encoder contains five first convolutional blocks, each containing a convolutional layer, a normalization layer, and a pooling layer. The decoder contains second convolutional blocks that correspond one-to-one with the first convolutional blocks of the encoder. Each second convolutional block contains an upsampling layer and a convolutional layer. Each second convolutional block receives a pooling index corresponding to the first convolutional block. The upsampling layer is used to upsample the enhanced feature map output by the attention module based on the pooling index, replacing the learnable upsampling kernel.
5. The method as described in claim 4, characterized in that, The attention module is located between the normalization layer and the pooling layer of the last first convolutional block of the encoder; The attention module includes a position attention submodule and a channel attention submodule, which are used to perform position attention enhancement and channel attention enhancement on the target feature map output by the last layer of the encoder, respectively, and then weighted and superimposed to output an enhanced feature map.
6. The method according to any one of claims 1-5, characterized in that, The category-balanced cross-entropy loss function is: In the formula, For category weights, For pixel-based true labels, represents the pixel class probability predicted by the model, where i=1 represents the lesion class and i=2 represents the background class.
7. A deep learning-based device for segmenting ovarian borderline tumor lesions, used to implement the steps of the method described in any one of claims 1-6, characterized in that, include: The acquisition unit is used to acquire a dataset consisting of MRI images of ovarian borderline tumor lesions from multiple hospitals; An enhancement unit is used to preprocess and randomly transform the samples in the dataset to obtain a dataset containing several target samples. The segmentation unit is used to input several target samples into the SegNet+ image segmentation network to train a tumor segmentation model; wherein the SegNet+ image segmentation network contains an encoder, an attention module, and a decoder.
8. A computer device, characterized in that, The computer device includes a memory and a processor. The memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory to implement the steps of the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the method described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1-6.