An Adaptive Segmentation and Measurement Method for Laryngeal Cancer-Hypopharyngeal Cancer Medical Images

The two-stage contrastive learning framework addresses the challenges of throat and hypopharyngeal cancer image segmentation by improving model adaptability across diverse imaging conditions, enhancing diagnostic accuracy and efficiency while reducing costs.

CN118674728BActive Publication Date: 2025-07-15NINGBO MEDICAL CENT LIHUILI HOSPITACL +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410657350.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-25
Publication Date
2025-07-15
Estimated Expiration
2044-05-25

AI Technical Summary

Technical Problem

The prior art has problems such as high recognition difficulty, high morphological variability, and poor adaptability of different imaging devices in the medical image segmentation and measurement of laryngeal and hypopharyngeal carcinoma, resulting in low diagnostic efficiency and high cost.

Method used

Adaptive segmentation and measurement method of laryngeal-hysterical carcinoma medical images is adopted. By constructing the comparative learning pre-training of the source domain and the target domain images, the generator is used to perform style transfer and random number enhancement, combining the comparison loss function and mixed segmentation loss, the network learning parameters are optimized, and the training process of the segmentation model is simplified.

Benefits of technology

It improves the generalization ability of the model in different medical imaging fields, improves the accuracy and practicality of the segmentation model, especially in the case of scarce labeling data, which significantly reduces medical costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118674728B_ABST
    Figure CN118674728B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of medical image processing, and particularly to a method for adaptive segmentation and measurement of laryngeal cancer - hypopharyngeal cancer medical images. The key steps of this method include contrastive learning pre-training and source domain fine-tuning. In contrastive learning pre-training, an encoder-decoder with shared parameters is used as the network backbone. By designing a non-parametric feature projection head, the pixel-level representation is mapped to the hypersphere space, and a contrastive loss function is used to learn the consistent representation between domains. Source domain fine-tuning is to transfer the pre-trained knowledge to the segmentation model to achieve accurate pixel-level segmentation. Using this method not only optimizes the accuracy and efficiency of the diagnosis process, but also has significant significance in reducing medical costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and particularly to a method for adaptive segmentation and measurement of laryngeal cancer - hypopharyngeal cancer medical images. Background Art

[0002] Laryngeal cancer and hypopharyngeal cancer, as relatively common types of malignant tumors in the fields of otolaryngology and head and neck surgery, have an extremely urgent need for early diagnosis. This early diagnosis plays a crucial role in improving the survival rate of patients and their quality of life. Unfortunately, due to the difficulty in directly observing the lesion sites of these cancers and the insignificance of early symptoms, the diagnosis and treatment processes are full of challenges. Currently, traditional medical imaging and pathological methods still show certain limitations in the early detection of laryngeal cancer and hypopharyngeal cancer, especially in accurately identifying and locating early lesions. Generally, patients need to undergo various different types of examinations, which not only reduces the diagnostic efficiency and the utilization effect of medical resources, but also increases the economic burden and mental stress of patients.

[0003] Although unsupervised domain adaptation methods can achieve cross - domain analysis in laryngeal cancer - hypopharyngeal cancer medical images, there are still many challenges in the process of medical image segmentation and measurement of laryngeal cancer and hypopharyngeal cancer: First, the identification of these tumor regions in medical images is difficult, mainly because the contrast between tumors and surrounding tissues in images is usually low. In addition, the morphology and size of tumors have a high degree of variability, which increases the difficulty of accurate image segmentation. At the same time, under different imaging devices and conditions, the generalization ability of segmentation algorithms is often limited, making it complex to apply in a clinical environment. In past research, although these problems have been solved to a certain extent by improving network structures and algorithms, these methods often rely on complex network modules and are usually only applicable to specific domain adaptation scenarios, and it is difficult to be widely applied to the diverse clinical environments of laryngeal cancer - hypopharyngeal cancer medical images. When these methods attempt to adapt to different imaging conditions and devices, a large number of customized parameter adjustments are required, which not only increases the complexity of the model but also results in the inability to widely adapt to imaging devices with different parameters. Summary of the Invention

[0004] The technical solution to be solved by the present invention is: to provide a method for adaptive segmentation and measurement of laryngeal cancer - hypopharyngeal cancer medical images, which not only optimizes the accuracy and efficiency of the diagnostic process but also has significant significance in reducing medical costs.

[0005] The technical solution adopted by the present invention is: a method for adaptive segmentation and measurement of laryngeal cancer - hypopharyngeal cancer medical images, characterized in that it includes the following steps:

[0006] S1. Select the dataset: Construct the source domain images and target domain images;

[0007] S2. Data preprocessing: Augment the source domain images and target domain images constructed in step S1;

[0008] S3. Domain connection construction:

[0009] Within the domain, obtain the global view of the target domain image by identity mapping of the target domain image preprocessed in step S2, and then perform a random number enhancement operation on the target domain image to obtain the local view of the target domain image; Between domains, input the source domain image preprocessed in step S2 into the trained generator for style transfer conversion to obtain the global view of the fake target domain image, and then perform a random number enhancement operation on the source domain image to obtain the local view of the source domain image;

[0010] Within the domain, use the global view and local view of the target domain image as positive sample pairs. Between domains, use the global view of the fake target domain image obtained by inputting the source domain image into the trained generator for style transfer conversion and the local view of the source domain image obtained by the random number enhancement operation as positive sample pairs; Within the domain, use the local views and global views of different target domain images as negative sample pairs. Between domains, use the local views of different source domain images and the global views of the fake target domain images obtained by the generator as negative sample pairs;

[0011] S4. Contrastive learning pre-training: Input the within-domain views and between-domain views obtained in step S3 into the encoder-decoder with shared parameters to obtain pixel-level representations, map the pixel-level representations to the hypersphere space through a designed non-parametric feature projection head, calculate the loss using the contrastive loss function, and then optimize the network learning parameters through backpropagation, making the distance between positive sample pairs between domains the smallest and the distance between negative sample pairs the largest. Then, continuously iterate and optimize until the contrastive loss function converges, and finally obtain the trained network learning parameters;

[0012] S5. Source domain fine-tuning: Construct a semantic segmenter, initialize the segmenter with the network learning parameters trained in step S4, and randomly initialize the segmentation head at the same time. Then, use the image pairs obtained in step S3 to adjust the segmenter to obtain a trained segmenter;

[0013] S6. Use the segmenter trained in step S5 to segment the medical images to be segmented.

[0014] Preferably, the trained generator in step S3 refers to the CycleGAN or CUT generator trained with the data preprocessed in step S2.

[0015] Preferably, the formula of the contrastive loss function in step S4 is:

[0016]

[0017] where L contrast is the contrastive loss function, sim(·) is used to calculate the cosine similarity of feature vectors, and f i represents a feature vector as an anchor point, and f i ′ represents the feature vector of the positive sample corresponding to the anchor point, and f j represents the feature vector of the negative sample, + means to pull closer the feature vector of the positive sample, - means to push away the feature vector of the negative sample, v is the temperature parameter, n is the number of feature vectors, and j represents the counting variable.

[0018] Preferably, the non-parametric feature projection head in step S4 refers to: inputting a feature map, then obtaining an intermediate layer feature map through a combination layer of several max-pooling layers and non-linear activation functions, and then performing a vector flattening operation on the intermediate layer feature map to obtain a feature vector.

[0019] Preferably, after fine-tuning the segmenter using the image obtained in step S3 in step S5, it is also necessary to use a hybrid segmentation loss to guide the segmenter, and the hybrid segmentation loss is composed of a weighted cross-entropy loss and a Dice loss.

[0020] Preferably, the formulas for the segmenter and the hybrid segmentation loss are:

[0021]

[0022]

[0023] where F(x) is the segmenter, Θ is the network learning parameter, and the superscript s represents the source domain image sample or the prediction result based on the source domain image sample; s→t means that the source domain sample has undergone a style transfer transformation from the source domain to the target domain by the generator G; is the output of the segmenter F, and y s is the ground truth of the source domain image sample, p represents the prediction; λ = 1.0 and η = 2.0 are the contribution coefficients of the hybrid segmentation loss function;

[0024]

[0025]

[0026] The superscript s represents the source domain image sample or the prediction result based on the source domain image sample; c represents the number of true semantic categories; β is the ratio vector of the pixel ratio of a specific category in a batch to the pixel ratio of all categories, and pixels represents a pixel-by-pixel traversal operation.

[0027] Compared with the prior art using the above method, the present invention has the following advantages:

[0028] (1) By means of contrastive learning pre-training technology, the present invention effectively utilizes unlabeled data and significantly improves the generalization ability of the model in different medical imaging fields, especially when dealing with medical images of multiple centers, multiple sites, multiple sequences, and multiple modalities.

[0029] (2) The two-stage contrastive learning training framework of the present invention simplifies the traditional unsupervised domain adaptation (UDA) process and improves the practicality and accuracy of the segmentation model in clinical applications, especially in the case of scarce labeled data.

[0030] (3) The method of the present invention demonstrates excellent performance in multiple medical image segmentation tasks, including but not limited to the segmentation of laryngeal cancer and hypopharyngeal cancer, etc., proving its potential in actual clinical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a model training and prediction model framework diagram of a method for self-adaptive segmentation and measurement of laryngeal cancer - hypopharyngeal cancer medical images according to the present invention.

[0032] Figure 2 It is a technical roadmap of a non-parametric feature projection head proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.

[0034] Embodiment 1:

[0035] A method for self-adaptive segmentation and measurement of laryngeal cancer - hypopharyngeal cancer medical images, as Figure 1 shown, the core steps of this method include contrastive learning pre-training and source domain fine-tuning, and the contrastive learning pre-training further includes domain connection construction, and it generally includes the following steps:

[0036] S1. Select a data set: This model constructs a relevant CT and MRI fusion data set for laryngeal cancer and hypopharyngeal cancer. Among them, the CT modal data and the MRI modal data respectively contain 3D grayscale images composed of 60 samples. 1600 2D plane slice images are obtained by slicing, with a resolution of 256×256. The training set, validation set, and test set are divided according to a ratio of 8:1:1, and the CT image S is the source domain image, and the MRI image T is the target domain image;

[0037] S2. Data preprocessing: To construct a more comprehensive training dataset, we merged all the data and adopted three rotation methods: horizontal rotation, vertical rotation, and horizontal-vertical rotation. By randomly applying these rotations, we successfully expanded the original 1,600 images to 3,200. In the expanded dataset, each sample contains three different rotated versions: the upper right is horizontally rotated, the lower left is vertically rotated, and the lower right is horizontally-vertically rotated. Such processing not only increases the diversity of the dataset but also helps improve the model's robustness to image rotation changes.

[0038] S3. Domain connection construction

[0039] Within the domain, by taking the view of the target-domain image after the identity mapping μ as the target-domain global view and the view of the target-domain image with random data augmentation τ as the target-domain local view, considering their scales can ensure that the semantics of these two views always match in the latent space, ensuring semantic consistency within the domain. Specifically, through contrastive learning, the target-domain local view and the target-domain global view are made as close as possible semantically, while pushing other views farther away, so as to learn a dense representation containing sufficient image context during the pre-training stage. This strategy not only avoids the semantic inconsistency problem caused by random augmentation τ but also provides a fine-grained representation for subsequent segmentation tasks.

[0040] Between domains, the source-domain images are used as the generator G of the unsupervised pre-trained CycleGAN or CUT for style transfer transformation to synthesize the global view of the fake target-domain images. Where S represents the source domain and T represents the target domain. S T represents the set of style transfer images based on the source domain. represents an image sample in S T The local view of the source domain is obtained through random augmentation τ to maintain semantic consistency. Specifically, by constructing positive pair views that have both target-domain features and maintain semantic consistency. Among them, represents the augmented view of the i-th source-domain image sample. represents the sample after the source-domain sample is transferred by the generator G. + represents pulling the positive sample features closer, and then it is used for contrastive pre-training to reduce the difference between the source domain and the target domain, so as to establish an effective connection between the two domains. Through this strategy, we can achieve a closer connection between domains and provide strong prior knowledge for subsequent segmentation tasks.

[0041] And it is also necessary to construct positive sample pairs and negative sample pairs, where:

[0042] Positive sample pairs, within the same domain: the global view of the target domain and its local view; across domains, the global view of the fake target domain image generated by the generator G from the global view of the source domain image and the local view obtained by performing a random number enhancement operation on the global view of the source domain image;

[0043] Negative sample pairs, combinations of enhanced views of different images, such as the global view of the target domain and the local view of the source domain;

[0044] S4. Contrastive learning pre-training: Initialize the network learning parameters, including the parameters of the encoder-decoder and the non-parametric feature projection head, which are randomly initialized; input the within-domain views and the across-domain views into the parameter-shared encoder-decoder respectively to obtain pixel-level representations, and then, through the non-parametric feature projection head, map the pixel-level representations to the hypersphere space, calculate the loss using the contrastive loss function, and then optimize the network learning parameters through backpropagation, so that the distance between positive sample pairs across domains is minimized and the distance between negative sample pairs is maximized. Then, continuously iterate and optimize until the contrastive loss function converges, and finally obtain the trained network learning parameters, where:

[0045] The parameter-shared encoder-decoder is a prior art. For example, common encoder-decoder structures include U-Net, Autoencoder, etc. Therefore, it is not elaborated in detail in this application. Its main function is to encode the input image into a low-dimensional latent representation, and the decoder decodes the latent representation back to the original image size, thereby retaining the pixel-level information of the image; the specific process is that the input image is x, and the output of the encoder-decoder is h = E(x). The pixel-level representation h is a multi-dimensional tensor, and each pixel point has a corresponding feature vector, representing the high-dimensional features of the pixel point; in this embodiment, there are two encoder-decoders, and the parameters of the two encoder-decoders are shared. Input the global view of the target domain and the local view of the target domain into one encoder-decoder, and input the global view of the fake target domain image and the local view of the source domain into the other encoder-decoder;

[0046] The non-parametric feature projection head mainly functions to map the pixel-level representation h to the hypersphere space, which is usually implemented using a linear transformation or a simple multi-layer perceptron (MLP). Specifically, it is: the input is a 2-channel feature map with a resolution of 256×256, and then through a combination layer of several max-pooling layers and non-linear activation functions to obtain an intermediate layer feature map representation. Then, the intermediate layer feature map is flattened into a feature vector of n×1 dimension. When the input feature map size is 256*2, n = 128. Among them, the max-pooling layer is followed by a non-linear activation function layer, and neither has learnable / trainable parameters. The purpose is to increase the sparse representation of features by the model to a certain extent and improve the sensitivity of the non-parametric feature projection head to important feature vectors;

[0047] Hypersphere space refers to the feature vector space on the unit hypersphere. The L2 norm of each feature vector is 1, which can prevent the size difference of the feature vector from affecting the contrastive learning. Through the non-parametric feature projection head, each pixel-level representation is normalized to ensure that all feature vectors fall on the unit hypersphere. The pixel-level representation is mapped to the hypersphere space mainly to calculate the similarity of the feature vectors in the contrastive learning process to ensure that the size difference of the feature vector does not affect the similarity calculation;

[0048] The contrast loss function is mainly used to learn the consistency representation between domains, ensuring that the corresponding views of the source domain and the target domain are close in the feature space, while the views of different images are far away. The contrast loss function imposes different constraints on positive sample pairs and negative sample pairs: for positive sample pairs, the contrast loss function hopes that the similarity is close to 1; for negative sample pairs, the contrast loss function hopes that the similarity is less than a set threshold. The formula is:

[0049]

[0050] Where L contrast is the contrast loss function, sim(·) is used to calculate the cosine similarity of the feature vector, f i represents a feature vector as an anchor point, f i ′ represents the feature vector of the positive sample corresponding to the anchor point, f j represents the feature vector of negative samples, + represents the feature vector that brings the positive samples closer, - represents the feature vector that pushes the negative samples away, v is the temperature parameter, n is the number of feature vectors, where the number of feature vectors is the same as the number of training batches, and j represents a counting variable;

[0051] The network learning parameters are optimized by minimizing the contrast loss function, and the parameters are updated using the gradient descent method. The gradient of the contrast loss function to the network learning parameters is calculated each time, and the parameters are adjusted according to the gradient to gradually reduce the loss function. This optimization process is iterated repeatedly, and the parameters are continuously updated until the loss function converges, that is, the loss value changes very little or reaches the preset number of iterations. Ultimately, the optimized network learning parameters can minimize the distance between sample pairs in the same field and maximize the distance between sample pairs in different fields in the feature space, achieving the goal of contrastive learning, thereby training an encoder-decoder model that can effectively distinguish samples from different fields;

[0052] S5, fine-tuning the source domain, constructing a semantic segmenter, initializing the segmenter using the network learning parameters trained in step S4, and randomly initializing the segmentation head, and then adjusting the segmenter using the image obtained in step S3 to obtain a trained segmenter; specifically, it includes:

[0053] Construct a semantic segmenter F to learn the anatomical categories of specific classes. Initialize the segmenter F(x) using the prior domain connection knowledge (i.e., the network learning parameters) Θ, and randomly initialize the segmentation head. Then, use the corresponding image labels (x s→t ∈S T , y s ∈Y s ) to fine-tune the segmenter F(x|Θ). The fine-tuning means that the number of adjustments is very small, within 10 times, to obtain the per-pixel semantics of the anatomical structure. Among them, the superscript s represents the prediction result based on the source domain image samples or the prediction results of the source domain image samples; s→t means that the source domain samples have undergone a style transfer transformation from the source domain to the target domain by the generator G; is the output of the segmenter F, y s is the ground truth of the source domain image samples, p represents prediction; λ = 1.0 and η = 2.0 are the contribution coefficients of the hybrid segmentation loss function,

[0054]

[0055]

[0056] Considering the uneven distribution of the semantic category of the anatomical structure in the laryngeal cancer - hypopharyngeal cancer medical images, we use the hybrid segmentation loss L seg to guide the segmenter F. The hybrid segmentation loss L seg is composed of the weighted cross-entropy loss L wce and the Dice loss L dice . Among them, the superscript s represents the prediction result based on the source domain image samples or the prediction results of the source domain image samples; c represents the number of true semantic categories; β is the ratio vector of the pixel ratio of a specific category in a batch to the pixel ratio of all categories; pixels represents the per-pixel traversal operation;

[0057]

[0058]

[0059] The segmenter F is an image segmentation network based on U-Net. The dimension of the input image data of this network is R W×H×3 , and the dimension of the output pixel-level classification probability is R W×H×C , where W, H, and C represent the length, width, and number of classes of the input image, respectively;

[0060] S6. Use the segmenter trained in step S5 to segment the medical images to be segmented.

[0061] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

[0062] For those skilled in the art, various changes and modifications will undoubtedly be obvious after reading the above description. Therefore, the appended claims should be regarded as covering all changes and modifications that embrace the true spirit and scope of the present invention. Any and all equivalent ranges and contents within the scope of the claims should be considered to still fall within the spirit and scope of the present invention.

Claims

1. An adaptive segmentation and measurement method for laryngeal cancer - hypopharyngeal cancer medical images, characterized in that It includes the following steps: S1. Select the dataset: Construct the source domain images and target domain images; S2. Data preprocessing: Augment the source domain images and target domain images constructed in step S1; S3. Domain connection construction: Within the domain, obtain the global view of the target domain image through the identity mapping of the target domain image preprocessed in step S2, and then perform a random number enhancement operation on the target domain image to obtain the local view of the target domain image; Between the domains, input the source domain image preprocessed in step S2 into the trained generator for style transfer conversion to obtain the global view of the fake target domain image, and then perform a random number enhancement operation on the source domain image to obtain the local view of the source domain image; Within the domain, use the global view and local view of the target domain image as a positive sample pair. Between the domains, use the global view of the fake target domain image obtained by inputting the source domain image into the trained generator for style transfer conversion and the local view of the source domain image obtained through the random number enhancement operation as a positive sample pair; Within the domain, use the local view and global view of different target domain images as negative sample pairs. Between the domains, use the local view of different source domain images and the global view of the fake target domain image obtained by the generator as negative sample pairs; S4. Contrastive learning pre-training: Input the intra-domain views and inter-domain views obtained in step S3 into the encoder-decoder with shared parameters to obtain pixel-level representations. Map the pixel-level representations to the hypersphere space through a designed non-parametric feature projection head, calculate the loss using the contrastive loss function, and then optimize the network learning parameters through backpropagation to minimize the distance between positive sample pairs between the domains and maximize the distance between negative sample pairs. Then, continuously iterate and optimize until the contrastive loss function converges, and finally obtain the trained network learning parameters; S5. Source domain fine-tuning: Construct a semantic segmenter, initialize the segmenter with the network learning parameters trained in step S4, and randomly initialize the segmentation head at the same time. Then, use the images obtained in step S3 to adjust the segmenter to obtain a trained segmenter; S6. Use the segmenter trained in step S5 to segment the medical images to be segmented.

2. The method for self-adaptive segmentation and measurement of laryngeal cancer - hypopharyngeal cancer medical images according to claim 1, wherein: The trained generator in step S3 refers to the CycleGAN or CUT generator trained with the data preprocessed in step S2.

3. A method for adaptive segmentation and measurement of laryngeal cancer - hypopharyngeal cancer medical images according to claim 1, characterized in that: The formula of the contrastive loss function in step S4 is: where L contrast is the contrastive loss function, sim(·) is used to calculate the cosine similarity of feature vectors, and f i represents a feature vector as an anchor, f i ′ represents the feature vector of the positive sample corresponding to the anchor, f j represents the feature vector of the negative sample, + means to pull closer the feature vector of the positive sample, - means to push away the feature vector of the negative sample, v is the temperature parameter, n is the number of feature vectors, and j represents the counting variable.

4. A method for adaptive segmentation and measurement of laryngeal cancer - hypopharyngeal cancer medical images according to claim 1, characterized in that: The non-parametric feature projection head in step S4 refers to: input the feature map, then obtain the intermediate layer feature map through a combination layer of several max pooling layers and non-linear activation functions, and then perform a vector flattening operation on the intermediate layer feature map to obtain the feature vector.

5. A method for adaptive segmentation and measurement of laryngeal cancer - hypopharyngeal cancer medical images according to claim 1, characterized in that: After fine-tuning the segmenter with the images obtained in step S3 in step S5, it is also necessary to use the hybrid segmentation loss to guide the segmenter, and the hybrid segmentation loss is composed of the weighted cross-entropy loss and the Dice loss.

6. A method for adaptive segmentation and measurement of laryngeal cancer - hypopharyngeal cancer medical images according to claim 5, characterized in that: The formulas of the segmenter and the hybrid segmentation loss are: Among them, F(x) is the segmenter, Θ is the network learning parameter, and the superscript s represents the source-domain image samples or the prediction results based on the source-domain image samples; s→t means that the source-domain samples have undergone a style transfer transformation from the source domain to the target domain by the generator G; is the output of the segmenter F, y s is the ground truth of the source-domain image samples, p represents prediction; λ = 1.0 and η = 2.0 are the contribution coefficients of the hybrid segmentation loss function; The superscript s represents the source-domain image samples or the prediction results based on the source-domain image samples; c represents the number of true semantic categories; β is the ratio vector of the pixel ratio of a specific category in a batch to the pixel ratio of all categories, and pixels represents a pixel-by-pixel traversal operation.

Citation Information

Patent Citations

  • Pathological image segmentation method based on domain adversarial self-supervised learning

    CN113379764A

  • Feature prototype-based semi-supervised domain adaptive semantic segmentation method and system

    CN114529900A