A method, system and medium for colorectal cancer image segmentation

Through student-teacher network architecture and pseudo-labeling technology, the problem of low tumor area segmentation accuracy in colorectal cancer CT images is solved, the dependence on labeled data is reduced, and the segmentation accuracy and generalization ability are improved.

CN119741499BActive Publication Date: 2025-06-17THE SECOND AFFILIATED HOSPITAL OF NAVAL MEDICAL UNIVERSITY PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510214030.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-17
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The prior art faces the characteristics of blurred boundaries, complex shapes, and different sizes when segmenting tumor areas in colorectal cancer CT images, resulting in low segmentation accuracy and relying on manual labeling of data, which is expensive and scarce labeling of data.

Method used

The student teacher network architecture is used in combination with pseudo-label technology, and the labeled data is used for supervising learning, and the pseudo-label generation is optimized by unlabeled data to reduce the dependence on labeled data. The teacher network updates parameters through exponential moving averages to improve the stability and quality of pseudo-label generation.

Benefits of technology

It significantly reduces the need for labeled data, makes full use of unlabeled data, improves the accuracy and generalization ability of colorectal cancer image segmentation, especially improves segmentation performance in complex areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741499B_ABST
    Figure CN119741499B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of image segmentation, and particularly relates to a method, system and medium for colorectal cancer image segmentation. The method includes the following steps: obtaining colorectal cancer image annotation data and unannotated data; initializing a student-teacher network architecture to obtain a student network model and a teacher model; performing supervised learning on the student network with the annotation data to obtain a preliminarily optimized student network model; inputting the unannotated data into the preliminarily optimized student network to generate prediction data, and performing pseudo-label calculation with the unannotated data to obtain pseudo-label data; optimizing the student network in combination with the pseudo-label data to obtain a secondary optimized student network model; using the secondary optimized student network to update the teacher model, and further optimizing the teacher model with the annotation data to obtain a colorectal cancer image segmentation model. The present invention utilizes limited annotation data and combines a large amount of unannotated data to realize a colorectal cancer image segmentation model with high precision and strong robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation, and in particular to a method, system and medium for colorectal cancer image segmentation. Background Art

[0002] With the rapid development of medical imaging technology, in the early screening and diagnosis of colorectal cancer, CT (Computed Tomography) images have become an important non-invasive means. Through CT images, the internal tissue structure and abnormal areas of the human body can be observed, especially the growth of tumors. However, due to the characteristics of blurred boundaries, complex shapes, and different sizes in the tumor area, accurately segmenting the tumor area in CT images remains a huge technical challenge.

[0003] With the rise of deep learning, Convolutional Neural Networks (CNNs) have been widely used in medical image segmentation tasks. Among them, U-Net and its variants have become the most commonly used segmentation architectures. The deep learning-based method has the following advantages: it can automatically learn multi-level features in the image, thus achieving high-precision segmentation. By introducing convolution operations, it can capture the global context and local details of the tumor area. However, the deep learning-based segmentation method also faces the following problems: such as dependence on labeled data. The annotation of medical images requires manual operation by clinical experts, which is a cumbersome and costly process. The scarcity of labeled data limits the performance of deep learning models; insufficient generalization ability. Since the model mainly relies on labeled data during training, its performance degrades when dealing with unseen data or when the data distribution changes. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention proposes a method, system and medium for colorectal cancer image segmentation to solve at least one of the above technical problems.

[0005] The present application provides a method for colorectal cancer image segmentation, including the following steps:

[0006] Step S1: Obtain colorectal cancer image data, where the colorectal cancer image data includes colorectal cancer image annotation data and colorectal cancer image unlabeled data;

[0007] Step S2: Initialize the student-teacher network architecture according to the colorectal cancer image data to obtain an initialized student network model and an initialized teacher model respectively;

[0008] Step S3: Perform supervised learning on the initialized student network model according to the colorectal cancer image annotation data to obtain a preliminarily optimized student network model; input the unannotated data of the colorectal cancer images into the preliminarily optimized student network model for calculation to obtain preliminary student model prediction data, and perform pseudo-label calculation on the unannotated data of the colorectal cancer images and the preliminary student model prediction data to obtain pseudo-label data; optimize the preliminarily optimized student network model according to the preliminary student model prediction data and the pseudo-label data to obtain a secondary optimized student network model;

[0009] Step S4: Update the parameters of the initialized teacher model according to the secondary optimized student network model to obtain an updated teacher model, and optimize the updated teacher model according to the colorectal cancer image annotation data to obtain a colorectal cancer image segmentation model.

[0010] In the present invention, by introducing the pseudo-label technology and the student-teacher network architecture, the unannotated data of colorectal cancer images is fully utilized, reducing the over-reliance on annotated data. The demand for annotated data is significantly reduced, and at the same time, the potential information contained in the unannotated data is mined. The teacher network is updated through exponential moving average (EMA), with smoother parameters and stronger robustness to noise, providing high-quality guidance for the generation of pseudo-labels for unannotated data. Based on supervised learning, the student network is further optimized through pseudo-labels, iteratively improving the segmentation performance. Pseudo-label generation combines the student model prediction with unannotated data, improving the credibility of pseudo-labels. Regional verification and dynamic optimization are performed on the pseudo-label data to eliminate low-confidence regions and incorrect labels, improving the quality of pseudo-labels. Under the combined action of annotated data and unannotated data, the model undergoes multiple optimizations (such as preliminary optimization, secondary optimization, teacher network update) to form a closed-loop improvement mechanism. The parameters of the student network and the teacher network gradually approach the optimal, making the performance of the generated image segmentation model more excellent.

[0011] Preferably, step S1 is specifically:

[0012] Interact with the multi-center hospital imaging database through the medical imaging data interface to obtain preliminary colorectal cancer CT imaging data;

[0013] Group the preliminary colorectal cancer CT imaging data to obtain annotated data and unannotated data respectively;

[0014] Extract annotation features from the annotated data to obtain annotation feature data;

[0015] Screen the unannotated data to obtain unannotated screened data;

[0016] Perform data augmentation on the annotation feature data and the unannotated screened data to obtain colorectal cancer image annotation data and colorectal cancer image unannotated data respectively.

[0017] In the present invention, colorectal cancer CT image data is obtained from multiple authoritative medical data sources, ensuring the diversity, representativeness, and true clinical value of the data. The grouping of labeled data and unlabeled data is the basis for subsequent supervised learning and pseudo-label generation. The labeled data is extracted separately to ensure the high quality of supervised learning; the unlabeled data is separated to facilitate pseudo-label generation and optimization. It improves the feature expression ability of the labeled data and reduces the redundant burden of model training. Focus on key regions (such as tumor boundaries) to improve the learning efficiency and segmentation accuracy of the model. Use quality assessment algorithms (such as clarity, noise level, integrity) to screen unlabeled data and eliminate low-quality data (such as samples with low resolution, high noise, or data corruption). By screening, the quality of the unlabeled data is guaranteed, providing high-quality input for pseudo-label generation. It expands the scale of labeled data and unlabeled data, alleviating the problem of insufficient labeled data. It enhances the adaptability of the model to different scenarios and different tumor morphologies.

[0018] Preferably, step S2 is specifically as follows:

[0019] Extract dimensional characteristic features based on the colorectal cancer image data to obtain dimensional characteristic feature data;

[0020] Generate a network structure based on the dimensional characteristic feature data to obtain network architecture metadata;

[0021] Initialize the student-teacher network architecture based on the network architecture metadata to obtain an initialized student network model and an initialized teacher model respectively.

[0022] In the present invention, dimensionality characteristics analysis is performed on CT image data of colorectal cancer, including features such as image resolution, spatial dimension (2D or 3D), signal intensity distribution, etc. These characteristics are extracted as an expression of the dimension and complexity of the input data, and are mapped through a preset parameter library to design an adapted network architecture. Images generated by different CT devices or scanning protocols have different dimensionality characteristics (such as slice thickness, resolution). By extracting dimensionality characteristics, the generated network architecture can adapt to different input data, enhancing the generality of the method. For images with different resolutions, the parameters of the input layer and intermediate layer of the network are adjusted, avoiding unnecessary computational overhead. Combining dimensionality characteristics, a network architecture is automatically generated, including key parameters such as the dimension of the input layer, the number of convolutional layers, the number of feature channels, the number of downsamplings, etc. An architecture generation method based on search (such as NAS, Neural Architecture Search) or a rule-driven method is used to generate an adapted network architecture. The network architecture generated based on data characteristics can better meet the actual data requirements, thereby improving the training efficiency and segmentation performance of the model. The student network is used for direct training, while the teacher network is smoothly updated through the exponential moving average (EMA) parameters of the student network. In the initial stage, the weights of the teacher network are the same as those of the student network, ensuring the stability of training. Through the EMA smoothing parameter update of the teacher network, the impact of parameter fluctuations of the student network on the model performance during the training process is avoided. The initialized student and teacher network parameters are the same, providing a reliable initial condition for pseudo-label generation.

[0023] Preferably, step S3 is specifically as follows:

[0024] Perform supervised learning on the initialized student network model according to the annotation data of colorectal cancer images to obtain a preliminary optimized student network model;

[0025] Input the unannotated data of colorectal cancer images into the preliminary optimized student network model for calculation to obtain preliminary student model prediction data;

[0026] Perform spatial consistency pseudo-label calculation on the unannotated data of colorectal cancer images and the preliminary student model prediction data to obtain spatial consistency pseudo-label data;

[0027] Perform contrastive learning pseudo-label calculation on the unannotated data of colorectal cancer images and the preliminary student model prediction data to obtain contrastive learning pseudo-label data;

[0028] Perform confidence screening according to the spatial consistency pseudo-label data and the contrastive learning pseudo-label data to obtain pseudo-label data;

[0029] Optimize the preliminary optimized student network model according to the preliminary student model prediction data and the pseudo-label data to obtain a secondary optimized student network model.

[0030] In the present invention, supervised learning is carried out using labeled data to improve the initial segmentation ability of the student network. At the same time, through the pseudo-label generation technology, the potential information of unlabeled data is fully exploited to improve the learning efficiency of the model. The spatial consistency constraint is adopted to generate pseudo-labels, and the spatial relationship between pixels is used to ensure the smoothness and continuity of the pseudo-labels. Techniques such as conditional random field (CRF) or joint histogram are introduced to make the pseudo-labels more accurate at the boundaries and in small regions. The contrastive learning method is used to generate pseudo-labels, and through feature similarity analysis, the discrimination ability of the pseudo-labels for complex backgrounds and fuzzy regions is enhanced. Pseudo-labels are generated through dynamic contrast target sampling and spatial perception characteristics, making the model more adaptable to complex scenarios. Through the confidence screening of pseudo-labels, high-quality labels are retained and low-quality labels are removed. A dynamic threshold (such as a dynamic screening strategy based on the mean and standard deviation) is used to ensure the flexibility of the screening results. Supervised learning is used for the preliminary optimization of the model, and pseudo-label generation and optimization are used to further improve the segmentation performance. The model gradually achieves a stable improvement in performance through preliminary optimization and secondary optimization. At the same time, spatial consistency pseudo-labels and contrastive learning pseudo-labels are introduced. By generating pseudo-labels from multiple angles, the diversity and adaptability of pseudo-label generation are enhanced. The two types of pseudo-labels are fused using confidence screening, retaining the high-quality part and removing the noise. The supervised learning of labeled data improves the initial segmentation performance, and the optimization of pseudo-labels of unlabeled data enhances the generalization ability of the model. In some special scenarios (such as the tumor boundary of colorectal cancer being blurred and the segmentation of small lesion regions), the model performs more excellently.

[0031] Preferably, step S4 is specifically as follows:

[0032] The parameter of the initialized teacher model is updated according to the secondary optimized student network model to obtain an updated teacher model;

[0033] The updated teacher model is optimized according to the colorectal cancer image annotation data to obtain a teacher annotation optimized model;

[0034] The teacher annotation optimized model is optimized according to the pseudo-label data to obtain a teacher pseudo-label optimized model;

[0035] The weights of the teacher pseudo-label optimized model are frozen to obtain a colorectal cancer image segmentation model.

[0036] In the present invention, the parameters of the teacher model are updated through the exponential moving average (EMA) of the student model, which smooths the noise and fluctuations in training and makes the teacher model more stable. The teacher model gradually approaches the optimal parameters through the EMA mechanism, ensuring that the generated pseudo-labels have high quality. The teacher model is optimized using the labeled data to make it more accurate in segmenting the tumor region (especially the region with complex boundaries). The optimized teacher model is more reliable when generating pseudo-labels, avoiding the negative impact of low-quality labels on model training. The teacher model is optimized using high-quality pseudo-labels to enable it to extract more potential information from the unlabeled data, further improving the generalization ability of the model. Through pseudo-label optimization, the teacher model can absorb the knowledge of both labeled data and unlabeled data, realizing the advantages of semi-supervised learning. After the teacher model is optimized, its weights are frozen to generate a segmentation model, avoiding parameter drift in further training. The frozen model can be directly used for inference, improving the efficiency and stability in clinical applications. The parameter update of the teacher model depends on the secondary optimization results of the student model, and the collaborative optimization of the two forms a closed loop, gradually improving the performance of the overall network. The student model focuses on dynamically learning the labeled data and pseudo-labels, while the teacher model steadily improves its performance through smooth parameter updates.

[0037] Preferably, the calculation of the spatially consistent pseudo-labels is specifically as follows:

[0038] Calculate the geological distance for the unlabeled data of colorectal cancer images and the prediction data of the preliminary student model to obtain the geological distance map data;

[0039] Calculate the confidence for the prediction data of the preliminary student model to obtain the confidence map data;

[0040] Perform multi-scale analysis based on the geological distance map data and the confidence map data to obtain the multi-scale data of the geological distance map and the multi-scale data of the confidence map respectively;

[0041] Perform spatial consistency constraints based on the multi-scale data of the geological distance map and the multi-scale data of the confidence map to obtain the consistent segmentation data;

[0042] Perform confidence screening on the prediction data of the preliminary student model based on the consistent segmentation data to obtain the spatially consistent pseudo-label data.

[0043] In the present invention, the geological distance calculation can effectively capture the spatial relationship between pixels, especially model the spatial continuity of the boundary of the tumor region. The geological distance map highlights the spatial separation between the target region (such as a tumor) and the background region, enabling the model to more accurately identify the tumor region. According to the prediction probability of the student model, confidence map data is generated to evaluate the classification credibility of each pixel. The pseudo-labels in the high-confidence regions are more credible, and the low-confidence regions are eliminated through a screening mechanism to ensure the high quality of the pseudo-labels. Multi-scale analysis processes the geological distance map and the confidence map at different resolutions to capture the global features and local details of the target region. The boundary details of the tumor are captured at a small scale, and the global characteristics are captured at a large scale, enhancing the model's segmentation ability for complex regions. Using the multi-scale data of geological distance and confidence, the spatial smoothness and coherence of the pseudo-labels are ensured through consistency constraints. The consistency segmentation data further improves the accuracy of the pseudo-labels by eliminating isolated pixels and discontinuous regions. The low-confidence regions in the consistency segmentation data are eliminated, and only the high-confidence pseudo-label data is retained to further improve the reliability of the pseudo-labels. Through the threshold screening mechanism, the flexibility and adaptability of pseudo-label screening are ensured.

[0044] Preferably, the calculation of the contrast learning pseudo-label is specifically as follows:

[0045] For the unlabeled data of colorectal cancer images, multi-level feature extraction is performed through preliminary optimization of the student network model to obtain multi-level feature data;

[0046] The multi-level features are processed in blocks to obtain local feature embedding data;

[0047] Feature alignment is performed based on the local feature embedding data to obtain globally aligned feature data;

[0048] Contrastive target sampling is performed on the globally aligned feature data to obtain contrastive target data;

[0049] Contrastive learning data augmentation is performed based on the contrastive target data to obtain enhanced contrastive target data;

[0050] Feature aggregation is performed based on the enhanced contrastive target data to obtain aggregated feature data;

[0051] Cross-node contrastive consistency optimization is performed based on the aggregated feature data to obtain consistency-optimized feature data;

[0052] Similarity pseudo-label generation is performed based on the consistency-optimized feature data to obtain preliminary pseudo-label data;

[0053] Pseudo-label region verification is performed based on the preliminary pseudo-label data and the preliminary student model prediction data to obtain pseudo-label verification data;

[0054] Perform Gaussian smoothing on the data verified according to the pseudo-labels to obtain the contrastive learning pseudo-label data;

[0055] Among them, the contrastive target sampling is specifically as follows:

[0056] Perform target region boundary analysis on the globally aligned feature data to obtain boundary feature data;

[0057] Perform target region confidence evaluation based on the boundary feature data to obtain target region confidence evaluation data;

[0058] When it is determined that the target region area data corresponding to the target region confidence evaluation data is greater than or equal to the preset target region area threshold data, then perform sparse sampling according to the target region confidence evaluation data to obtain contrastive target sampling data;

[0059] When it is determined that the target region area data corresponding to the target region confidence evaluation data is less than the preset target region area threshold data, then perform dense sampling according to the target region confidence evaluation data to obtain contrastive target sampling data;

[0060] Perform confidence-guided target screening according to the contrastive target sampling data to obtain confidence screening data;

[0061] Perform spatial perception target generation according to the confidence screening data to obtain spatial perception target data;

[0062] Perform region and background contrastive target sampling according to the spatial perception target data to obtain contrastive target data.

[0063] In the present invention, multi-level feature data capturing shallow features (such as texture and edge information) and deep features (such as semantic information) is ensured to comprehensively capture features of different scales and complexities by the model. The computational complexity is reduced through the chunking and embedding mechanisms, while enhancing the coherence of the local region feature expression. The alignment process eliminates the distribution differences of features between different scales and regions, providing high-quality basic features for subsequent pseudo-label generation. The sampling strategy is dynamically adjusted to generate contrastive target data based on confidence, spatial perception, and boundary shape, ensuring the pertinence and diversity of the sampling data. The sampling strategy ensures that the sampling focuses on the tumor core region and the key boundary regions, avoiding excessive attention to non-related background regions. Data augmentation makes the model more adaptable to tumor regions of different morphologies, scales, and orientations, reducing the dependence on specific data distributions. Aggregation enhances the features of the contrastive target data, and consistency optimization is used to ensure the coherence of feature expression globally and locally. The similarity pseudo-label generation strategy effectively captures the feature differences between the target region and the background, avoiding the generation of incorrect labels. Gaussian smoothing processing reduces the isolated noise points in the pseudo-labels, enhancing the spatial continuity of the pseudo-labels.

[0064] Extract the boundary information of the target area using boundary detection algorithms (such as Canny or Sobel operators), and optimize the boundary features by combining morphological operations. The extracted boundary feature data defines the shape, size, and boundary complexity of the target area. Calculate the average confidence of the target area based on the prediction results of the student network to evaluate the credibility of the model's classification of the target area. High-confidence areas are marked as reliable targets, and low-confidence areas are further processed through subsequent screening mechanisms. Sparse sampling is used for larger target areas to avoid redundant data affecting the calculation efficiency. Dense sampling is used for smaller target areas to capture more detailed information and ensure sufficient learning of small target areas. Based on the confidence ranking, eliminate the sampling points with lower confidence and only retain the target data with high confidence. The screened target data is more accurate and helps with subsequent feature processing and contrast learning. Combine spatial distribution information (such as distance, density, and adjacency relationships) to generate a spatial perception matrix to clarify the spatial relationships of the target areas. The generation of spatially aware targets makes the target data more conform to the actual spatial distribution characteristics. Sample the high-confidence pixels within the target area and simultaneously compare the pixels with the greatest feature differences from the target area in the background area. The contrast sampling between the area and the background enables the model to better distinguish the target area from the background area.

[0065] Preferably, the multi-scale analysis specifically is:

[0066] Perform preliminary multi-scale partitioning based on the geological distance map data and the confidence map data to obtain the preliminary multi-scale data of the geological distance map and the preliminary multi-scale data of the confidence map;

[0067] Perform global-local entropy calculation based on the preliminary multi-scale data of the geological distance map and the preliminary multi-scale data of the confidence map to obtain the multi-scale entropy data of the geological distance map and the multi-scale entropy data of the confidence map;

[0068] Perform multi-scale entropy similarity calculation based on the multi-scale entropy data of the geological distance map and the multi-scale entropy data of the confidence map to obtain the multi-scale entropy similarity data;

[0069] Perform non-similar entropy screening on the preliminary multi-scale data of the geological distance map and the preliminary multi-scale data of the confidence map based on the multi-scale entropy similarity data to obtain the multi-scale screening data of the geological distance map and the multi-scale screening data of the confidence map respectively;

[0070] Perform consistency distribution calculation based on the multi-scale screening data of the geological distance map and the multi-scale screening data of the confidence map to obtain the multi-scale data of the geological distance map and the multi-scale data of the confidence map.

[0071] In the present invention, the input data is divided into multiple scales, and the geological distance map and the confidence map are divided into different scale versions to adapt to different resolutions and target sizes. The initial division captures global and local features, enabling the data to exhibit diversity at different scales. Entropy calculation is used to measure the complexity and uncertainty of information in the image, distinguishing structured regions (such as tumors) from unstructured regions (such as the background). Global entropy calculation evaluates the overall information distribution, while local entropy calculation captures the information complexity within a small range. Entropy similarity is used to quantify the similarity between the two at multiple scales, capturing the mutual relationship between geological distance and confidence. Similarity calculation determines the distribution consistency between multi-scale data and eliminates inconsistent regions. Inconsistent or abnormal regions in the multi-scale data are screened out, and regions with high noise or low confidence are removed. The regions to be removed are determined by threshold methods (such as screening strategies based on mean and standard deviation), ensuring the flexibility and robustness of the screening. The consistent distribution calculation ensures the distribution coherence of multi-scale data at different resolutions, providing high-quality input for pseudo-label generation and model optimization. The combination of multi-scale data of the geological distance map and the confidence map provides more spatial and semantic information for pseudo-label generation. After removing abnormal regions, pseudo-label generation is more accurate, reducing the interference of incorrect labels.

[0072] Preferably, the present application also provides a colorectal cancer image segmentation system for performing the colorectal cancer image segmentation method as described above. The system includes:

[0073] A colorectal cancer image data acquisition module for acquiring colorectal cancer image data, where the colorectal cancer image data includes colorectal cancer image annotation data and colorectal cancer image unannotated data;

[0074] A student-teacher network architecture initialization module for initializing the student-teacher network architecture according to the colorectal cancer image data to obtain an initialized student network model and an initialized teacher model respectively;

[0075] A secondary optimized student network model construction module for performing supervised learning on the initialized student network model according to the colorectal cancer image annotation data to obtain a preliminary optimized student network model; inputting the colorectal cancer image unannotated data into the preliminary optimized student network model for calculation to obtain preliminary student model prediction data, and performing pseudo-label calculation on the colorectal cancer image unannotated data and the preliminary student model prediction data to obtain pseudo-label data; optimizing the preliminary optimized student network model according to the preliminary student model prediction data and the pseudo-label data to obtain a secondary optimized student network model;

[0076] A colorectal cancer image segmentation model generation module, configured to update the parameters of the initialized teacher model according to the secondary optimized student network model to obtain an updated teacher model, and optimize the updated teacher model according to the colorectal cancer image annotation data to obtain a colorectal cancer image segmentation model.

[0077] Preferably, the present application further provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the method described in any one of the above when running.

[0078] The beneficial effects of the present invention are as follows: By combining labeled data and unlabeled data, the present invention uses the labeled data to supervise and train the student network, and at the same time fully mines the unlabeled data through the pseudo-label generation technology. The utilization rate of unlabeled data is significantly improved, reducing the dependence on manually labeled data, thereby reducing the labeling cost. The pseudo-label generation combines the spatial consistency pseudo-label calculation and the contrastive learning pseudo-label calculation, and generates accurate and reliable pseudo-labels from the perspectives of spatial consistency and feature contrast respectively. The spatial consistency pseudo-label uses the geodesic distance and the confidence map, combines multi-scale analysis and consistency distribution calculation, and improves the spatial coherence of the pseudo-label. The contrastive learning pseudo-label generates high-quality pseudo-labels through feature alignment, dynamic sampling and similarity calculation, significantly improving the effective utilization of unlabeled data. The teacher network updates the parameters through the exponential moving average (EMA) of the student network, and the parameters are smoother and more robust to noise, improving the stability of pseudo-label generation. The teacher network gradually improves the segmentation ability for complex regions (such as regions with blurred boundaries and small lesions) through multiple optimizations of the labeled data and the pseudo-label data. The student network forms a closed loop through supervised learning and semi-supervised pseudo-label optimization, and continuously improves the segmentation performance during training. By combining the supervised learning of the labeled data and the semi-supervised pseudo-label optimization of the unlabeled data, the present invention makes the model training more efficient. The teacher network parameters are updated through EMA, further smoothing the parameter changes and avoiding oscillations and instabilities during the training process. The present invention is applicable to scenarios where the labeled data is insufficient but the unlabeled data is rich, and significantly reduces the dependence on the labeled data through the pseudo-label generation technology. In medical image segmentation tasks, the labeled data is often scarce and expensive, so the present invention has high application value. Description of the Drawings

[0079] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects and advantages of the present application will become more obvious:

[0080] Figure 1 Shows a step flow chart of a method for colorectal cancer image segmentation according to an embodiment;

[0081] Figure 2The flowchart of the steps of a method for collecting colorectal cancer image data according to an embodiment is shown;

[0082] Figure 3 The flowchart of the steps of a method for initializing a student-teacher network architecture according to an embodiment is shown;

[0083] Figure 4 The flowchart of the steps of a method for constructing a secondary optimized student network model according to an embodiment is shown;

[0084] Figure 5 The flowchart of the steps of a method for generating a colorectal cancer image segmentation model according to an embodiment is shown. Detailed implementation manners

[0085] The technical method of the present invention patent will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0086] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.

[0087] It should be understood that although the terms "first", "second", etc. may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.

[0088] The labeled dataset is 100,000 colorectal cancer CT images (labeled, each image size is 512×512 pixels). The unlabeled dataset is 500,000 colorectal cancer CT images (unlabeled). The model architecture is that the student network is based on the U-Net segmentation architecture of ResNet-50. The teacher network has the same architecture as the student network, and the parameters are updated through EMA. The student network is trained using labeled data and unlabeled data, and the utilization efficiency of unlabeled data is improved through pseudo-label generation. The goal is to improve the segmentation performance (measured by the Dice coefficient, the higher the better).

[0089] Obtain labeled data and unlabeled data. The labeled data is 100,000 CT images, and each image is labeled with the pixel mask of the tumor area. The unlabeled data is 500,000 CT images without any annotations. Normalize all images and adjust the pixel value range to [0,1]. Flip, rotate (random angle) and scale the labeled data to expand the labeled data set to 300,000 images.

[0090] The student network is a U-Net architecture based on ResNet-50, with randomly initialized weights. The initial parameters of the teacher network are the same as those of the student network. The input image size is 512×512. The output segmentation map size is 512×512. The initial convolutional layer output channels are 64, which increase layer by layer. The labeled data (300,000 images) is used as input. The supervised learning uses Dice loss. , where P is the predicted mask and G is the true mask. After preliminary training on the labeled data, the Dice coefficient of the student network on the validation set increased from the initial 0.55 to 0.72.

[0091] 500,000 unlabeled data were input into the student network to generate preliminary prediction results. In a certain unlabeled image, the student network predicted that the mask size of the tumor area was 2000 pixels. Through distance transformation, the distance from the tumor boundary to the target pixel was obtained to obtain a geological distance map. Based on the classification probability predicted by the student network, the confidence distribution of each pixel was generated to obtain a confidence map. The geological distance map and confidence map were divided into blocks, and the local entropy and global entropy were calculated. For example, the local entropy distribution of a certain image showed that the entropy value of the boundary area was 0.85, and the entropy value of the central area was 0.65. Pixels with confidence values ​​lower than 0.6 were removed to obtain spatially consistent pseudo labels.

[0092] Dynamic sampling is performed on the target area, with dense sampling for small tumor areas (area < 1500 pixels) and sparse sampling for large tumor areas. The similarity of features in different regions is calculated by cosine similarity, and pseudo labels for comparative learning are screened out.

[0093] The student network is optimized using pseudo-label training, using the loss function: ,in It is the Dice loss based on pseudo-labels. After optimization with pseudo-labels, the Dice coefficient of the student network increased from 0.72 to 0.79.

[0094] Use the student network to update the teacher network. The EMA update formula: , where are the model weights of the current teacher network, including convolutional kernel weights, biases, etc. are the model weights of the current student network, including convolutional kernel weights, biases, etc., and λ = 0.99. Further optimize the teacher network using the labeled data to enhance the segmentation ability for high-confidence regions. Use the pseudo-labeled data for secondary optimization to improve the generalization ability of the teacher network. After optimization, freeze the weights of the teacher network as the segmentation model.

[0095] On the test set (500,000 unlabeled CT images), the Dice coefficient of the segmentation model reached 0.84, a 12% improvement compared to the model trained only with labeled data (Dice coefficient 0.72). The training time for the labeled data was 500 hours. The time for generating pseudo-labels for the unlabeled data was 100 hours. The total training time was 600 hours.

[0096] Please refer to Figures 1 to 5 , this application provides a method for colorectal cancer image segmentation, including the following steps:

[0097] Step S1: Obtain colorectal cancer image data, where the colorectal cancer image data includes colorectal cancer image annotation data and colorectal cancer image unlabeled data;

[0098] Specifically, by cooperating with multiple medical institutions, use the medical imaging data interface to obtain the CT image data of colorectal cancer patients. This data includes colorectal cancer image annotation data with expert annotations and unlabeled data. Among the obtained image data, first group the annotation data and unlabeled data. The annotation data is used for supervised learning training, and the unlabeled data is used to generate pseudo-labels. In the annotation data, extract the regional feature information of colorectal cancer, such as the size, shape, boundary, etc. of the tumor. Perform data augmentation processing on the annotation data and unlabeled data respectively, including rotation, translation, scaling, flipping, cropping, etc., with the aim of increasing the data diversity and thus improving the generalization ability of the model.

[0099] Step S2: Initialize the student-teacher network architecture according to the colorectal cancer image data to obtain the initialized student network model and the initialized teacher model respectively;

[0100] Specifically, based on colorectal cancer image data, the dimensional characteristic features of the image are extracted through a deep learning model, such as the texture, shape, and structural information of the image. According to the extracted dimensional characteristic feature data, a network structure suitable for colorectal cancer image segmentation is generated. This network structure includes a student network and a teacher network. The student network is used to learn the target segmentation task, while the teacher network is used to guide the learning direction of the student network. The generated network architecture metadata is used to initialize the student network model and the teacher model. The parameters of the initial student network and teacher network are usually randomly initialized, or initialized using parameters pre-trained from other image segmentation tasks to accelerate the convergence speed.

[0101] Step S3: Perform supervised learning on the initialized student network model according to the colorectal cancer image annotation data to obtain a preliminarily optimized student network model; input the unannotated colorectal cancer image data into the preliminarily optimized student network model for calculation to obtain preliminary student model prediction data, and perform pseudo-label calculation on the unannotated colorectal cancer image data and the preliminary student model prediction data to obtain pseudo-label data; optimize the preliminarily optimized student network model according to the preliminary student model prediction data and the pseudo-label data to obtain a secondary optimized student network model.

[0102] Specifically, input the colorectal cancer image annotation data into the initialized student network for supervised learning training. During the training process, use the cross-entropy loss function to optimize the student network so that it can accurately predict the colorectal cancer region. After several rounds of training, a preliminarily optimized student network model is obtained. Input the unannotated colorectal cancer image data into the preliminarily optimized student network for prediction and generate the prediction data of the preliminary student model. Perform pseudo-label calculation based on the prediction data of the student model and the data of the unannotated image. The pseudo-label refers to the inference result of the student model for the unannotated data. The pseudo-label will be regarded as a "virtual label" for further training the model. Combine the prediction data of the preliminary student model and the pseudo-label data as the training set to continue optimizing the student network model. Through the backpropagation algorithm, the student network adjusts according to the difference between the true label and the pseudo-label, thereby obtaining a secondary optimized student network model.

[0103] Step S4: Update the parameters of the initialized teacher model according to the secondary optimized student network model to obtain an updated teacher model, and optimize the updated teacher model according to the colorectal cancer image annotation data to obtain a colorectal cancer image segmentation model.

[0104] Specifically, according to the secondary optimized student network model, the parameters of the initialized teacher model are updated. The knowledge gradually learned by the student network during the training process will be used to guide the teacher network, thereby improving the accuracy of the teacher model. The update of the teacher model uses a relatively slow learning rate to avoid over-adjustment. The updated teacher model is further optimized using colorectal cancer image annotation data. Through supervised learning, the teacher model is trained on the annotation data and its segmentation accuracy is further improved. The optimization process includes fine-tuning the parameters of the teacher model to better match the annotation data. In addition to using annotation data, pseudo-label data can also be used to optimize the teacher model, increasing the model's learning ability on unlabeled data and improving the model's robustness. After the teacher model is fully optimized, its weight parameters are frozen to avoid further training. At this time, the teacher model already has strong colorectal cancer image segmentation ability, and the obtained teacher model is the colorectal cancer image segmentation model, which can automatically segment the colorectal cancer area on new images.

[0105] Preferably, step S1 is specifically as follows:

[0106] Step S11: Interact with the multi-center hospital imaging database through the medical imaging data interface to obtain preliminary colorectal cancer CT imaging data;

[0107] Specifically, interact with the imaging database of the multi-center hospital through the medical imaging data interface to obtain the CT imaging data of colorectal cancer patients. This database contains the CT scan images of multiple colorectal cancer patients and their related pathological information.

[0108] Step S12: Group the preliminary colorectal cancer CT imaging data to obtain annotation data and unlabeled data respectively;

[0109] Specifically, the obtained preliminary colorectal cancer CT imaging data is divided into two groups according to whether it contains expert annotations: annotation data and unlabeled data. Annotation data refers to the CT images of colorectal cancer tumors that have been annotated by medical experts. Usually, these data contain information such as the location, size, and shape of the tumor; unlabeled data has no expert annotations and contains unrecognized colorectal cancer CT images.

[0110] Step S13: Extract annotation features from the annotation data to obtain annotation feature data;

[0111] Specifically, feature extraction is performed on the labeled data, aiming to extract representative image features from CT images of colorectal cancer. The labeled data is processed through common image processing methods (such as edge detection, texture analysis, shape analysis, etc.) to obtain labeled feature data. The labeled feature data will include boundary information, texture features, structural features, etc. of the colorectal cancer tumor, providing the target information required for the training of the student network.

[0112] Step S14: Screen the unlabeled data to obtain unlabeled screened data;

[0113] Specifically, the unlabeled data is screened to remove irrelevant or low-quality images, ensuring that the screened data can effectively participate in pseudo-label generation and model training. The screening process can be carried out by checking factors such as image quality, resolution, and noise in the image. For example, overly blurred or incomplete CT images are removed, and those image data with clear structures and high image quality are retained, ensuring that only high-quality data is included in the unlabeled data, and this data will be used to generate pseudo-labels.

[0114] Step S15: Perform data augmentation on the labeled feature data and the unlabeled screened data to obtain labeled data of colorectal cancer images and unlabeled data of colorectal cancer images respectively.

[0115] Specifically, data augmentation processing is performed on the labeled data and the screened unlabeled data to increase data diversity and improve the generalization ability of the model, including rotation and flipping, randomly rotating or flipping the image to simulate images at different angles; scaling and cropping, scaling the image or cropping different regions from the image to enhance the model's adaptability to objects of different sizes; translation and mirroring, performing translation and mirroring on the image to simulate spatial changes in the image; adding noise, adding a certain degree of noise to the image to make the model more robust during the training process.

[0116] Two datasets are obtained: labeled data of colorectal cancer images and unlabeled data of colorectal cancer images. The labeled data contains information on the tumor regions labeled by experts and can be directly used for supervised learning; while the unlabeled data, after screening and data augmentation, can be used for pseudo-label generation and unsupervised learning. The generated datasets have high-quality labeled data and unlabeled data required for training deep learning models, providing an ample basis for network training.

[0117] Preferably, step S2 is specifically as follows:

[0118] Step S21: Extract dimensional characteristic features based on the colorectal cancer image data to obtain dimensional characteristic feature data;

[0119] Specifically, in this step, first, the dimensional characteristic features of colorectal cancer image data are extracted. Dimensional characteristic features refer to the information describing aspects such as the shape, texture, and structure of the image, and the model captures the unique manifestations of colorectal cancer. The image is processed through a deep learning model or traditional image processing methods (such as texture analysis, shape analysis, etc.) to extract the dimensional characteristic features of the image. Using an edge detection algorithm, the tumor edge features in the image are extracted. This can help the network accurately locate the boundary of the tumor during segmentation. The texture analysis method (such as gray-level co-occurrence matrix or Gabor filter) is used to extract the texture features of the image to help the model understand the detailed structure of the tumor area. Through morphological analysis (such as dilation, erosion, etc.), the shape information of the image is extracted to identify the shape features of colorectal cancer tumors.

[0120] Step S22: Generate a network structure based on the dimensional characteristic feature data to obtain network architecture metadata;

[0121] Specifically, based on the extracted dimensional characteristic feature data, network architecture metadata is generated. Network architecture metadata refers to the structural information of the network layers, including the input and output sizes of each layer, the number of convolutional kernels, the type of activation function, etc. According to the complexity of the dimensional characteristic features, the number of network layers and the number of neurons in each layer are designed. For example, if the colorectal cancer image has strong texture features, more convolutional layers are required to extract high-level features. According to the spatial characteristics of the features, the configuration of convolutional layers and pooling layers is determined. Convolutional layers are used to extract local features, and pooling layers are used to reduce the spatial dimension while retaining important features. An activation function is selected according to the nature of the network. For example, the ReLU activation function is selected because it can effectively accelerate convergence in most cases and handle the non-linear features in colorectal cancer image data.

[0122] Dimensional characteristic features (such as the texture, edge, contrast, etc. of the image) play a crucial role in designing the network architecture. The number of network layers and the number of neurons in each layer are adjusted according to the characteristics of the image. When the image has strong texture features (for example, the tumor edge and cell tissue in the colorectal cancer image have complex textures), a deeper network structure is required to extract these high-order features. The number of convolutional layers should be increased, and the number of neurons between layers is increased to capture more complex texture information. If the texture of the image is relatively simple, the network depth is appropriately reduced, and fewer convolutional layers, such as 16 layers, can meet the feature extraction requirements. When the image contains more complex structural information (for example, the morphological changes in the tumor area are large), a deeper network, such as 128 layers, is required to gradually extract spatial information at different scales. If the structure of the image is relatively simple and only basic shape or edge features need to be extracted, a shallower network is used.

[0123] The convolutional layer is responsible for extracting local features from the image. When designing the convolutional layer, if the colorectal cancer image has rich and complex texture features, it is recommended to use multiple convolutional layers, and each convolutional layer gradually extracts higher-order features. For example, the first few convolutional layers are used to extract low-order features such as edges and colors, while the subsequent layers extract higher-order features of the tumor region or specific structures. The size of the convolutional kernel determines the range of feature extraction. Smaller convolutional kernels (such as 3x3 or 5x5) are suitable for images with rich details and can effectively capture small spatial variations. For complex texture images, multiple convolutional layers need to be stacked, such as 32 layers or 64 layers, to gradually aggregate local information.

[0124] The pooling layer is used to reduce the spatial dimension, reduce the computational amount and retain the important features of the image. The pooling operation uses max pooling or average pooling. The choice of the pooling layer involves the feature distribution of the image. A pooling layer is inserted after every two convolutional layers to gradually reduce the spatial dimension while retaining the key features of the image. During the processing of colorectal cancer images, the pooling layer can extract features with spatial stability (such as the shape and distribution of tumors) and remove some unimportant details.

[0125] For simple image data, a small number of convolutional layers (such as 2 - 3 layers) and a small number of neurons (such as 64, 128) can perform effective feature extraction and classification. For complex colorectal cancer image data, based on the result of the weight calculation of the texture features of the image, a threshold judgment is made, and more convolutional layers (such as 4 - 6 layers) are generated through preset parameters. The number of neurons in each layer increases layer by layer. There are fewer layers in the front and more layers in the back, and the number of neurons gradually increases. For example, the first convolutional layer uses 32 neurons, the second uses 64 neurons, the third uses 128 neurons, and the last layer requires more neurons (such as 256 or 512) to capture more complex features.

[0126] Step S23: Initialize the student-teacher network architecture according to the network architecture metadata to obtain the initialized student network model and the initialized teacher model respectively.

[0127] Specifically, in this step, the student network model is initialized according to the generated network architecture metadata. The student network is typically used to learn the specific features and task objectives of images and continuously adjusts its weight parameters during the training process. The initialization of the student network usually adopts a random initialization method or uses a pre-trained model (such as a model pre-trained on ImageNet) for initialization, and the specific choice depends on the complexity of the task and the characteristics of the data. The teacher network is used to guide the student network to learn and is more complex or more powerful than the student network. The initialization process of the teacher network is similar to that of the student network. First, the network structure of the teacher model is initialized according to the network architecture metadata, and the network weights are set using pre-trained parameters or random initialization. The parameter update speed of the teacher model is usually slower to avoid over-adjustment. The student and teacher networks can have the same structure or can be different at certain levels. The teacher network contains more layers or larger convolutional kernels to learn the global features of colorectal cancer images at a higher level; while the student network focuses on low-level feature learning and is gradually optimized through collaborative training with the teacher network.

[0128] Preferably, step S3 is specifically as follows:

[0129] Step S31: Perform supervised learning on the initialized student network model according to the colorectal cancer image annotation data to obtain a preliminarily optimized student network model;

[0130] Specifically, use the colorectal cancer image annotation data to perform supervised learning on the initialized student network model. The goal of supervised learning is to train the student network to learn the features of colorectal cancer images through the annotated image data. By inputting the annotated image data, calculate the loss between the prediction result and the actual label, and adjust the parameters in the network through backpropagation to minimize the loss function. Adopt the cross-entropy loss function or the Dice coefficient as the loss function, and update the network weights through an optimization algorithm (such as the Adam optimizer) until the loss value converges. Obtain a preliminarily optimized student network model that can perform relatively accurate image segmentation on the annotated data.

[0131] Step S32: Input the unannotated data of colorectal cancer images into the preliminarily optimized student network model for calculation to obtain preliminary student model prediction data;

[0132] Specifically, input the unannotated data of colorectal cancer images into the preliminarily optimized student network model for prediction. The unannotated data is used as input and, after being processed by the student network, prediction data is obtained. These prediction data represent the inference results of the student network on the unannotated images. Usually, the prediction results will include the probability or label of each pixel belonging to a certain category (such as tumor and non-tumor).

[0133] Step S33: Calculate the spatial consistency pseudo-labels for the unlabeled data of colorectal cancer images and the prediction data of the preliminary student model to obtain spatial consistency pseudo-label data;

[0134] Specifically, based on the unlabeled data of colorectal cancer images and the prediction data of the preliminary student model, calculate the spatial consistency pseudo-labels. Utilize the spatial consistency between different regions in the image to generate pseudo-labels. In medical imaging, tumors often have specific spatial structures, that is, tumor tissues will show a certain coherence or consistency in space. Calculate the geological distance between each pixel point in the image to identify the degree of regional change in the image. According to the prediction results of the preliminary student model, evaluate the confidence of each pixel point to identify the certainty of the model's prediction of a certain region. Analyze the spatial consistency information at different scales to capture the manifestations of the tumor region at different levels. Through spatial consistency analysis, generate a pseudo-label data with high spatial consistency, that is, by combining the model prediction data and spatial consistency features, determine which regions can be regarded as pseudo-labels.

[0135] Step S34: Calculate the contrastive learning pseudo-labels for the unlabeled data of colorectal cancer images and the prediction data of the preliminary student model to obtain contrastive learning pseudo-label data;

[0136] Specifically, based on the unlabeled data of colorectal cancer images and the prediction data of the preliminary student model, further generate pseudo-labels through the contrastive learning method. The goal of contrastive learning is to optimize the model by learning the similarity between data. Extract the multi-level features of the image and perform block processing on the feature maps to obtain local feature embedding data. Align the extracted local features and convert them into globally aligned feature data. Through the contrastive learning method, select contrastive targets from the globally aligned feature data. The contrastive target refers to the image region with similar features, which is used to train the model to distinguish different classes of features. Perform data augmentation on the selected contrastive target data to further increase the diversity of samples. Aggregate the enhanced feature data to generate a more representative feature expression and optimize its consistency so that the model can better distinguish the tumor region from other regions. Generate preliminary pseudo-labels according to the consistency-optimized feature data and perform regional verification to ensure the accuracy of the pseudo-labels.

[0137] Step S35: Perform confidence filtering based on the spatial consistency pseudo-label data and the contrastive learning pseudo-label data to obtain pseudo-label data;

[0138] Specifically, combine the spatial consistency pseudo-label data with the contrastive learning pseudo-label data and further screen out reliable pseudo-label data through confidence filtering. Calculate the confidence of each pseudo-label and screen out the pseudo-labels with higher confidence. These high-confidence pseudo-label data will be used to further optimize the student network.

[0139] Step S36: Optimize the preliminary optimized student network model according to the preliminary student model prediction data and the pseudo-label data to obtain a secondary optimized student network model.

[0140] Specifically, according to the prediction data of the preliminary student model and the filtered pseudo-label data, retrain the preliminarily optimized student network. During the training process, use the labeled data, prediction data, and pseudo-label data to jointly optimize the model. While continuously learning the labeled data, the student network can gradually obtain effective learning signals from the pseudo-labels. The secondary student network model optimized by the pseudo-labels will be able to better perform colorectal cancer image segmentation and further improve the accuracy of image segmentation.

[0141] Preferably, step S4 is specifically as follows:

[0142] Step S41: Update the parameters of the initialized teacher model according to the secondary optimized student network model to obtain an updated teacher model;

[0143] Specifically, after obtaining the secondary optimized student network model, first use this student model as the guidance for the teacher model. By feeding back on the learning effect of the student network, update the parameters of the teacher model. By comparing the outputs of the student model and the teacher model, adjust the parameters of the teacher model to a state more suitable for the current data. The purpose of parameter update is to make the teacher model gradually approach the learning results of the student model, thereby improving the performance of the teacher model on unlabeled data. The parameter update of the teacher model includes using the weights of the student network as the initial values for fine-tuning the teacher model. During the update process, the learning rate of the teacher model is small to ensure that it does not deviate from the effect of the student model and gradually aligns with the learning results of the student model. An updated teacher model is obtained, whose performance is closer to the student model, can handle more diverse input data, and has strong generalization ability.

[0144] Step S42: Optimize the updated teacher model according to the colorectal cancer image labeled data to obtain a teacher annotation optimized model;

[0145] Specifically, input the colorectal cancer image labeled data into the updated teacher model for further optimization. Through the method of supervised learning, the teacher model is trained based on the labeled data, and its parameters are continuously adjusted to reduce the difference between the prediction result and the actual label. Calculate the loss function (such as cross-entropy loss or Dice coefficient) using the labeled data, and update the weights of the teacher model through backpropagation. The optimization process of the teacher model is similar to that of the student model, but since the teacher model usually has stronger reasoning ability, its optimization process pays more attention to details and high-order features in the data.

[0146] Through the training of labeled data, the teacher model will further improve the accuracy of the image segmentation task, especially when dealing with high-quality labeled data.

[0147] Step S43: Optimize the teacher annotation optimization model according to the pseudo-label data to obtain the teacher pseudo-label optimization model;

[0148] Specifically, after the teacher model is optimized by the labeled data, it is further optimized using the generated pseudo-label data. The pseudo-label data provides a large amount of unlabeled image information. The teacher model is trained on this pseudo-label data to enable the model to adapt to more diverse image features and improve its performance on unlabeled data. Training the teacher model with pseudo-label data enables it to recognize and learn the features contained in the pseudo-labels. The pseudo-label data supplements the deficiencies of the labeled data to a certain extent. Through the self-learning of the model, the teacher model can understand more potential structures in the images. During the training process, the quality of the pseudo-labels is crucial for the optimization effect. Through confidence filtering and other mechanisms, it can be ensured that the model focuses on relatively reliable pseudo-labels, thus avoiding the negative impact of low-quality parts of the pseudo-labels on model training. Through the training with pseudo-labels, the teacher model will gradually acquire more segmentation capabilities, especially the processing ability for unlabeled images has been significantly improved.

[0149] Step S44: Freeze the weights of the teacher pseudo-label optimization model to obtain the colorectal cancer image segmentation model.

[0150] Specifically, after the teacher model is optimized by the labeled data and the pseudo-label data, it enters the weight freezing stage. The main purpose of freezing the weights is to avoid overfitting or learning unnecessary noise in the further training of the teacher model. After freezing, the parameters of all layers of the teacher model will no longer be updated, thus ensuring the stability of the model in subsequent applications. Based on the teacher pseudo-label optimization model, selectively freeze the weights of certain layers, freeze the lower layers of the network (such as convolutional layers), and keep the higher layers (such as fully connected layers) for fine-tuning. The strategy of freezing weights is usually based on the structure of the model and the training process. The frozen layers should include those parts of the network that have converged on the labeled data and the pseudo-label data to ensure stability. The teacher model after freezing the weights is the final colorectal cancer image segmentation model. It can not only process labeled data but also has strong ability to process unlabeled data, and has high robustness and accuracy in practical applications.

[0151] Preferably, the calculation of the spatial consistency pseudo-label is specifically as follows:

[0152] Calculate the geological distance for the unlabeled data of colorectal cancer images and the prediction data of the preliminary student model to obtain the geological distance map data;

[0153] Specifically, in medical images, the pixel gradients in the boundary regions are relatively large, while those within the regions are relatively small. To capture the structural information of the images more precisely, image gradient information should be introduced in the calculation of geological distance. The gradient of an image is represented by the local variation of the image. The calculation formula: , where is the gradient value of pixel point , is the image in the th pixel, is the grayscale value of the image, is the gradient value of the image in the direction, is the gradient value of the image in the direction. After calculating the gradient of each pixel, the gradient information is used as a weighting factor to affect the calculation of geological distance. , where is the weighted geological distance, is the and the spatial distance between them, calculated by the Euclidean distance, is the weighting coefficient of the gradient, is the gradient value of pixel , is the gradient value of pixel . These calculation results are stored in the form of a matrix to form geological distance map data. In the geological distance map, each element represents the weighting of the spatial distance and gradient information between pixel and pixel .

[0154] Calculate the confidence of the preliminary student model prediction data to obtain confidence map data;

[0155] Specifically, obtain the class probability distribution of each pixel point from the output of the preliminary student network model. The output of the model is a tensor , where and represent the height and width of the image respectively, represents the number of classes, is the data type space in the output probability distribution tensor, which represents the real number space where the data is located in the output layer of the neural network. Each element represents the prediction probability that the pixel at position in the image belongs to class . Calculate the confidence value of each pixel, and select the maximum probability value from the class probability distribution output by the model as the confidence of this pixel. The calculation formula: , is the maximum index of the category, that is, each pixel of the image has possible categories, and select the pixel The maximum predicted probability of the corresponding category is used as the confidence value of the pixel. This confidence value represents the degree of certainty of the model's prediction for the pixel. The higher the value, the more confident the model is in predicting the pixel category. Different weighting coefficients are assigned to different regions by calculating the gradient of the image. Calculate the local confidence weighting coefficient for each pixel, and this coefficient is proportional to the gradient magnitude. , where is the local confidence weighting coefficient, is the image in the row and column pixel point corresponding gradient value. The weighted confidence value of the boundary region is calculated as: , and the confidence map is a matrix of size , and each element represents the confidence of the pixel at position in the image.

[0156] Perform multi-scale analysis based on the geological distance map data and the confidence map data to obtain the geological distance map multi-scale data and the confidence map multi-scale data respectively;

[0157] Specifically, the geological distance map data and the confidence map data have different spatial characteristics. In the preliminary analysis, the scale division can be performed on these two types of data respectively. In order to process the multiple scale features in the image, the original image is decomposed according to different scale ranges, such as selecting large scale, medium scale, and small scale for processing. The geological distance map is divided into multiple scales. This can be achieved by performing Gaussian blurring of different sizes on the original image. For example, using Gaussian convolution kernels of different sizes (such as , , ) ( represents different scales, such as , , ), and the image is blurred respectively to obtain three different scale geological distance map data , , . For the confidence map data, similarly, it is processed through convolutions of different scales. Use convolution kernels of different sizes (such as , , , such as , , )( indicating different scales), convolve the confidence map data to obtain confidence maps of three scales , , . Obtain geological distance map data and confidence map data of different scales, and fuse these scale data. By means of weighted average or convolution kernel, etc., the information at different scales is combined to obtain a more comprehensive feature representation. Or, fuse the information of different scales through a convolutional layer. Through a preset multi-scale convolutional kernel, perform a convolution operation on the image data of multiple scales to obtain the fused features. The feature representation is (small-scale data, medium-scale data, large-scale data, multi-scale fusion data).

[0158] Perform spatial consistency constraints based on the multi-scale data of the geological distance map and the multi-scale data of the confidence map to obtain consistent segmentation data;

[0159] Specifically, for each pixel point, its surrounding neighborhood pixels can be examined, and based on the geological distance and confidence, determine whether these neighborhood pixels should be classified into the same category through threshold judgment. Cross-scale consistency ensures that pixels or regions in the image maintain consistent labels at different scales. Smaller scales focus on detailed information, while larger scales focus on global features. The geological distance maps and confidence maps obtained from different scales (such as large, medium, and small scales) will provide image features at different levels. At each scale, use the aforementioned adjacent region consistency strategy to label the region. For each pixel point, it is necessary to check the label consistency at multiple scales. If the labels are inconsistent at multiple scales, smooth and adjust the labels, such as through over-smoothing operations, that is, at the places where the labels are inconsistent, adjust the labels through weighted smoothing or majority voting mechanisms to make the labels at each scale tend to be consistent. For each pixel point, adjust its label by calculating the weighted average in the multi-scale and neighborhood. Specifically, different weights are assigned to each scale, and weighted adjustment is combined with the geological distance and confidence to ensure that the label information at different scales is fused and avoid the excessive influence of the label of a certain scale. Further optimize the label consistency through optimization algorithms (such as gradient descent or conditional random field, etc.). By introducing global constraint conditions, optimize the label of each pixel point to make the labels of adjacent regions or pixels more consistent, thus obtaining a consistent segmentation result.

[0160] Perform confidence screening on the preliminary student model prediction data according to the consistent segmentation data to obtain spatially consistent pseudo-label data.

[0161] Specifically, the confidence of the prediction data of the preliminary student model is screened using the consistent segmentation data to obtain spatially consistent pseudo-label data. The consistent segmentation data is screened through a preset confidence threshold data, and the regions with higher confidence are selected as pseudo-labels. The higher the confidence, the stronger the reliability of the pseudo-labels. For the region verification of the pseudo-labels, the selected regions are further verified to ensure that the segmentation results of these regions are accurate and stable. If large errors are found in the regions of the pseudo-labels, corrections will be made or the regions will be ignored. To eliminate the noise at the boundaries of the pseudo-labels, Gaussian smoothing is performed to make the pseudo-label data smoother and more continuous, thereby improving the segmentation quality.

[0162] Preferably, the calculation of the contrastive learning pseudo-label is specifically as follows:

[0163] The unlabeled data of colorectal cancer images is subjected to multi-level feature extraction through a preliminarily optimized student network model to obtain multi-level feature data;

[0164] Specifically, in the first layer of the convolutional neural network, multiple convolutional kernels (filters) are used to perform convolutional operations on the image to extract low-level features of the image. For example, the convolutional layer extracts basic features such as edges, textures, and lines in the image. For colorectal cancer images, the low-level features capture the boundaries between tumor tissues and normal tissues. The input image is calculated with multiple convolutional kernels through convolutional operations to obtain a set of feature maps, which represent the activation of local regions in the image. Each convolutional kernel will learn different features, such as edges, corners, textures, etc. The pooling layer follows the convolutional layer and is used to reduce the spatial dimension of the feature maps and retain the main information, including max pooling and average pooling. Pooling operations are performed on each convolutional feature map to obtain the pooled feature maps. For example, in max pooling, it takes the maximum value within a small region, thereby retaining the most important information. As the depth of the network increases, the network can learn more complex and abstract features. The shallow features (feature maps of the first and second layers) include the edges, corners, and simple texture information of the image, such as the basic shapes of regions such as the contours of tumors and blood vessels. The middle-level features (feature maps of the middle layers) represent more complex structures, such as local regions of tumors and texture patterns between tissues. The deep features (feature maps of the deeper layers) represent the global information of the image, such as tumor regions, normal tissues, and backgrounds.

[0165] The multi-level features are processed in blocks to obtain local feature embedding data;

[0166] Specifically, the image is divided into local regions of a fixed size or feature-based region partitioning (for example, partitioning based on saliency or changing regions). The features of each local region are transformed into vector form through an embedding layer for subsequent feature alignment and learning.

[0167] Align features based on locally-embedded data to obtain globally-aligned feature data;

[0168] Specifically, align features through locally-embedded data to obtain globally-aligned feature data. The purpose of feature alignment is to make the features extracted in different regions and at different scales consistent. By calculating the similarity or distance metric between feature vectors, global feature alignment is performed, using metrics such as cosine similarity or Euclidean distance to align the features in different regions to a unified scale or space. According to the results of global alignment, further optimize the feature vectors to make the features in different regions more consistent.

[0169] Perform contrastive target sampling on the globally-aligned feature data to obtain contrastive target data;

[0170] Specifically, based on the globally-aligned feature data, perform contrastive target sampling to obtain contrastive target data. The purpose of contrastive target sampling is to select the most representative regions from the globally-aligned features for further learning, identify the target regions and background regions of the image in the globally-aligned features, calculate the similarity between regions, and select the target regions as the objects for comparison. According to the data after feature alignment, select the regions closest in the feature space as the contrastive targets, and the target regions contain tumors or other abnormal regions of interest to the model.

[0171] Perform contrastive learning data augmentation based on the contrastive target data to obtain augmented contrastive target data;

[0172] Specifically, after obtaining the contrastive target data, perform data augmentation for contrastive learning, aiming to improve the diversity of the data and help the model generalize better. Perform operations such as rotation, translation, and scaling on the selected contrastive target data to simulate changes under different conditions, so that the model can learn richer features. Through this augmentation, the model can learn the performance of the target region under different transformations, thereby improving its robustness.

[0173] Perform feature aggregation based on the augmented contrastive target data to obtain aggregated feature data;

[0174] Specifically, perform feature aggregation on the augmented contrastive target data to obtain aggregated feature data. The purpose of feature aggregation is to fuse multiple augmented contrastive target data into an overall feature representation for the model to train and optimize. Combine the feature data generated by different augmentation methods, and weighted average, concatenation, or other aggregation methods can be used to fuse multiple feature information into an overall. Through feature aggregation, the model can obtain more comprehensive and representative target region features.

[0175] Perform cross-node contrastive consistency optimization based on the aggregated feature data to obtain consistency-optimized feature data;

[0176] Specifically, based on the aggregated feature data, cross-node comparison consistency optimization is performed. Ensure the consistency of features between different nodes, and reduce the errors caused by differences between nodes. Check whether the feature data between different nodes is consistent, and optimize it by measuring the differences in feature distributions between different nodes. Through the method of contrastive learning, align the features between different nodes so that similar regions can maintain consistency between different nodes.

[0177] During the optimization process, first, measure the feature consistency by calculating the differences in feature distributions between different nodes (such as mean squared error, KL divergence, etc.). When the feature data is inconsistent, use the method of contrastive learning. By calculating the contrastive loss between positive and negative sample pairs, adjust the weights of the network so that the feature vectors from similar regions are closer in the feature space, while the distance between the feature vectors of dissimilar regions increases. Through this optimization process, the network can align the feature data between different nodes (hospitals or devices), making the representations of similar regions more consistent between different nodes. The optimization process can effectively reduce the errors caused by differences between different nodes, thereby improving the segmentation performance of the model in a multi-center, multi-device environment. Especially in the segmentation of tumor regions in colorectal cancer images, there is a significant improvement in effect.

[0178] Through the contrastive learning optimization process, the representations in the feature space are gradually aligned, so that the features from similar regions, regardless of which node they come from, have similar feature representations. Through the optimization process, the tumor regions in colorectal cancer images from different nodes (such as different hospitals or devices) have approximate vector representations in the feature space. Through the optimization of the contrastive loss, the features from different regions (such as tumor and normal tissue regions) have a large distance, thus maintaining the distinctiveness of the features.

[0179] Generate similarity pseudo-labels based on the consistency-optimized feature data to obtain preliminary pseudo-label data;

[0180] Specifically, based on the feature data optimized for consistency, generate similarity pseudo-labels to obtain preliminary pseudo-label data. The generation of pseudo-labels mainly depends on the similarity between features, that is, predict labels through similar regions. Calculate the feature similarity between different regions, and select regions with high similarity to generate pseudo-labels. Assign labels to each pixel or region according to the similarity.

[0181] Verify the pseudo-label regions based on the preliminary pseudo-label data and the preliminary student model prediction data to obtain pseudo-label verification data;

[0182] Specifically, verify the pseudo-labeled regions of the preliminary pseudo-labeled data to ensure that the generated pseudo-labels are accurate and effective. The verification steps include verifying whether the boundaries of the pseudo-labeled regions are accurate to ensure the matching of the pseudo-labels with the true regions of the images. If there are obvious errors or deviations in the pseudo-labeled regions, correct or discard them.

[0183] Perform Gaussian smoothing on the pseudo-label verification data to obtain contrastive learning pseudo-label data;

[0184] Specifically, use Gaussian smoothing to smooth the pseudo-label verification data to remove boundary noise and make the pseudo-labels smoother. Gaussian smoothing processes each pixel point through weighted averaging, making the labeled regions more continuous and smooth.

[0185] Among them, the contrastive target sampling is specifically as follows:

[0186] Perform target region boundary analysis on the globally aligned feature data to obtain boundary feature data;

[0187] Specifically, by analyzing the globally aligned feature data, identify the boundaries of the target regions in the images. The goal of boundary analysis is to determine the regions in the images with obvious structural changes, which represent tumors, lesions, or other important tissue structures. Use image processing techniques (such as edge detection algorithms) to identify the obvious boundaries in the images, and use information such as gradient changes and texture changes to identify these regions. By analyzing the characteristics of the boundaries such as shape, size, and complexity, obtain the boundary feature data of the target regions. These features can include the length, curvature, density, etc. of the boundaries.

[0188] Perform target region confidence evaluation based on the boundary feature data to obtain target region confidence evaluation data;

[0189] Specifically, based on the boundary feature data, perform the confidence evaluation of the target regions to determine which regions contain higher information content and which regions are noise or irrelevant. Evaluate the confidence of each target region according to the globally aligned feature data and boundary features, and achieve this by calculating indicators such as the intensity of the pixels within the region, the clarity of the boundary, and the regional consistency. Generate a confidence score for each target region. Regions with high confidence indicate that the region contains important tumor or lesion information, while regions with low confidence are background or noise.

[0190] When it is determined that the target region area data corresponding to the target region confidence evaluation data is greater than or equal to the preset target region area threshold data, then perform sparse sampling according to the target region confidence evaluation data to obtain contrastive target sampling data;

[0191] When it is determined that the target area data corresponding to the target area confidence evaluation data is less than the preset target area threshold data, dense sampling is performed according to the target area confidence evaluation data to obtain comparison target sampling data;

[0192] Specifically, according to the area size of the target area and the confidence evaluation data, the sampling density is determined: when the area of the target area is greater than or equal to the preset target area threshold (for example, the target area is large), sparse sampling is selected, which means sampling only at the key parts (such as the edge or the center) of the target area to reduce the computational burden and the risk of overfitting. When the area of the target area is less than the preset target area threshold (for example, the target area is small), dense sampling is performed, which means sampling evenly throughout the target area to ensure that all relevant features are captured.

[0193] Confidence-guided target screening is performed according to the comparison target sampling data to obtain confidence screening data;

[0194] Specifically, based on the comparison target data obtained by sampling, further confidence-guided target screening is performed to ensure that the finally selected target data has high representativeness and reliability. A confidence threshold is set, and only those sampling data with confidence scores higher than the threshold will be retained. Low-confidence target data will be removed to reduce the interference of noise on subsequent learning. All sampling data are sorted according to the confidence scores, and the target data with higher confidence is selected as the final comparison target data.

[0195] Spatial perception target generation is performed according to the confidence screening data to obtain spatial perception target data;

[0196] Specifically, according to the target data after confidence screening, spatial perception target data is generated. Spatial perception refers to generating features that can effectively distinguish the target from the background by analyzing the spatial relationships between target data. Analyze the spatial relationships between the target area and the background area, such as the geometric shape and distribution characteristics of the target area. By modeling these spatial information, spatial perception target data is obtained.

[0197] Region-to-background comparison target sampling is performed according to the spatial perception target data to obtain comparison target data.

[0198] Specifically, after generating the spatial perception target data, region-to-background comparison target sampling is performed. According to the spatial perception target data, the target area and the background area are clearly distinguished, which is achieved through image segmentation techniques or threshold-based segmentation methods. Comparison target sampling is performed between the target area and the background area. Comparison sampling helps to improve the robustness of the model, especially when dealing with complex backgrounds and blurred regions.

[0199] Preferably, the multi-scale analysis specifically includes:

[0200] Perform a preliminary multi-scale division based on the geological distance map data and the confidence map data to obtain the preliminary multi-scale data of the geological distance map and the preliminary multi-scale data of the confidence map;

[0201] Specifically, perform a preliminary multi-scale division on the geological distance map data and the confidence map data. The geological distance map is segmented according to different scales, and the division accuracy can be controlled by defining different scale factors. For example, the geological distance map is segmented using a fixed window size to obtain regional features at different scales. Similarly, the confidence map data is divided according to the same scale. The regions at each scale represent the confidence distribution of different regions in the image.

[0202] Calculate the global-local entropy based on the preliminary multi-scale data of the geological distance map and the preliminary multi-scale data of the confidence map to obtain the multi-scale entropy data of the geological distance map and the multi-scale entropy data of the confidence map;

[0203] Specifically, based on the preliminary multi-scale data, calculate the global-local entropy to measure the complexity of the data distribution at different scales. The purpose of entropy calculation is to evaluate the amount of information in different regions of the image. Calculate the global entropy values of the entire geological distance map and the confidence map at different scales. The global entropy is obtained by statistically analyzing the data of the entire image. Within each sub-block region, calculate its local entropy value, which reflects the degree of change of the features within the region. The local entropy value can reflect the local detailed information of the image.

[0204] Calculate the multi-scale entropy similarity based on the multi-scale entropy data of the geological distance map and the multi-scale entropy data of the confidence map to obtain the multi-scale entropy similarity data;

[0205] Specifically, by comparing the multi-scale entropy data of the geological distance map and the confidence map, calculate the similarity between them. Adopt the entropy similarity calculation method, based on the entropy data of the geological distance map and the confidence map, evaluate the similarity between image regions at different scales, and calculate the correlation or distance between the entropy data. Regions with higher similarity are considered regions with similar features in the image, and these regions can provide more information for screening and optimization.

[0206] Perform non-similar entropy screening on the preliminary multi-scale data of the geological distance map and the preliminary multi-scale data of the confidence map according to the multi-scale entropy similarity data to obtain the multi-scale screening data of the geological distance map and the multi-scale screening data of the confidence map respectively;

[0207] Specifically, based on the multi-scale entropy similarity data, non-similar entropy screening is performed on the preliminary multi-scale data of the geological distance map and the preliminary multi-scale data of the confidence map. According to the similarity calculation results, those regions with lower entropy similarity are removed. That is, those regions with lower information content and belonging to noise are selected for elimination, and regions with higher similarity are retained. A similarity threshold is set, and regions below this threshold will be screened out, only those regions with significant features are retained.

[0208] Based on the multi-scale screening data of the geological distance map and the multi-scale screening data of the confidence map, consistency distribution calculation is performed to obtain the multi-scale data of the geological distance map and the multi-scale data of the confidence map.

[0209] Specifically, after multi-scale screening, consistency distribution calculation is performed on the geological distance map and the confidence map. Check whether the screened geological distance map and confidence map have consistent characteristic distributions at different scales. Specifically, it is necessary to evaluate whether the same region has consistent confidence and geological characteristics at multiple scales. According to the screened data, calculate their distribution characteristics at different scales, such as position, morphology, regional consistency, etc., to further optimize image segmentation.

[0210] Preferably, the present application also provides a colorectal cancer image segmentation system for performing the colorectal cancer image segmentation method as described above. The system includes:

[0211] A colorectal cancer image data acquisition module for acquiring colorectal cancer image data, where the colorectal cancer image data includes colorectal cancer image annotation data and colorectal cancer image unannotated data;

[0212] A student-teacher network architecture initialization module for initializing the student-teacher network architecture according to the colorectal cancer image data to obtain an initialized student network model and an initialized teacher model respectively;

[0213] A secondary optimized student network model construction module for performing supervised learning on the initialized student network model according to the colorectal cancer image annotation data to obtain a preliminary optimized student network model; inputting the colorectal cancer image unannotated data into the preliminary optimized student network model for calculation to obtain preliminary student model prediction data, and performing pseudo-label calculation on the colorectal cancer image unannotated data and the preliminary student model prediction data to obtain pseudo-label data; optimizing the preliminary optimized student network model according to the preliminary student model prediction data and the pseudo-label data to obtain a secondary optimized student network model;

[0214] A colorectal cancer image segmentation model generation module is used to update the parameters of the initialized teacher model according to the secondary optimized student network model to obtain an updated teacher model, and optimize the updated teacher model according to the colorectal cancer image annotation data to obtain a colorectal cancer image segmentation model.

[0215] Preferably, the present application also provides a computer-readable storage medium, in which a computer program is stored, and wherein the computer program is configured to execute the method described in any one of the above when running.

[0216] The beneficial effects of the present invention are as follows: By combining labeled data and unlabeled data, the present invention uses the labeled data to supervise the training of the student network, and at the same time fully excavates the unlabeled data through the pseudo-label generation technology. The utilization rate of the unlabeled data is significantly improved, the dependence on manually labeled data is reduced, and thus the labeling cost is reduced. The pseudo-label generation combines the calculation of spatially consistent pseudo-labels and the calculation of contrastive learning pseudo-labels, and generates accurate and reliable pseudo-labels from the perspectives of spatial consistency and feature contrast respectively. The spatially consistent pseudo-labels utilize the geodesic distance and the confidence map, and combine multi-scale analysis and consistency distribution calculation to improve the spatial coherence of the pseudo-labels. The contrastive learning pseudo-labels generate high-quality pseudo-labels through feature alignment, dynamic sampling and similarity calculation, significantly improving the effective utilization of the unlabeled data. The parameters of the teacher network are updated by the exponential moving average (EMA) of the student network, the parameters are smoother and more robust to noise, and the stability of the pseudo-label generation is improved. The teacher network gradually improves its segmentation ability for complex regions (such as regions with blurred boundaries and small lesions) through multiple optimizations of the labeled data and the pseudo-label data. The student network forms a closed loop through supervised learning and semi-supervised pseudo-label optimization, and continuously improves its segmentation performance during training. By combining the supervised learning of the labeled data and the semi-supervised pseudo-label optimization of the unlabeled data, the present invention makes the model training more efficient. The parameters of the teacher network are updated by EMA, further smoothing the parameter changes and avoiding oscillations and instabilities during the training process. The present invention is applicable to scenarios where the labeled data is insufficient but the unlabeled data is abundant, and significantly reduces the dependence on the labeled data through the pseudo-label generation technology. In medical image segmentation tasks, the labeled data is often scarce and expensive, so the present invention has high application value.

[0217] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended application documents rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be included in the present invention.

[0218] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for colorectal cancer image segmentation, characterized in that: The following steps are involved: Step S1: acquiring colorectal cancer image data, wherein the colorectal cancer image data includes colorectal cancer image annotated data and colorectal cancer image unannotated data; Step S2: Initializing the student-teacher network architecture according to the colorectal cancer image data to obtain an initialized student network model and an initialized teacher model; Step S3: Performing supervised learning on the initialized student network model according to the annotated data of colorectal cancer images to obtain a preliminary optimized student network model; inputting the unlabeled data of colorectal cancer images into the preliminary optimized student network model for calculation to obtain preliminary student model prediction data, and performing pseudo-label calculation on the unlabeled data of colorectal cancer images and the preliminary student model prediction data to obtain pseudo-label data; optimizing the preliminary optimized student network model according to the preliminary student model prediction data and the pseudo-label data to obtain a secondary optimized student network model; Step S4: Update the parameters of the initialized teacher model according to the secondary optimized student network model to obtain an updated teacher model, and optimize the updated teacher model according to the colorectal cancer image annotation data to obtain a colorectal cancer image segmentation model.

2. The method according to claim 1, characterized in that Step S1 is specifically as follows: Interact with multi-center hospital image databases through medical image data interfaces to obtain preliminary colorectal cancer CT image data; The preliminary colorectal cancer CT image data are grouped to obtain labeled data and unlabeled data; Extracting annotation features from the annotation data to obtain annotation feature data; Perform non-labeled data screening on the non-labeled data to obtain non-labeled screened data; Data enhancement is performed on the labeled feature data and the unlabeled screened data to obtain colorectal cancer image labeled data and colorectal cancer image unlabeled data, respectively.

3. The method according to claim 1, characterized in that Step S2 is specifically as follows: Extracting dimensionality characteristics based on the colorectal cancer image data to obtain dimensionality characteristics data; Generate network structure based on dimensionality characteristic feature data to obtain network architecture metadata; The student-teacher network architecture is initialized according to the network architecture metadata to obtain an initialized student network model and an initialized teacher model respectively.

4. The method according to claim 1, characterized in that: Step S3 is specifically as follows: The initialized student network model was supervised and learned based on the annotated data of colorectal cancer images to obtain a preliminary optimized student network model; The unlabeled data of colorectal cancer images are input into the preliminary optimized student network model for calculation to obtain the preliminary student model prediction data; Calculate spatially consistent pseudo labels for unlabeled colorectal cancer image data and preliminary student model prediction data to obtain spatially consistent pseudo label data; Comparative learning pseudo-label calculations are performed on the unlabeled data of colorectal cancer images and the data predicted by the preliminary student model to obtain comparative learning pseudo-label data; Confidence screening is performed based on spatial consistency pseudo-label data and comparative learning pseudo-label data to obtain pseudo-label data; The preliminary optimized student network model is optimized according to the preliminary student model prediction data and pseudo-label data to obtain a secondary optimized student network model.

5. The method according to claim 4, characterized in that Step S4 is specifically as follows: Update the parameters of the initialized teacher model according to the secondary optimized student network model to obtain an updated teacher model; The updated teacher model is optimized according to the annotated data of colorectal cancer images to obtain the teacher annotation optimization model; The teacher annotation optimization model is optimized according to the pseudo-label data to obtain the teacher pseudo-label optimization model; The weights of the teacher pseudo-label optimization model were frozen to obtain the colorectal cancer image segmentation model.

6. The method according to claim 4, characterized in that The spatial consistency pseudo label calculation is as follows: The geological distance was calculated for the unlabeled data of colorectal cancer images and the preliminary student model prediction data to obtain the geological distance map data; Calculate the confidence of the preliminary student model prediction data to obtain confidence map data; Perform multi-scale analysis based on the geological distance map data and the confidence map data to obtain geological distance map multi-scale data and confidence map multi-scale data respectively; According to the multi-scale data of the geological distance map and the multi-scale data of the confidence map, spatial consistency constraints are performed to obtain consistent segmentation data; The confidence of the preliminary student model prediction data is screened according to the consistent segmentation data to obtain spatially consistent pseudo-label data.

7. The method according to claim 4, characterized in that The calculation of pseudo labels for contrastive learning is as follows: The unlabeled data of colorectal cancer images are extracted by preliminarily optimizing the student network model to obtain multi-level feature data; Multi-level features are processed in blocks to obtain local feature embedding data; Perform feature alignment based on local feature embedding data to obtain global alignment feature data; Performing comparison target sampling on the global alignment feature data to obtain comparison target data; Perform contrast learning data enhancement according to contrast target data to obtain enhanced contrast target data; Perform feature aggregation according to the enhanced contrast target data to obtain aggregated feature data; Perform cross-node comparison and consistency optimization based on aggregated feature data to obtain consistent optimized feature data; Generate similarity pseudo labels based on the consistent optimized feature data to obtain preliminary pseudo label data; Perform pseudo-label region verification based on preliminary pseudo-label data and preliminary student model prediction data to obtain pseudo-label verification data; Gaussian smoothing is performed on the pseudo-label verification data to obtain the pseudo-label data for comparative learning; The comparison target sampling is as follows: Performing target area boundary analysis on the global alignment feature data to obtain boundary feature data; Performing target area confidence assessment according to the boundary feature data to obtain target area confidence assessment data; When it is determined that the target area area data corresponding to the target area confidence evaluation data is greater than or equal to the preset target area area threshold data, sparse sampling is performed according to the target area confidence evaluation data to obtain comparison target sampling data; When it is determined that the target area area data corresponding to the target area confidence evaluation data is less than the preset target area area threshold data, dense sampling is performed according to the target area confidence evaluation data to obtain comparative target sampling data; Conduct confidence-guided target screening according to the comparison target sampling data to obtain confidence screening data; The data is filtered according to the confidence level to generate a spatial perception target, thereby obtaining spatial perception target data; According to the spatial perception target data, regional and background contrast target sampling is performed to obtain contrast target data.

8. The method according to claim 6, characterized in that The multi-scale analysis is as follows: Perform preliminary multi-scale division according to the geological distance map data and the confidence map data to obtain preliminary multi-scale data of the geological distance map and preliminary multi-scale data of the confidence map; Perform global local entropy calculation based on preliminary multi-scale data of geological distance map and preliminary multi-scale data of confidence map to obtain multi-scale entropy data of geological distance map and multi-scale entropy data of confidence map; Multi-scale entropy similarity calculation is performed based on the multi-scale entropy data of the geological distance map and the multi-scale entropy data of the confidence map to obtain multi-scale entropy similarity data; According to the multi-scale entropy similarity data, the preliminary multi-scale data of the geological distance map and the preliminary multi-scale data of the confidence map are screened for non-similar entropy, and the multi-scale screening data of the geological distance map and the multi-scale screening data of the confidence map are obtained respectively; The consistency distribution calculation is performed based on the multi-scale screening data of the geological distance map and the multi-scale screening data of the confidence map to obtain the multi-scale data of the geological distance map and the multi-scale data of the confidence map.

9. A colorectal cancer image segmentation system, characterized in that: For executing the method for colorectal cancer image segmentation according to claim 1, the system comprises: A colorectal cancer image data acquisition module, used to acquire colorectal cancer image data, wherein the colorectal cancer image data includes colorectal cancer image annotated data and colorectal cancer image unannotated data; A student-teacher network architecture initialization module is used to initialize the student-teacher network architecture according to colorectal cancer image data, and obtain an initialized student network model and an initialized teacher model respectively; The secondary optimized student network model construction module is used to perform supervised learning on the initialized student network model according to the annotated data of colorectal cancer images to obtain a preliminary optimized student network model; input the unlabeled data of colorectal cancer images into the preliminary optimized student network model for calculation to obtain preliminary student model prediction data, and perform pseudo-label calculation on the unlabeled data of colorectal cancer images and the preliminary student model prediction data to obtain pseudo-label data; optimize the preliminary optimized student network model according to the preliminary student model prediction data and the pseudo-label data to obtain a secondary optimized student network model; The colorectal cancer image segmentation model generation module is used to update the parameters of the initialized teacher model according to the secondary optimized student network model to obtain an updated teacher model, and optimize the updated teacher model according to the colorectal cancer image annotation data to obtain the colorectal cancer image segmentation model.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 8 when executed.

Citation Information

Patent Citations

  • MR medical image colorectal cancer segmentation method and system based on semi-supervised learning

    CN117409201A

  • Colposcope image segmentation model construction method, image classification method and device

    CN118334336A