Colorectal cancer automatic segmentation method and device based on multi-phase feature alignment fusion model

By employing a cross-period attention mechanism and multi-scale feature fusion, the spatial misalignment and scale difference issues in multi-phase CT image segmentation were resolved, enabling high-precision automatic segmentation of colorectal cancer lesions and reducing the annotation burden on doctors.

CN120976255APending Publication Date: 2025-11-18GUANGDONG GENERAL HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511087717.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively utilize the complementary information of multi-phase CT images in colorectal cancer CT image segmentation. Furthermore, due to spatial misalignment issues and a lack of multi-scale processing capabilities, it is difficult to accurately segment colorectal cancer lesions of different sizes.

Method used

We employ a cross-phase attention mechanism to adaptively align CT image features from the arterial and venous phases, integrate multi-scale features through the attention mechanism, expand the receptive field by combining depthwise separable convolution and Transformer structure, and optimize the model using a dual-path depth supervision strategy.

Benefits of technology

It significantly improves the accuracy and robustness of colorectal cancer lesion segmentation, reduces the annotation burden on doctors, and enables precise segmentation of lesions of different sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976255A_ABST
    Figure CN120976255A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic colorectal cancer segmentation method and device based on a multi-phase feature alignment fusion model. The method comprises the following steps: respectively extracting multi-scale features of an arterial phase CT image and a venous phase CT image through parallel encoders; inputting the arterial phase features and the venous phase features under the same scale into a corresponding multi-phase alignment fusion module, dynamically distributing weights through an inter-phase attention mechanism, and generating aligned and fused features; after the aligned and fused features of adjacent scales are spliced through an attention-guided multi-scale fusion module, effective expression of the features is guided on different scales through an attention map, and therefore features with detailed information are obtained; respectively applying a depth supervision strategy to the original feature path and the multi-scale refined feature path of the decoder, and optimizing the model in combination with cross entropy loss and Dice loss; and performing automatic segmentation on a to-be-processed colorectal cancer image based on the optimized model. According to the method, the robustness and the accuracy of automatic segmentation of the colorectal cancer can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of medical image processing, and particularly relates to a colorectal cancer automatic segmentation method and device based on a multi-phase feature alignment fusion model. BACKGROUND

[0002] Colorectal cancer (CRC) is one of the most common malignant tumors. Due to the advantages of easy acquisition and high imaging quality, computed tomography (CT) has been widely used in the examination and diagnosis of various clinical diseases. In particular, accurate delineation of CRC on CT images can guide individualized surgical and radiotherapy planning. However, manual delineation of colorectal cancer is time-consuming and laborious, and can only be completed by radiologists with professional knowledge. Therefore, it is urgent to develop an automatic and accurate colorectal cancer segmentation technology. With the rapid development of deep learning, a large number of CRC automatic segmentation algorithms have been proposed. However, current researches only use venous phase CT images for tumor segmentation, ignoring the rich features contained in multi-phase CT images. Specifically, venous phase images can show the size, shape and infiltration range of colorectal cancer, and are the main phase for colorectal cancer identification; while arterial phase images can provide more rich blood supply features to assist in the identification and segmentation of colorectal cancer.

[0003] In recent years, multi-phase or multi-modal fusion technology has received more and more attention in medical image segmentation. A simple strategy is to add or splice multi-phase information at the image layer or feature layer. In addition, various multi-phase fusion modules based on attention mechanism (MAML, PA-Net, F2Net) have also been proposed. However, these methods rely on good alignment between multi-phase inputs, while in abdominal CT images containing colorectal cancer, there is a problem of spatial misalignment between multi-phase CT images due to the influence of respiratory motion and intestinal peristalsis, which limits the effectiveness of CRC multi-phase feature fusion. In addition, the size of colorectal cancer lesions varies significantly, and existing models (such as UNet++, DeepLabv3+) are limited by convolution operations and downsampling mechanisms, making it difficult to capture the details of lesions of different sizes, especially small lesions.

[0004] Therefore, how to realize efficient alignment and fusion of multi-phase features and improve the segmentation accuracy of the model for colorectal cancer lesions of different sizes is a technical problem that needs to be solved in the field. SUMMARY

[0005] The main purpose of the present application is to overcome the shortcomings and deficiencies of the prior art, provide a multi-phase feature alignment fusion model based automatic segmentation method and device for colorectal cancer, the present application aligns the arterial phase and venous phase CT image features adaptively through the cross-phase attention mechanism, alleviates the multi-phase feature space misalignment problem, and effectively integrates multi-scale features through the attention mechanism, enhances the modeling ability of the model for tumors of different sizes.

[0006] In order to achieve the above purpose, the present application adopts the following technical scheme:

[0007] In the first aspect, the present application provides a multi-phase feature alignment fusion model based automatic segmentation method for colorectal cancer, comprising the following steps:

[0008] The multi-scale features of the arterial phase CT image and the venous phase CT image are extracted by parallel encoders respectively, and the multi-scale arterial phase features and the multi-scale venous phase features are obtained.

[0009] The arterial phase features and the venous phase features under the same scale are input into the corresponding multi-phase alignment fusion module, the weight is dynamically allocated through the cross-phase attention mechanism, and the aligned and fused features are generated; the cross-phase attention mechanism realizes the receptive field expansion by combining the inverted bottleneck structure similar to the Transformer structure with the depth separable convolution.

[0010] In the decoding stage, the aligned and fused features of adjacent scales are spliced through the attention guided multi-scale fusion module, and the effective expression of the features is guided on different scales through the attention map, so that the features with detailed information are obtained.

[0011] The depth supervision strategy is applied to the original feature path and the multi-scale refined feature path of the decoder respectively, and the cross-entropy loss and the Dice loss are combined in each path to optimize the model.

[0012] Based on the optimized model, the colorectal cancer image to be processed is automatically segmented.

[0013] As a preferred technical scheme, each encoder includes a plurality of double convolution layers and a normalization layer, and the combination of double convolution and normalization is performed once in each level of the encoder until the features on all channels are extracted.

[0014] As a preferred technical scheme, the specific implementation of the multi-phase alignment fusion module includes:

[0015] The arterial phase features and the venous phase features under the same scale are input into the corresponding multi-phase alignment fusion module;

[0016] The arterial phase features and the venous phase features are merged in the channel dimension, and the merged features are subjected to convolution, standardization and linear correction processing, and the corrected features are input into a TransConv block, the TransConv block combines an inverted bottleneck structure with a depth separable convolution, and realizes feature alignment by expanding the receptive field;

[0017] The dynamic weights of the arterial phase features and the venous phase features are generated by an activation function;

[0018] The arterial phase features and the venous phase features are weighted and summed according to the dynamic weights to obtain the fusion features.

[0019] As a preferred technical solution, the TransConv block is defined as follows:

[0020] TransConv(·)=C1(GN(GL(C1(BN(DWC(x))))))+x

[0021] Wherein, DWC(·) is a depth separable convolution with a 5*5*5 convolution kernel, GN(·) is group normalization, GL(·) is a GELU activation function, and C1(·) is a 1*1*1 convolution.

[0022] As a preferred technical solution, the attention-guided multi-scale fusion module is represented as follows:

[0023]

[0024] Wherein represents features with detailed information, UP(·) represents an up-sampling operation, and Concat(·) represents a concatenation operation, represents matrix multiplication, and Ch(·) represents a channel convolution using a 3*3*3 convolution kernel; σ is a Sigmoid activation function, which is used to obtain the attention map of the current scale, f c is a combination operation composed of a 1*1*1 convolution kernel, batch normalization and an activation function applied to the output to generate multi-scale refined features.

[0025] As a preferred technical solution, the dual-path deep supervision strategy includes an original feature path supervision and a multi-scale refined feature supervision;

[0026] The original feature path supervision directly supervises the features of each layer of the decoder;

[0027] The multi-scale refined feature supervision supervises the refined features output by the attention-guided multi-scale fusion module.

[0028] As a preferred technical solution, the model is optimized by combining a cross-entropy loss and a Dice loss, specifically:

[0029] The implementation of the Dice loss is as follows:

[0030]

[0031] where x is the softmax output of the network, y is the one-hot encoding of the label segmentation map, the shape of x and y is I x K, v e I represents the number of voxels in the training batch, and k e K represents the number of classes;

[0032] The implementation of the cross-entropy loss is as follows:

[0033]

[0034] where N is the number of input CT images, p j e [0, 1] represents the predicted probability of the jth body data, and q j e [0, 1] represents the real segmentation label corresponding to the jth body data.

[0035] In a second aspect, the present application provides a colorectal cancer automatic segmentation system based on a multi-phase feature alignment fusion model, which is applied to the colorectal cancer automatic segmentation method based on the multi-phase feature alignment fusion model, and includes a feature extraction module, a multi-phase alignment fusion module, an attention-guided multi-scale fusion module, a deep supervision module, and an image segmentation module.

[0036] The feature extraction module is configured to extract multi-scale features of the arterial phase CT image and the venous phase CT image through parallel encoders respectively, to obtain multi-scale arterial phase features and venous phase features.

[0037] The multi-phase alignment fusion module is configured to input the arterial phase features and the venous phase features under the same scale into corresponding multi-phase alignment fusion modules, to generate features after alignment and fusion by dynamically allocating weights through a cross-phase attention mechanism; the cross-phase attention mechanism realizes receptive field expansion by combining an inverted bottleneck structure similar to a Transformer structure with a depth separable convolution.

[0038] The attention-guided multi-scale fusion module is configured to, in a decoding stage, splice the aligned and fused features of adjacent scales through the attention-guided multi-scale fusion module, and guide the effective expression of the features at different scales through an attention map, so as to obtain features with detailed information.

[0039] The deep supervision module is configured to apply a deep supervision strategy to the original feature path and the multi-scale refined feature path of the decoder respectively, and combine a cross-entropy loss and a Dice loss to optimize the model in each path.

[0040] The image segmentation module is configured to automatically segment the colorectal cancer image to be processed based on the optimized model.

[0041] In a third aspect, the present application provides an electronic device, comprising:

[0042] at least one processor; and

[0043] a memory in communication with the at least one processor; wherein

[0044] The memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor to enable the at least one processor to execute the colorectal cancer automatic segmentation method based on the multi-phase feature alignment fusion model.

[0045] In a fourth aspect, the present application provides a computer readable storage medium storing a program, and the program is executed by a processor to implement the colorectal cancer automatic segmentation method based on the multi-phase feature alignment fusion model.

[0046] Compared with the prior art, the present application has the following advantages and beneficial effects:

[0047] (1) Due to the particularity of medical images, generally only radiologists and people with medical clinical knowledge can judge CRC and perform pixel-level annotation based on images, and the present application greatly reduces the burden of doctors on image annotation layer by layer, and realizes automatic CRC segmentation.

[0048] (2) The present application makes full use of the complementary information in multi-phase CT images, and through the multi-phase feature alignment fusion module, the arterial phase and venous phase features are adaptively aligned and fused, effectively alleviating the spatial misalignment problem and significantly improving the segmentation accuracy. Compared with the simple feature stacking method relying on image alignment in the prior art, the multi-phase feature alignment fusion module can dynamically adjust the contribution of different phase features to ensure accurate integration of information.

[0049] (3) The present application introduces an attention-guided multi-scale fusion module to fuse feature information of different resolutions, dynamically focuses on key areas using the attention mechanism, and significantly improves the adaptability of the model to colorectal cancer lesions with complex size and morphology. Compared with the scheme lacking multi-scale processing in the prior art, the present application performs more robustly in complex scenes.

[0050] (4) The application uses a double-path deep supervision strategy to supervise the traditional decoder output and multi-scale refined features respectively, maintains the integrity of multi-scale information, and avoids feature loss. Compared with the single-path supervision method in the prior art, the double-path deep supervision strategy enhances the model's ability to capture details, making the segmentation result more consistent and delicate. BRIEF DESCRIPTION OF DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0052] Figure 1 The flow chart of the colon cancer automatic segmentation method based on the multi-phase feature alignment fusion model of the embodiment of the present application;

[0053] Figure 2 The block diagram of the colon cancer automatic segmentation system based on the multi-phase feature alignment fusion model of the embodiment of the present application.

[0054] Figure 3 The structural diagram of the electronic device of the embodiment of the present application. DETAILED DESCRIPTION

[0055] In order to make the person skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0056] In the present application, the phrase "embodiment" means that the specific features, structures or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase in the specification does not necessarily mean the same embodiment, nor is it an independent or alternative embodiment to other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments.

[0057] Multi-phase features refer to features extracted from images acquired at different time points (phases) from the same medical image examination. For example, in contrast-enhanced CT (CECT), after the patient is injected with contrast agent, scanning is performed at the arterial phase and the venous phase respectively, obtaining images of different hemodynamic states. The images of each phase reflect the enhancement characteristics of the tissue at different time points, and the multi-phase features are a mathematical representation of these differentiated information. Fusing multi-phase features can comprehensively utilize the vascular heterogeneity information in the arterial phase and the morphological information in the venous phase, avoiding the one-sidedness of single-phase information.

[0058] As shown in the embodiment, the multi-phase feature alignment fusion model-based automatic colorectal cancer segmentation method provided by the embodiment comprises the following steps: Figure 1

[0059] S1, multi-phase feature extraction;

[0060] The acquired arterial phase CT image and the venous phase CT image are respectively input into two groups of encoders to extract multi-scale features, thereby obtaining multi-scale arterial phase features and venous phase features

[0061] The multi-phase feature extraction can be expressed by the following formula:

[0062]

[0063] It can be understood that in step S1, the two groups of encoders are arranged in parallel and perform feature extraction on the arterial phase CT image and the venous phase CT image respectively. The encoder comprises a plurality of double convolution layers and a standardization layer, and is used to extract multi-scale arterial phase features and venous phase features.

[0064] S2, using a multi-phase alignment fusion module to perform fusion processing on the multi-scale arterial phase features and venous phase features to generate aligned and fused features Specifically,

[0065] The multi-phase alignment fusion module uses and as inputs to generate The overall dual-phase feature aggregation can be expressed as:

[0066]

[0067] The fused features are obtained by weighted sum of and

[0068] where W i phase ​​may be expressed as:

[0069]

[0070] where Concat(·) represents a concatenation operation, i.e., combining the arterial and venous phase features in the channel dimension, C(·) represents using a 3x3x3 convolutional layer, then passing through batch normalization BN(·) and ReLu activation function ReLU(·), finally, through the Sigmoid function σ(·) to ensure W i ap +W i vp = 1.

[0071] It can be understood that the current multi-phase fusion module may be affected by the problem of spatial misalignment, thereby hindering the effective integration of multi-phase information. In order to solve this problem, the present application introduces a module inspired by Transformer-TransConv. The TransConv module combines an inverted bottleneck structure with a deep separable convolution (DWC) to achieve feature alignment by expanding the receptive field. Thus, the attention weight can be further optimized as:

[0072]

[0073] where TransConv(·) is defined as:

[0074] TransConv(·) = C1(GN(GL(C1(BN(DWC(x)))))) + x

[0075] In this formula, the deep separable convolution (DWC(·)) uses a 5x5x5 convolution kernel size to expand the receptive field of the neural network while maintaining computational efficiency.

[0076] S3, obtaining detailed features through an attention-guided multi-scale fusion module, specifically:

[0077] In order to enable the network to fully learn the context information in the decoding stage, the present application proposes an attention-guided multi-scale fusion module (AGMF). Specifically, the features D i and After concatenation, the attention map guides the effective expression of the features at different scales, thereby obtaining refined features Specifically, it can be expressed as:

[0078]

[0079] wherein UP(·) represents an up-sampling operation, Concat(·) represents a concatenation operation, and represents a matrix multiplication operation.Ch(·) performs a channel convolution operation using a 3x3x3 convolution kernel, and sigma is a sigmoid activation function used to obtain an attention map at a current scale. c is a combination operation composed of a 1x1x1 convolution kernel, batch normalization and an activation function applied to the output, thereby generating features with rich detailed information.

[0080] S4, a depth supervision strategy is applied to the original feature path and the multi-scale refined feature path of the decoder respectively;

[0081] To effectively capture the original features and multi-scale features, the application proposes a double-path depth supervision strategy, so that the total loss L total is expressed as:

[0082]

[0083] wherein, and are calculated in a similar manner and can be expressed as:

[0084]

[0085]

[0086] In order to keep the network consistent globally and put the main optimization pressure on the last layer, so as to ensure that the final output has the best pixel-level accuracy, the application takes gamma, delta, epsilon, xi and zeta as 0.0625, 0.125, 0.25, 0.5 and 1 respectively. and They use a combination of cross-entropy loss and Dice loss as the segmentation loss, that is:

[0087] L seg =L dice +L CE

[0088] The implementation of the Dice loss is as follows:

[0089]

[0090] wherein x is the softmax output of the network, and y is the one-hot encoding of the label segmentation map. The shapes of x and y are IxK, vI represents the number of voxels in the training batch, and kK represents the number of classes.

[0091] The cross-entropy loss is as follows:

[0092]

[0093] where N is the number of input CT images, p j ∈ [0, 1] represents the prediction probability of the j-th individual data, q j ∈ [0, 1] represents the true segmentation label corresponding to the j-th individual data.

[0094] It can be understood that in the embodiment, the cross-entropy loss focuses on the confidence of the pixel-by-pixel classification in tumor segmentation. The Dice loss focuses on the overall overlap of the predicted result and the true tumor region. Dice directly measures the intersection union ratio of two sets.

[0095] Therefore, the cross-entropy is responsible for the pixel-level fine classification and detail extraction, and the Dice loss is responsible for the region-level complete coverage and overall consideration. The two complement each other, so as to simultaneously consider the boundary accuracy and the overall tumor volume overlap, and improve the segmentation quality.

[0096] S5, based on the optimized model, the colorectal cancer image to be processed is automatically segmented.

[0097] Further, in the optimization process, the test process of the model is included, specifically:

[0098] The input picture is calculated in the segmentation network to obtain a prediction result:

[0099] P pred =f seg (AP,VP)

[0100] In the test stage, the trained network model only needs to input the arterial phase CT image AP and the venous phase CT image VP as input, and simply retains the segmentation prediction P pred as the final output.

[0101] In a more specific embodiment, the above-mentioned colorectal cancer automatic segmentation method based on multi-phase feature alignment fusion model includes the following steps:

[0102] (1) Input multi-phase CT images: resample the original input CT images AP and VP to a space with a size of (0.746 mm, 0.746 mm, 5 mm), and then randomly crop them into patches X with a size of 256x256x32 voxels. The corresponding pixel-level annotation label y and the coordinate graph e are two matrices with the same size as X. Y contains three values 0 and 1, where different values represent the semantic class to which the corresponding pixel belongs (0 represents background and 1 represents colorectal cancer), and e ranges between [0, 1].

[0103] (2) Feature encoding: two independent encoders are used to extract features from the arterial phase and venous phase images respectively to obtain preliminary feature representations of the multi-phase. Each encoder includes multiple convolutional layers for extracting features of different scales.

[0104] (3) Multi-phase feature alignment and fusion: The multi-phase feature alignment and fusion module is used to align and fuse the arterial and venous phase features. First, the cross-phase attention mechanism is used to dynamically calculate the weights of the arterial and venous phase features. Next, the Transformer-like module is integrated to expand the receptive field to alleviate the spatial misalignment problem, and finally the multi-phase fusion features after alignment are output.

[0105] (4) Multi-scale feature fusion: The aligned multi-phase features and the features of the traditional decoder are input into the attention-guided multi-scale fusion module to improve the model's segmentation ability for complex lesion regions by fusing features of different resolutions.

[0106] (5) Dual-path deep supervision strategy: One path supervises the traditional decoder features, and the other path supervises the multi-scale refined features to ensure information integrity and consistency of the segmentation results.

[0107] (6) The overall optimization process uses the RAdam optimizer with an initial learning rate of 0.001, and the specific learning rate decay strategy is polynomial decay. The batch size of the training process is 2, and the iteration training is 600 times.

[0108] The present application conducts ablation research to evaluate the impact of steps (3), (4), and (5) components on model performance. Table 1 below shows the specific results.

[0109] Table 1

[0110]

[0111] Data with * indicates p-value<0.05compared with baseline method -Multi-phase(cat).

[0112] Effectiveness of multi-phase information: The present application extends the network, using dual encoders for arterial and venous phase CT images, and connecting multi-phase features on multiple layers. As shown in Table 1, integrating multi-phase information improves the DSC score by 2.38% compared with the single arterial phase method and by 1.33% compared with the single venous phase method. These results confirm that using complementary information from AP and VP enhances CRC segmentation.

[0113] Effectiveness of the multi-phase feature alignment and fusion module: To evaluate the effectiveness of the multi-phase feature alignment and fusion module of the present application, the present application replaces the cascading operation in the multi-phase network with MPFA. As shown in the table above, MPFA achieves more comprehensive feature fusion, and the DSC performance is improved by 0.35% compared with simple cascading. The simulation results verify the effectiveness of the proposed method in the integration and alignment of multi-phase image feature information.

[0114] Effectiveness of multi-scale feature fusion: To comprehensively evaluate the attention-guided multi-scale fusion (AGMF) module, the present application conducts two comparative experiments: 1) "Multi-phase+AGMF" without MPFA and 2) "Multi-phase+MPFA+AGMF" with the complete architecture. As shown in the table above, the AGMF module produces DSC improvements of 0.41% and 1.11%, respectively, compared with its "Multi-phase(cat)" configuration. It is worth noting that the enhanced performance in the second experiment indicates that AGMF can better utilize multi-scale features for accurate tumor boundary delineation when combined with the feature alignment capability of MPFA. This synergistic effect shows that MPFA provides well-aligned feature representations, laying a solid foundation for the AGMF module's cross-scale attention mechanism.

[0115] Effectiveness of the dual-path deep supervision strategy: To study the effect of the proposed dual-path deep supervision strategy (DDS), the present application conducts "Multi-phase+AGMF+DDS" and "Multi-phase+MPFA+AGMF+DDS" experiments. As shown in the table, these results show that the introduction of the DDS strategy into the network helps to retain rich multi-scale features during model optimization, thereby improving sensitivity to target features and improving segmentation accuracy.

[0116] The application systematically optimizes the medical image segmentation performance through a multi-phase relative alignment fusion module, an attention-guided multi-scale fusion module and a double-path deep supervision strategy. The multi-phase relative alignment fusion module adopts a Transformer structure based on a cross-phase attention mechanism to dynamically align the feature distribution of the arterial phase and the venous phase CT images and expand the receptive field, effectively alleviating the spatial misalignment of multi-phase images caused by scanning time delay and the insufficient complementary utilization of features, and significantly improving the liver metastasis lesion segmentation accuracy. The attention-guided multi-scale fusion module combines a channel-space double attention mechanism through a pyramid feature integration network to adaptively focus on key areas, strengthen the modeling ability of the shape diversity of colorectal cancer lesions, and greatly improve the detection sensitivity of small lesions and heterogeneous areas. The double-path deep supervision strategy introduces a double supervision loss of a traditional feature branch and a multi-scale refined feature branch in parallel in the decoder, maintains the integrity of multi-scale information through double constraints, ensures the stability and boundary continuity of the segmentation results, and finally realizes robust and accurate lesion segmentation.

[0117] It should be noted that, for the foregoing method embodiments, in order to facilitate description, they are all expressed as a series of action combinations, but those skilled in the art should know that the application is not limited by the order of the described actions, because according to the application, certain steps can be performed in other orders or simultaneously.

[0118] Based on the same idea as the automatic colorectal cancer segmentation method based on the multi-phase feature alignment fusion model in the above embodiment, the application also provides an automatic colorectal cancer segmentation system based on a multi-phase feature alignment fusion model, which can be used to execute the automatic colorectal cancer segmentation method based on the multi-phase feature alignment fusion model described above. For the sake of convenience, in the structural schematic diagram of the embodiment of the automatic colorectal cancer segmentation system based on the multi-phase feature alignment fusion model, only the parts related to the embodiments of the application are shown, and those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and can include more or fewer components than the illustrated, or combine certain components, or different component arrangements.

[0119] Please refer to Figure 2 In another embodiment of the present application, an automatic colorectal cancer segmentation system 100 based on a multi-phase feature alignment fusion model is provided, which includes a feature extraction module 101, a multi-phase relative alignment fusion module 102, an attention-guided multi-scale fusion module 103, a deep supervision module 104 and an image segmentation module 105.

[0120] The feature extraction module 101 is configured to extract multi-scale features of the arterial phase CT image and the venous phase CT image through parallel encoders respectively, to obtain multi-scale arterial phase features and venous phase features.

[0121] The multi-phase alignment fusion module 102 is configured to input arterial phase features and venous phase features under the same scale into corresponding multi-phase alignment fusion modules, and generate features after alignment and fusion by dynamically allocating weights through a cross-phase attention mechanism; the cross-phase attention mechanism realizes receptive field expansion by combining an inverted bottleneck structure similar to a Transformer structure with a deep separable convolution;

[0122] The attention-guided multi-scale fusion module 103 is configured to splice features after alignment and fusion of adjacent scales through the attention-guided multi-scale fusion module in the decoding stage, and guide effective expression of the features on different scales through an attention map, so as to obtain features with detailed information.

[0123] The deep supervision module 104 is configured to apply a deep supervision strategy to an original feature path and a multi-scale refined feature path of the decoder respectively, and combine a cross-entropy loss and a Dice loss to optimize the model in each path.

[0124] The image segmentation module 105 is configured to automatically segment a colorectal cancer image to be processed based on the optimized model.

[0125] It should be noted that the colorectal cancer automatic segmentation system based on the multi-phase feature alignment fusion model of the present application corresponds to the colorectal cancer automatic segmentation method based on the multi-phase feature alignment fusion model of the present application. The technical features and advantages described in the embodiment of the colorectal cancer automatic segmentation method based on the multi-phase feature alignment fusion model are applicable to the embodiment of the colorectal cancer automatic segmentation based on the multi-phase feature alignment fusion model. For specific content, please refer to the description in the method embodiment, which will not be repeated here. It is hereby declared.

[0126] In addition, in the embodiment of the colorectal cancer automatic segmentation system based on the multi-phase feature alignment fusion model of the above embodiment, the logical division of each program module is only an example. In actual application, the above functions can be completed by different program modules according to needs, for example, considering the configuration requirements of the corresponding hardware or the convenience of software implementation, that is, the internal structure of the colorectal cancer automatic segmentation system based on the multi-phase feature alignment fusion model is divided into different program modules to complete all or part of the functions described above.

[0127] Please refer to Figure 3In one embodiment, an electronic device implementing a method for automatic segmentation of colorectal cancer based on multi-phase feature alignment fusion model is provided. The electronic device 200 can include a first processor 201, a first memory 202, and a bus. The electronic device 200 can further include a computer program stored in the first memory 202 and executable on the first processor 201, such as an automatic segmentation program of colorectal cancer based on multi-phase feature alignment fusion model 203.

[0128] The first memory 202 includes at least one type of readable storage medium, such as flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the first memory 202 can be an internal storage unit of the electronic device 200, such as a mobile hard disk of the electronic device 200. In other embodiments, the first memory 202 can also be an external storage device of the electronic device 200, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the first memory 202 can include both an internal storage unit and an external storage device of the electronic device 200. The first memory 202 can be used to store application software and various data installed in the electronic device 200, such as the code of the automatic segmentation program of colorectal cancer based on multi-phase feature alignment fusion model 203, and can also be used to temporarily store data that has been output or will be output.

[0129] The first processor 201 can be composed of an integrated circuit in some embodiments, such as a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, combinations of various control chips, etc. The first processor 201 is the control unit of the electronic device, which connects various components of the electronic device through various interfaces and lines, and executes or runs programs or modules stored in the first memory 202 and calls data stored in the first memory 202 to perform various functions and process data of the electronic device 200.

[0130] Figure 3 Only an electronic device with components is shown, and those skilled in the art can understand that, Figure 3The illustrated structure does not constitute a limitation on the electronic device 200, and can include fewer or more components than illustrated, or combine certain components, or different component arrangements.

[0131] The colorectal cancer automatic segmentation program 203 stored in the first memory 202 in the electronic device 200 is a combination of multiple instructions, which, when running in the first processor 201, can implement:

[0132] Multi-scale features of the arterial phase CT image and the venous phase CT image are extracted by parallel encoders respectively, to obtain multi-scale arterial phase features and venous phase features;

[0133] The arterial phase features and the venous phase features under the same scale are input into corresponding multi-phase alignment fusion modules, and through a cross-period attention mechanism, weights are dynamically allocated to generate aligned and fused features; the cross-period attention mechanism uses an inverted bottleneck structure similar to the Transformer structure combined with a depth separable convolution to realize the expansion of the receptive field;

[0134] In the decoding stage, through the attention-guided multi-scale fusion module, the aligned and fused features of adjacent scales are spliced, and then through the attention map, the effective expression of the features is guided at different scales, so as to obtain features with detailed information;

[0135] A depth supervision strategy is applied to the original feature path and the multi-scale refined feature path of the decoder respectively, and in each path, a cross-entropy loss and a Dice loss are combined to optimize the model;

[0136] Based on the optimized model, the colorectal cancer image to be processed is automatically segmented.

[0137] Further, the modules / units of the electronic device 200, if implemented in the form of software function units and sold or used as independent products, can be stored in a non-volatile computer readable storage medium. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM).

[0138] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing relevant hardware. The program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0139] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.

[0140] The above embodiments are the preferred embodiments of the present application, but the embodiments of the present application are not limited to the above embodiments, and any changes, modifications, substitutions, combinations and simplifications made without departing from the spirit and principles of the present application should be equivalent replacement methods and should be included in the protection scope of the present application.

Claims

1. An automatic segmentation method for colorectal cancer based on a multi-phase feature alignment and fusion model, characterized in that, Includes the following steps: Multi-scale features of arterial phase CT images and venous phase CT images are extracted separately by parallel encoders to obtain multi-scale arterial phase features and venous phase features. Arterial and venous phase features at the same scale are input into the corresponding multi-phase relatively aligned fusion module. Weights are dynamically allocated through a cross-phase attention mechanism to generate aligned and fused features. The cross-phase attention mechanism uses an inverted bottleneck structure similar to Transformer combined with depthwise separable convolution to expand the receptive field. During the decoding stage, the attention-guided multi-scale fusion module stitches together the aligned and fused features from adjacent scales, and then guides the effective expression of features at different scales through attention maps, thereby obtaining features with detailed information. A deep supervision strategy is applied to the original feature path and the multi-scale refined feature path of the decoder, and the model is optimized by combining cross-entropy loss and Dice loss in each path. The optimized model is used to automatically segment the colorectal cancer images to be processed.

2. The automatic segmentation method for colorectal cancer based on a multi-phase feature alignment and fusion model according to claim 1, characterized in that, Each encoder consists of multiple double convolutional layers and normalization layers. In each stage of the encoder, a combination of double convolution and normalization is performed until features on all channels have been extracted.

3. The automatic segmentation method for colorectal cancer based on a multi-phase feature alignment and fusion model according to claim 1, characterized in that, The specific implementation of the multi-phase relative homogeneity fusion module includes: Arterial and venous phase features at the same scale are used as inputs to the corresponding multi-phase relatively homogeneous fusion module; Arterial and venous phase features are merged along the channel dimension. The merged features are then processed by convolution, normalization, and linear correction. The corrected features are then input into the TransConv block, which combines an inverted bottleneck structure with depthwise separable convolution to achieve feature alignment by expanding the receptive field. Dynamic weights for arterial and venous phase features are generated using activation functions; The fused features are obtained by weighting and summing the arterial phase features and venous phase features according to dynamic weights.

4. The automatic segmentation method for colorectal cancer based on a multi-phase feature alignment and fusion model according to claim 3, characterized in that, The TransConv block is defined as follows: TransConv(·)=C1(GN(GL(C1(BN(DWC(x))))))+x Where DWC(·) is a depthwise separable convolution with a 5×5×5 kernel, GN(·) is group normalization, GL(·) is the GELU activation function, and C1(·) is a 1×1×1 convolution.

5. The automatic segmentation method for colorectal cancer based on a multi-phase feature alignment and fusion model according to claim 1, characterized in that, The attention-guided multi-scale fusion module is represented as follows: in These represent features with detailed information; UP(·) indicates an upsampling operation, and Concat(·) indicates a concatenation operation. Ch(·) represents matrix multiplication, and Ch(·) indicates channel convolution using a 3×3×3 kernel; σ is the sigmoid activation function used to obtain the attention map at the current scale, and f c This involves applying a combination of operations consisting of a 1×1×1 convolution kernel, batch normalization, and activation functions to the output to generate multi-scale refined features.

6. The automatic segmentation method for colorectal cancer based on a multi-phase feature alignment and fusion model according to claim 1, characterized in that, The dual-path deep supervision strategy includes original feature path supervision and multi-scale refined feature supervision; The original feature path supervision directly applies supervision to the features of each layer of the decoder; The multi-scale refinement feature supervision applies supervision to the refinement features output by the attention-guided multi-scale fusion module.

7. The automatic segmentation method for colorectal cancer based on a multi-phase feature alignment and fusion model according to claim 1, characterized in that, The optimization model combining cross-entropy loss and Dice loss is as follows: The Dice loss is implemented as follows: Where x is the softmax output of the network, y is the one-hot encoding of the label segmentation map, the shape of x and y is I×K, v∈I represents the number of voxels in the training batch, and k∈K represents the number of classes; The implementation of the cross-entropy loss is as follows: Where N is the number of input CT images, p j ∈[0,1] represents the predicted probability of the j-th individual data, q j ∈[0,1] represents the true segmentation label corresponding to the j-th individual data.

8. An automatic segmentation system for colorectal cancer based on a multi-phase feature alignment and fusion model, characterized in that, The automatic segmentation method for colorectal cancer based on the multi-phase feature alignment fusion model applied to any one of claims 1-7 includes a feature extraction module, a multi-phase alignment fusion module, an attention-guided multi-scale fusion module, a depth supervision module, and an image segmentation module. The feature extraction module is used to extract multi-scale features from arterial phase CT images and venous phase CT images respectively through parallel encoders, so as to obtain multi-scale arterial phase features and venous phase features. The multi-phase relative homogeneous fusion module is used to input arterial phase features and venous phase features at the same scale into the corresponding multi-phase relative homogeneous fusion module, and dynamically allocate weights through a cross-phase attention mechanism to generate aligned and fused features; the cross-phase attention mechanism adopts an inverted bottleneck structure similar to the Transformer structure combined with depthwise separable convolution to achieve receptive field expansion. The attention-guided multi-scale fusion module is used in the decoding stage to stitch together the aligned and fused features of adjacent scales, and then guide the effective expression of features at different scales through the attention map, thereby obtaining features with detailed information. The deep supervision module is used to apply deep supervision strategies to the original feature path and the multi-scale refined feature path of the decoder, respectively. Each path combines cross-entropy loss and Dice loss to optimize the model. The image segmentation module is used to automatically segment the colorectal cancer image to be processed based on the optimized model.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions that can be executed by the at least one processor, the computer program instructions being executed by the at least one processor to enable the at least one processor to perform the automatic segmentation method for colorectal cancer based on a multi-phase feature alignment and fusion model as described in any one of claims 1-7.

10. A computer-readable storage medium storing a program, characterized in that, When the program is executed by the processor, it implements the automatic segmentation method for colorectal cancer based on the multi-phase feature alignment and fusion model as described in any one of claims 1-7.