Image Segmentation Method, Apparatus and Device Based on Multimodal Task-Specific Fusion

Through the multimodal task-specific fusion method, the problems of multimodal feature fusion and task adaptive modulation in medical image segmentation are solved, and the high accuracy and robustness of tumor and lymph node segmentation are achieved, adapting to changes in different data sets and patient groups, and improving the reliability of clinical applications.

CN120125829BActive Publication Date: 2025-07-22XIAMEN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510607867.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-22
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The existing medical image segmentation technology has insufficient in multimodal feature fusion, task adaptive modulation and anatomical prior knowledge utilization, resulting in low accuracy of tumor and lymph node segmentation, poor robustness, and difficulty in adapting to changes in different data sets and patient populations.

Method used

Through multimodal task-specific fusion methods, including feature alignment, task-driven modulation, selective inhibition and neuron clustering processing, the feature fusion and segmentation process is optimized by combining anatomical area priors and graph neural network propagation mechanisms.

Benefits of technology

It significantly improves the accuracy, robustness and adaptability of tumor and lymph node segmentation, and provides more reliable clinical application support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125829B_ABST
    Figure CN120125829B_ABST
Patent Text Reader

Abstract

The image segmentation method, device and equipment based on multi-modal task-specific fusion provided by the present invention relate to the technical field of image segmentation. In the present invention, the obtained medical images of different modalities are respectively mapped to a shared feature space for feature alignment; then, according to the task requirements, the contributions of the modality features after feature alignment in different tasks are dynamically adjusted to obtain a task-driven adjusted feature map; and a selective suppression mechanism is adopted for dynamic filtering to obtain a suppressed modality feature map; simulating the biological neuron cluster mechanism, features related to the task requirements in the suppressed modality feature map are respectively extracted and enhanced through different convolutional structures to obtain neuron cluster features, and their spatial features and channel features are respectively extracted and fused and reconstructed to obtain fused features; finally, the fused features are input into the segmentation space and combined with an activation function to output a segmentation result. The present invention improves the accuracy, robustness of medical image segmentation and the adaptability to multi-modal medical images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of neural networks and image segmentation, and in particular, to an image segmentation method, device, equipment and medium based on multi-modal task-specific fusion. Background Art

[0002] In the field of medical image analysis, the accurate segmentation of laryngeal cancer and lymph nodes is of great significance. The medical image segmentation task needs to accurately extract the target area from complex anatomical structures, which is crucial for disease diagnosis, staging and treatment planning. Traditional methods mainly rely on manual feature extraction and rule-based algorithms, but these methods often show problems of insufficient accuracy and poor robustness when dealing with multi-modal data and complex anatomical structures, and it is difficult to meet the clinical needs. With the development of deep learning technology, the automatic segmentation of nasopharyngeal carcinoma has gradually turned to methods based on single-modal or multi-modal medical images. Single-modal segmentation methods usually rely on medical images of T1, T1c or T2 modalities and perform segmentation through convolutional neural networks (CNNs). Among them, the T1 modality is mainly used for the identification of anatomical structures, the T1c modality focuses on the detection of tumor enhancement signals, and the T2 modality is used for the identification of edema and fluid. However, single-modal methods are limited by the limitations of their information and are difficult to comprehensively capture the complex features of tumors and lymph nodes.

[0003] In recent years, more and more studies have begun to attempt multi-modal information fusion in order to improve the segmentation accuracy by using the complementarity between different modalities. Existing multi-modal methods usually fuse the features of T1, T1c and T2 modalities by means of splicing, weighting or joint learning, and perform joint training based on deep learning models (such as U-Net and its variants). Although these methods have improved the segmentation effect to a certain extent, there are still significant deficiencies. First of all, multi-modal feature fusion methods usually rely on simple splicing or weighted average operations, and fail to fully optimize and effectively fuse the features of different modalities, resulting in the complementarity between modalities not being fully exerted. Secondly, existing methods generally ignore the extraction of task-specific features and fail to adaptively adjust the modal contributions according to the different requirements of T staging (tumor segmentation) and N staging (lymph node segmentation), making the model unstable when dealing with complex anatomical structures. In addition, existing technologies often fail to make full use of the anatomical prior knowledge in medical images, and it is difficult to achieve an ideal segmentation accuracy when the boundary between tumors and normal tissues is blurred. Finally, existing technologies have weak generalization ability when dealing with changes in different datasets, devices and patient groups, which limits their wide application in actual clinical environments.

[0004] In view of this, the applicant specifically proposes this application. Summary of the Invention

[0005] The present invention aims to provide an image segmentation method, device, equipment and medium based on multi-modal task-specific fusion, so as to solve the problems existing in the existing medical image segmentation technology in aspects such as multi-modal feature fusion, task adaptive modulation, and utilization of anatomical prior knowledge.

[0006] To solve the above technical problems, the present invention is realized through the following technical solutions:

[0007] An image segmentation method based on multi-modal task-specific fusion, including:

[0008] S1, obtaining medical images of different modalities;

[0009] S2, mapping the medical images of different modalities to a shared feature space respectively for feature alignment;

[0010] S3, dynamically adjusting the contributions of the modality features after feature alignment in different tasks according to task requirements to obtain a task-driven adjusted feature map;

[0011] S4, using a selective suppression mechanism to dynamically filter the task-driven adjusted feature map to obtain a suppressed modality feature map;

[0012] S5, simulating the biological neuron cluster mechanism, and respectively extracting and enhancing the features related to task requirements in the suppressed modality feature map through different convolutional structures to obtain neuron cluster features;

[0013] S6, respectively extracting the spatial features and channel features of the neuron cluster features and performing fusion reconstruction to obtain fusion features;

[0014] S7, inputting the fusion features into a segmentation space and outputting a segmentation result in combination with an activation function.

[0015] Preferably, the task requirements include a T staging task and an N staging task; wherein, the T staging task is to highlight the spatial details of anatomical regions by strengthening the spatial information of features; the N staging task is to enhance the contrast information between modality features to highlight the contrast in features.

[0016] Preferably, when dynamically adjusting the contributions of the modality features after feature alignment in different tasks according to task requirements:

[0017] Performing global compression of the features spatially through average pooling to generate modulation weights for the T staging task;

[0018] Highlighting the contrast in features through max pooling to generate modulation weights for the N staging task.

[0019] Preferably, the selective suppression mechanism is used to dynamically filter different modality information according to different tasks, and S4 is specifically:

[0020] First, calculate the modality difference degree between the current modality and other modalities. The formula is:

[0021] ;

[0022] Wherein, represents the modality difference degree of the -th modality at the pixel position ; represents the feature map of the -th modality after task modulation, that is, the task-driven adjusted feature map of the -th modality; represents the L1 norm; represents other modalities different from the modality ;

[0023] Then, introduce task-related modulation weights to weight the modality difference degree and generate a difference score to highlight the modality changes concerned by the current task. The formula is:

[0024] ;

[0025] Wherein, represents the difference score of the -th modality at the pixel position under the current task H; represents the task-sensitive weight of the -th modality under the current task H;

[0026] Finally, map the difference score through the Sigmoid activation function to obtain an inhibition weight, so as to obtain the final inhibited modality feature map. The formula is:

[0027] ;

[0028] ;

[0029] Wherein, represents the inhibition coefficient; represents the Sigmoid activation function, which is used to map the score into a differentiable weight; is the final inhibited modality feature map.

[0030] Preferably, when simulating the mechanism of biological neuron clusters to extract features of different task requirements, use the depthwise separable convolution structure to extract the features of the T staging task to strengthen the spatial information; use the 1x1 convolution structure to capture the global contrast information of the N staging task to enhance the perception of lymph nodes. The formulas are respectively:

[0031] ;

[0032] ;

[0033] Among them, and respectively represent the neuron cluster features of the T staging task and the N staging task; and respectively represent the neuron cluster processing modules of the T staging task and the N staging task, and extract features related to the tumor boundary and lymph node distribution respectively; and respectively represent the inhibitory mode features of the T staging task and the N staging task.

[0034] Preferably, the S6 is specifically:

[0035] Extract the spatial feature and channel feature of the neuron cluster feature through convolution operation respectively, and the formula is:

[0036] ;

[0037] ;

[0038] Among them, and respectively represent the extracted spatial feature and channel feature; and respectively represent the decoupling operations of the spatial feature and the channel feature; represents the neuron cluster feature of the T staging task;

[0039] Fuse the spatial feature and the channel feature through a reconstruction operation to obtain the final fused feature, and the formula is:

[0040] ;

[0041] Among them, represents the fused feature; represents the reconstruction operation.

[0042] Preferably, the S7 is specifically: Map the fused feature to the segmentation spaces of the T staging task and the N staging task respectively, and output the segmentation result of each task through the Sigmoid activation function, and the expression:

[0043] ;

[0044] ;

[0045] Among them, and The segmentation predictions for the T-stage task and the N-stage task respectively; is the fused feature; is the Sigmoid activation function; 、 are the weights for the T-stage task and the N-stage task respectively; 、 are the biases for the T-stage task and the N-stage task respectively.

[0046] Preferably, it further includes optimizing the fused feature by using anatomical region priors and the graph neural network propagation mechanism, specifically:

[0047] First, the medical image is divided into multiple anatomical regions according to the anatomical structure;

[0048] Next, the global semantic average representation within each anatomical region is calculated based on the fused feature map as the feature input for the nodes of the anatomical region graph, and the structural similarity between adjacent anatomical regions is used as the connection strength of the edges of the anatomical region graph to construct the global anatomical region graph. The formula is:

[0049] ;

[0050] ;

[0051] Among them, represents the graph node feature vector of anatomical region r; R represents the reconstructed data; represents all the pixel sets of anatomical region r; represents the feature representation at the pixel position in the fused feature map;

[0052] represents anatomical region and the edge connection strength between; 、 respectively represent the geometric center coordinates of anatomical regions 、 ; represents the square of the Euclidean distance; is the scale parameter of the Gaussian distance kernel;

[0053] Then, the medical image is divided into several patch blocks, and a local patch adjacency graph is constructed according to the spatial adjacency between the patches, and the features of each patch are extracted and updated through a standard GCN;

[0054] The node embedding information of the global anatomical region graph is fed back to the local patch features for cross-layer fusion propagation. The expression is:

[0055] ;

[0056] Among them, is the feature of the i-th patch; is the feature of the updated i-th patch, which incorporates the high-level structural information of the anatomical region map;

[0057] is a learnable cross-image projection weight matrix for mapping structural semantics from the regional domain to the patch domain;

[0058] represents the belonging probability that the i-th patch belongs to region r, which is calculated based on the distance between the patch's position and the region's center, and the formula is:

[0059] ;

[0060] Among them, is the center coordinate of the i-th patch; is the center coordinate of region r; is the center coordinate of region ; represents the Euclidean distance.

[0061] The present invention also provides a nasopharyngeal carcinoma image segmentation device based on multi-modal task-specific fusion, including:

[0062] An acquisition unit for acquiring medical images of different modalities;

[0063] A multi-modal feature alignment unit for mapping medical images of different modalities to a shared feature space for feature alignment;

[0064] A task-driven modulation unit for dynamically adjusting the contributions of the feature-aligned multi-modal features in different tasks according to task requirements to obtain a task-driven adjusted feature map;

[0065] A dynamic suppression unit for dynamically filtering the task-driven adjusted feature map using a selective suppression mechanism to obtain a suppressed modal feature map;

[0066] A neuron cluster unit for simulating the biological neuron cluster mechanism, and respectively extracting and enhancing the features related to task requirements in the suppressed modal feature map through different convolutional structures to obtain neuron cluster features;

[0067] A decoupling and reconstruction unit for respectively extracting the spatial features and channel features of the neuron cluster features and performing fusion and reconstruction to obtain fusion features;

[0068] A segmentation prediction unit, configured to input the fused features into a segmentation space and output a segmentation result by combining with an activation function.

[0069] The present invention also provides a nasopharyngeal carcinoma image segmentation device based on multi-modal task-specific fusion, including a processor and a memory. A computer program is stored in the memory, and the computer program can be executed by the processor to implement a method for image segmentation based on multi-modal task-specific fusion as described above.

[0070] The present invention also provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, a method for image segmentation based on multi-modal task-specific fusion as described above is implemented.

[0071] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0072] The present invention aims to combine innovative technologies such as multi-modal feature projection, task-driven modulation, neuron cluster processing, and spatial-channel decoupled fusion to improve the accuracy, robustness, and adaptability to multi-modal medical images of medical image segmentation of tumors and lymph nodes, thereby providing more reliable technical support for clinical applications.

[0073] The present invention solves the problem that multi-modal features in the prior art cannot be fully optimized and effectively fused through multi-modal feature alignment and fusion technology. Among them, independent projection mapping is used to align the features of multi-modal medical images (such as T1, T1c, and T2 modalities), and a task-driven modulation mechanism is used to dynamically adjust the contributions of each modality, giving full play to the complementarity between modalities.

[0074] Furthermore, the present invention improves the distribution modeling ability of medical images such as tumor boundaries and lymph nodes by task-adaptive modulation technology, which dynamically enhances or suppresses task-related features according to different task requirements. In particular, the modality-sensitive neuron dynamic suppression technology effectively suppresses task-irrelevant or ambiguous modality information through a selective suppression mechanism, highlighting the significant expression of the dominant modality.

[0075] In addition, the present invention uses spatial-channel decoupled fusion technology to extract spatial features and channel features respectively, and optimizes the feature fusion process through a reconstruction operation, further improving the segmentation accuracy.

[0076] Finally, the present invention uses a structure partition-guided double-layer graph optimization technology, combines anatomical region priors with a graph neural propagation mechanism, and realizes the modeling of the spatial structure correlation and context consistency between tumors and lymph nodes in medical images, significantly improving the structural consistency of segmentation.

[0077] The present invention can effectively fuse multi-modal image information, adaptively adjust the contributions of image modalities, and at the same time make full use of anatomical prior knowledge and the propagation mechanism of graph neural networks to optimize the fused features, significantly improving the performance of medical image segmentation such as nasopharyngeal carcinoma, and showing obvious advantages in terms of segmentation accuracy, robustness, and generalization ability, providing strong support for clinical diagnosis and treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0079] Figure 1 Schematic flow diagram of an image segmentation method based on multi-modal task-specific fusion provided for Example 1.

[0080] Figure 2 Schematic overall framework diagram of an image segmentation method based on multi-modal task-specific fusion provided for Example 1.

[0081] Figure 3 Schematic diagram of a nasopharyngeal carcinoma image segmentation device based on multi-modal task-specific fusion provided for Example 2.

[0082] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0083] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the present invention claimed, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0084] Example 1

[0085] Embodiment 1 of the present invention provides an image segmentation method based on multi-modal task-specific fusion, which can be implemented by an image segmentation device based on multi-modal task-specific fusion (hereinafter referred to as the segmentation device), and particularly, is executed by one or more processors in the segmentation device.

[0086] In this embodiment, the segmentation device can be an electronic device equipped with a processor, and the processor has a computer program of the image segmentation method based on multi-modal task-specific fusion and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited herein.

[0087] In medical image analysis, especially in neuroimaging (such as brain MRI) and tumor diagnosis, T1-weighted (T1), T1 Contrast-enhanced (T1c), and T2-weighted (T2) are three of the most commonly used MRI modalities. Each modality highlights different physical characteristics of tissues through different imaging parameters, providing complementary information for clinical diagnosis and machine learning tasks (such as segmentation, classification).

[0088] In this embodiment, T1 refers to T1-weighted imaging, which mainly reflects the differences in the longitudinal relaxation time of tissues. T1C is enhanced T1-weighted imaging, which is performed after intravenous injection of a gadolinium-based contrast agent (such as gadopentetate dimeglumine) on the basis of T1-weighted imaging. T2 refers to T2-weighted imaging, which highlights the differences in the transverse relaxation time (T2) through long TR and long TE parameters, reflecting the degree of freedom of water molecules in tissues (such as extracellular edema). T1 provides an anatomical reference, T1c enhances specificity, and T2 sensitively captures pathological changes. The combination of the three can significantly improve the diagnostic accuracy.

[0089] As Figure 1 shown, an image segmentation method based on multi-modal task-specific fusion includes steps S1 to S7.

[0090] S1, Obtain medical images of different modalities.

[0091] Obtain multi-modal medical images of the same anatomical region (such as T1, T1c, T2 MRI, or CT, PET, etc.).

[0092] S2, Map medical images of different modalities to a shared feature space respectively for feature alignment.

[0093] The purpose of this step is to map the input multi-modal images (such as T1, T1c, T2) to a shared feature space, solve the problem of modality heterogeneity, make the features of different modalities comparable, and facilitate subsequent fusion operations.

[0094] Since the features of different modalities vary in dimension and distribution, it is necessary to align the features of each modality to a unified space through projection operations. In this embodiment, for each modality, a 1x1 convolution operation is used to map it to a hidden space with the same number of channels (e.g., 256 dimensions). This operation ensures that the features of different modalities are fused in the same dimension, avoiding dimensional mismatches between modalities.

[0095] Specifically, taking the images of three modalities, T1, T1c, and T2, as an example, the expression of the projection operation is:

[0096] ;

[0097] where, is the projection operation; are the modality features after being mapped to the hidden space; are the input modality features; t represents the image modality.

[0098] S3. Dynamically adjust the contributions of the modality features after feature alignment in different tasks according to the task requirements to obtain a task-driven adjusted feature map.

[0099] After completing the alignment of multi-modal features, the next step is to enhance the contribution of each modality in different tasks through a task-driven modulation mechanism.

[0100] Specifically, in this embodiment, the task requirements include the T staging task and the N staging task; among them, the T staging task is to highlight the spatial details of the anatomical region (such as the edema boundary of the T2 modality) by strengthening the spatial information of the features; the N staging task is to enhance the contrast information between modality features to highlight the contrast in the features (such as the enhanced signal of the T1c modality).

[0101] In this embodiment, the three modality features are modulated respectively to dynamically enhance or suppress the task-related features.

[0102] T staging task modulation: Globally compress the features spatially through average pooling to generate the modulation weights for the T staging task. This operation helps to highlight spatial details such as tumor boundaries.

[0103] N staging task modulation: Highlight the contrast in the features through max pooling to generate the modulation weights for the N staging task, enhancing the perception of important structures such as the location of lymph nodes.

[0104] The formula for the modulated features is as follows:

[0105] ;

[0106] where, 、 Respectively represent the features after task-driven modulation of T staging and N staging; 、 、 、 Respectively represent the features of T, T1, T1c, and T2 modalities; 、 Respectively represent the task-driven modulation operations of T staging and N staging tasks, aiming to enhance task-related features.

[0107] S4, adopt a selective inhibition mechanism to dynamically filter the task-driven adjusted feature map to obtain an inhibited modality feature map.

[0108] To further highlight the dominant modalities and suppress redundant modalities, this module is designed based on the "selective inhibition mechanism" in neuroscience, and a task-sensitive dynamic inhibition mechanism is proposed. The selective inhibition mechanism is used to dynamically filter different modality information according to different tasks, including modality difference perception, task modulation weighting, and inhibition weight generation.

[0109] Modality difference perception: Calculate the modality difference degree between the current modality and other modalities. The formula is:

[0110] ;

[0111] Among them, represents the modality difference degree of the th modality at the pixel position ; represents the feature map of the th modality after task modulation, that is, the task-driven adjusted feature map of the th modality; represents the L1 norm, that is, the sum of the absolute differences between modalities at this position; represents other modalities different from modality ; , represents the difference between this modality and all other modalities.

[0112] This step is used to measure whether a certain modality has information uniqueness in a local area. The greater the difference, the more likely it is that the modality feature conflicts or is redundant with other modalities.

[0113] Task modulation weighting: Introduce task-related modulation weights to weight the modality difference degree and generate a difference score to highlight the modality changes concerned by the current task. The formula is:

[0114] ;

[0115] Among them, represents the The difference score of a modality at the pixel position ; denotes the task-sensitive weight of the -th modality under the current task H.

[0116] This step combines the task requirements and modality characteristics, re-weights the modality difference degree, and is used to highlight the modality changes that the current task pays more attention to.

[0117] Inhibitory weight generation: The inhibitory weight is obtained by mapping the difference score through the Sigmoid activation function, so as to obtain the final inhibitory modality feature map. The formula is:

[0118] ;

[0119] ;

[0120] where denotes the inhibition coefficient, and its range is between 0 and 1; denotes the Sigmoid activation function, which is used to map the score into a differentiable weight; is the final inhibitory modality feature map for subsequent fusion;

[0121] is the task-driven adjusted feature map of the -th modality.

[0122] When the score is larger, it indicates that there may be redundant information or ambiguous expressions in this modality. The smaller the inhibitory coefficient output by the Sigmoid, the more the contribution of this modality in this region is dynamically suppressed. On the contrary, when this modality performs well in this region, approaches 1, so as to retain its strong expression.

[0123] S5 simulates the mechanism of biological neuron clusters, and extracts and enhances the features related to the task requirements in the inhibitory modality feature map through different convolutional structures to obtain neuron cluster features.

[0124] This step adopts the mechanism of simulating biological neuron clusters, and extracts the features related to T staging (such as tumor segmentation) and N staging (such as lymph node segmentation) through different convolutional structures respectively. T staging mainly focuses on the boundary of the tumor, so depthwise separable convolution is used to strengthen the extraction of spatial information; N staging uses 1x1 convolution to capture global contrast information to enhance the perception of lymph nodes. The specific formulas are respectively:

[0125] ;

[0126] ;

[0127] Among them, and respectively represent the neuron cluster features of the T staging task and the N staging task; and respectively represent the neuron cluster processing modules of the T staging task and the N staging task, and respectively extract features related to the tumor boundary and lymph node distribution; and respectively represent the inhibitory mode features of the T staging task and the N staging task.

[0128] S6, respectively extract the spatial features and channel features of the neuron cluster features and perform fusion and reconstruction to obtain the fusion features.

[0129] In order to further improve the segmentation accuracy, the method of the present invention proposes a spatial-channel decoupled fusion strategy. In this step, first, the spatial features and channel features are respectively extracted through convolution operations. The spatial features mainly reflect the local structural information in the image (such as the tumor boundary), while the channel features reflect the contrast information between modalities (such as the enhanced signal of T1c). Then, the spatial features and channel features are fused through a reconstruction operation to obtain the final fusion features.

[0130] Specifically, the spatial features and channel features of the neuron cluster features are respectively extracted through convolution operations, and the formula is:

[0131] ;

[0132] ;

[0133] Among them, and respectively represent the extracted spatial features and channel features; and respectively represent the decoupling operations of the spatial features and channel features; represents the neuron cluster features of the T staging task;

[0134] The spatial features and channel features are fused through a reconstruction operation to obtain the final fusion features, and the formula is:

[0135] ;

[0136] Among them, represents the fusion feature; represents the reconstruction operation.

[0137] S7, input the fusion feature into the segmentation space and combine the activation function to output the segmentation result.

[0138] After feature fusion is completed, the final segmentation output is performed through a segmentation head module. This module maps the fused features to the segmentation spaces of T staging and N staging, and outputs the segmentation results of each task through the Sigmoid activation function.

[0139] Map the fused features to the segmentation spaces of the T staging task and the N staging task respectively, and output the segmentation results of each task through the Sigmoid activation function. The expression is:

[0140] ;

[0141] ;

[0142] where and are the segmentation predictions of the T staging task and the N staging task respectively; is the fused feature; is the Sigmoid activation function; and are the weights of the T staging task and the N staging task respectively; and are the biases of the T staging task and the N staging task respectively.

[0143] In another preferred embodiment, before the segmentation prediction step of the method of the present invention, the fused features can also be optimized for the segmentation structure by using anatomical region priors and graph neural network propagation mechanisms.

[0144] Anatomical Region Priors: In medical image analysis, different anatomical regions (such as the heart, lungs, etc.) have specific structures and functions. Utilizing this prior knowledge can help the model better understand and segment medical images.

[0145] Graph Neural Networks (GNNs): A deep learning model for processing graph-structured data that can capture the relationships between nodes and the topological structure of the graph.

[0146] Global anatomical region graph: Represent the anatomical regions in the entire medical image as nodes in the graph, and use the structural similarity between regions as the weights of the edges.

[0147] Patch: Divide the medical image into small local regions (patches), and each patch can be regarded as a node in the graph.

[0148] Standard GCN (Graph Convolutional Network): A commonly used graph neural network for performing convolutional operations on graph-structured data.

[0149] This module aims at structure perception and proposes a two-layer graph structure modeling method that combines anatomical region priors with a graph neural propagation mechanism. By establishing a global anatomical region graph and a local patch-level graph and introducing a cross-layer fusion mechanism, it realizes the modeling of the spatial structure correlation and context consistency between tumors and lymph nodes in medical images, effectively improving the model's structural understanding and segmentation accuracy.

[0150] Specifically, first, the image is divided into multiple anatomical regions according to the anatomical structure of the medical image (such as the nasopharynx in CT images). This can be achieved through a predefined anatomical atlas or unsupervised clustering.

[0151] Next, the global semantic average representation within each anatomical region is calculated based on the fused feature map and used as the feature input for the nodes of the anatomical region graph (for example, using the central points of anatomical regions such as tumors, left lymph nodes, and right lymph nodes as nodes, and the node feature being the average value of the pixels within the region). The structural similarity (such as spatial distance, morphological similarity, etc.) between adjacent anatomical regions is used as the connection strength of the edges of the anatomical region graph, and the global anatomical region graph is constructed. The formula is:

[0152] ;

[0153] ;

[0154] where represents the graph node feature vector of anatomical region r; R represents the reconstructed data; represents all the pixel sets of anatomical region r; represents the feature representation at pixel position in the fused feature map;

[0155] represents anatomical region and the edge connection strength between; , respectively represent the geometric center coordinates of anatomical regions , ; represents the square of the Euclidean distance; is the scale parameter of the Gaussian distance kernel, which can be set to a fixed value or be learnable. For example, it can be set to 3.2 (obtained by grid search and optimized on the validation set).

[0156] This edge weight function represents the structural similarity or "propagation intimacy" between adjacent anatomical regions, with strong connections for close neighbors and weak connections for long distances, which is in line with the real anatomical structure.

[0157] Then, the medical image is divided into several patches, and each patch serves as a node in the graph. A local patch adjacency graph is constructed based on the spatial adjacency or feature similarity between patches, and the features of each patch are updated by extracting through a standard GCN;

[0158] The node embedding information of the global anatomical region graph is fed back to the local patch features for cross-layer fusion propagation to achieve the fusion of global and local information. The expression is:

[0159] ;

[0160] where, is the feature of the i-th patch; is the updated feature of the i-th patch, which incorporates the high-level structural information of the anatomical region graph;

[0161] is a learnable cross-graph projection weight matrix for mapping the structural semantics from the region domain to the patch domain;

[0162] represents the belonging probability that the i-th patch belongs to region r, which is calculated from the distance between the patch's position and the region center. The formula is:

[0163] ;

[0164] where, is the center coordinate of the i-th patch; is the center coordinate of region r; is the center coordinate of region ; represents the Euclidean distance.

[0165] The fused features are used for downstream tasks (such as segmentation, classification, etc.), and the parameters of the global graph and the local graph are jointly optimized through backpropagation.

[0166] As Figure 2 shown, it presents a schematic diagram of the overall network structure of the present invention, which includes multiple key modules and their connection relationships. The input data includes T1-modal images, T1c-modal images, and T2-modal images. These modalities respectively correspond to different biological information: the T1 modality is used to identify the anatomical structure of the nasopharynx, the T1c modality highlights the tumor enhancement signal, and the T2 modality detects the tissue edema distribution.

[0167] These images are preprocessed and then input into the multi-modal projection and feature alignment module to perform feature alignment operations. In practical applications, first, an independent convolutional projector maps the features of each modality into a shared hidden feature space. For each modality, a 1x1 convolutional operation is used to map its features into a hidden space with the same number of channels. This operation ensures that features of different modalities are fused in the same dimension, avoiding the problem of dimensional mismatch between modalities. The convolutional projector is designed following the lightweight principle, and the feature mapping is completed only through 1x1 convolutional kernels, thereby reducing the computational complexity and improving the model running efficiency. Specifically, in the Pytorch framework, this module is implemented by defining a sub-network containing three independent 1x1 convolutional layers, and each convolutional layer processes the input features of T1, T1c, and T2 modalities respectively.

[0168] After completing multi-modal feature alignment, it enters the task-driven modulation module. The goal of this module is to dynamically adjust the contributions of each modality according to different task requirements. For the T staging task, global average pooling is used to generate modulation weights to highlight spatial details such as tumor boundaries; for the N staging task, max pooling is used to generate modulation weights to enhance the perception of important structures such as lymph node positions. The generation process of modulation weights extracts feature distribution information from a global perspective through pooling operations, thereby realizing the dynamic adjustment of modality contributions. For example, when the model needs to focus on segmenting tumor boundaries, the system will strengthen the spatial information of the T2 modality according to the requirements of the T staging task, while suppressing the expression of irrelevant features in other modalities.

[0169] Next, the modality-sensitive neuron dynamic inhibition module further filters out task-irrelevant or ambiguous modality information. This module is designed based on the selective inhibition mechanism in neuroscience, including three stages: modality difference perception, task modulation weighting, and inhibition weight generation. In the specific implementation, in the modality difference perception stage, the differences between modalities are calculated by comparing the activation values of different modality features pixel by pixel; in the task modulation weighting stage, difference scores are generated in combination with task requirements. For example, in the N staging task, more attention is paid to the enhanced signal of the T1c modality; in the inhibition weight generation stage, the difference scores are mapped to a value between 0 and 1 through the Sigmoid function to ensure that their value ranges are reasonable and can be precisely controlled. In the code implementation, this module is completed through a series of tensor operations and non-linear activation functions.

[0170] Subsequently, the neuron cluster feature extraction module extracts the features related to T staging and N staging respectively, and enhances the feature expression of each task. For the T staging task, depthwise separable convolution is used to strengthen the extraction of spatial information, focusing on the tumor boundary; for the N staging task, 1x1 convolution is used to capture global contrast information and enhance the perception of lymph nodes. The design of depthwise separable convolution separates the processing of spatial and channel information through depthwise convolution and pointwise convolution, thereby reducing the computational overhead and improving the feature extraction efficiency. In practical applications, this module is implemented by defining two parallel convolutional branches, where one branch uses depthwise separable convolution to extract T staging-related features, and the other branch uses 1x1 convolution to extract N staging-related features. The results of the two branches are merged through a concatenation operation to form a feature representation containing rich task-specific information.

[0171] Next, the spatial-channel decoupled fusion module extracts spatial features and channel features respectively and fuses them through a reconstruction operation. Spatial features mainly reflect the local structural information in the image, such as the tumor boundary; channel features reflect the contrast information between modalities, such as the enhanced signal of T1c. In specific implementation, first, spatial features and channel features are extracted through convolution operations respectively, and then a reconstruction operation is used to fuse the two to obtain the final fused features. The reconstruction operation combines spatial features and channel features through weighted summation, thereby realizing the comprehensive modeling of multimodal information. In code implementation, this module extracts spatial and channel features by defining two independent convolutional layers respectively, and realizes feature fusion through weighted summation operation. The weighting coefficients are automatically learned during the training process to ensure that the fusion result can maximize the segmentation performance.

[0172] On this basis, the structure partition-guided double-layer graph optimization module combines the anatomical region prior and the graph neural propagation mechanism to optimize the structural consistency of the segmentation. This module includes two parts: a global anatomical region graph and a local patch graph. The global anatomical region graph takes the center points of anatomical regions such as tumors, left lymph nodes, and right lymph nodes as nodes, and the node features are the pixel means within the regions; the edge weights between regions are defined in the form of Gaussian kernels. The local patch graph divides the image into several patches, constructs an adjacency graph, and updates the patch representation through a standard graph convolutional network. Cross-layer fusion propagation feeds back the region graph embedding information to the local patch features, thereby realizing the modeling of the spatial structure correlation and context consistency between tumors and lymph nodes in medical images. In practical applications, this module is implemented by defining a double-layer graph neural network, where the global anatomical region graph calculates the edge weights using Gaussian kernels, and the local patch graph updates the node features through a standard GCN. Cross-layer fusion propagation feeds back the region graph embedding information to the local patch features through matrix multiplication, thereby improving the structural consistency of the segmentation.

[0173] Finally, the segmentation head module outputs the segmentation prediction results for T staging and N staging. This module maps the fused features to the segmentation spaces of T staging and N staging, and outputs the segmentation results of each task through the Sigmoid activation function. In the specific implementation, the segmentation head module adopts a lightweight fully convolutional network structure, and defines two parallel convolutional branches to output the T staging segmentation result and the N staging segmentation result respectively. The Sigmoid activation function ensures that the output results are between 0 and 1, facilitating subsequent binarization processing. The final segmentation result can be converted into a binary mask by setting a threshold for the probability map, generating detailed tumor and lymph node segmentation results.

[0174] Through the collaborative work of the above modules, the present invention realizes a comprehensive optimization of image segmentation tasks such as nasopharyngeal carcinoma. In actual application scenarios, this method can be applied to clinical diagnosis and treatment plan formulation. For example, in radiotherapy planning, doctors can quickly locate the positions of tumors and lymph nodes through the three-dimensional reconstruction model generated by this method, so as to formulate accurate radiotherapy plans. In addition, due to the innovative designs in feature alignment, task modulation, modality suppression, feature extraction, feature fusion, and structural optimization of this method, its segmentation performance shows obvious advantages in terms of accuracy, robustness, and generalization ability, providing reliable technical support for clinical practice.

[0175] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0176] The present invention significantly improves the segmentation accuracy of image segmentation tasks such as tumors (T staging) and lymph nodes (N staging) by combining innovative technologies such as multi-modal feature projection, task-adaptive modulation, neuron cluster processing, and spatial-channel decoupled fusion. This method aligns the features of T1, T1c, and T2 modalities through independent projection, and uses a task-driven modulation mechanism to dynamically adjust the modality contributions, enhancing task-related features in an adaptive manner. The neuron cluster module effectively extracts task-specific features, and the spatial-channel decoupled fusion strategy further optimizes the feature fusion process, improving the segmentation accuracy and task performance.

[0177] Embodiment 2

[0178] As Figure 3 shown, the second embodiment of the present invention also provides a nasopharyngeal carcinoma image segmentation device based on multi-modal task-specific fusion, including:

[0179] An acquisition unit for acquiring medical images of different modalities;

[0180] A multi-modal feature alignment unit for mapping medical images of different modalities to a shared feature space for feature alignment;

[0181] A task-driven modulation unit, configured to dynamically adjust the contributions of the modality features after feature alignment in different tasks according to task requirements, and obtain a task-driven adjusted feature map;

[0182] A dynamic suppression unit, configured to dynamically filter the task-driven adjusted feature map by using a selective suppression mechanism to obtain a suppressed modality feature map;

[0183] A neuron cluster unit, configured to simulate a biological neuron cluster mechanism, and respectively extract and enhance the features related to task requirements in the suppressed modality feature map through different convolutional structures to obtain neuron cluster features;

[0184] A decoupling and reconstruction unit, configured to respectively extract the spatial features and channel features of the neuron cluster features and perform fusion and reconstruction to obtain fusion features;

[0185] A segmentation prediction unit, configured to input the fusion features into a segmentation space and output a segmentation result by combining an activation function.

[0186] Embodiment III

[0187] The third embodiment of the present invention further provides a nasopharyngeal carcinoma image segmentation device based on multi-modal task-specific fusion, which includes a memory and a processor. A computer program is stored in the memory, and the computer program can be executed by the processor to implement the image segmentation method based on multi-modal task-specific fusion as described above.

[0188] Embodiment IV

[0189] The fourth embodiment of the present invention further provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, the image segmentation method based on multi-modal task-specific fusion as described above is implemented.

[0190] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the part of the module, program segment, or code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0191] In addition, the functional modules in each embodiment of the present invention can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.

[0192] If the above functions are implemented in the form of software functional modules and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes. It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements that are not explicitly listed, or also includes elements inherent to such a process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.

[0193] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "said", and "the" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.

[0194] It should be understood that the term " / and" used herein is merely a description of the associated relationship of associated objects, indicating that three relationships may exist. For example, A / and B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0195] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0196] The "first / second" mentioned in the embodiments is only to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when permitted. It should be understood that the objects distinguished by the "first / second" can be interchanged appropriately so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0197] The foregoing is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image segmentation method based on multi-modal task-specific fusion, characterized in that Including: S1. Obtain medical images of different modalities; S2. Map medical images of different modalities to a shared feature space respectively for feature alignment; S3. Dynamically adjust the contributions of the features of each modality after feature alignment in different tasks according to task requirements to obtain a task-driven adjusted feature map; S4. Use a selective suppression mechanism to dynamically filter the task-driven adjusted feature map to obtain a suppressed modality feature map; specifically: First, calculate the modality difference degree between the current modality and other modalities, and the formula is: ; Among them, represents the modal difference degree of the -th modality at the pixel position ; represents the feature map of the -th modality after task modulation, that is, the task-driven adjusted feature map of the -th modality; represents the L1 norm; represents other modalities different from the modality ; Then, introduce a task-related modulation weight to weight the modality difference degree to generate a difference score to highlight the modality changes concerned by the current task, and the formula is: ; Among them, represents the difference score of the -th modality at the pixel position under the current task H; represents the task-sensitive weight of the -th modality under the current task H; Finally, map the difference score through a Sigmoid activation function to obtain a suppression weight, so as to obtain the final suppressed modality feature map, and the formula is: ; ; Among them, represents an inhibition coefficient; represents the Sigmoid activation function, which is used to map the score into a differentiable weight; is the final inhibition mode feature map; S5. Simulate the biological neuron cluster mechanism, and respectively extract and enhance the features related to task requirements in the suppressed modality feature map through different convolutional structures to obtain neuron cluster features; S6. Respectively extract the spatial features and channel features of the neuron cluster features and perform fusion reconstruction to obtain fusion features; S7. Input the fusion features into a segmentation space and combine with an activation function to output a segmentation result.

2. The image segmentation method based on multi-modal task-specific fusion according to claim 1, wherein , The task requirements include a T staging task and an N staging task; among them, the T staging task is to highlight the spatial details of the anatomical region by strengthening the spatial information of the features; the N staging task is to enhance the contrast information between modality features to highlight the contrast in the features.

3. The image segmentation method based on multi-modal task-specific fusion according to claim 2, wherein , When dynamically adjusting the contributions of the features of each modality after feature alignment in different tasks according to task requirements: Globally compress the features spatially through average pooling to generate a modulation weight for the T staging task; Highlight the contrast in the features through max pooling to generate a modulation weight for the N staging task.

4. The image segmentation method based on multi-modal task-specific fusion according to claim 2, wherein , When simulating the biological neuron cluster mechanism to extract features of different task requirements, use a depthwise separable convolutional structure to extract the features of the T staging task to strengthen the spatial information; use a 1x1 convolutional structure to capture the global contrast information of the N staging task to enhance the perception of lymph nodes, and the formulas are respectively: ; ; Among them, , respectively represent the neuron cluster features of the T staging task and the N staging task; , respectively represent the neuron cluster processing modules of the T staging task and the N staging task, and extract features related to tumor boundaries and lymph node distributions respectively; , respectively represent the inhibitory mode features of the T staging task and the N staging task.

5. The image segmentation method based on multi-modal task-specific fusion according to claim 2, wherein , The specific content of S6 is: Respectively extract the spatial features and channel features of the neuron cluster features through convolutional operations, and the formula is: ; ; Among them, and respectively represent the extracted spatial features and channel features; and respectively represent the decoupling operations of spatial features and channel features; represents the neuron cluster features for the T staging task; Fuse the spatial features and channel features through a reconstruction operation to obtain the final fusion features, and the formula is: ; Among them, represents the fusion feature; represents the reconstruction operation.

6. The image segmentation method based on multi-modal task-specific fusion according to claim 2, wherein , The specific content of S7 is: Map the fusion features to the segmentation spaces of the T staging task and the N staging task respectively, and output the segmentation results of each task through a Sigmoid activation function, and the expression is: ; ; Among them, and are the segmentation predictions of the T-staging task and the N-staging task respectively; is the fused feature; is the Sigmoid activation function; and are the weights of the T-staging task and the N-staging task respectively; and are the biases of the T-staging task and the N-staging task respectively.

7. A method for image segmentation based on multi-modal task-specific fusion according to claim 2, characterized in that , It also includes using anatomical region priors and graph neural network propagation mechanisms to optimize the fusion features, specifically: First, divide the image into multiple anatomical regions according to the anatomical structure of the medical image; Then, calculate the global semantic average representation within each anatomical region according to the fusion feature map as the feature input of the anatomical region graph nodes, and use the structural similarity between adjacent anatomical regions as the connection strength of the anatomical region graph edges to construct a global anatomical region graph, and the formula is: ; ; Among them, represents the graph node feature vector of the anatomical region r; R represents the reconstructed data; represents the set of all pixels in the anatomical region r; represents the pixel position in the fused feature map feature representation at; Represents an anatomical region and the edge connection strength between; , respectively represent the geometric center coordinates of the anatomical regions , ; Represents the square of the Euclidean distance; is the scale parameter of the Gaussian distance kernel; Then, the medical image is divided into several patches, a local patch adjacency graph is constructed according to the spatial adjacency between the patches, and the features of each patch are extracted and updated through a standard GCN; The node embedding information of the global anatomical region graph is fed back to the local patch features for cross-layer fusion propagation, and the expression is: ; Among them, is the feature of the i-th patch; is the feature of the updated i-th patch, which integrates the high-level structural information of the anatomical region map; is a learnable cross-image projection weight matrix for mapping structural semantics from the region domain to the patch domain; Indicates the belonging probability that the i-th patch belongs to region r, which is calculated from the distance between the position of the patch and the center of the region. The formula is: ; Among them, is the center coordinate of the i-th patch; is the center coordinate of the region r; is the center coordinate of the region ; represents the Euclidean distance.

8. An image segmentation device for nasopharyngeal carcinoma based on multi-modal task-specific fusion, characterized in that, Including: An acquisition unit for acquiring medical images of different modalities; A multi-modal feature alignment unit for mapping medical images of different modalities to a shared feature space for feature alignment; A task-driven modulation unit for dynamically adjusting the contributions of the modality features after feature alignment in different tasks according to task requirements to obtain a task-driven adjusted feature map; A dynamic suppression unit for dynamically filtering the task-driven adjusted feature map by using a selective suppression mechanism to obtain a suppressed modality feature map; specifically: First, calculate the modality difference degree between the current modality and other modalities, and the formula is: ; Among them, represents the modal difference degree of the -th modality at the pixel position ; represents the feature map of the -th modality after task modulation, that is, the task-driven adjusted feature map of the -th modality; represents the L1 norm; represents other modalities different from the modality . Then, introduce a task-related modulation weight to weight the modality difference degree to generate a difference score to highlight the modality changes concerned by the current task, and the formula is: ; Among them, represents the difference score of the -th modality at the pixel position ; represents the task-sensitive weight of the -th modality under the current task H; Finally, map the difference score through a Sigmoid activation function to obtain a suppression weight, so as to obtain the final suppressed modality feature map, and the formula is: ; ; Among them, represents the inhibition coefficient; represents the Sigmoid activation function, which is used to map the score to a differentiable weight; is the final inhibition modal feature map; A neuron cluster unit for simulating the biological neuron cluster mechanism, and respectively extracting and enhancing the features related to the task requirements in the suppressed modality feature map through different convolutional structures to obtain neuron cluster features; A decoupling and reconstruction unit for respectively extracting the spatial features and channel features of the neuron cluster features and performing fusion reconstruction to obtain fusion features; A segmentation prediction unit for inputting the fusion features into a segmentation space and outputting a segmentation result in combination with an activation function.

9. A nasopharyngeal carcinoma image segmentation device based on multi-modal task-specific fusion, characterized in that, Including a processor and a memory, and a computer program is stored in the memory, and the computer program can be executed by the processor to implement an image segmentation method based on multi-modal task-specific fusion as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method and system for processing medical image data

    CN119624978A

  • Efficient segmentation of tumours from lung ct

    US20250061682A1