Image segmentation method, device and equipment based on multi-modal task specific fusion

By adopting multimodal task-specific fusion methods in medical image segmentation technology, including feature alignment, dynamic modulation, selective inhibition and neuron cluster processing, the shortcomings of multimodal feature fusion and task adaptive modulation in the existing technology are solved, and the segmentation accuracy and robustness are significantly improved, providing more reliable clinical support.

CN120125829AActive Publication Date: 2025-06-10XIAMEN UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510607867.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-10
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The existing medical image segmentation technology has shortcomings in multimodal feature fusion, task adaptive modulation, and the utilization of anatomical prior knowledge, resulting in insufficient segmentation accuracy and robustness, which is difficult to meet clinical needs.

Method used

A method of image segmentation based on multimodal task-specific fusion is proposed. By acquiring medical images of different modalities, feature alignment and dynamic modulation are performed, task-related features are extracted using selective inhibition mechanism and neuron cluster mechanism, and segmentation results are optimized priori through space-channel decoupling and fusion and anatomical region.

Benefits of technology

It significantly improves the accuracy, robustness of medical image segmentation such as tumors and lymph nodes, and its adaptability to multimodal medical images, providing more reliable technical support for clinical diagnosis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125829A_ABST
    Figure CN120125829A_ABST
Patent Text Reader

Abstract

The invention provides an image segmentation method, device and equipment based on multi-modal task specific fusion, and relates to the technical field of image segmentation. According to the invention, the obtained medical images of different modalities are respectively mapped to the shared feature space for feature alignment; according to task requirements, contributions of modal features after feature alignment in different tasks are dynamically adjusted, and a task-driven adjustment feature map is obtained; carrying out dynamic filtering by adopting a selective inhibition mechanism to obtain an inhibition mode characteristic pattern; simulating a biological neuron cluster mechanism, respectively extracting features related to task requirements in the suppression modal feature map through different convolution structures, enhancing the features to obtain neuron cluster features, respectively extracting spatial features and channel features of the neuron cluster features, and performing fusion reconstruction to obtain fusion features; and finally, inputting the fusion features into a segmentation space and outputting a segmentation result in combination with an activation function. According to the method, the accuracy and robustness of medical image segmentation and the adaptive capacity to the multi-modal medical image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of neural networks and image segmentation, and in particular, to an image segmentation method, device, equipment and medium based on multi-modal task-specific fusion. Background Art

[0002] In the field of medical image analysis, the accurate segmentation of laryngeal cancer and lymph nodes is of great significance. The medical image segmentation task requires accurately extracting the target area from complex anatomical structures, which is crucial for disease diagnosis, staging, and treatment planning. Traditional methods mainly rely on manual feature extraction and rule-based algorithms, but these methods often show problems of insufficient accuracy and poor robustness when dealing with multi-modal data and complex anatomical structures, and are difficult to meet clinical needs. With the development of deep learning technology, the automatic segmentation of nasopharyngeal carcinoma has gradually turned to methods based on single-modal or multi-modal medical images. Single-modal segmentation methods usually rely on medical images of T1, T1c, or T2 modalities and perform segmentation through a convolutional neural network (CNN). Among them, the T1 modality is mainly used for the identification of anatomical structures, the T1c modality focuses on the detection of tumor enhancement signals, and the T2 modality is used for the identification of edema and fluid. However, single-modal methods are limited by the limitations of their information and are difficult to comprehensively capture the complex features of tumors and lymph nodes.

[0003] In recent years, more and more research has begun to attempt multi-modal information fusion in order to improve the segmentation accuracy by using the complementarity between different modalities. Existing multi-modal methods usually fuse the features of T1, T1c, and T2 modalities by means of splicing, weighting, or joint learning, and perform joint training based on deep learning models (such as U-Net and its variants). Although these methods have improved the segmentation effect to a certain extent, there are still significant deficiencies. First, multi-modal feature fusion methods usually rely on simple splicing or weighted average operations, and fail to fully optimize and effectively fuse the features of different modalities, resulting in the complementarity between modalities not being fully exploited. Second, existing methods generally ignore the extraction of task-specific features and fail to adaptively adjust the modality contributions according to the different requirements of T staging (tumor segmentation) and N staging (lymph node segmentation), making the model unstable when dealing with complex anatomical structures. In addition, existing technologies often fail to make full use of the anatomical prior knowledge in medical images, and it is difficult to achieve an ideal segmentation accuracy when the boundary between tumors and normal tissues is blurred. Finally, existing technologies have weak generalization ability when dealing with changes in different datasets, devices, and patient groups, which limits their wide application in actual clinical environments.

[0004] In view of this, the applicant specifically proposes this application. Summary of the Invention

[0005] The present invention aims to provide an image segmentation method, device, equipment and medium based on multi-modal task-specific fusion, so as to solve the problems existing in the existing medical image segmentation technology in aspects such as multi-modal feature fusion, task adaptive modulation, and utilization of anatomical prior knowledge.

[0006] To solve the above technical problems, the present invention is implemented through the following technical solutions: An image segmentation method based on multi-modal task-specific fusion, comprising: S1, obtaining medical images of different modalities; S2, mapping the medical images of different modalities to a shared feature space respectively for feature alignment; S3, dynamically adjusting the contributions of the feature-aligned features of each modality in different tasks according to task requirements to obtain a task-driven adjusted feature map; S4, using a selective suppression mechanism to dynamically filter the task-driven adjusted feature map to obtain a suppressed modality feature map; S5, simulating the biological neuron cluster mechanism, and respectively extracting and enhancing the features related to task requirements in the suppressed modality feature map through different convolutional structures to obtain neuron cluster features; S6, respectively extracting the spatial features and channel features of the neuron cluster features and performing fusion and reconstruction to obtain fusion features; S7, inputting the fusion features into a segmentation space and combining with an activation function to output a segmentation result.

[0007] Preferably, the task requirements include a T staging task and an N staging task; wherein, the T staging task is to highlight the spatial details of anatomical regions by strengthening the spatial information of features; the N staging task is to enhance the contrast information between modality features to highlight the contrast in features.

[0008] Preferably, when dynamically adjusting the contributions of the feature-aligned features of each modality in different tasks according to task requirements: Performing global compression of the features spatially through average pooling to generate a modulation weight for the T staging task; Highlighting the contrast in features through max pooling to generate a modulation weight for the N staging task.

[0009] Preferably, the selective suppression mechanism is used to dynamically filter different modality information according to different tasks, and S4 is specifically: First, calculating the modality difference degree between the current modality and other modalities, and the formula is: ; Wherein, represents the th modality at the pixel position Modal difference degree at Indicates the feature map of the -th modality after task modulation, that is, the task-driven adjusted feature map of the -th modality; Indicates the L1 norm; Indicates other modalities different from modality ; Then, introduce task-related modulation weights to weight the modal difference degree to generate a difference score to highlight the modal changes concerned by the current task. The formula is: ; Wherein, Indicates the difference score of the -th modality at pixel position under the current task H; Indicates the task-sensitive weight of the -th modality under the current task H; Finally, map the difference score through the Sigmoid activation function to obtain an inhibition weight, thereby obtaining the final inhibited modal feature map. The formula is: ; ; Wherein, Indicates the inhibition coefficient; Indicates the Sigmoid activation function, which is used to map the score into a differentiable weight; is the final inhibited modal feature map.

[0010] Preferably, when simulating the mechanism of biological neuron clusters to extract features of different task requirements, use a depthwise separable convolution structure to extract the features of the T staging task to strengthen spatial information; use a 1x1 convolution structure to capture the global contrast information of the N staging task to enhance the perception of lymph nodes. The formulas are respectively: ; ; Wherein, , respectively indicate the neuron cluster features of the T staging task and the N staging task; , respectively indicate the neuron cluster processing modules of the T staging task and the N staging task, and respectively extract the features related to the tumor boundary and lymph node distribution; , respectively indicate the inhibited modal features of the T staging task and the N staging task.

[0011] Preferably, S6 is specifically as follows: Extract the spatial feature and channel feature of the neuron cluster feature respectively through convolution operation, and the formula is: ; ; Wherein, , respectively represent the extracted spatial feature and channel feature; , respectively represent the decoupling operations of the spatial feature and channel feature; represents the neuron cluster feature of the T staging task; Fuse the spatial feature and channel feature through a reconstruction operation to obtain the final fused feature, and the formula is: ; Wherein, represents the fused feature; represents the reconstruction operation.

[0012] Preferably, S7 is specifically as follows: Map the fused feature to the segmentation spaces of the T staging task and the N staging task respectively, and output the segmentation results of each task through the Sigmoid activation function. The expression is: ; ; Wherein, , are the segmentation predictions of the T staging task and the N staging task respectively; is the fused feature; is the Sigmoid activation function; , are the weights of the T staging task and the N staging task respectively; , are the biases of the T staging task and the N staging task respectively.

[0013] Preferably, it further includes optimizing the fused feature by using the anatomical region prior and the graph neural network propagation mechanism, specifically as follows: First, divide the medical image into multiple anatomical regions according to the anatomical structure of the medical image; Next, calculate the global semantic average representation within each anatomical region according to the fused feature map, use it as the feature input of the anatomical region graph node, and use the structural similarity between adjacent anatomical regions as the connection strength of the anatomical region graph edge to construct a global anatomical region graph. The formula is: ; ; Among them, represents the graph node feature vector of the anatomical region r; R represents the reconstructed data; represents all pixel sets of the anatomical region r; represents the feature representation at the pixel position in the fused feature map; represents the anatomical region and the edge connection strength between; , respectively represent the geometric center coordinates of the anatomical regions , ; represents the squared Euclidean distance; is the scale parameter of the Gaussian distance kernel; Then, the medical image is divided into several patch blocks, a local patch adjacency graph is constructed according to the spatial adjacency between patches, and the features of each patch are extracted and updated through a standard GCN; The node embedding information of the global anatomical region graph is fed back to the local patch features for cross-layer fusion propagation, and the expression is: ; Among them, is the feature of the i-th patch; is the updated feature of the i-th patch, which fuses the high-level structure information of the anatomical region graph; is a learnable cross-graph projection weight matrix for mapping the structural semantics from the region domain to the patch domain; represents the belonging probability that the i-th patch block belongs to the region r, which is calculated from the distance between the position of the patch block and the region center, and the formula is: ; Among them, is the center coordinate of the i-th patch; is the center coordinate of the region r; is the center coordinate of the region ; represents the Euclidean distance.

[0014] The present invention also provides a nasopharyngeal carcinoma image segmentation device based on multi-modal task-specific fusion, including: An acquisition unit for acquiring medical images of different modalities; A multi-modal feature alignment unit for mapping medical images of different modalities to a shared feature space for feature alignment; A task-driven modulation unit, configured to dynamically adjust the contributions of each modality feature after feature alignment in different tasks according to task requirements, so as to obtain a task-driven adjusted feature map; A dynamic suppression unit, configured to dynamically filter the task-driven adjusted feature map by using a selective suppression mechanism to obtain a suppressed modality feature map; A neuron cluster unit, configured to simulate a biological neuron cluster mechanism, and respectively extract and enhance the features related to task requirements in the suppressed modality feature map through different convolutional structures to obtain neuron cluster features; A decoupling and reconstruction unit, configured to respectively extract the spatial features and channel features of the neuron cluster features and perform fusion and reconstruction to obtain fused features; A segmentation prediction unit, configured to input the fused features into a segmentation space and combine an activation function to output a segmentation result.

[0015] The present invention also provides a nasopharyngeal carcinoma image segmentation device based on multi-modal task-specific fusion, including a processor and a memory. A computer program is stored in the memory, and the computer program can be executed by the processor to implement an image segmentation method based on multi-modal task-specific fusion as described above.

[0016] The present invention also provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, an image segmentation method based on multi-modal task-specific fusion as described above is implemented.

[0017] In summary, compared with the prior art, the present invention has the following beneficial effects: The present invention aims to combine innovative technologies such as multi-modal feature projection, task-driven modulation, neuron cluster processing, and spatial-channel decoupling fusion to improve the accuracy, robustness, and adaptability to multi-modal medical images of medical image segmentation of tumors and lymph nodes, so as to provide more reliable technical support for clinical applications.

[0018] Through multi-modal feature alignment and fusion technologies, the present invention solves the problem that multi-modal features in the prior art cannot be fully optimized and effectively fused. Among them, independent projection mapping is used to perform feature alignment on multi-modal medical images (such as T1, T1c, and T2 modalities), and a task-driven modulation mechanism is used to dynamically adjust the contributions of each modality, giving full play to the complementarity between modalities.

[0019] Furthermore, through the task adaptive modulation technology, the present invention dynamically enhances or suppresses task-related features according to different task requirements, improving the distribution modeling ability for medical images such as tumor boundaries and lymph nodes. In particular, the modal-sensitive neuron dynamic suppression technology effectively suppresses task-irrelevant or ambiguous modal information through a selective suppression mechanism, highlighting the significant expression of the dominant modality.

[0020] In addition, the present invention uses the spatial-channel decoupled fusion technology to extract spatial features and channel features respectively, and optimizes the feature fusion process through a reconstruction operation, further improving the segmentation accuracy.

[0021] Finally, the present invention adopts the structure partition-guided double-layer graph optimization technology, combines the anatomical region prior with the graph neural propagation mechanism, realizes the modeling of the spatial structure correlation and context consistency between tumors and lymph nodes in medical images, and significantly improves the structural consistency of segmentation.

[0022] The present invention can effectively fuse multi-modal image information, adaptively adjust the contribution of image modalities, and at the same time make full use of anatomical prior knowledge and the graph neural network propagation mechanism to optimize the fused features, significantly improving the performance of medical image segmentation such as nasopharyngeal carcinoma, showing obvious advantages in terms of segmentation accuracy, robustness, and generalization ability, and providing strong support for clinical diagnosis and treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 FIG. 18 is a schematic flowchart of an image segmentation method based on multi-modal task-specific fusion provided for Embodiment 1.

[0025] Figure 2 FIG. 22 is a schematic overall framework diagram of an image segmentation method based on multi-modal task-specific fusion provided for Embodiment 1.

[0026] Figure 3 FIG. 26 is a schematic diagram of a nasopharyngeal carcinoma image segmentation device based on multi-modal task-specific fusion provided for Embodiment 2.

[0027] The following further details the present invention in conjunction with the drawings and specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0029] Example 1 Example 1 of the present invention provides an image segmentation method based on multi-modal task-specific fusion, which can be implemented by an image segmentation device based on multi-modal task-specific fusion (hereinafter referred to as the segmentation device), and particularly, is executed by one or more processors in the segmentation device.

[0030] In this embodiment, the segmentation device may be an electronic device equipped with a processor, and the processor has a computer program for the image segmentation method based on multi-modal task-specific fusion and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited herein.

[0031] In medical image analysis, especially in neuroimaging (such as brain MRI) and tumor diagnosis, T1-weighted (T1), T1 Contrast-enhanced (T1c), and T2-weighted (T2) are three of the most commonly used MRI modalities. Each modality highlights different physical characteristics of tissues through different imaging parameters, providing complementary information for clinical diagnosis and machine learning tasks (such as segmentation and classification).

[0032] In this embodiment, T1 refers to T1-weighted imaging, which mainly reflects the differences in the longitudinal relaxation time of tissues. T1C is enhanced T1-weighted imaging, which is performed after intravenous injection of a gadolinium-based contrast agent (such as gadopentetate dimeglumine) on the basis of T1-weighted imaging. T2 refers to T2-weighted imaging, which highlights the differences in the transverse relaxation time (T2) through long TR and long TE parameters, reflecting the degree of freedom of water molecules in tissues (such as extracellular edema). T1 provides an anatomical reference, T1c enhances specificity, and T2 sensitively captures pathological changes. The combination of the three can significantly improve the diagnostic accuracy.

[0033] Such as Figure 1As shown, an image segmentation method based on multi-modal task-specific fusion, which includes steps S1 to S7.

[0034] S1. Obtain medical images of different modalities.

[0035] Obtain multi-modal medical images of the same anatomical region (such as T1, T1c, T2 MRI, or CT, PET, etc.).

[0036] S2. Map the medical images of different modalities to a shared feature space respectively for feature alignment.

[0037] This step aims to map the input multi-modal images (such as T1, T1c, T2) to a shared feature space, solve the problem of modality heterogeneity, make the features of different modalities comparable, and facilitate subsequent fusion operations.

[0038] Since the features of different modalities vary in dimension and distribution, it is necessary to align the features of each modality to a unified space through a projection operation. In this embodiment, for each modality, a 1x1 convolution operation is used to map it to a hidden space with the same number of channels (such as 256 dimensions). This operation ensures that the features of different modalities are fused in the same dimension and avoids dimension mismatch between modalities.

[0039] Specifically, taking the images of three modalities T1, T1c, and T2 as an example, the expression of the projection operation is: ; Among them, is the projection operation; is the modality feature after mapping to the hidden space; is the input modality feature; t represents the image modality.

[0040] S3. Dynamically adjust the contributions of the modality features after feature alignment in different tasks according to the task requirements to obtain a task-driven adjusted feature map.

[0041] After completing the alignment of multi-modal features, the next step is to enhance the contribution of each modality in different tasks through a task-driven modulation mechanism.

[0042] Specifically, in this embodiment, the task requirements include the T staging task and the N staging task; among them, the T staging task is to highlight the spatial details of the anatomical region by strengthening the spatial information of the features (such as the edema boundary of the T2 modality); the N staging task is to enhance the contrast information between modality features to highlight the contrast in the features (such as the enhanced signal of the T1c modality).

[0043] In this embodiment, three types of modal features are modulated respectively to dynamically enhance or suppress task-related features.

[0044] T-stage task modulation: The features are globally compressed spatially through average pooling to generate the modulation weights for the T-stage task. This operation helps to highlight spatial details such as tumor boundaries.

[0045] N-stage task modulation: The contrast in the features is highlighted through max pooling to generate the modulation weights for the N-stage task, enhancing the perception of important structures such as the location of lymph nodes.

[0046] The formula for the modulated features is as follows: ; where and represent the features after being modulated by the T-stage and N-stage task-driven respectively; and and and represent the features of the T, T1, T1c, and T2 modalities respectively; and represent the task-driven modulation operations for the T-stage and N-stage tasks respectively, aiming to enhance task-related features.

[0047] S4. A selective inhibition mechanism is used to dynamically filter the task-driven adjusted feature map to obtain the suppressed modal feature map.

[0048] To further highlight the dominant modalities and suppress redundant modalities, this module is designed based on the "selective inhibition mechanism" in neuroscience, and a task-sensitive dynamic inhibition mechanism is proposed. The selective inhibition mechanism is used to dynamically filter different modal information according to different tasks, including modal difference perception, task modulation weighting, and suppression weight generation.

[0049] Modal difference perception: Calculate the modal difference degree between the current modality and other modalities. The formula is: ; where represents the modal difference degree of the th modality at the pixel position ; represents the feature map of the th modality after task modulation, that is, the task-driven adjusted feature map of the th modality; represents the L1 norm, that is, the sum of the absolute differences between modalities at this position; represents other modalities different from modality ; , indicating the difference between this modality and all other modalities.

[0050] This step is used to measure whether a modality has information uniqueness in a local area. The greater the difference, the more likely it is that the modality features conflict or are redundant with other modalities.

[0051] Task modulation weighting: Introduce task-related modulation weights to weight the modality difference degree and generate a difference score to highlight the modality changes concerned by the current task. The formula is: ; Where, represents the difference score of the th modality at the pixel position under the current task H; represents the task-sensitive weight of the th modality under the current task H.

[0052] This step combines the task requirements and modality characteristics to re-weight the modality difference degree, which is used to highlight the modality changes that the current task pays more attention to.

[0053] Inhibitory weight generation: Map the difference score through the Sigmoid activation function to obtain the inhibitory weight, so as to obtain the final inhibitory modality feature map. The formula is: ; ; Where, represents the inhibition coefficient, and the range is between 0 and 1; represents the Sigmoid activation function, which is used to map the score into a differentiable weight; is the final inhibitory modality feature map for subsequent fusion; is the task-driven adjustment feature map of the th modality.

[0054] When the score is larger, it indicates that there may be redundant information or ambiguous expressions in this modality. The smaller the inhibitory coefficient output by Sigmoid, the more the contribution of this modality in this area is dynamically suppressed. On the contrary, when this modality performs well in this area, tends to 1, so as to retain its strong expression.

[0055] S5, simulating the mechanism of biological neuron clusters, extracts and enhances the features related to the task requirements in the inhibitory modality feature map through different convolutional structures to obtain neuron cluster features.

[0056] This step adopts the mechanism of simulating biological neuron clusters, and extracts features related to T staging (such as tumor segmentation) and N staging (such as lymph node segmentation) through different convolutional structures respectively. T staging mainly focuses on the boundary of the tumor, so depthwise separable convolution is used to strengthen the extraction of spatial information; N staging uses 1x1 convolution to capture global contrast information to enhance the perception of lymph nodes. The specific formulas are as follows: ; ; Among them, and respectively represent the neuron cluster features of the T staging task and the N staging task; and respectively represent the neuron cluster processing modules of the T staging task and the N staging task, and extract features related to the tumor boundary and lymph node distribution respectively; and respectively represent the inhibitory mode features of the T staging task and the N staging task.

[0057] S6 extracts the spatial features and channel features of the neuron cluster features respectively and performs fusion and reconstruction to obtain the fusion features.

[0058] To further improve the segmentation accuracy, the method of the present invention proposes a spatial-channel decoupled fusion strategy. In this step, spatial features and channel features are first extracted through convolutional operations respectively. The spatial features mainly reflect the local structural information in the image (such as the tumor boundary), while the channel features reflect the contrast information between modalities (such as the enhanced signal of T1c). Then, the spatial features and channel features are fused through a reconstruction operation to obtain the final fusion features.

[0059] Specifically, the spatial features and channel features of the neuron cluster features are extracted through convolutional operations respectively. The formulas are as follows: ; ; Among them, and respectively represent the extracted spatial features and channel features; and respectively represent the decoupling operations of the spatial features and channel features; represents the neuron cluster features of the T staging task; The spatial features and channel features are fused through a reconstruction operation to obtain the final fusion features. The formula is as follows: ; Among them, represents the fusion features; Indicates a reconstruction operation.

[0060] S7, input the fused feature into the segmentation space and output the segmentation result by combining with the activation function.

[0061] After feature fusion is completed, the final segmentation output is performed through the segmentation head module. This module maps the fused feature to the segmentation space of T staging and N staging, and outputs the segmentation result of each task through the Sigmoid activation function.

[0062] Map the fused feature to the segmentation space of the T staging task and the N staging task respectively, and output the segmentation result of each task through the Sigmoid activation function. The expression is: ; ; where , are the segmentation predictions of the T staging task and the N staging task respectively; is the fused feature; is the Sigmoid activation function; , are the weights of the T staging task and the N staging task respectively; , are the biases of the T staging task and the N staging task respectively.

[0063] In another preferred embodiment, before the segmentation prediction step of the method of the present invention, the fused feature can also be optimized in terms of segmentation structure by using anatomical region priors and graph neural network propagation mechanism.

[0064] Anatomical Region Priors: In medical image analysis, different anatomical regions (such as the heart, lungs, etc.) have specific structures and functions. Utilizing this prior knowledge can help the model better understand and segment medical images.

[0065] Graph Neural Networks (GNNs): A deep learning model for processing graph-structured data, which can capture the relationships between nodes and the topological structure of the graph.

[0066] Global anatomical region graph: Represent the anatomical regions in the entire medical image as nodes in the graph, and use the structural similarity between regions as the weight of the edge.

[0067] Patch: Divide the medical image into small local regions (patches), and each patch can be regarded as a node in the graph.

[0068] Standard GCN (Graph Convolutional Network): A commonly used graph neural network for performing convolutional operations on graph-structured data.

[0069] This module aims at structure perception and proposes a two-layer graph structure modeling method that combines anatomical region priors with graph neural propagation mechanisms. By establishing a global anatomical region graph and a local patch-level graph and introducing a cross-layer fusion mechanism, it realizes the modeling of the spatial structure correlation and context consistency between tumors and lymph nodes in medical images, effectively improving the model's structural understanding and segmentation accuracy.

[0070] Specifically, first, the image is divided into multiple anatomical regions according to the anatomical structure of the medical image (such as the nasopharynx in CT images). This can be achieved through a predefined anatomical atlas or unsupervised clustering.

[0071] Next, the global semantic average representation within each anatomical region is calculated based on the fused feature map and used as the feature input for the nodes of the anatomical region graph (for example, using the central points of anatomical regions such as tumors, left lymph nodes, and right lymph nodes as nodes, and the pixel mean within the region as the node feature). The structural similarity (such as spatial distance, morphological similarity, etc.) between adjacent anatomical regions is used as the connection strength of the edges of the anatomical region graph, and the global anatomical region graph is constructed. The formula is: ; ; Among them, represents the graph node feature vector of anatomical region r; R represents the reconstructed data; represents the set of all pixels in anatomical region r; represents the feature representation at pixel position in the fused feature map; represents anatomical region and ; , respectively represent the geometric center coordinates of anatomical regions , ; represents the squared Euclidean distance; is the scale parameter of the Gaussian distance kernel, which can be set to a fixed value or be learnable. For example, it can be set to 3.2 (obtained by grid search and optimized on the validation set).

[0072] This edge weight function represents the structural similarity or "propagation intimacy" between adjacent anatomical regions. The connection is strong for nearby neighbors and weak for long distances, which is in line with the real anatomical structure.

[0073] Then, the medical image is divided into several patches, and each patch serves as a node in the graph. A local patch adjacency graph is constructed based on the spatial adjacency or feature similarity between patches, and the features of each patch are updated by extracting through the standard GCN; The node embedding information of the global anatomical region graph is fed back to the local patch features for cross-layer fusion propagation to achieve the fusion of global and local information. The expression is: ; where, is the feature of the i-th patch; is the updated feature of the i-th patch, which fuses the high-level structural information of the anatomical region graph; is a learnable cross-graph projection weight matrix for mapping the structural semantics from the regional domain to the patch domain; represents the membership probability that the i-th patch belongs to region r, which is calculated from the distance between the position of the patch and the center of the region. The formula is: ; where, is the center coordinate of the i-th patch; is the center coordinate of region r; is the center coordinate of region ; represents the Euclidean distance.

[0074] The fused features are used for downstream tasks (such as segmentation, classification, etc.), and the parameters of the global graph and the local graph are jointly optimized through backpropagation.

[0075] As Figure 2 shown, it shows the overall network structure diagram of the present invention, which includes multiple key modules and their connection relationships. The input data includes T1-modal images, T1c-modal images, and T2-modal images. These modalities correspond to different biological information: the T1 modality is used to identify the anatomical structure of the nasopharynx, the T1c modality highlights the tumor enhancement signal, and the T2 modality detects the tissue edema distribution.

[0076] These images are preprocessed and then input into the multi-modal projection and feature alignment module to perform feature alignment operations. In practical applications, first, the features of each modality are mapped to a shared hidden feature space through an independent convolutional projector. For each modality, a 1x1 convolutional operation is used to map its features to a hidden space with the same number of channels. This operation ensures that features of different modalities are fused in the same dimension, avoiding the problem of dimensional mismatch between modalities. The convolutional projector is designed following the lightweight principle, and the feature mapping is completed only through 1x1 convolutional kernels, thereby reducing the computational complexity and improving the model running efficiency. Specifically, in the Pytorch framework, this module is implemented by defining a sub-network containing three independent 1x1 convolutional layers, and each convolutional layer processes the input features of T1, T1c, and T2 modalities respectively.

[0077] After completing the multi-modal feature alignment, it enters the task-driven modulation module. The goal of this module is to dynamically adjust the contributions of each modality according to different task requirements. For the T staging task, modulation weights are generated through global average pooling to highlight spatial details such as tumor boundaries; for the N staging task, modulation weights are generated through max pooling to enhance the perception of important structures such as lymph node positions. The generation process of modulation weights extracts feature distribution information from a global perspective through pooling operations, thereby realizing the dynamic adjustment of the contributions of modalities. For example, when the model needs to focus on segmenting tumor boundaries, the system will strengthen the spatial information of the T2 modality according to the requirements of the T staging task, while suppressing the expression of irrelevant features in other modalities.

[0078] Next, the modality-sensitive neuron dynamic inhibition module further filters task-irrelevant or ambiguous modality information. This module is designed based on the selective inhibition mechanism in neuroscience and includes three stages: modality difference perception, task modulation weighting, and inhibition weight generation. In the specific implementation, in the modality difference perception stage, the differences between modalities are calculated by comparing the activation values of different modality features pixel by pixel; in the task modulation weighting stage, difference scores are generated in combination with task requirements. For example, in the N staging task, more attention is paid to the enhanced signal of the T1c modality; in the inhibition weight generation stage, the difference scores are mapped to a value between 0 and 1 through the Sigmoid function to ensure that their value ranges are reasonable and can be precisely controlled. In the code implementation, this module is completed through a series of tensor operations and non-linear activation functions.

[0079] Subsequently, the neuron cluster feature extraction module extracts features related to T staging and N staging respectively, and enhances the feature representation of each task. For the T staging task, depthwise separable convolution is used to strengthen the extraction of spatial information, with a focus on the tumor boundary; for the N staging task, 1x1 convolution is used to capture global contrast information and enhance the perception of lymph nodes. The design of depthwise separable convolution separates the processing of spatial and channel information through depthwise convolution and pointwise convolution, thereby reducing the computational overhead and improving the feature extraction efficiency. In practical applications, this module is implemented by defining two parallel convolutional branches. One branch uses depthwise separable convolution to extract features related to T staging, and the other branch uses 1x1 convolution to extract features related to N staging. The results of the two branches are merged through a concatenation operation to form a feature representation containing rich task-specific information.

[0080] Next, the spatial-channel decoupled fusion module extracts spatial features and channel features respectively and fuses them through a reconstruction operation. Spatial features mainly reflect the local structural information in the image, such as the tumor boundary; channel features reflect the contrast information between modalities, such as the enhanced signal of T1c. In specific implementation, first, spatial features and channel features are extracted through convolution operations respectively, and then a reconstruction operation is used to fuse the two to obtain the final fused features. The reconstruction operation combines spatial features and channel features through weighted summation, thereby achieving comprehensive modeling of multi-modal information. In code implementation, this module extracts spatial and channel features by defining two independent convolutional layers respectively, and realizes feature fusion through weighted summation operation. The weighting coefficients are automatically learned during the training process to ensure that the fusion result can maximize the segmentation performance.

[0081] On this basis, the structure partition-guided double-layer graph optimization module combines the anatomical region prior and the graph neural propagation mechanism to optimize the structural consistency of the segmentation. This module includes two parts: a global anatomical region graph and a local patch graph. The global anatomical region graph uses the central points of anatomical regions such as tumors, left lymph nodes, and right lymph nodes as nodes, and the node features are the pixel means within the regions; the edge weights between regions are defined in the form of Gaussian kernels. The local patch graph divides the image into several patches, constructs an adjacency graph, and updates the patch representation through a standard graph convolutional network. Cross-layer fusion propagation feeds back the region graph embedding information to the local patch features, thereby realizing the modeling of the spatial structure correlation and context consistency between tumors and lymph nodes in medical images. In practical applications, this module is implemented by defining a double-layer graph neural network. The global anatomical region graph calculates the edge weights using Gaussian kernels, and the local patch graph updates the node features through a standard GCN. Cross-layer fusion propagation feeds back the region graph embedding information to the local patch features through matrix multiplication, thereby improving the structural consistency of the segmentation.

[0082] Finally, the segmentation head module outputs the segmentation prediction results for T staging and N staging. This module maps the fused features to the segmentation spaces of T staging and N staging, and outputs the segmentation results for each task through the Sigmoid activation function. In the specific implementation, the segmentation head module adopts a lightweight fully convolutional network structure, and defines two parallel convolutional branches to output the T staging segmentation result and the N staging segmentation result respectively. The Sigmoid activation function ensures that the output results are between 0 and 1, facilitating subsequent binarization processing. The final segmentation result can convert the probability map into a binary mask by setting a threshold, generating detailed tumor and lymph node segmentation results.

[0083] Through the collaborative work of the above modules, the present invention realizes the comprehensive optimization of image segmentation tasks such as nasopharyngeal carcinoma. In actual application scenarios, this method can be applied to clinical diagnosis and treatment plan formulation. For example, in radiotherapy planning, doctors can quickly locate the positions of tumors and lymph nodes through the three-dimensional reconstruction model generated by this method, so as to formulate a precise radiotherapy plan. In addition, due to the innovative designs of this method in feature alignment, task modulation, modality suppression, feature extraction, feature fusion, and structural optimization, its segmentation performance shows obvious advantages in terms of accuracy, robustness, and generalization ability, providing reliable technical support for clinical practice.

[0084] In summary, compared with the prior art, the present invention has the following beneficial effects: The present invention significantly improves the segmentation accuracy of image segmentation tasks such as tumors (T staging) and lymph nodes (N staging) by combining innovative technologies such as multi-modal feature projection, task-adaptive modulation, neuron cluster processing, and spatial-channel decoupled fusion. This method aligns the features of T1, T1c, and T2 modalities through independent projection, and uses a task-driven modulation mechanism to dynamically adjust the modality contributions, enhancing the task-related features in an adaptive manner. The neuron cluster module effectively extracts task-specific features, and the spatial-channel decoupled fusion strategy further optimizes the feature fusion process, improving the segmentation accuracy and task performance.

[0085] Embodiment 2 As Figure 3 shown, the second embodiment of the present invention also provides a nasopharyngeal carcinoma image segmentation device based on multi-modal task-specific fusion, including: An acquisition unit for acquiring medical images of different modalities; A multi-modal feature alignment unit for mapping medical images of different modalities to a shared feature space for feature alignment; A task-driven modulation unit for dynamically adjusting the contributions of the modality features after feature alignment in different tasks according to task requirements to obtain a task-driven adjusted feature map; A dynamic suppression unit, which is used to dynamically filter the task-driven adjusted feature map by adopting a selective suppression mechanism to obtain a suppressed modal feature map; A neuron cluster unit, which is used to simulate the biological neuron cluster mechanism, extract and enhance the features related to task requirements in the suppressed modal feature map respectively through different convolutional structures to obtain neuron cluster features; A decoupling and reconstruction unit, which is used to extract the spatial features and channel features of the neuron cluster features respectively and perform fusion and reconstruction to obtain fusion features; A segmentation prediction unit, which is used to input the fusion features into a segmentation space and combine an activation function to output a segmentation result.

[0086] Embodiment III The third embodiment of the present invention further provides a nasopharyngeal carcinoma image segmentation device based on multi-modal task-specific fusion, which includes a memory and a processor. A computer program is stored in the memory, and the computer program can be executed by the processor to implement the image segmentation method based on multi-modal task-specific fusion as described above.

[0087] Embodiment IV The fourth embodiment of the present invention further provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, the image segmentation method based on multi-modal task-specific fusion as described above is implemented.

[0088] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are only illustrative. For example, the flowcharts in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0089] In addition, in each embodiment of the present invention, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0090] If the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs. It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements that are not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.

[0091] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0092] It should be understood that the term " / and" used herein is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0093] Depending on the context, as used herein, the word "if" can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0094] The "first / second" mentioned in the embodiments is only used to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when permitted. It should be understood that the objects distinguished by the "first / second" can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0095] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image segmentation method based on multimodal task-specific fusion, characterized in that: include: S1, obtain medical images of different modalities; S2, mapping medical images of different modalities to the shared feature space for feature alignment; S3, dynamically adjust the contribution of each modal feature after feature alignment in different tasks according to task requirements to obtain the task-driven adjusted feature map; S4, dynamically filtering the task-driven adjustment feature map using a selective inhibition mechanism to obtain an inhibition modal feature map; S5, simulating the biological neuron cluster mechanism, extracting and enhancing the features related to the task requirements in the inhibitory modal feature map through different convolution structures, and obtaining neuron cluster features; S6, respectively extracting the spatial features and channel features of the neuron cluster features and fusing and reconstructing them to obtain fused features; S7, input the fusion feature into the segmentation space and combine it with the activation function to output the segmentation result.

2. The image segmentation method based on multimodal task-specific fusion according to claim 1 is characterized in that The task requirements include T staging tasks and N staging tasks; wherein, the T staging task is to highlight the spatial details of the anatomical area by strengthening the spatial information of the features; the N staging task is to enhance the contrast information between modal features to highlight the contrast in the features.

3. The image segmentation method based on multimodal task-specific fusion according to claim 2 is characterized in that ,When dynamically adjusting the contribution of each modal feature after feature alignment in different tasks according to task requirements: The features are globally compressed in space through average pooling to generate the modulation weights of the T-stage task; The modulation weights for the N-stage task are generated by max-pooling the contrast in the salient features.

4. The image segmentation method based on multimodal task-specific fusion according to claim 1, characterized in that ,The selective suppression mechanism is used to dynamically filter different modal information according to different tasks, and S4 is specifically: First, calculate the modal difference between the current mode and other modes. The formula is: ; in, Indicates The mode is at pixel position The modal difference at ; Indicates The feature map of the modality after task modulation, i.e. Task-driven adjustment feature maps of each modality; represents the L1 norm; Indicates that it is different from modal Other modes of Then, the task-related modulation weight is introduced to weight the modal difference and generate a difference score to highlight the modal changes that the current task focuses on. The formula is: ; in, Indicates that under the current task H The mode is at pixel position Difference scores at Indicates the current task H. Task-sensitive weights of each modality; Finally, the difference score is mapped through the Sigmoid activation function to obtain the inhibition weight, thereby obtaining the final inhibition modal feature map, the formula is: ; ; in, represents the suppression coefficient; represents the Sigmoid activation function, which is used to map the score to a differentiable weight; is the final suppressed mode feature map.

5. The image segmentation method based on multimodal task-specific fusion according to claim 2 is characterized in that ,When simulating the biological neuron cluster mechanism to extract features of different task requirements, the deep separable convolution structure is used to extract the features of the T staging task to enhance the spatial information; the 1x1 convolution structure is used to capture the global contrast information of the N staging task to enhance the perception of the lymph nodes. The formulas are: ; ; in, , Respectively represent the neuron cluster characteristics of T-stage task and N-stage task; , The neuron cluster processing modules represent the T staging task and the N staging task, respectively, extracting features related to the tumor boundary and lymph node distribution; , Represent the inhibition modal characteristics of T-stage task and N-stage task respectively.

6. The image segmentation method based on multimodal task-specific fusion according to claim 2 is characterized in that , the S6 is specifically: The spatial features and channel features of the neuron cluster features are extracted respectively through convolution operations, and the formula is: ; ; in, , Represent the extracted spatial features and channel features respectively; , Respectively represent the decoupling operations of spatial features and channel features; Represents the characteristics of neuronal clusters in the T-staging task; The spatial features and channel features are fused through reconstruction operations to obtain the final fusion features. The formula is: ; in, Indicates fusion features; Represents a refactoring operation.

7. The image segmentation method based on multimodal task-specific fusion according to claim 2 is characterized in that ,S7 specifically includes: mapping the fusion features to the segmentation space of T-stage tasks and N-stage tasks respectively, and outputting the segmentation result of each task through the Sigmoid activation function, the expression is: ; ; in, , They are the segmentation predictions for T-stage tasks and N-stage tasks respectively; For fusion features; is the Sigmoid activation function; , are the weights of T-stage tasks and N-stage tasks respectively; , are the biases of T-stage tasks and N-stage tasks respectively.

8. The image segmentation method based on multimodal task-specific fusion according to claim 2 is characterized in that ,It also includes,optimizing the fusion features by using anatomical region prior and graph neural network propagation mechanism, specifically: First, the image is divided into multiple anatomical regions according to the anatomical structure of the medical image; Next, the global semantic average representation in each anatomical region is calculated based on the fusion feature map as the feature input of the anatomical region map node, and the structural similarity between adjacent anatomical regions is used as the connection strength of the anatomical region map edge to construct the global anatomical region map. The formula is: ; ; in, represents the feature vector of the graph node of the anatomical region r; R represents the reconstructed data; represents the set of all pixels of the anatomical region r; Represents the pixel position in the fused feature map Feature representation at; Indicates anatomical region and The edge connection strength between them; , Represents anatomical regions , The geometric center coordinates of; represents the square of the Euclidean distance; is the scale parameter of the Gaussian distance kernel; Then, the medical image is divided into several patches, a local patch adjacency graph is constructed based on the spatial adjacency between patches, and the features of each patch are extracted and updated through the standard GCN. The node embedding information of the global anatomical region map is fed back to the local patch feature for cross-layer fusion propagation, and the expression is: ; in, is the feature of the i-th patch; The updated features of the ith patch incorporate the high-level structural information of the anatomical region map; is a learnable cross-graph projection weight matrix used to map structural semantics from the region domain to the patch domain; It represents the probability that the ith patch belongs to region r, which is calculated by the distance between the position of the patch and the center of the region. The formula is: ; in, is the center coordinate of the i-th patch; is the center coordinate of region r; For Region The center coordinates of represents the Euclidean distance.

9. A nasopharyngeal carcinoma image segmentation device based on multimodal task-specific fusion, characterized in that: include: An acquisition unit, used for acquiring medical images of different modalities; A multimodal feature alignment unit, used to map medical images of different modalities to a shared feature space for feature alignment; The task-driven modulation unit is used to dynamically adjust the contribution of each modal feature after feature alignment in different tasks according to task requirements to obtain a task-driven adjustment feature map; A dynamic suppression unit, used for dynamically filtering the task-driven adjustment feature map by using a selective suppression mechanism to obtain a suppression modal feature map; A neuron cluster unit is used to simulate the biological neuron cluster mechanism, and extracts and enhances the features related to the task requirements in the inhibitory modal feature map through different convolution structures to obtain neuron cluster features; A decoupling and reconstruction unit, used to extract the spatial features and channel features of the neuron cluster features respectively and perform fusion reconstruction to obtain fusion features; The segmentation prediction unit is used to input the fusion feature into the segmentation space and combine it with the activation function to output the segmentation result.

10. A nasopharyngeal carcinoma image segmentation device based on multimodal task-specific fusion, characterized in that: It includes a processor and a memory, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement an image segmentation method based on multimodal task-specific fusion as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Multimodal MR image segmentation method and device based on 3D-Ghost network

    CN115496769A

  • Method and system for processing medical image data

    CN119624978A

  • Efficient segmentation of tumours from lung ct

    US20250061682A1

  • Medical image segmentation method and system, terminal, and storage medium

    WO2023087300A1