Method and system for identifying lymph node metastasis in mediastinal region based on double-control routing hybrid expert model

The method for identifying mediastinal lymph node metastasis based on a dual-control routing hybrid expert model is proposed. The task is split into detection and prediction sub-tasks. By using a dual-control routing gating mechanism and gradient balancing algorithm, the problems of feature interference and gradient conflict in the identification of mediastinal lymph node metastasis are solved, and a high-accuracy identification effect is achieved.

CN121435014APending Publication Date: 2026-01-30NORTHEAST FORESTRY UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511600728.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Existing technologies suffer from low accuracy and poor identification results in the task of identifying lymph node metastasis in the mediastinal region. Furthermore, the hybrid expert model neglects the allocation and calculation of weights for similar features of related tasks in its gating mechanism design, leading to conflicts between tasks and gradient imbalance.

Method used

A hybrid expert model based on dual-control routing is adopted. The task of identifying mediastinal lymph node metastasis is divided into two auxiliary subtasks: mediastinal detection and lymph node metastasis prediction. The dual-control routing gating mechanism and the two-dimensional gradient balancing algorithm are used to extract task similarity and unique features respectively, and feature fusion and gradient optimization are performed.

Benefits of technology

It effectively reduces interference from multi-dimensional features, improves task performance, realizes gradient balance and feature sharing among tasks in a multi-task learning system, and enhances the accuracy and stability of mediastinal lymph node metastasis identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121435014A_ABST
    Figure CN121435014A_ABST
Patent Text Reader

Abstract

The invention discloses a mediastinal region lymph node metastasis identification method and system based on a double-control routing hybrid expert model, and belongs to the technical field of computer intelligent auxiliary diagnosis. The invention aims to solve the problems that the existing mediastinal region lymph node metastasis identification method is low in accuracy and poor in identification effect due to the fact that the mediastinal region lymph node metastasis identification task is complicated and includes multi-dimensional information identification. According to the technical key points, a mediastinal region lymph node metastasis identification task is split into two auxiliary sub-tasks of mediastinal region detection and lymph node metastasis prediction, and the three tasks jointly form a multi-task learning system. Knowledge transmission and complementation are realized through feature sharing between tasks and gradient-based collaborative optimization, so that the feature distinguishability of each task is improved. The method is characterized in that a double-control routing gating mechanism is constructed, similar features and unique features between different tasks are extracted through a feature routing branch and a task routing branch respectively, and correlation and difference between the tasks are balanced; meanwhile, a two-dimensional gradient balance algorithm is designed, multi-task gradient optimization is carried out from two angles of gradient direction alignment and amplitude dynamic balance, and the problems of gradient conflicts and task dominance are solved while gradient feature distribution between tasks is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer-aided diagnosis technology, specifically relating to a method and system for identifying mediastinal lymph node metastasis based on a dual-control routing hybrid expert model. Background Technology

[0002] Lung cancer is a tumor with a high incidence and mortality rate. The staging diagnosis of lung adenocarcinoma is one of the most critical clinical factors affecting the treatment outcome and prognosis of patients with non-small cell lung cancer (NSCLC). According to the international lung cancer TNM staging system (T represents tumor size, N represents lymph node metastasis, and M represents distant metastasis) [1], whether the mediastinal lymph nodes have metastasized is a key indicator for determining the stage of lung adenocarcinoma patients. Manual diagnostic methods mainly rely on assessing the size of lymph nodes on CT images to determine whether metastasis has occurred, but lymph node enlargement is not necessarily caused by malignant tumors, and non-enlarged lymph nodes may also be carriers of metastasis [2]. This indicates that the identification of lymph node metastasis involves many unknown features, and there is no unified standard for diagnosis based on visual assessment by doctors [3]. Therefore, it is crucial to introduce intelligent auxiliary diagnostic methods. Intelligent auxiliary diagnostic methods must extract mediastinal information and lymph node status from a single image [4] to determine whether lymph node metastasis exists in different mediastinal regions. This indicates that the task of identifying mediastinal lymph node metastasis is complex and involves multi-dimensional information identification. Typical single-task image classification methods are usually designed for natural image recognition [5,6,7,8]. They face challenges in medical imaging, such as the presence of a large amount of noise and significant structural similarities between organs, and difficulties in extracting fine-grained features across multiple dimensions from the same image, leading to information interference across different dimensions. Therefore, this invention decomposes and decouples the mediastinal lymph node metastasis prediction task, taking it as the main task and constructing mediastinal region detection and lymph node metastasis prediction as auxiliary sub-tasks, forming a multi-task learning system to solve complex multidimensional classification problems.

[0003] In the identification of mediastinal lymph node metastasis and its auxiliary subtasks, significant similarities were observed between features of different tasks, but the importance of these similar features varied considerably across tasks. Therefore, distinguishing between similar and unique features across different tasks is crucial. However, some hybrid expert models neglect the allocation and calculation of weights for similar features across related tasks in their gating mechanisms, resulting in a lack of fine-grained control over expert selection and feature sharing. Furthermore, some multi-task learning methods [9, 10, 11] rely too heavily on shared feature extraction while neglecting task characteristic modeling. Some hybrid expert (MoE) models [12, 13, 14] ignore the allocation and weights of similar features across related tasks in their gating mechanisms, leading to a lack of fine-grained control over expert selection and feature sharing.

[0004] Basic experiments revealed significant performance gaps between different tasks. This reflects substantial differences in features and data distribution across tasks. Furthermore, the mediastinal lymph node metastasis identification task is more complex, resulting in significantly lower detection performance compared to the other two tasks. This leads to larger loss and gradient values ​​in multi-task learning systems, potentially causing task dominance. Therefore, mitigating potential conflicts and task dominance while maintaining differences in task gradient distribution is a crucial problem. Some multi-task gradient balancing methods ensure loss function convergence by modifying gradients [12, 13, 14] or balancing tasks [15, 16]. However, these methods struggle to achieve multi-angle and fine-grained gradient adjustments, making them unsuitable for adapting to the gradient variation characteristics of different tasks.

[0005] In recent years, there has been an increasing amount of research on the detection of mediastinal lymph nodes and the identification of lymph node metastasis, for example,

[0006] Existing patent document CN118229645A discloses a method, system, device, and medium for detecting mediastinal lymph nodes in the lungs. Its purpose is to address the technical problems of lymph node center offset errors following a Gaussian distribution, the dynamic variation in detection difficulty due to lymph node center movement, and poor detection results. The constructed lymph node segmentation model includes a backbone network, a prediction module, and a correction module. The output of the backbone network serves as the input to the prediction module. The output of the prediction module is then weighted with the outputs of each downsampling module in the backbone network and used as the input to the correction module. The prediction module outputs the bounding box information and a first probability of the target nodule, while the correction module outputs a second probability of the target nodule. The loss is calculated by combining the bounding box information, the average of the first and second probabilities, and the corresponding label data, and the parameters of the lymph node segmentation model are updated accordingly.

[0007] For example, existing patent document CN117893462A discloses a method and system for identifying cervical lymph node metastasis in thyroid cancer based on SwinTransformer. This method includes preprocessing a color Doppler ultrasound image to determine the proportion of red channel values ​​and blue channel values ​​of each pixel in the overall RGB value; segmenting the preprocessed ultrasound image into multiple patches and numbering each patch; adding the patch number to a set if a patch contains pixels with a ratio greater than a threshold; using the segmented patches as input to the SwinTransformer; when the SW-MSA layer is executed, obtaining multiple sliding distances based on the window size, calculating the difference in the number of patches within the same window number in the set relative to when there is no sliding; determining the sliding distance used by the SW-MSA layer based on the difference in sliding distances; and using the output of the last SwinTransformerBlock as input to the MLP, with the MLP output serving as the identification result. This invention not only identifies lymph node metastasis but also has high accuracy.

[0008] It can be seen that, in the existing technology, no one has proposed a technical solution to identify mediastinal lymph node metastasis using a dual-control routing hybrid expert model. Summary of the Invention

[0009] The technical problem to be solved by this invention is:

[0010] To address the issues of poor accuracy and recognition performance of existing methods for identifying mediastinal lymph node metastasis due to the complexity of the task and the inclusion of multi-dimensional information, this invention proposes a method and system for identifying mediastinal lymph node metastasis based on a dual-control routing hybrid expert model.

[0011] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:

[0012] A method for identifying mediastinal lymph node metastasis based on a dual-control routing hybrid expert model, the method comprising:

[0013] Improvement of the Hybrid Expert Model (MMoE): Add a dual-control routing gating mechanism consistent with the number of multi-tasks to the input of the Hybrid Expert Model (MMoE) to obtain a dual-control routing hybrid expert model;

[0014] First, the data of the mediastinal lymph node metastasis identification task and its different dimension identification auxiliary subtasks (the mediastinal lymph node metastasis identification task and its different dimension identification auxiliary subtasks represent multiple tasks) are input into the expert network to perform expert selection and activation, and a two-dimensional gradient balancing algorithm is designed to realize the gradient update of multiple tasks.

[0015] Simultaneously, a dual-control routing gating mechanism is employed for feature calculation and allocation for each task. This mechanism consists of a feature routing branch and a task routing branch. The feature routing branch extracts similar features between mediastinal images, including color, brightness, and frequency domain features. The task routing branch extracts unique features based on specific label signals from different tasks. The unique features for each task are as follows: the unique features for the mediastinal lymph node metastasis identification task include the morphological features of the mediastinal region and the environmental information of the lungs; the unique features for the mediastinal region detection task are its location and edge features; and the unique features for the lymph node metastasis prediction task are the structural, color, and texture features of the lymph nodes.

[0016] Then, the task similarity and unique features extracted from the feature routing branch and the task routing branch are weighted and fused to achieve a comprehensive description of the features of each task;

[0017] Furthermore, the feature computation of the expert network is integrated with the feature extraction of the dual-control routing gating mechanism to fully balance the feature requirements of shared and independent tasks;

[0018] Finally, the decoding branch of each task completes the recognition process based on the fusion coding results of the expert network and the dual-control routing gating mechanism.

[0019] The technical concept of this invention lies in splitting the task of identifying mediastinal lymph node metastasis into two auxiliary sub-tasks: mediastinal region detection and lymph node metastasis prediction. These three tasks together constitute a multi-task learning system. Through feature sharing and gradient-based collaborative optimization among tasks, knowledge transfer and complementarity are achieved, thereby improving the feature distinguishability of each task. The key to this invention is the construction of a dual-control routing gating mechanism. This mechanism extracts similar and unique features between different tasks through feature routing branches and task routing branches, balancing the correlation and differences between tasks. Simultaneously, a two-dimensional gradient balancing algorithm is designed. While maintaining the gradient feature distribution among tasks, it optimizes multi-task gradients from the perspectives of gradient direction alignment and magnitude balance, resolving the problems of gradient conflict and task dominance.

[0020] The present invention has the following beneficial technical effects:

[0021] This invention decomposes complex multidimensional classification tasks into multiple auxiliary subtasks, utilizing multi-task collaboration to achieve fine-grained representation of complex classification task features, effectively reducing interference between different classification dimensions. This invention extracts similar and unique features between different tasks through feature routing branches and task routing branches, balancing the correlation and differences between tasks. This invention achieves multi-task gradient update and adjustment from two perspectives: gradient direction alignment and amplitude balancing. This invention establishes a multi-task collaborative framework based on a dual-control routing hybrid expert model for the complex problem of mediastinal lymph node metastasis identification. This is a novel design approach, with mediastinal lymph node metastasis identification as the main task and mediastinal detection and lymph node metastasis prediction as auxiliary subtasks, enhancing the performance of each task, especially the main task, through a multi-task feature sharing mechanism. The dual-control routing gating mechanism constructed in this invention optimizes multi-task learning performance in scenarios with a large number of similar features by balancing the weights of task similarity features and task unique features. The dual-dimensional gradient balancing algorithm designed in this invention optimizes multi-task gradients from two perspectives: gradient direction alignment and amplitude balancing, while maintaining the gradient feature distribution between tasks. The method of this invention effectively solves the problems of mutual interference of multi-dimensional features in complex recognition tasks, the difficulty in balancing task relevance and task characteristic weight calculation in multi-task learning systems, and gradient conflict and task dominance among multiple tasks. Attached Figure Description

[0022] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0023] Figure 1 This is an overall flowchart of the mediastinal lymph node metastasis identification method based on a dual-control routing hybrid expert model as described in this invention;

[0024] Figure 2 It is a detailed flowchart of the calculation of each task in the dual-control routing hybrid expert model;

[0025] Figure 3 The confusion matrix for different classification methods on the task of mediastinal lymph node metastasis is shown. The numbers in the matrix represent: "1-3" represent no lymph node metastasis in the upper, middle and lower mediastinal regions, respectively, and "4-6" represent lymph node metastasis in the upper, middle and lower mediastinal regions, respectively. In the figure: (a) is ConvNeXt, (b) is FasterNet, (c) is MobileVit, (d) is SwinTransformer, (e) is VisionTransformer, and (f) is the method of this invention.

[0026] Figure 4The columns represent the core attention areas under different routing mechanisms, with each column representing the core attention area of ​​a different gating mechanism. Detailed Implementation

[0027] Combined with appendix Figure 1-4 The implementation of the mediastinal lymph node metastasis identification method based on a dual-control routing hybrid expert model described in this invention is explained as follows:

[0028] Overall Methodology and Flow Figure 1 As shown: First, the data from the mediastinal lymph node metastasis identification task and its various auxiliary sub-tasks are input into an expert network for expert selection and activation. A two-dimensional gradient balancing algorithm is then used to update the gradients for multiple tasks. Simultaneously, a dual-control routing gating mechanism is employed for feature calculation and allocation for each task. The feature routing branch extracts similar features between mediastinal images, including color, brightness, and frequency domain features. The task routing branch extracts unique features based on specific label signals from different tasks. (The unique features of the mediastinal lymph node metastasis identification task include the morphological features of the mediastinal region and lung environmental information; the unique features of the mediastinal region detection task lie in its location and edge features; and the unique features of the lymph node metastasis prediction task lie in the structure, color, and texture features of the lymph nodes.) Then, the similar and unique features extracted by the feature routing branch and the task routing branch are weighted and fused to achieve a comprehensive description of the features for each task. Finally, the feature calculation of the expert network is fused with the feature extraction of the dual-control routing gating mechanism to fully balance the shared and independent feature requirements between tasks. Finally, the decoding branch of each task completes the recognition process based on the fusion coding results of the expert network and the dual-control routing gating mechanism.

[0029] The following example, using the processing of any task in a multi-tasking system, illustrates the detailed process and steps of the method of this invention. Figure 2 As shown:

[0030] First, the input image data for each task is divided into image patches. Then, a sample feature generator processes these image patches to obtain sample input features. Input the sample into the feature The inputs are respectively fed into the expert network and the dual-control routing gating mechanism.

[0031] (1) Sample input features Processing in expert networks:

[0032] First, apply the regularization function. Input samples as features Transform into regularized features :

[0033]

[0034] The shared parameters of the expert network are represented as follows:

[0035]

[0036] Each Representing the An expert analyzed the sample input features. The encoding result, . Number of representative experts. It is a shared representation matrix generated by all expert networks. Then, based on the regularization features... Construct a binary expert choice matrix for each task. This matrix is ​​composed of A binary single expert selector Composition, defined as:

[0037]

[0038] Each Indicates whether the first [function name] has been activated for the current task. There are [number] experts. There are a total of [number] experts here. indivual Setting it to 1 indicates that the task is active. One expert was selected, while the rest remained inactive. Here... The value is a manually set parameter; the optimal value is determined based on ablation experiments. .

[0039] (2) Sample input features Handling in dual-control routing gating mechanism:

[0040] First, apply the flattening function. Input samples as features Transform into sample flattening features :

[0041]

[0042] The dual-control routing gating mechanism consists of a feature routing branch and a task routing branch, and the feature weights are calculated through the feature routing branch. The process is as follows:

[0043]

[0044] Indicates sample flattening features The selection weights for each expert incorporate similar sample features across multiple tasks, helping the model capture the correlations between different tasks. The weight matrix represents a linearly changing weight matrix. The bias vector representing the linear transformation is a learnable parameter whose value is automatically learned by the gradient descent algorithm based on the characteristics of multi-task data and model properties. The function is applied to the row vector of each sample.

[0045] Furthermore, keyword feature filters are used. Separate the flattening features of each task , , This represents the total number of tasks.

[0046]

[0047] Calculate the weight of each task using task routing branches :

[0048]

[0049] Indicates the first The weight matrix that varies linearly for each task. Indicates the first Bias vectors for linear transformations of each task , This represents the total number of tasks. This represents the selection weights of the expert network for each task, reflecting the model's understanding of task characteristics. It is combined with feature weights. and task weight Generate the routing weights for each task to the expert network. :

[0050]

[0051] in, The hyperparameters set during the experiment were optimized based on the ablation experiments. .

[0052] (3) Finally, by combining the outputs of the expert network and the dual-control routing gating mechanism, the encoder output of each task is obtained. :

[0053]

[0054] in, This represents the dot product operation. This represents the Hadamard operation.

[0055] Furthermore, in the process of multi-task gradient update, this invention employs a two-dimensional gradient balancing algorithm, the formula of which is:

[0056]

[0057] This process involves four variables that need to be calculated. , , and :

[0058] (1) The above The calculation process is as follows:

[0059] First, calculate the gradient for each task based on the loss value:

[0060]

[0061] in, Representative model parameters gradient, Representative task loss function, This represents the number of tasks; therefore, the gradient matrix of the model... It can be represented as:

[0062]

[0063] And thus obtain The iterative calculation formula, with an initial value of 0:

[0064]

[0065] in, It is the attenuation factor. It increases with the number of iterations. and Based on empirical values, they were set to 0.8 and 0.2 respectively; As a batch buffer gradient, it can reflect the latest task information while retaining historical gradient trends, thus achieving a dynamic balance in gradient updates.

[0066] (2) The above The calculation process is as follows:

[0067] Based on the gradient matrix of the model Calculate the gradient autocorrelation matrix in the task space. The gradient similarity between tasks is obtained using the following formula:

[0068]

[0069] matrix It contains statistical relationships between feature maps in the task space, which is crucial for the model to understand the style and content of the task image; matrix The eigenvalues ​​and eigenvectors are calculated as follows:

[0070]

[0071] set up It is a diagonal matrix, and the diagonal elements are matrices. eigenvalues ; It is an orthogonal matrix whose column vectors are matrices. eigenvectors; then based on the matrix Given the rank, construct a new diagonal matrix. :

[0072]

[0073] In the task space, perform gradient decomposition and compute the transformation matrix. :

[0074]

[0075] Transformation matrix Used to adjust the magnitude and direction of the gradient; for the transformation matrix Summing the columns to compute the gradient matrix vector :

[0076]

[0077] in Indicates to No. The sum of all elements in the list; This is used to align gradient directions for different tasks, thereby avoiding gradient conflicts between tasks.

[0078] (3) The above The calculation process is as follows:

[0079] First, calculate the magnitude of each gradient vector to form a norm vector. As shown below:

[0080]

[0081]

[0082] Then, the normalized norm vector of the gradient is calculated. :

[0083]

[0084]

[0085] It is used to dynamically adjust the size of the task gradient to prevent task dominance caused by an excessively large gradient and task omission caused by an excessively small gradient.

[0086] (4) The process for determining the value is as follows:

[0087] As a weight adjustment factor, the two weights are convexly combined. This allows for flexible adjustment of the weight proportions based on task characteristics and data features, achieving complementary adjustment of gradient deficiencies and improving the balance and stability of training across tasks. In the described model, The optimal value was determined to be 0.5 through the ablation experiment.

[0088] Example:

[0089] This invention is applied to the task of identifying lymph node metastasis in the mediastinum. First, the task is broken down into a mediastinal region detection task and a lymph node metastasis prediction task. Then, a hybrid expert model based on dual-control routing is constructed to achieve predictions for the above three tasks.

[0090] The following example, using the processing of any task in a multi-task system, demonstrates the detailed steps of a hybrid expert model based on dual-control routing in the multi-task detection process:

[0091] First, the input image data for the task is divided into image patches. Then, these image patches are processed by a sample feature generator to obtain sample input features. Input the sample into the feature The inputs are respectively fed into the expert network and the dual-control routing gating mechanism.

[0092] (1) Sample input features Processing in expert networks:

[0093] First, apply the regularization function. Input features Transform into regularized features :

[0094]

[0095] In a multi-task system for identifying mediastinal lymph node metastasis, the regularization function is used. The definition of is:

[0096]

[0097] It is a three-segment function, representing that... The difference is smoothed within a range, with constants 0 and 1 at the two ends.

[0098] The shared parameters of the expert network are represented as follows:

[0099]

[0100] Each Representing the An expert analyzed the sample input features. The encoding result, . It is a shared representation matrix generated by all expert networks. Then, based on the regularization features... Construct a binary expert choice matrix for each task. This matrix consists of 4 binary single expert selectors. Composition, defined as:

[0101]

[0102] Each Indicates whether the first [function name] has been activated for the current task. Two experts were involved in the testing of this task. When set to 1, it means that the task has activated 2 experts, while the remaining experts remain inactive.

[0103] (2) Sample input features Processing in dual-control routing gating

[0104] First, apply the flattening function. Input samples as features Transform into sample flattening features :

[0105]

[0106] Here Indicates input features for samples A one-dimensional flattening operation. The dual-control routing gating mechanism consists of a feature routing branch and a task routing branch, with feature weights calculated through the feature routing branch. The process is as follows:

[0107]

[0108] Indicates sample flattening features The selection weights for each expert incorporate similar sample features across multiple tasks, helping the model capture the correlations between different tasks. The weight matrix represents a linearly changing weight matrix. The bias vector representing the linear transformation is a learnable parameter whose value is automatically learned by the gradient descent algorithm based on the characteristics of multi-task data and model properties. The function is applied to the row vector of each sample.

[0109] Furthermore, keyword feature filters are used. Separate the flattening features of each task , Here It is a linear layer set according to each task tag.

[0110]

[0111] Calculate the weight of each task using task routing branches :

[0112]

[0113] Indicates the first The weight matrix that varies linearly for each task. Indicates the first Bias vectors for linear transformations of each task . This represents the selection weights of the expert network for each task, reflecting the model's understanding of task characteristics. It is combined with feature weights. and task weight Generate the routing weights for each task to the expert network. :

[0114]

[0115] in, Based on the ablation experiment results, The model performs best when the target time is reached.

[0116] (3) Finally, by combining the outputs of the expert network and the dual-control routing gating mechanism, the encoder output of each task is obtained. :

[0117]

[0118] in, This represents the dot product operation. This represents the Hadamard operation.

[0119] Furthermore, in the multi-task gradient update process, a two-dimensional gradient balancing algorithm is adopted, the formula of which is:

[0120]

[0121] This process involves four variables that need to be calculated. , , and :

[0122] (1) The above The calculation process is as follows:

[0123] First, calculate the gradient for each task based on the loss value:

[0124]

[0125] in, Representative model parameters gradient, Representative task loss function, This represents the number of tasks; therefore, the gradient matrix of the model... It can be represented as:

[0126]

[0127] And thus obtain The iterative calculation formula, with an initial value of 0:

[0128]

[0129] in, It is the attenuation factor. It increases with the number of iterations. and Based on empirical values, they were set to 0.8 and 0.2 respectively; As a batch buffer gradient, it can reflect the latest task information while retaining historical gradient trends, thus achieving a dynamic balance in gradient updates.

[0130] (2) The above The calculation process is as follows:

[0131] Based on the gradient matrix of the model Calculate the gradient autocorrelation matrix in the task space. The gradient similarity between tasks is obtained using the following formula:

[0132]

[0133] matrix It contains statistical relationships between feature maps in the task space, which is crucial for the model to understand the style and content of the task image; matrix The eigenvalues ​​and eigenvectors are calculated as follows:

[0134]

[0135] set up It is a diagonal matrix, and the diagonal elements are matrices. eigenvalues ; It is an orthogonal matrix whose column vectors are matrices. eigenvectors; then based on the matrix Given the rank, construct a new diagonal matrix. :

[0136]

[0137] In the task space, perform gradient decomposition and compute the transformation matrix. :

[0138]

[0139] Transformation matrix Used to adjust the magnitude and direction of the gradient; for the transformation matrix Summing the columns to compute the gradient matrix vector :

[0140]

[0141] in Indicates to No. The sum of all elements in the list; This is used to align gradient directions for different tasks, thereby avoiding gradient conflicts between tasks.

[0142] (3) The above The calculation process is as follows:

[0143] First, calculate the magnitude of each gradient vector to form a norm vector. As shown below:

[0144]

[0145]

[0146] Then, the normalized norm vector of the gradient is calculated. :

[0147]

[0148]

[0149] It is used to dynamically adjust the size of the task gradient to prevent task dominance caused by an excessively large gradient and task omission caused by an excessively small gradient.

[0150] (4) The process for determining the value is as follows:

[0151] As a weight adjustment factor, the two weights are convexly combined. This allows for flexible adjustment of the weight proportions based on task characteristics and data features, achieving complementary adjustment of gradient deficiencies and improving the balance and stability of training across tasks. In the described model, The optimal value was determined to be 0.5 through the ablation experiment.

[0152] The technical effectiveness and advantages (including key performance indicators) of this invention are verified as follows:

[0153] (1) Advantages of a multi-task collaboration framework:

[0154] To verify the advantages of the hybrid expert model based on dual-control routing in the task of identifying mediastinal lymph node metastasis, this invention was compared with advanced image classification algorithms. The experimental data are shown in Table 1. The low F1 scores of TransNext and FasterNet indicate their difficulty in handling the classification challenges posed by imbalanced medical image datasets. The remaining single-task classification methods also struggle to achieve satisfactory classification performance. In contrast, the ACC of this invention exceeds 80%, and it achieves the best performance across all three evaluation metrics.

[0155] Table 1. Comparison of detection performance between the single-task classification algorithm and the present invention for identifying mediastinal lymph node metastasis.

[0156]

[0157] Furthermore, we calculated the confusion matrix of various methods in the task of identifying mediastinal lymph node metastasis, such as... Figure 3 As shown in the figure, the results reveal significant misclassification issues across different methods, misclassifying category 1 as category 4, category 2 as category 5, and category 3 as category 6. This indicates that different methods struggle to accurately determine the lymph node metastasis status in different mediastinal regions. In contrast, this invention leverages the advantages of multi-task learning to fully acquire the locational features of the mediastinal region and lymph node metastasis characteristics, significantly enhancing the model's ability to understand features across different dimensions.

[0158] (2) Advantages of the dual-control routing gating mechanism:

[0159] To verify the advantages of this invention, we compared it with other multi-task learning frameworks, including the classic multi-task learning methods HPS

[17] , MTAN[9] and the classic hybrid expert models CGC

[14] , MMoE

[12] and DSelect-k

[13] . The experimental results are shown in Table 2. Analysis of the data shows that the HPS shared encoder is not sensitive to task-specific features and has difficulty distinguishing the key features of different tasks, resulting in poor detection performance for each task. In contrast, the detection performance of the MTAN framework is significantly improved, but its detection capability is much lower than that of this invention. The CGC method has too many parameters, resulting in long model training time. Its core application scenario is recommendation system, which is not suitable for lightweight deployment of medical image models. The hybrid expert models MMoE and DSelect-k achieve a balance between computation and detection performance, but the routing gating mechanisms of different methods are quite different, resulting in inconsistent detection performance of different models for the identification of mediastinal lymph node metastasis and related tasks. This invention demonstrates significant performance advantages among all compared multi-task learning methods. It achieves an accuracy of 80.53%, an AUC of 94.96%, and an F1 score of 73.03% in the task of identifying mediastinal lymph node metastasis (Task 1), significantly improving the performance of mediastinal lymph node metastasis identification. Meanwhile, the accuracy of the two auxiliary sub-tasks (Task 2 and Task 3) is improved to 81.38% and 95.86%, respectively, indicating that the overall performance of this invention is good and stable.

[0160] Table 2. Comparative experiments between different multi-task learning methods. Task 1 represents the mediastinal lymph node metastasis identification task; Task 2 represents the mediastinal region detection task; Task 3 represents the lymph node metastasis prediction task.

[0161]

[0162] (3) Advantages of the two-dimensional gradient balancing algorithm

[0163] To verify the advantages of this invention, we compared it with other multi-task gradient balancing algorithms [15,18-23], and the experimental results are shown in Table 3. Analysis of the accuracy of different methods on different tasks reveals that this invention exhibits significant advantages in multiple evaluation metrics. In terms of accuracy, this invention achieves 80.53% accuracy for Task 1 and 81.38% accuracy for Task 2, both higher than other methods. This indicates that this invention can more accurately distinguish complex main task samples and also has a good balance for auxiliary sub-tasks, demonstrating a balanced optimization capability between tasks. In the AUC metric, this invention performs excellently in detection for each task, especially far exceeding other methods in the main task, indicating that the model has stronger discriminative ability and robustness. In the F1 metric, this invention has the most significant advantage in the main task, indicating its superior ability in handling class imbalance and complex decision boundaries. In summary, the overall performance, stability, and generalization ability of this invention in multi-task joint detection are superior to existing mainstream algorithms. This shows that this invention can effectively alleviate gradient conflicts between multiple tasks, stably optimize the shared feature space, and efficiently realize the transfer of information between tasks.

[0164] Table 3. Comparative experiments between different multi-task gradient balancing algorithms. Task 1 represents the mediastinal lymph node metastasis identification task; Task 2 represents the mediastinal region detection task; Task 3 represents the lymph node metastasis prediction task.

[0165]

[0166] Features of this invention (including key performance indicators):

[0167] The dual-control routing gating mechanism consists of a task routing branch and a feature-based routing branch. Its effectiveness is demonstrated through ablation experiments, and the experimental data are shown in Table 4.

[0168] Table 4. Effectiveness experiments of each component of the dual-control routing gating mechanism. Task 1 represents the mediastinal lymph node metastasis identification task; Task 2 represents the mediastinal region detection task; Task 3 represents the lymph node metastasis prediction task.

[0169]

[0170] When using only feature-based routing branches, the detection performance for all tasks is not optimal. This indicates that different tasks require different feature weights, and relying solely on feature-based routing branches cannot capture task-specific features, resulting in poor system learning ability. When using only task-based routing branches, the ACC and F1 scores for Task 1 significantly improve, while the AUC slightly decreases. This suggests that while task-based routing branches can capture task-specific information, they may overlook key similarity features between tasks. In contrast, the dual-control routing gating mechanism considers both task similarity and characteristics, enabling the model to capture task-specific features while leveraging inter-task relationships, thus enhancing its ability to extract similar task features in multi-task learning.

[0171] Figure 4 Attention focus areas under different routing mechanisms. Each column represents the core attention area of ​​different gating mechanisms. Figure 4 This demonstrates the differences in the regions of interest for image features among different gating mechanisms. The feature routing branch primarily focuses on the mediastinal region and its surrounding bright areas, which contain rich overall mediastinal features. The task routing branch, on the other hand, focuses on identifying organ features within the mediastinum. The dual-control routing gating mechanism, however, concentrates on the darker areas of the mediastinum containing lymph nodes, providing a crucial basis for identifying lymph node metastasis in the mediastinum.

[0172] The number of experts was verified through ablation experiments. The optimal values ​​for the parameters are shown in Table 5. When the number of parameters increased from 26.07M to 52.13M, both accuracy and F1 score reached their peak. When the number of experts was further increased to 78.20M and 103.36M, these two key indicators showed a fluctuating decline, but the AUC score gradually increased. This phenomenon indicates that increasing the number of experts can improve the model's ability to distinguish between positive and negative samples. However, due to the large number of similar features between the auxiliary subtasks and the main task, experts may only need to change feature weights in different task calculations. Increasing the number of experts will cause redundancy in feature extraction, leading to a decrease in the model's ACC and F1 scores. Therefore, At that time, the expert network achieved the best balance between performance and efficiency with only 52.13M parameters.

[0173] Table 5 Optimal values ​​for ablation experiment table

[0174]

[0175] Verified through ablation experiments The optimal value is shown in Table 6, and the experimental data is as follows. When the model only depends on the feature routing branch or the task routing branch ( or Its detection capability across multiple tasks is limited, indicating that a single routing branch encounters a bottleneck in feature representation. Under the specified configuration, the model exhibits the best overall performance, with an accuracy of 85.92%, an F1 score of 81.85%, and a high AUC value. This indicates that the model demonstrates strong performance in feature extraction and weight calculation for various tasks. With this configuration, the model achieves a high AUC, but the accuracy and F1 score decrease significantly. This indicates that both the combination of the two feature routes and a higher weighting of task routes can enhance the learning ability of multi-task models. However, the weighting of shared and unique parameters for tasks and the feature extraction of multi-task systems are not achieved through a simple averaging of the two route gating methods or by relying solely on task routes.

[0176] Table 6 Optimal values ​​for ablation experiment table

[0177]

[0178] Verified through ablation experiments The optimal value of is shown in Table 7, and the experimental data are as follows. Higher values... This will lead to stronger alignment between task gradients, while lower... This emphasizes balancing gradient magnitudes to reduce conflicts. Too high ( When the model overemphasizes gradient consistency, it weakens learning for specific tasks and leads to a decline in the three evaluation metrics. In this scenario, the optimization process is entirely dominated by gradient alignment, with task gradients pointing in similar directions. This setup improves inter-task consistency but limits task-specific discriminative power, leading to poor accuracy and F1 performance. Too low ( When the optimization process causes a loss of coordination between tasks, it leads to unstable convergence and a decrease in overall accuracy. In contrast, when At this stage, the model focuses solely on balancing gradient magnitudes without considering alignment. The model exhibits the worst overall performance in this process, indicating weak sharing between tasks and unstable gradient dynamics. The optimal approach... A value of 0.5 achieves the optimal trade-off between gradient direction and magnitude optimization, improving the model's stability and generalization ability. This further demonstrates that dynamic adjustment... It is crucial for achieving stable and effective gradient adjustment in multi-task learning.

[0179] Table 7 Optimal values ​​for ablation experiment table

[0180]

[0181] Verification has shown that the method proposed in this invention solves the technical problem raised in this invention. Simulation experiments and practical applications have both verified the technical effects and practicality claimed in this invention.

[0182] The proposed method (algorithm) for identifying mediastinal lymph node metastasis based on a dual-control routing hybrid expert model is the core technology of this invention, and various products can be derived from this algorithm. Based on the algorithm (method) proposed in this invention, a mediastinal lymph node metastasis identification system based on a dual-control routing hybrid expert model is developed using a programming language. This system has program modules corresponding to the steps of the identification method, and executes the steps of the aforementioned method for identifying mediastinal lymph node metastasis based on a dual-control routing hybrid expert model during runtime. The developed system (software) computer program is stored on a computer-readable storage medium, and the computer program is configured to implement the steps of the aforementioned method for identifying mediastinal lymph node metastasis based on a dual-control routing hybrid expert model when called by a processor. In other words, this invention is materialized on a carrier, becoming a computer program product used to identify mediastinal lymph node metastasis.

[0183] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, application-specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0184] The computational programs (also referred to as programs, software, software applications, or code) of this invention include machine instructions of a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device PLD) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0185] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, they are all within the protection scope of this invention.

[0186] The following is a list of references cited in this invention:

[0187] [1] Zhao Kejia, Liu Chengwu, Liu Lunxu. Interpretation of the IASLC Ninth Edition TNM Staging for Lung Cancer [J]. Chinese Journal of Thoracic and Cardiovascular Surgery, 2024, 31(04):489-497.

[0188] [2]El-Sherief AH, Lau CT, Carter BW, et al. Staging lung cancer: regional lymph node classification[J]. Radiologic Clinics, 2018, 56(3): 399-409.

[0189] [3]de Sousa IP, Vellasco MMBR, da Silva E C. Evolved explainable classifications for lymph node metastases[J]. Neural Networks, 2022, 148: 1-12.

[0190] [4]Gao Z, Luo Y, Wang M, et al. Seeking multi-view commonality and peculiarity: A novel decoupling method for lung cancer subtype classification[J]. Expert Systems with Applications, 2025, 260: 125397.

[0191] [5]Woo S, Debnath S, Hu R, et al. Convnext v2: Co-designing andscaling convnets with masked autoencoders[C] / / Proceedings of the IEEE / CVFconference on computer vision and pattern recognition. 2023: 16133-16142.

[0192] [6]Chen J, Kao S, He H, et al. Run, don't walk: chasing higher FLOPSfor faster neural networks[C] / / Proceedings of the IEEE / CVF conference oncomputer vision and pattern recognition. 2023: 12021-12031.

[0193] [7]Mehta S, Rastegari M. Mobilevit: light-weight, general-purpose,and mobile-friendly vision transformer[J]. arXiv preprint arXiv:2110.02178,2021.

[0194] [8]Shi D. Transnext: Robust foveal visual perception for visiontransformers[C] / / Proceedings of the IEEE / CVF conference on computer visionand pattern recognition. 2024: 17773-17783.

[0195] [9]Liu S, Johns E, Davison A J. End-to-end multi-task learning withattention[C] / / Proceedings of the IEEE / CVF conference on computer vision andpattern recognition. 2019: 1871-1880.

[0196]

[10] Misra I, Shrivastava A, Gupta A, et al. Cross-stitch networks formulti-task learning[C] / / Proceedings of the IEEE conference on computer visionand pattern recognition. 2016: 3994-4003.

[0197]

[11] Pfeiffer J, Kamath A, Rücklé A, et al. Adapterfusion: Non-destructive task composition for transfer learning[J]. arXiv preprint arXiv:2005.00247, 2020.

[0198]

[12] Ma J, Zhao Z, Yi X, et al. Modeling task relationships in multi-task learning with multi-gate mixture-of-experts[C] / / Proceedings of the 24thACM SIGKDD international conference on knowledge discovery & data mining.2018: 1930-1939.

[0199]

[13] Hazimeh H, Zhao Z, Chowdhery A, et al. Dselect-k: Differentiableselection in the mixture of experts with applications to multi-task learning[J]. Advances in Neural Information Processing Systems, 2021, 34: 29335-29347.

[0200]

[14] Tang H, Liu J, Zhao M, et al. Progressive layered extraction(ple): A novel multi-task learning (mtl) model for personalizedrecommendations[C] / / Proceedings of the 14th ACM conference on recommendersystems. 2020: 269-278。

[0201]

[15] Liu L, Li Y, Kuang Z, et al. Towards impartial multi-tasklearning[C] / / International conference on learning representations. 2021.

[0202]

[16] Sener O, Koltun V. Multi-task learning as multi-objectiveoptimization[J]. Advances in neural information processing systems, 2018, 31.

[0203]

[17] Caruana R. Multitask learning: A knowledge-based source ofinductive bias1[C] / / Proceedings of the Tenth International Conference onMachine Learning. 1993: 41-48.

[0204]

[18] Navon A, Shamsian A, Achituve I, et al. Multi-task learning as abargaining game[J]. arXiv preprint arXiv:2202.01017, 2022.

[0205]

[19] Lin B, Jiang W, Ye F, et al. Dual-balancing for multi-tasklearning[J]. arXiv preprint arXiv:2308.12029, 2023.

[0206]

[20] Senushkin D, Patakin N, Kuznetsov A, et al. Independent componentalignment for multi-task learning[C] / / Proceedings of the IEEE / CVF Conferenceon Computer Vision and Pattern Recognition. 2023: 20083-20093.

[0207]

[21] Liu B, Liu X, Jin X, et al. Conflict-averse gradient descent formulti-task learning[J]. Advances in Neural Information Processing Systems,2021, 34: 18878-18890.

[0208]

[22] Lin X, Zhang X, Yang Z, et al. Smooth Tchebycheff Scalarizationfor Multi-Objective Optimization[C] / / Forty-first International Conference onMachine Learning.

[0209]

[23] Ban H, Ji K. Fair Resource Allocation in Multi-Task Learning[C] / / International Conference on Machine Learning. PMLR, 2024: 2715-2731.

Claims

1. A method for mediastinal lymph node metastasis identification based on a double-control routing hybrid expert model, characterized in that, The method comprises: The mixed expert model MMoE is improved: a double-control routing gating mechanism consistent with the number of multi-tasks is added to the input end of the mixed expert model MMoE to obtain a double-control routing mixed expert model; Firstly, the data of the mediastinal lymph node metastasis identification task and different dimension identification auxiliary sub-tasks are respectively input into the expert network for expert selection and activation, and a double-dimension gradient balancing algorithm is designed to realize multi-task gradient updating; Meanwhile, a double-control routing gating mechanism is used for feature calculation and distribution for each task; the double-control routing gating mechanism is composed of a feature routing branch and a task routing branch; wherein the feature routing branch is used for extracting similar features between mediastinal images, including color, brightness and frequency domain features; the task routing branch extracts task-specific features based on specific label signals in different tasks, and the specific features of each task are as follows: the specific features of the mediastinal lymph node metastasis identification task include the morphological features of the mediastinal region and the environmental information of the lung; the specific features of the mediastinal region detection task are position and edge features; the specific features of the lymph node metastasis prediction task are lymph node structure, color and texture features; Then, the similar and specific features of the tasks extracted by the feature routing branch and the task routing branch are weighted and fused to realize comprehensive description of the features of each task; Then, the feature calculation of the expert network and the feature extraction of the double-control routing gating mechanism are fused to fully balance the shared and independent feature requirements between tasks; Finally, the decoding branch of each task completes the identification process according to the fusion coding result of the expert network and the double-control routing gating mechanism.

2. The mediastinal lymph node metastasis identification method based on the dual-control routing hybrid expert model according to claim 1, characterized in that, The method divides input image data of each task into image blocks, processes the image blocks through a sample feature generator to obtain sample input features ; and inputs the sample input features into an expert network and a double-control routing gate mechanism respectively to obtain an encoder output of each task.

3. The mediastinal lymph node metastasis identification method based on the dual-control routing hybrid expert model according to claim 2, characterized in that, The specific implementation process of the method is: (1) Sample input features Processing in the expert network: First, a regularization function is applied The sample input features are transformed into regularization features :​ The shared parameters of the expert network are represented as: each represents the number of experts; represents the encoding result of the sample input features by the i-th expert, , , represents the number of experts; is a shared representation matrix generated by the network of all experts; Then, according to the regularized features a binary expert selection matrix is constructed for each task which consists of binary single-expert selectors defined as: Each indicates whether the task is currently active for the expert; there are is set to 1, indicating that the task is active for the expert, while the remaining experts remain inactive, value is a parameter set manually by a person;​ (2) sample input features Processing in double control routing gating mechanism: First apply a flattening function input features to sample flattened features : The double control routing and gating mechanism is composed of a feature routing branch and a task routing branch, and the feature weight is calculated through the feature routing branch The process is as follows: representing sample flattening features The selection weights for each expert contain similar sample features across multiple tasks, helping the model to capture the correlation between different tasks, representing a weight matrix for a linear transformation, representing a bias vector for a linear transformation, The function is applied to the row vector of each sample; Further, the keyword feature filter is utilized Separate out each task flat feature , , Represents the total number of tasks; Computing a weight for each task using task routing branches : a weight matrix representing linear variation of the th task, a bias vector representing linear variation of the th task, , a selection weight of each task to the expert network, which contains the understanding of the model to the characteristics of the task, combined with the feature weight and the task weight , to generate the routing weight of each task to the expert network : wherein, is a hyperparameter set for the experimental process; (3) Finally, combine the outputs of the expert network and the double-control routing gate mechanism to obtain the encoder output of each task : wherein, denotes a dot product operation, denotes a Hadamard operation.

4. The mediastinal lymph node metastasis identification method based on the dual-control routing hybrid expert model according to claim 3, characterized in that, Bias vector representing a linear transformation The determination process is that the learnable parameters The values are automatically learned by gradient descent algorithm according to multi-task data characteristics and model characteristics.

5. The mediastinal lymph node metastasis identification method based on the dual-control routing hybrid expert model according to claim 4, characterized in that, The hyperparameters set during the experiment process The determination process of the value of the parameter is as follows: The value range of the parameter is: According to the ablation experiment results, the best model comprehensive average performance index value is selected as the final selection of the present application. value as the final selection of the present application.

6. The mediastinal lymph node metastasis identification method based on the dual-control routing hybrid expert model according to claim 5, characterized in that, Hyperparameters set during the experiment The determination process is: determined according to the ablation experiment result.

7. The mediastinal lymph node metastasis identification method based on the dual-control routing hybrid expert model according to claim 6, characterized in that, The implementation process of the double-dimension gradient balancing algorithm is: The process contains four variables that need to be calculated , , and : (1) the The calculation procedure is as follows: Firstly, the gradient of each task is calculated according to the loss value: wherein, representing model parameters the gradient of, representing the loss function for a task, representing the number of tasks; thus the gradient matrix of the model can be expressed as: Further, the iteration formula of is obtained, whose initial value is 0: wherein, is a decay factor, increases with the number of iterations, and are set to 0.8 and 0.2 respectively according to empirical values; As a batch buffer gradient, it can reflect the latest task information while preserving the trend of historical gradients, achieving dynamic balance of gradient update. (2) the The calculation process is: According to the gradient matrix of the model , a gradient autocorrelation matrix in the task space is calculated , the gradient similarity between tasks is obtained, and the calculation formula is as follows: Matrix contains the statistical relationship of feature maps in the task space, which is crucial for the model to understand the style and content of the task image; Matrix The eigenvalues and eigenvectors of the matrix are calculated as follows: Let be a diagonal matrix with diagonal elements being the eigenvalues of the matrix ; ; be an orthogonal matrix whose column vectors are the eigenvectors of the matrix ; then construct a new diagonal matrix according to the rank of the matrix : In the task space, perform gradient decomposition and compute the transformation matrix : transformation matrix for adjusting the size and direction of the gradient; summing the columns of the transformation matrix to compute the gradient matrix vector : wherein denotes the sum over The summing all elements in a column; align the gradient directions of different tasks to avoid gradient conflicts between tasks; (3) the The calculation process is: First, the magnitude of each gradient vector is computed to form a norm vector As follows: Then, the normalized norm vector of the gradient is calculated : for dynamically adjusting the size of the task gradient, preventing task domination caused by too large gradient and task omission caused by too small gradient; (4) the The determination flow of the value is: The two weights are convexly combined as a weight adjustment factor; it can flexibly adjust the proportion of each weight according to the task characteristics and data characteristics, realize the complementary adjustment of the gradient defects, and improve the balance and stability of the training between tasks; in the model, The optimal value of the weight adjustment factor is determined through an ablation experiment process.

8. A mediastinal lymph node metastasis identification method system based on a double-control routing hybrid expert model, characterized in that: The system has program modules corresponding to the steps of any one of claims 1-7, and when running, the steps in the mediastinal lymph node metastasis identification method based on the double-control routing mixed expert model are executed.

9. A computer-readable storage medium, characterized in that: The computer readable storage medium stores a computer program, and the computer program is configured to realize the steps of the mediastinal lymph node metastasis identification method based on the double-control routing mixed expert model in any one of claims 1-7 when called by the processor.

Citation Information

Patent Citations

  • Thyroid cancer neck lymph node metastasis identification method and system based on Swin Transformer

    CN117893462A

  • Lung mediastinal lymph node detection method, system, equipment and medium

    CN118229645A