An Esophageal Cancer Lymph Node Metastasis Prediction System Based on Clinical Dynamic Modulation

CN122575745APending Publication Date: 2026-08-14ZHEJIANG CANCER HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

然而,此种基于单一几何尺寸的评估范式存在显著的诊断局限性:其不仅难以有效辨识食管癌中常见的、未引起淋巴结显著增大的微小转移灶,亦无法准确区分由炎性增生反应所导致的淋巴结假阳性增大

Benefits of technology

(1)通过构建包含多层感知机的权重生成网络,低维稀疏的临床特征向量被映射至富含语义表达能力的隐空间表征,并基于该表征并行导出具有通道特异性的缩放调制参数与偏置校准参数。该机制突破了传统后融合方法中两种模态特征仅在决策层进行简单拼接的局限,实现了临床知识对影像特征提取过程的深层介入。具体而言,缩放参数能够依据患者个体临床风险背景对特定特征通道的信息流通强度进行门控调节,而偏置参数则可对特征响应基线进行适应性校准,二者协同作用使得影像特征提取网络能够依据患者特异性的病理生理状态动态调整其对不同纹理模式的敏感度,从而有效增强了模型对微小转移灶相关影像表征的响应能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575745A_ABST
    Figure CN122575745A_ABST
Patent Text Reader

Abstract

This invention discloses a system for predicting esophageal cancer lymph node metastasis based on clinical dynamic modulation, relating to the field of medical image processing. The system includes: a data acquisition module for acquiring enhanced CT image data and structured clinical data of the object to be predicted, and performing standardized preprocessing on the data; a clinical feature-driven weight generation network for transforming low-dimensional clinical feature vectors to a high-dimensional latent space through nonlinear mapping, and generating dynamic modulation parameters in parallel; an image feature extraction and dynamic modulation network for extracting image feature maps and performing channel-by-channel feature recalibration to generate fusion features that embed clinical prior knowledge into the image representation process; and a metastasis prediction classification module for performing spatial dimensionality reduction and probability mapping on the fusion features, generating and outputting prediction results. This application's solution, through clinical prior dynamic modulation of image weights, can significantly improve the prediction accuracy of small lymph node metastasis in esophageal cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing, and in particular to an esophageal cancer lymph node metastasis prediction system based on clinical dynamic modulation. Background Technology

[0002] Esophageal cancer, a highly lethal malignant tumor of the digestive tract, relies heavily on accurate assessment of lymph node metastasis. This assessment is crucial for determining whether neoadjuvant chemoradiotherapy is necessary and for ascertaining the extent of intraoperative lymph node dissection. Currently, preoperative clinical assessment primarily relies on enhanced computed tomography (CT) images, and lymph node short diameter exceeding a specific millimeter threshold is commonly used as a morphological criterion for metastasis. However, this single-geometric-size-based assessment paradigm has significant diagnostic limitations: it struggles to effectively identify common micrometastases in esophageal cancer that do not cause significant lymph node enlargement, and it also fails to accurately distinguish false-positive lymph node enlargement caused by inflammatory hyperplasia. These inherent defects in morphological diagnostic criteria can easily lead to systematic biases in preoperative staging, resulting in inappropriate treatment strategies and ultimately impacting patient survival and quality of life. Therefore, exploring intelligent auxiliary diagnostic technologies that transcend traditional morphological frameworks and integrate multidimensional information to achieve accurate lymph node metastasis identification has become a critical technical challenge in this field.

[0003] To overcome the limitations of single-image morphological assessment, existing technologies have begun to incorporate deep learning and radiomics methods to integrate CT images and clinical indicators by constructing multimodal fusion models. However, current mainstream fusion paradigms mostly employ post-fusion strategies, which involve simply concatenating or linearly combining independently extracted image feature vectors with clinical feature vectors at the model's decision-making stage. These methods suffer from two inherent technical drawbacks: First, the intermodal interactions are superficial and lack temporal sequence. The image feature extraction network is completely isolated from prior clinical knowledge during computation, making it unable to dynamically adjust the weights of attention to specific texture patterns or subtle heterogeneous regions in the images based on clinical high-risk cues such as abnormal tumor markers or low differentiation. This makes it easy to lose crucial micrometastatic information during the layer-by-layer abstraction process of deep convolution. Second, there is a significant modal semantic gap. Imaging data is represented as a high-dimensional, dense spatial pixel matrix, while clinical data is a low-dimensional, sparse structured vector. Simple vector concatenation operations cannot effectively align and synergistically gain information from these two heterogeneous modalities within a unified semantic space, thus limiting the model's generalization performance in complex clinical scenarios. Therefore, based on the above challenges, this invention proposes an esophageal cancer lymph node metastasis prediction system based on clinical dynamic modulation. Summary of the Invention

[0004] To address the aforementioned issues, the present invention aims to provide an esophageal cancer lymph node metastasis prediction system based on clinical dynamic modulation. By constructing a nonlinear mapping mechanism from a low-dimensional clinical feature space to a high-dimensional imaging feature channel, the system enables deep intervention and guidance control of patient-specific clinical information in the feature extraction process of deep convolutional networks. This effectively enhances the model's ability to characterize micrometastases and occult lesions, and significantly improves the accuracy of esophageal cancer lymph node metastasis risk assessment.

[0005] To achieve the above objectives, this invention provides an esophageal cancer lymph node metastasis prediction system based on clinical dynamic modulation. Its core architecture consists of a data acquisition module, a clinical feature-driven weight generation network, an image feature extraction and dynamic modulation network, and a metastasis prediction classification module. The data acquisition module acquires enhanced CT images and structured clinical data and performs standardized preprocessing. The weight generation network adopts a hypernetwork architecture, transforming the standardized low-dimensional clinical feature vectors to a high-dimensional latent space through nonlinear mapping, and generating scaling modulation parameters and bias calibration parameters corresponding to the image feature channels in parallel based on this latent space representation. The image feature extraction and dynamic modulation network uses a deep residual convolutional network as its backbone, embedding a channel-by-channel feature recalibration mechanism based on the modulation parameters only in the high-level feature extraction stage corresponding to tumor heterogeneity semantic parsing, thereby organically integrating clinical prior knowledge into the dynamic construction process of image feature expression. Finally, the metastasis prediction classification module performs global spatial dimensionality reduction and probability mapping on the fused features modulated by clinical priors, outputting the final lymph node metastasis risk prediction result. The above modules work together to form an end-to-end, jointly optimizable framework, enabling a paradigm shift from general image recognition to clinical knowledge-oriented recognition.

[0006] In a first aspect, the present invention provides an esophageal cancer lymph node metastasis prediction system based on clinical dynamic modulation, comprising: The data acquisition module is used to acquire enhanced CT image data and structured clinical data of the object to be predicted, and to perform standardized preprocessing on the data; A clinical feature-driven weight generation network is used to transform the standardized low-dimensional clinical feature vectors to a high-dimensional latent space through a nonlinear mapping in the hypernetwork architecture, and to generate two sets of channel-specific dynamic modulation parameters in parallel based on the high-dimensional latent space. An image feature extraction and dynamic modulation network is used to extract image feature maps with spatial hierarchical structure from standardized image data, and only at the deep feature extraction level corresponding to the tumor heterogeneity semantic parsing stage, the dynamic modulation parameters are used to perform channel-by-channel feature recalibration operation on the corresponding image feature maps to generate fusion features that embed clinical prior knowledge into the image representation process. The metastasis prediction and classification module is used to perform spatial dimensionality reduction and probability mapping on the fused features, and generate and output prediction results characterizing the risk of esophageal cancer lymph node metastasis.

[0007] Furthermore, the structured clinical data includes continuous and discrete variables describing the patient's physiological state and tumor pathological attributes. The continuous variables are statistically standardized to eliminate dimensional differences, and the discrete variables are binarized and encoded to transform them into high-dimensional sparse vectors. This provides numerically stable input conditions for the subsequent weight generation network while preserving the integrity of the pathophysiological meaning of the clinical variables.

[0008] Furthermore, the nonlinear mapping transformation constructed by the weight generation network, through the introduction of a learnable parameter matrix and a normalization mechanism, achieves adaptive adjustment of the response intensity of each feature channel in the deep feature extraction layer and baseline bias correction based on the differences in the individual patient's pathophysiological state.

[0009] Furthermore, the dynamic modulation parameters include scaling parameters for controlling the intensity of information flow in the feature channels and bias parameters for adjusting the reference response of the feature channels; when generating the scaling parameters, the weight generation network introduces a nonlinear activation constraint mechanism to limit the range of its numerical fluctuations, while retaining the numerical degrees of freedom of the linear mapping when generating the bias parameters.

[0010] Furthermore, the image feature extraction and dynamic modulation network uses a deep convolutional network with a bottleneck residual structure as the feature extraction backbone. The deep feature extraction layer corresponds to the high-level semantic bottleneck residual block group in the deep convolutional network, which is responsible for parsing the heterogeneous texture and necrotic region features inside the tumor. The shallow bottleneck residual block group for extracting general visual features is excluded from the intervention range of the dynamic modulation parameters. This avoids noise interference from clinical prior knowledge on the general visual feature extraction process and ensures that modal interaction occurs at the optimal semantic target where clinical knowledge plays a guiding role.

[0011] Furthermore, the feature recalibration operation is configured at the end of each residual unit in the high-level semantic bottleneck residual block group, and the specific embedding position is after the last convolution operation and normalization operation in the unit and before the nonlinear activation operation and residual path superposition operation. In this way, while maintaining the residual path identity mapping advantage, the precise channel-by-channel intervention of the clinical modulation signal on the normalized feature response is achieved.

[0012] Furthermore, the weight generation network also includes a clinical uncertainty measurement branch, used to generate a scalarized clinical risk indicator based on the high-dimensional latent space representation; the image feature extraction and dynamic modulation network also includes a spatial attention generation unit, which receives the scalarized clinical risk indicator as a dynamic adjustment signal, and generates a spatial dimension attention-weighted mask for the feature map output by the deep feature extraction layer based on the signal, so as to perform secondary enhancement modulation of the spatial dimension on the feature map after channel recalibration.

[0013] Furthermore, the weight generation network also includes a clinical information complexity assessment branch, used to calculate the complexity assessment value of the completeness and determinism of the clinical features based on the high-dimensional latent space representation; the image feature extraction and dynamic modulation network adaptively determines the intervention initiation depth of the dynamic modulation parameters in the feature extraction level based on the complexity assessment value, so that the guiding strength and intervention range of clinical prior knowledge can be dynamically adjusted according to the quality differences of the input clinical information.

[0014] Furthermore, the metastasis prediction classification module receives the top-level image feature map after the feature recalibration operation, and uses a global spatial average pooling mechanism to compress the feature map into a one-dimensional feature descriptor. The feature descriptor condenses the deep image semantic information after being weighted and enhanced by the dynamic modulation parameters, so that the feature representation used for classification decision is highly condensed with clinical prior-guided imaging evidence related to tumor heterogeneity.

[0015] Furthermore, in the parameter optimization stage, the system employs an adaptive weighting strategy based on class distribution priors to construct a loss function. This strategy determines the corresponding cost-sensitive weight factors based on the inverse frequency of positive and negative classes in the training samples, in order to correct the learning bias of the model under class imbalance conditions.

[0016] Furthermore, during the parameter optimization phase, the system applies random spatial transformation augmentation operations to the input enhanced CT image data. These augmentation operations include random mirror flipping around a spatial axis, random rotation transformation within a defined angle range, and random scaling transformation within a defined scale range, in order to enhance the model's ability to generalize and adapt to spatial variations in the images.

[0017] Furthermore, during the parameter optimization phase, the system introduces a cross-modal contrastive regularization constraint mechanism. This mechanism constructs positive and negative constraint pairs at the image feature level based on the similarity measure of clinical feature vectors among samples within the training batch, and constructs an auxiliary loss term based on the criteria of maximizing the mutual information of positive constraint pairs and minimizing the mutual information of negative constraint pairs. The auxiliary loss term, together with the classification supervision loss, is used in the iterative update process of network parameters to force the dynamically modulated high-level semantic feature representation of the image and the corresponding clinical state to maintain structured semantic alignment in the embedding space.

[0018] Secondly, a method for predicting esophageal cancer lymph node metastasis based on clinical dynamic modulation is also provided. This method is based on the system described in the first aspect and includes: Acquire enhanced CT images and structured clinical data of the subjects to be predicted, and perform standardized preprocessing on each. Standardized clinical data is input into a weighted supernetwork, which generates a high-dimensional latent space representation from the low-dimensional clinical feature space through a nonlinear mapping mechanism, and derives the scaling and bias calibration values ​​for the image feature channels in parallel based on the representation. Standardized image data is input into a deep feature extraction network to extract image feature map sequences with spatial hierarchical structure. Only in the deep feature extraction stage corresponding to tumor heterogeneity semantic parsing, the scaling control amount and bias calibration amount are used to perform channel-wise feature response recalibration operation on the extracted feature maps to obtain fused feature maps modulated by clinical prior depth. Global spatial dimensionality reduction and nonlinear decision mapping are performed on the fused feature map to generate and output predicted values ​​representing the probability of esophageal cancer lymph node metastasis.

[0019] This invention proposes a clinically dynamic modulation-based esophageal cancer lymph node metastasis prediction system. The system performs standardized preprocessing to normalize the representation of enhanced CT images and structured clinical data. Subsequently, a weighted generation network based on a hypernetwork architecture is introduced to transform low-dimensional clinical feature vectors into a high-dimensional latent space via nonlinear mapping. Based on this latent space representation, scaling modulation parameters and bias calibration parameters corresponding to the image feature channels are derived in parallel. Building upon this, an image feature extraction and dynamic modulation network with a deep residual convolutional network as its backbone performs channel-by-channel feature recalibration only at the high-level feature extraction stage corresponding to tumor heterogeneity semantic parsing. This organically embeds patient-specific clinical prior knowledge into the dynamic construction of the image feature representation process. Finally, the modulated and enhanced fused features are reduced in dimensionality globally and mapped probabilistically to output the lymph node metastasis risk prediction result. These modules collaboratively constitute an end-to-end jointly optimizable framework, achieving a fundamental shift from a general image recognition paradigm to a clinical knowledge-oriented recognition paradigm.

[0020] Regarding the depth of modal interaction, the system overcomes the limitations of superficial and fragmented intermodal interactions in traditional methods by constructing a nonlinear mapping mechanism from clinical features to imaging features. This allows for deep integration of clinical prior knowledge into the imaging feature extraction process, enabling dynamic adjustments to the representation of imaging features based on the individual patient's pathophysiological background. In terms of key feature capture capabilities, the system strategically limits dynamic modulation operations to the high-level semantic feature layer responsible for parsing tumor heterogeneity textures, avoiding undue interference from clinical information on superficial general visual features. This significantly enhances the model's sensitivity and robustness in capturing subtle imaging features related to micrometastases and occult lesions. Regarding its clinical decision support value, the system effectively improves the accuracy of lymph node metastasis risk assessment without introducing additional examination costs or procedural burdens. It provides clinicians with intelligent auxiliary decision-making support that combines high sensitivity and high specificity for developing individualized lymph node dissection ranges and neoadjuvant therapy strategies, demonstrating significant clinical application value for optimizing treatment pathways for esophageal cancer patients.

[0021] Beneficial effects By implementing the esophageal cancer lymph node metastasis prediction system based on clinical dynamic modulation provided by the present invention, the following technical effects are achieved: (1) By constructing a weighted generation network containing a multilayer perceptron, low-dimensional sparse clinical feature vectors are mapped to a latent space representation rich in semantic expression, and channel-specific scaling and modulation parameters and bias calibration parameters are derived in parallel based on this representation. This mechanism breaks through the limitation of traditional post-fusion methods that simply splice the two modal features at the decision layer, and realizes the deep intervention of clinical knowledge in the image feature extraction process. Specifically, the scaling parameter can gate the information flow intensity of specific feature channels according to the individual clinical risk background of the patient, while the bias parameter can adaptively calibrate the feature response baseline. The synergistic effect of the two enables the image feature extraction network to dynamically adjust its sensitivity to different texture patterns according to the patient's specific pathophysiological state, thereby effectively enhancing the model's response capability to image representations related to small metastases.

[0022] (2) By introducing clinical modulation only in the high-level semantic feature extraction stage, which is responsible for parsing tumor heterogeneity and necrotic region features, rather than intervening in the shallow network that extracts general visual features such as edge texture, the destructive interference of clinical priors on the continuity of basic visual feature extraction is effectively avoided. This design allows the image feature extraction network to maintain unbiased representation of general image patterns in the shallow layer, while receiving directional guidance from clinical knowledge in the critical stage of semantic abstraction, thereby achieving the optimal selection of modality fusion timing. Compared with full-layer modulation or shallow-layer modulation strategies, this selective modulation scheme significantly enhances the sensitivity to capturing subtle imaging signs related to lymph node metastasis, while effectively suppressing the noise amplification effect that may be caused by prematurely introducing clinical information.

[0023] (3) A spatial attention co-modulation branch guided by clinical risk entropy is introduced on the basis of channel modulation. On the one hand, clinical risk entropy, as a scalar uncertainty measure extracted from clinical feature vectors, provides patient-specific dynamic adjustment signals for the generation of spatial attention masks, enabling the spatial focus intensity to be adaptively scaled according to the level of clinical risk. On the other hand, the synergistic effect of spatial attention masks and channel modulation parameters further clarifies the spatial positioning information of "where to focus" on the feature map, based on "what features to focus on". This mechanism significantly improves the network's focusing accuracy on key anatomical regions in images associated with clinical risk backgrounds, especially enhancing the spatial positioning ability of small lesions located in atypical locations or areas with blurred boundaries, effectively compensating for the guidance blind spot of simple channel modulation at the spatial resolution level.

[0024] (4) By constructing positive and negative sample constraint pairs based on clinical feature similarity during the training phase and applying contrastive learning auxiliary loss, the high-level semantic feature vectors of images and their corresponding clinical states are forced to maintain a structured semantic alignment in the embedding space. This mechanism ensures that the image feature representations of patients with similar clinicopathological backgrounds are close to each other in the embedding space, while the image representations of patients with different clinical states are far apart. This constraint effectively improves the clinical discriminative power and inter-class separability of image feature representations, while significantly enhancing the robustness of the model under conditions of sparse or missing clinical data, enabling the dynamically modulated fused features to more accurately reflect the true pathophysiological state of patients.

[0025] (5) By adding a clinical information complexity evaluation branch to the weight generation network, the completeness and certainty of the input clinical feature vector are quantitatively evaluated in real time, and the intervention depth of the clinical modulation operation is dynamically decided based on the evaluation results. For highly certain samples with rich and clear clinical information, the modulation operation remains at the deep semantic level to reduce unnecessary computational overhead; for low-certainty samples with sparse clinical information or contradictory indicators, the modulation operation is extended to a shallower level to provide supplementary prior constraints. This adaptive strategy ensures sufficient clinical guidance for complex and difficult cases on the one hand, and avoids redundant calculations for simple cases on the other. While improving diagnostic accuracy, it effectively controls the average computational resource consumption in the inference stage, achieving the optimal trade-off between model capacity utilization and clinical adaptability. Attached Figure Description

[0026] To make the above-described esophageal cancer lymph node metastasis prediction system based on clinical dynamic modulation of the present invention more obvious and understandable, the accompanying drawings used in the specific embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0027] Figure 1 This is a flowchart illustrating the method described in this application; Figure 2 A schematic diagram illustrating the network architecture for image feature modulation and classification based on dynamic weights; Figure 3 ROC curves for different models. Detailed Implementation

[0028] Example 1: This embodiment proposes an esophageal cancer lymph node metastasis prediction system based on dynamically modulated image weights according to clinical features. Its core lies in breaking through the shallow interaction paradigm of traditional multimodal fusion, where clinical information and image features are simply spliced ​​together at the decision level. Instead, it constructs a dynamic modulation mechanism for image features driven by prior clinical knowledge. This allows the deep convolutional network to adaptively adjust the response intensity to different texture patterns and semantic features based on the patient's individualized pathophysiological background during CT image feature extraction, thereby significantly improving the prediction accuracy for small lymph node metastases and occult lesions in esophageal cancer. The overall workflow of this embodiment is as follows: Figure 1 As shown, the details are as follows.

[0029] In the multimodal data acquisition and standardized preprocessing stage, the system first acquires the preoperative venous phase enhanced CT image data of the subject to be predicted, which is typically stored in DICOM format. Centered on the primary lesion and mediastinal region as labeled by the radiologist, a region of interest (ROI) cropping operation is performed on the image to obtain a local image block containing the lesion and surrounding background tissue. To further eliminate interference from irrelevant tissues such as bone and air in subsequent feature extraction, a soft tissue window parameter (window width 400 HU, window level 40 HU) is used to truncate the CT values ​​of the cropped image. Simultaneously, the system acquires the patient's corresponding structured clinical data and constructs an original clinical feature vector based on this, denoted as […]. The features contained in this vector can be divided into two categories: continuous variables and discrete variables. Continuous variables include patient age, carcinoembryonic antigen (CEA) value, and squamous cell carcinoma antigen (SCCA) value. These variables are processed using the Z-Score statistical standardization method to eliminate numerical differences caused by different dimensions. Discrete variables include gender, tumor location, tumor differentiation degree, and clinical T stage. These variables are converted into binary vector representations using one-hot encoding. Finally, the standardized continuous variables and the encoded discrete variables are concatenated along the feature dimension to form a unified high-dimensional clinical feature input vector, denoted as [vector name missing]. This provides numerically stable input conditions for the subsequent weight generation network.

[0030] In the clinical feature-driven dynamic image weight generation stage, the system introduces a weight generation network based on a hypernetwork architecture, aiming to establish a nonlinear mapping relationship from the low-dimensional clinical feature space to the high-dimensional image feature channel space. The main structure of this weight generation network is a multilayer perceptron containing two fully connected layers. Batch normalization layers and linear unit activation functions with leakage correction are sequentially embedded between each fully connected layer to enhance the network's nonlinear expressive power and stabilize gradient propagation during training. The high-dimensional clinical feature input vector obtained in the preceding steps is then used as the input. The data is fed into the multilayer perceptron and, after layer-by-layer nonlinear transformation, a latent feature vector containing the patient's individual pathophysiological information is extracted, denoted as... This latent feature vector can be viewed as a compact representation of the patient's clinical state in a high-dimensional semantic space, encapsulating the interaction relationships of multi-dimensional clinical information such as age, tumor marker levels, differentiation degree, and stage. Subsequently, based on this latent feature vector... The network generates two sets of dynamic modulation parameters through two parallel linear transformation layers, each strictly corresponding to the number of feature channels in the modulation layer of the image feature extraction network: one is the channel scaling weight, denoted as... The second is the channel bias weight, denoted as... For the first image to be modulated in the image extraction network A convolutional layer with a set number of feature channels. The generated weight dimensions and All with Strict correspondence. Scaling weights. The generation process is represented as:

[0031] In the formula, It is a learnable generator matrix; For learnable bias terms; The hyperbolic tangent activation function restricts the scaling weights to a symmetrical range of -1 to 1. This effectively suppresses abnormal gradient amplification during modulation, ensuring the numerical stability of the model training. Bias weights The generation process is represented as:

[0032] In the formula, It is a learnable generator matrix; This is a learnable bias term; the numerical degrees of freedom of the linear mapping are preserved here without imposing interval constraints, thus giving it sufficient flexibility to adaptively calibrate the feature response baseline. The scaling weights mentioned above... With bias weights Together, they constitute the control signal for dynamically modulating the image feature extraction network.

[0033] In the image feature modulation and classification stage based on dynamic weights, the system adopts a deep convolutional network with a bottleneck residual structure as the backbone architecture for image feature extraction. Its network architecture is as follows: Figure 2 As shown, the backbone network sequentially comprises multiple bottleneck residual block groups along the forward propagation direction. The shallow bottleneck residual block groups (i.e., the first and second groups) are primarily responsible for extracting general visual features such as edges, corners, and local textures. The deep bottleneck residual block groups (i.e., the third and fourth groups) are responsible for parsing high-level abstract semantic features closely related to tumor heterogeneity, such as the morphology of necrotic areas within the tumor, heterogeneous enhancement patterns, and signs of marginal infiltration. The key lies in strategically limiting dynamic modulation operations to the deep bottleneck residual block groups. Specifically, dynamic modulation layers are embedded only in each residual unit contained in the third and fourth bottleneck residual block groups, while the shallow bottleneck residual block groups retain their original structure and are unaffected by clinical modulation signals. The starting point of this selective modulation strategy is that the extraction of shallow general visual features should maintain the ability to represent the basic pattern of the image without bias. Prematurely introducing clinical prior intervention may introduce noise and disrupt the continuity of shallow feature extraction. High-level semantic features are an abstract condensation of tumor pathological attributes, which are the best semantic targets for clinical prior knowledge to play a directional guiding role.

[0034] Within each selected residual unit for dynamic modulation, the embedding position of the dynamic modulation layer is strictly set after the last 1×1 convolutional layer and its batch normalization layer within that unit, and before the addition operation of the ReLU nonlinear activation function and the residual path. Let the target stage... The first in The feature map output by each residual unit is ,in The values ​​are limited to the third and fourth bottleneck residual block groups. The system uses the aforementioned weight generation network to generate the corresponding channel scaling weights for this unit. With channel bias weight This involves performing a channel-by-channel affine transformation on the input feature map. Specifically, for a feature map with channel indices of... Spatial location coordinates are and The modulation calculation process for each element is expressed as follows:

[0035] In the formula, Input feature map; Output feature map; For channel indexing; Spatial location index in the height direction; Spatial position index in the width direction; For the corresponding to the first Phase 1 The th residual unit Scaling weights for each channel; These are the corresponding bias weights. In this affine transformation mechanism, the scaling weights... Through gating factors The form of the bias weights acts on the original feature response, enabling the network to adaptively adjust the information flow intensity of specific feature channels based on the patient's individualized clinical risk background; An adaptive offset correction is then applied to the baseline level of the feature response. The feature map output after this modulation is the fused feature after recalibrating the image representation in a high-level semantic space using clinical prior knowledge.

[0036] The dynamically modulated and enhanced top-level feature map output from the backbone network is then fed into a global average pooling layer to perform spatial dimension reduction and compression, resulting in a fixed-dimensional one-dimensional feature descriptor. This feature descriptor highly condenses the deep image semantic information guided by clinical prior weighting, preserving both the morphological and textural features in the image that are discriminative for lymph node metastasis and embedding the modulating effect of the patient-specific clinical background on the expression of these features. Finally, this one-dimensional feature descriptor is input into a fully connected classification layer, mapped to a probability interval between zero and one by a sigmoid activation function, and outputting a predicted probability value representing the patient's esophageal cancer lymph node metastasis. The system determines the positive or negative result of lymph node metastasis based on the relative relationship between the predicted probability value and a preset decision threshold (usually 0.5). It is worth noting that the above feature modulation process and classification prediction process constitute an end-to-end holistic differentiable framework. The parameters of each component of the network can be directly and jointly optimized for the lymph node metastasis prediction task without the need for staged training or manual intervention.

[0037] In the model training and evaluation phases, to objectively measure the system's generalization performance and prevent overfitting, this embodiment employs a strict data partitioning strategy. All collected multimodal paired data from esophageal cancer patients meeting the inclusion criteria are randomly divided into training, validation, and test sets in a 3:1:1 ratio. During partitioning, stratified sampling is used to ensure a consistent distribution ratio of lymph node metastasis-positive to negative samples across all subsets, thereby eliminating evaluation biases that may be introduced by uneven data distribution. The training set is used for iterative updates of the model's learnable parameters, the validation set is used for hyperparameter selection and monitoring of the training process, and the test set is only used for final performance evaluation after all training is complete.

[0038] To address the practical constraints of high cost and relatively limited sample size for medical image data annotation, this embodiment introduces online data augmentation technology during the training phase. In each iteration of reading training samples, spatial transformation operations are randomly applied to the input CT image data, including random mirror flips around a spatial axis, random rotations within a range of ±15 degrees, and random scaling transformations within a range of 0.8 to 1.2 times. These transformations can simulate subtle variations in patient positioning in clinical practice, enabling the model to learn robust image features with spatial transformation invariance, rather than mechanically memorizing specific pixel positions in the training samples. It should be noted that data augmentation operations are only applied to image modal data; clinical structured data retains its original values ​​during augmentation to strictly maintain the accuracy and consistency of its pathophysiological meaning.

[0039] In the core step of model optimization, considering the objective reality that the number of positive samples is usually significantly less than the number of negative samples in esophageal cancer lymph node metastasis prediction tasks, this embodiment adopts an adaptive weighting strategy based on class distribution priors to construct a classification loss function. This loss function introduces class weight factors that are correlated with the inverse of the frequency of each class sample in the training set, denoted as... and The contribution of positive and negative samples to the total loss is rebalanced. Let the batch size be... , No. The true label for each sample is The model predicts the transition probability as Then the weighted cross-entropy loss function The expression is:

[0040] During backpropagation, this weighting mechanism assigns greater penalty weights to minority class samples, effectively correcting learning bias caused by class imbalance and significantly improving the model's sensitivity in detecting metastatic positive cases. The entire training process employs an end-to-end joint optimization paradigm, meaning the parameters of the image feature extraction network and the clinical weight generation network are updated synchronously using gradients. The optimizer used is the AdamW algorithm, with an initial learning rate set to... The model employs a cosine annealing strategy for dynamic adjustment to finely converge to a better local extremum in the later stages of training. An early stopping monitoring mechanism is implemented during training; training automatically terminates when the loss function value on the validation set does not show a decreasing trend within ten consecutive epochs, and the model parameter state corresponding to the optimal performance on the validation set is saved. All experiments are based on the PyTorch deep learning framework and deployed and accelerated on a high-performance computing platform equipped with NVIDIA GPUs.

[0041] To fully verify the effectiveness of the technical solution proposed in this embodiment, multi-dimensional quantitative evaluation and comparative experiments were conducted on an independent test set. First, the proposed method was compared horizontally with a single-modal baseline model that did not incorporate any prior clinical knowledge and a multi-modal model employing a traditional clinical-image post-fusion strategy. Experimental results show that the single-modal baseline model, limited by the singularity of the information source, exhibits significantly insufficient sensitivity in detecting metastatic cases; while the traditional post-fusion model incorporates clinical information, its performance gain is limited due to the shallow level of intermodal interactions. In contrast, the dynamic modulation network proposed in this embodiment achieves significant improvements in key indicators such as area under the receiver operating characteristic (ROC) curve, accuracy, sensitivity, specificity, and F1 score, with a more balanced distribution of sensitivity and specificity. Specific performance indicators on the test set are shown in Table 1, and the corresponding ROC curves are shown in Table 2. Figure 3As shown, the ROC curve of this method is significantly better than that of the baseline model and the post-fusion model, which strongly demonstrates the unique advantages of the dynamic weight modulation mechanism in capturing occult micro lesions.

[0042] Table 1. Performance comparison of different models on the task of predicting lymph node metastasis in esophageal cancer

[0043] Secondly, to explore the optimal timing for clinical prior knowledge to intervene in the image feature extraction process, this embodiment conducted longitudinal ablation experiments at different modulation levels. The experiments tested three strategies: applying modulation only to the shallow bottleneck residual block group, applying full-layer modulation to all bottleneck residual block groups, and applying modulation only to the deep bottleneck residual block group as recommended in this invention. Experimental data showed that the shallow modulation strategy offered the most limited performance improvement. While the full-layer modulation strategy improved upon shallow modulation, it still did not reach the optimal level and introduced more additional learnable parameters. The deep modulation strategy adopted in this embodiment achieved the best performance across all indicators, and the increase in the number of parameters was less than that of the full-layer modulation scheme. This result fully confirms the following design insight: the high-level semantic feature layer, closely related to tumor heterogeneity, is the optimal target for clinical prior knowledge to play a directional guiding role. Premature or excessive clinical intervention not only fails to improve performance but may also introduce noise interference and increase unnecessary computational burden.

[0044] Example 2: Traditional dynamic modulation methods only operate on the feature channel dimension, neglecting the guiding role of the heterogeneity of clinical risk factors in the spatial distribution of images. This embodiment proposes a multi-scale channel-spatial co-modulation mechanism based on clinical risk entropy. First, a scalarized clinical risk entropy index is derived from the clinical feature vector using a weighted generation network. This index quantitatively represents the degree of uncertainty in lymph node metastasis implied by the patient's comprehensive clinical characteristics. Second, this risk entropy is used as a regulatory factor to generate an adaptive attention mask in the spatial dimension of the feature map. This mask guides the network to focus on key anatomical regions in the image that are associated with the clinical risk background. Finally, the spatial attention mask is synergistically combined with the original channel modulation parameters to form a channel-spatial joint modulation strategy. Through this mechanism, the network can not only adjust "what features to focus on" based on clinical characteristics, but also adjust "where to focus" based on the dynamic changes in risk entropy, thereby achieving precise localization and feature enhancement of occult metastases.

[0045] In the clinical feature-driven weight generation network, in addition to the original scaling and bias parameter generation branches, a parallel branch is added to calculate the clinical risk entropy. The standardized clinical feature input vector is fed into the multilayer perceptron of the weight generation network, and a latent feature vector is obtained through nonlinear mapping. Subsequently, an additional fully connected layer maps the latent feature vector to a scalar value, and the scalar value is compressed to an open interval using a sigmoid activation function to obtain the clinical risk entropy. This risk entropy reflects the degree of uncertainty in the metastasis risk inferred based on clinical information such as patient age, tumor marker levels, differentiation degree, and T stage. Its calculation process is expressed by the following formula:

[0046] In the formula, Clinical risk entropy is used to quantitatively characterize the degree of uncertainty in the risk of lymph node metastasis inferred from a comprehensive understanding of a patient's clinical characteristics. This is the Sigmoid activation function, used to map input values ​​to the (0,1) interval; It is a learnable weight matrix used to linearly map the normalized latent feature vectors to scalar values; For layer normalization operations, the latent feature vectors are normalized. The mean and variance are normalized to stabilize the feature distribution; For learnable bias terms, and Together they constitute the linear mapping parameters.

[0047] The calculated clinical risk entropy As a control signal, the spatial response of the high-level semantic feature map in the image feature extraction network is adaptively weighted. Specifically, let the feature map output by a bottleneck residual block in the deep feature extraction stage be... First, max pooling and average pooling operations are performed on the feature map along the channel dimension to obtain two two-dimensional spatial descriptors. These two descriptors are then concatenated along the channel dimension and input into a convolutional layer composed of small-sized kernels to generate an initial spatial attention map. Subsequently, clinical risk entropy is applied as a dynamic scaling factor to this initial spatial attention map, and then normalized using the sigmoid function to obtain the final spatial attention mask.

[0048] In the formula, The resulting spatial attention mask has the same spatial dimension as the feature map to be modulated, and is used to perform element-wise weighting of the spatial dimension on the feature map; The initial spatial attention map is generated by concatenating the feature maps after performing max pooling and average pooling along the channel dimension, and then passing them through a convolutional layer with a small kernel. Its spatial dimension is consistent with the feature map. Figure 1To.

[0049] This design allows the response amplitude of the spatial attention mask to be enhanced when the clinical risk entropy is high, prompting the network to search more actively for potential small lesion areas.

[0050] After the original channel modulation operation is completed, the spatial attention mask is applied to the channel-modulated feature map in an element-wise multiplication manner. Let the feature map modulated by the channel scaling parameters and bias parameters be... The final output feature map after co-modulation for:

[0051] In the formula, For broadcast element-wise multiplication.

[0052] Thus, based on the direct modulation of clinical priors in the channel dimension, the feature map is further weighted by attention guided by clinical risk entropy in the spatial dimension, realizing the synergistic driving of clinical information on "what to focus on" and "where to focus on".

[0053] To verify the effectiveness of the multi-scale channel-spatial co-modulation mechanism based on clinical risk entropy, this embodiment compares it with a baseline dynamic modulation scheme that only uses channel modulation. The comparative experiment was conducted on the same independent test set, which included multimodal data from pathologically confirmed esophageal cancer patients, where the ratio of positive to negative lymph node metastasis samples remained consistent with the actual clinical distribution. During the experiment, both schemes used the same ResNet-50 backbone network architecture, the same training hyperparameter settings, and the same data preprocessing procedures. The only variable was whether or not a spatial attention co-modulation branch guided by clinical risk entropy was introduced. The comparative results show that after introducing the spatial co-modulation mechanism, the area under the receiver operating characteristic (AUC) curve of the model increased from 0.875 in the baseline scheme to 0.913; the sensitivity index increased from 84.3% to 89.7%; and the specificity index also increased from 86.1% to 90.2%. These performance gains were statistically significant. Further analysis revealed that the gain effect of the spatial co-modulation mechanism was particularly pronounced in the subgroup of small metastatic lesions with a short diameter of less than 10 mm in lymph nodes, increasing the detection rate from 67.5% in the baseline protocol to 78.9%. This result confirms that the clinical risk entropy-guided spatial attention mechanism effectively directs the network to focus on key anatomical regions in images associated with the clinical risk context, significantly enhancing the model's spatial localization accuracy and feature capture ability for occult lesions, thereby achieving a substantial improvement in diagnostic performance without increasing model complexity.

[0054] Example 3: Existing methods for training multimodal fusion models typically rely solely on the supervision signal from the final classification task, neglecting the structured correspondence between clinical features and image features in the representation space. This embodiment introduces a clinical state-guided cross-modal contrastive regularization constraint mechanism. During the training phase, for samples within a training batch, positive and negative sample pairs are constructed based on the similarity of their clinical feature vectors. Furthermore, the high-level semantic feature vectors generated from dynamically modulated images of the same patient are required to remain aligned with their own clinical state representation in the embedding space, while simultaneously being far removed from the image feature representations of patients with different clinical states. By imposing this contrastive constraint, the image feature extraction network is forced to learn more discriminative and modally semantically aligned robust representations under the guidance of clinical priors. This eliminates the residual effects of the modality gap at the feature level and improves the model's generalization ability under sparse clinical data conditions.

[0055] In each training batch of the model, let the batch contain a total of One sample, each sample This corresponds to a standardized clinical feature vector and a high-level image semantic feature vector extracted by a dynamic modulation network. First, for any two samples within a batch... and Calculate the cosine similarity between their clinical feature vectors. Subsequently, for each anchor point sample Define its positive sample set and negative sample set: If If the similarity exceeds a preset threshold, the sample will be... If it is included in the positive sample set; If the sample is less than another preset threshold, then the sample will be... It is included in the negative sample set.

[0056] Design a contrast loss function based on clinical similarity weighting This loss function aims to maximize the lower bound of the mutual information between the feature vector of the anchor sample image and its positive sample image, while minimizing the similarity between the feature vector of the anchor sample image and its negative sample image. The specific loss calculation expression is as follows:

[0057] In the formula, To optimize network parameters, the cross-modal comparison loss value is used as an auxiliary loss term and jointly optimized with the classification supervision loss. This represents the number of samples in the current training batch. This is the index number of the anchor sample within the batch; For anchor point samples The set of positive samples, consisting of samples within the batch. Samples whose clinical feature vector cosine similarity is greater than a preset upper threshold are constituted; For anchor point samples Semantic feature vectors of high-level images extracted by a dynamic modulation network; Positive sample The corresponding high-level image semantic feature vector; The temperature coefficient for clinical state perception is used to adjust the smoothness of the similarity distribution and control the sensitivity of contrastive learning to differences in similarity. It is the index of any sample in the union of the positive and negative sample sets; For anchor point samples The negative sample set, which consists of samples within the batch and those in the sample set. The sample group consists of all samples whose cosine similarity to the clinical feature vector is less than a preset threshold, representing a sample group whose clinical status is significantly different from that of the anchor sample. For the sample The corresponding high-level image semantic feature vector; This is a clinical similarity enhancement factor function; For anchor point samples Compared with positive samples The cosine similarity of the clinical feature vectors between the two is taken from -1 to 1. The larger the value, the more similar the clinical conditions of the two are. For anchor point samples With any sample The cosine similarity of clinical feature vectors between them.

[0058] The aforementioned cross-modal contrastive loss is jointly optimized with the original weighted cross-entropy classification loss. During backpropagation, the image feature extraction network and the dynamic modulation network are updated simultaneously, enabling them to collaboratively learn feature representations that possess both clinical discriminative power and cross-modal semantic consistency.

[0059] To evaluate the actual contribution of the clinical state-guided cross-modal contrastive regularization constraint mechanism, this embodiment designed an ablation comparison experiment. The experiment used a model version incorporating the clinical risk entropy co-modulation mechanism as a baseline, and tested the performance differences before and after adding the cross-modal contrastive regularization constraint. Regarding experimental settings, the partition ratios of the training, validation, and test sets, as well as the stratified sampling strategy, remained unchanged; the clinical state-perceived temperature coefficient in the contrastive loss function was set to 0.1, the amplification factor of the clinical similarity enhancement factor was set to 2.0, and the total loss balance hyperparameter was determined to be 0.3 through grid search. The comparison process recorded the complete performance metrics of the two models on the same test set. Experimental data showed that after introducing the cross-modal contrastive regularization constraint, the area under the receiver operating characteristic curve (AUC) of the model further increased from 0.913 to 0.927; the F1 score increased from 0.868 to 0.885; and, notably, the robustness of the model on a sparse subset of samples with incomplete clinical feature information was significantly enhanced, with the sensitivity index on this subset increasing from 81.2% to 86.4%. Furthermore, qualitative observation of the embedding space distribution of high-level semantic feature vectors through t-SNE dimensionality reduction visualization analysis revealed that, after adding contrast regularization constraints, the image feature vectors of samples from different lymph node metastasis states exhibited clearer inter-class separation boundaries and more compact intra-class aggregation characteristics in the embedding space. These quantitative and qualitative results together demonstrate that the clinical state-guided cross-modal contrast regularization mechanism can effectively force the image feature extraction network to learn robust representations consistent with the clinical semantic space, eliminating residual semantic gaps between modalities, and thus significantly improving the predictive reliability of the model in real-world application scenarios where clinical data is incomplete.

[0060] Example 4: In the aforementioned embodiments, the modulation operation of clinical priors was fixedly applied during the deep semantic feature extraction stage. However, the completeness and information complexity of clinical features vary significantly among different patients: for high-risk patients with rich and clearly defined clinical information, deep modulation is sufficient to guide the network to focus on key regions; while for complex cases with sparse clinical information or contradictory indications, relying solely on deep modulation may result in insufficient guidance strength of the clinical priors, necessitating the introduction of clinical modulation at a shallower feature extraction stage to provide supplementary prior constraints. Therefore, this embodiment proposes an adaptive modulation depth selection strategy based on clinical feature complexity. This strategy constructs a lightweight complexity evaluation subnetwork to evaluate the information completeness and determinism of the input clinical feature vector in real time, and dynamically determines the initial depth of the feature extraction network for clinical modulation operation intervention based on the evaluation results. Through this mechanism, the system can adaptively adjust the timing and intensity of modality fusion intervention according to the quality differences of individual patient clinical information, thereby maximizing the guiding effectiveness of clinical priors while ensuring computational efficiency.

[0061] Within the weight generation network, a clinical complexity evaluation branch is added in parallel with the main branch. This branch consists of a fully connected layer with a small number of neurons and a sigmoid activation function. The latent feature vector is input into this branch, and it outputs a complexity score within an open interval. This score reflects the degree of certainty of the information contained in the current input clinical feature vector: a higher score indicates more explicit and targeted clinical information; a lower score indicates sparser clinical information or contradictory indicators, requiring stronger prior guidance.

[0062] Based on the complexity score, a discretized modulation depth hierarchy decision logic is established. Specifically, three modulation depth levels are preset: deep modulation only, medium-deep modulation, and medium-shallow to deep modulation. The complexity score is divided into three corresponding numerical intervals, from high to low, corresponding to the three levels mentioned above. During each forward propagation, the starting bottleneck residual block group number for activation of the dynamic modulation module in the image feature extraction network is dynamically determined based on the interval into which the calculated complexity score falls. For example, when the complexity score is high, the modulation operation is activated only in the third and fourth bottleneck residual blocks, consistent with the aforementioned deep modulation strategy; when the complexity score is medium, the modulation operation extends to the second bottleneck residual block; when the complexity score is low, the modulation operation further extends to the first bottleneck residual block, to achieve comprehensive clinical guidance starting from a shallower layer.

[0063] To ensure the differentiability of the discrete decision-making process during backpropagation, the Gumbel-Softmax reparameterization technique is used to relax the modulation depth selection. During training, each modulation depth level is treated as a classification option, and a soft-weight vector with approximate one-hot encoding is generated through Gumbel-Softmax sampling. This vector is then used to weight and fuse the modulation outputs corresponding to each level. During inference, a hard decision is used to select the modulation depth level with the highest probability. This design ensures both the differentiability of end-to-end training and efficient execution during inference.

[0064] After determining the modulation initiation depth, the image feature extraction network performs channel-by-channel recalibration of the output of each bottleneck residual block group from that specified level up to the top layer, applying dynamic modulation parameters generated by the weight generation network. For shallow feature maps before the modulation initiation depth, the original output is retained to avoid unnecessary or misleading clinical interventions. Finally, the adaptively depth-modulated feature maps are fed into the metastasis prediction and classification module to predict the probability of lymph node metastasis.

Claims

1. A clinically dynamic modulation-based esophageal cancer lymph node metastasis prediction system, characterized in that, include: The data acquisition module is used to acquire enhanced CT image data and structured clinical data of the object to be predicted, and to perform standardized preprocessing on the data; A clinical feature-driven weight generation network is used to transform the standardized low-dimensional clinical feature vectors to a high-dimensional latent space through a nonlinear mapping in the hypernetwork architecture, and to generate two sets of channel-specific dynamic modulation parameters in parallel based on the high-dimensional latent space. An image feature extraction and dynamic modulation network is used to extract image feature maps with spatial hierarchical structure from standardized image data, and only at the deep feature extraction level corresponding to the tumor heterogeneity semantic parsing stage, the dynamic modulation parameters are used to perform channel-by-channel feature recalibration operation on the corresponding image feature maps to generate fusion features that embed clinical prior knowledge into the image representation process. The metastasis prediction and classification module is used to perform spatial dimensionality reduction and probability mapping on the fused features, and generate and output prediction results characterizing the risk of esophageal cancer lymph node metastasis.

2. The system according to claim 1, characterized in that: The structured clinical data includes continuous and discrete variables describing the patient's physiological state and tumor pathological attributes; the continuous variables are statistically standardized to eliminate dimensional differences, and the discrete variables are binarized and encoded to transform them into high-dimensional sparse vectors.

3. The system according to claim 1, characterized in that: The nonlinear mapping transformation constructed by the weight generation network, through the introduction of a learnable parameter matrix and a normalization mechanism, achieves adaptive adjustment of the response intensity of each feature channel in the deep feature extraction layer and baseline bias correction based on the differences in the individual patient's pathophysiological state.

4. The system according to claim 1, characterized in that: The dynamic modulation parameters include scaling parameters for controlling the intensity of information flow in the feature channels and bias parameters for adjusting the response benchmark of the feature channels. When generating the scaling parameters, the weight generation network introduces a nonlinear activation constraint mechanism to limit the range of numerical fluctuations, while retaining the numerical degrees of freedom of the linear mapping when generating the bias parameters.

5. The system according to claim 1, characterized in that: The image feature extraction and dynamic modulation network uses a deep convolutional network with a bottleneck residual structure as the feature extraction backbone. The deep feature extraction layer corresponds to the high-level semantic bottleneck residual block group in the deep convolutional network, which is responsible for parsing the heterogeneous texture and necrotic region features inside the tumor. The shallow bottleneck residual block group, which extracts general visual features, is excluded from the intervention range of the dynamic modulation parameters.

6. The system according to claim 5, characterized in that: The feature recalibration operation is configured at the end of each residual unit in the high-level semantic bottleneck residual block group, and the specific embedding position is after the last convolution operation and normalization operation in the unit, and before the nonlinear activation operation and residual path superposition operation.

7. The system according to claim 1, characterized in that: The weight generation network further includes a clinical uncertainty measurement branch, used to generate a scalarized clinical risk indicator based on the high-dimensional latent space representation; the image feature extraction and dynamic modulation network further includes a spatial attention generation unit, which receives the scalarized clinical risk indicator as a dynamic adjustment signal, and generates a spatial dimension attention-weighted mask for the feature map output by the deep feature extraction layer based on the signal, so as to perform secondary enhancement modulation of the spatial dimension on the feature map after channel recalibration.

8. The system according to claim 1, characterized in that: The transfer prediction and classification module receives the top-level image feature map after the feature recalibration operation, and uses the global spatial average pooling mechanism to compress the feature map into a one-dimensional feature descriptor. The feature descriptor condenses the deep image semantic information after being weighted and enhanced by the dynamic modulation parameters.

9. The system according to claim 1, characterized in that: During the parameter optimization phase, the system employs an adaptive weighting strategy based on prior class distribution to construct a loss function. This strategy determines the cost-sensitive weight factors corresponding to the positive and negative classes based on their inverse frequencies in the training samples.

10. A method for predicting esophageal cancer lymph node metastasis based on clinical dynamic modulation, characterized in that: The method is implemented based on the system described in any one of claims 1-9: The method includes: Acquire enhanced CT images and structured clinical data of the subjects to be predicted, and perform standardized preprocessing on each. Standardized clinical data is input into a weighted supernetwork, which generates a high-dimensional latent space representation from the low-dimensional clinical feature space through a nonlinear mapping mechanism, and derives the scaling and bias calibration values ​​for the image feature channels in parallel based on the representation. Standardized image data is input into a deep feature extraction network to extract image feature map sequences with spatial hierarchical structure. Only in the deep feature extraction stage corresponding to tumor heterogeneity semantic parsing, the scaling control amount and bias calibration amount are used to perform channel-wise feature response recalibration operation on the extracted feature maps to obtain fused feature maps modulated by clinical prior depth. Global spatial dimensionality reduction and nonlinear decision mapping are performed on the fused feature map to generate and output predicted values ​​representing the probability of esophageal cancer lymph node metastasis.