A medical image prediction model construction system adaptive to physiological characteristics of children
Patent Information
- Application Number
- CN202610743196.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-27
- Publication Date
- 2026-08-21
AI Technical Summary
首先,儿童生理发育尚未成熟,器官形态、代谢水平、生理功能等与成人存在显著差异,而现有医学影像预测模型多基于成人数据训练,缺乏对儿童生理特征的针对性适配,导致模型在儿童影像分析中预测准确性不足
1、本系统针对儿童器官发育、代谢水平等生理特性,通过数据获取、预处理及特征提取模块的针对性设计,实现了对儿童核医学影像数据的精准分析。结合浅层残差卷积网络与深层非局部操作、多头自注意力机制,有效捕捉儿童影像中局部代谢特征与全局空间依赖关系,使复杂病理特征的识别能力显著提升,大幅提高了儿童疾病预测的准确率。
Smart Images

Figure CN122619286A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical imaging system technology, specifically to a medical image prediction model construction system adapted to children's physiological characteristics. Background Technology
[0002] Nuclear medicine imaging technology has significant application value in the diagnosis, monitoring, and treatment evaluation of pediatric diseases, but its clinical application still faces several technical bottlenecks. First, children's physiological development is not yet mature, and their organ morphology, metabolic levels, and physiological functions differ significantly from adults. Existing medical imaging prediction models are mostly trained on adult data, lacking specific adaptation to children's physiological characteristics, resulting in insufficient predictive accuracy in pediatric image analysis. Second, the fusion of multimodal nuclear medicine images (such as PET and SPECT) and clinical data has limitations. Traditional fusion methods use fixed weight allocation strategies, failing to dynamically adjust weights based on the correlation between modal features and the prediction target. This leads to key modal information being masked by redundant information, affecting the quality of feature representation. Third, traditional convolutional neural networks have shortcomings in feature capture, struggling to simultaneously and accurately extract local metabolic features (such as lesion edge details) and global spatial dependencies (such as cross-organ functional correlations), limiting their ability to identify complex pathological features. Furthermore, cross-center medical data collaborative modeling faces a contradiction between privacy protection and data sharing. Data from different medical centers is scattered and involves patient privacy; directly sharing data poses ethical and security risks. On the other hand, individual modeling is limited by insufficient sample size, resulting in weak model generalization ability. Finally, existing models have slow inference speeds and offer limited presentation of prediction results, lacking dynamic visualization and structured diagnostic suggestions, making it difficult to meet the needs of real-time clinical auxiliary diagnosis and hindering the clinical translation and application of the technology. Summary of the Invention
[0003] To address the shortcomings of existing technologies, this invention provides a medical image prediction model construction system adapted to children's physiological characteristics, thus solving the problems mentioned in the background technology.
[0004] To achieve the above objectives, the present invention provides the following technical solution: a medical image prediction model construction system adapted to children's physiological characteristics, comprising: The data acquisition module is used to acquire multimodal nuclear medicine imaging data and clinical data, including an intelligent quality control unit and an image acquisition unit; The data preprocessing module is used to perform image calibration, denoising, normalization and multimodal temporal registration on the nuclear medicine image data, including a temporal registration unit and a modal alignment unit. The feature extraction module is used to extract multimodal features from preprocessed image data, including a functional connectivity dynamic analysis unit, a spatial feature extraction unit, and a metabolic trajectory analysis unit. The feature fusion module is used to fuse extracted multimodal features through weighted combination, including cross-modal gating mechanisms; The model building module is used to build predictive models using fused features and corresponding clinical data, including a hierarchical feature interaction network. The loss function module is used to calculate the loss of the prediction model, including the adversarial regularization term and the temporal smoothing constraint term; The model evaluation module is used to evaluate the performance of the prediction model, including a time-dependent evaluation unit and a cross-center generalization evaluation unit. The model application module is used to apply the prediction model to the nuclear medicine imaging data of new patients, and to provide prediction results and auxiliary diagnostic suggestions, including a diagnostic suggestion unit, a real-time inference acceleration unit and a result visualization unit. The Federated Learning Collaboration Module supports joint modeling of nuclear medicine data across centers, employing secure multi-party computation and homomorphic encryption to protect privacy data.
[0005] Preferably, the intelligent quality control unit automatically detects image artifacts using a convolutional neural network and calculates the artifact detection loss using a triplet loss function: ; in, For anchor samples, For artifact-free positive samples, For negative samples with artifacts, For interval parameters; The intelligent quality control unit also includes a data cleaning subunit, which automatically marks image data that is abnormal in artifact detection and triggers a manual review process.
[0006] Preferably, the time-series registration unit of the data preprocessing module adopts a time-series registration algorithm based on mutual information, which achieves spatiotemporal alignment by maximizing the mutual information entropy between time series. The formula for calculating the time-series mutual information is: ; in, , The floating image and reference image sequence at time t. For time frames; and Representing single images and The information entropy reflects the uncertainty of pixel grayscale distribution; Representing two images and The joint information entropy reflects the common uncertainty of the gray-scale distribution of both. The modal alignment unit of the data preprocessing module is used to spatially register PET and SPECT images, achieving intermodal coordinate alignment by minimizing the Euclidean distance loss function. ; in, , For the first The coordinates of the registration points in PET and SPECT images This represents the total number of registration points. The denoising unit of the data preprocessing module employs a statistical noise correction algorithm for PET images and an iterative reconstruction denoising algorithm for SPECT images.
[0007] Preferably, the functional connectivity dynamic analysis unit extracts functional connectivity features by calculating the dynamic correlation coefficients of metabolic time series from different brain regions. The formula for the dynamic correlation coefficients is: ; The core dependent variable, i.e. Time of the first The region and the first The dynamic correlation coefficient of each region is directly output; the dynamic correlation strength between two regions is directly output, with a value range of [-1, 1]. For the covariance term, the core calculation unit, measures the first... Lagging in some areas The signal at time and the first each region The degree of linear correlation between signals at different times; For the first Each image area in The time-series signal value at time 10:00. Represents the signal strength of the region. It is the area code. It is the current moment. It is the time lag; For the first Each image area in The time-series signal value at time 10:00. Is with Different area codes, the rest are the same. ; as a reference signal, and Paired calculation of covariance constitutes the core analytical object for the dynamic correlation between two regions; For the first The standard deviation of the time-series signal in each region; measuring the standard deviation of the first region's time-series signal; The degree of dispersion of all time-series signals in each region is used to normalize the covariance, eliminate the interference caused by differences in signal amplitude, and make the correlation coefficients comparable. For the first The standard deviation of the time series signal in each region; and Consistency is used to normalize covariance and ensure The values are uniformly set in [-1, 1] to avoid deviations in correlation calculation due to different signal amplitudes in different regions; It represents the time lag; it adjusts the lag correlation of timing signals. The cross-modal gating mechanism uses a gating vector. The fusion weights of different modal features are dynamically adjusted, and the gating vector is calculated as follows: ; in, For the first Modal features, It is the Sigmoid activation function. , These are trainable parameters; The cross-modal gating mechanism also includes an attention mechanism, which dynamically adjusts the intermodal mutual information gain by calculating the intermodal gating gain. The weight value is calculated using the following formula: ; in, For the first Modal features and predicted labels Mutual information value; The feature fusion module also includes a feature normalization unit, which is used to standardize the fused features, using the Z-score normalization method. ; in, As a feature of fusion, The characteristic mean, The characteristic standard deviation is denoted as .
[0008] Preferably, the hierarchical feature interaction network includes a shallow local feature extraction layer and a deep global interaction layer: The shallow layer uses 3×3×3 convolutional kernels to capture local metabolic features and introduces residual connection structures: ; Deep learning models global spatial dependencies through non-local operations, as shown in the formula: ; in, For the first The feature vectors of each spatial location are updated by nonlocal operations and used to fuse the original local features and global spatial dependency information to achieve a collaborative representation of local metabolic features and global association information. For the first The original input feature vectors of each spatial location; These are the normalization coefficients; The number of sampling points for non-local operations; This is a similarity function used to calculate the reference position. eigenvectors With the target location eigenvectors The strength of the correlation between them; The deep global interaction layer employs a multi-head self-attention mechanism, which computes multiple attention heads in parallel to capture feature dependencies in different subspaces.
[0009] Preferably, the adversarial regularization term enhances the model's generalization ability through a generative adversarial network, and the adversarial loss formula is: ; in, For discriminator, For generator, For the true data distribution, Noise distribution; The temporal smoothing constraint term uses a first-order difference loss to force the prediction results of adjacent time points to be continuous, and the formula is as follows: ; in, The predicted value at time t; The total loss function of the loss function module is a weighted sum of adversarial loss, temporal smoothing loss, and task loss: ; in, , For weight parameters, A task-specific loss function.
[0010] Preferably, the global model update formula of the federated learning collaborative module is: ; in, For the first Local model parameters at the center, This represents the number of local samples. The total number of global samples supports joint modeling of heterogeneous data sources and aligns differences in image acquisition parameters between different centers through domain adaptation technology. The federated learning collaboration module supports the dynamic addition of new central nodes and achieves incremental updates of the global model through a secure aggregation protocol.
[0011] Preferably, the model evaluation module further includes a feature importance analysis unit, which quantifies the contribution of each modal feature to the prediction result based on the SHAP value; The time-dependent evaluation unit uses a time-dependent ROC curve to evaluate dynamic prediction performance. It calculates the AUC value by integrating the prediction probabilities at different time points, as shown in the following formula: ; in, express The area under the ROC curve at time t; The cross-center generalization evaluation unit uses leave-one-center cross-validation to evaluate the model's predictive performance on data with unknown centers within the federated learning framework.
[0012] Preferably, the real-time inference acceleration unit employs knowledge distillation technology to compress complex models into lightweight models, with the distillation loss function being: ; Where CE is the cross-entropy loss and MSE is the mean squared error loss. and These are the predicted probabilities for the student model and the teacher model, respectively. and These are the outputs of the feature layer; The real-time inference acceleration unit also employs model quantization compression technology and a hardware acceleration engine to control the image prediction latency to within 200ms while maintaining a prediction accuracy loss of ≤2%.
[0013] Preferably, the result visualization unit supports dynamic time-series curve display function, which can draw the changing trend of the indicators of interest over time in real time and overlay clinical event markers for reference; The auxiliary diagnostic suggestion unit automatically generates a structured report based on the prediction results. The structured report includes lesion location information, prediction probability, and related treatment suggestion links, and supports seamless integration with the hospital information system.
[0014] This invention provides a medical image prediction model construction system adapted to children's physiological characteristics, which has the following beneficial effects: Through innovative multi-module design, it significantly improves the accuracy, efficiency, and cross-center collaboration capabilities of nuclear medicine image prediction, with the following specific advantages: 1. This system, tailored to the physiological characteristics of children's organ development and metabolic levels, achieves precise analysis of pediatric nuclear medicine imaging data through targeted design of data acquisition, preprocessing, and feature extraction modules. By combining shallow residual convolutional networks with deep nonlocal operations and multi-head self-attention mechanisms, it effectively captures local metabolic features and global spatial dependencies in pediatric images, significantly improving the ability to identify complex pathological features and greatly enhancing the accuracy of pediatric disease prediction.
[0015] 2. A dynamic weighted fusion strategy combining cross-modal gating and attention mechanisms is adopted. By adjusting the gating vector and modality-label mutual information gain, the fusion weights of different modal features are dynamically allocated. This solves the problem that traditional fixed-weight fusion cannot highlight the contribution of key modalities, significantly improves the contribution of core features, and enhances the effectiveness and relevance of feature representation.
[0016] 3. A federated learning collaboration module is introduced, employing secure multi-party computation, homomorphic encryption, and domain adaptation techniques to achieve joint modeling of heterogeneous data sources while ensuring the privacy of data from each medical center is not leaked. It supports the dynamic addition of new center nodes and incremental updates to the global model, effectively solving the problem of insufficient sample size in single centers and significantly improving the model's generalization performance in cross-center scenarios.
[0017] 4. By integrating multiple technologies such as knowledge distillation, model quantization compression, and hardware acceleration engines, the image prediction latency is controlled within 200ms, while ensuring that the prediction accuracy loss is ≤2%, meeting the needs of real-time clinical diagnosis. The results visualization unit supports dynamic time-series curve display and clinical event marking, and the auxiliary diagnostic suggestion unit automatically generates structured reports containing lesion localization, prediction probability, and treatment suggestion links. It can also seamlessly integrate with hospital information systems, greatly improving the convenience and practicality of clinical applications.
[0018] 5. The intelligent quality control unit automatically detects image artifacts and triggers manual review through convolutional neural networks to ensure the quality of input data; the model evaluation module combines time-dependent ROC curves, leave-one-out cross-validation across centers, and SHAP value feature importance analysis to evaluate the model from multiple dimensions such as dynamic prediction performance, generalization ability, and feature contribution, thus comprehensively ensuring the reliability and stability of the model. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the principle of a medical image prediction model construction system adapted to children's physiological characteristics, as described in this invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] like Figure 1 As shown, the present invention provides a technical solution: a medical image prediction model construction system adapted to children's physiological characteristics, comprising: a data acquisition module, a data preprocessing module, a feature extraction module, a feature fusion module, a model construction module, a loss function module, a model evaluation module, and a federated learning collaboration module; The data acquisition module acquires multimodal nuclear medicine imaging data and clinical data, including an intelligent quality control unit and an image acquisition unit. The data preprocessing module performs image calibration, denoising, normalization, and multimodal temporal registration on the nuclear medicine imaging data, including a temporal registration unit and a modality alignment unit. The feature extraction module extracts multimodal features from the preprocessed image data, including a functional connectivity dynamic analysis unit, a spatial feature extraction unit, and a metabolic trajectory analysis unit. The feature fusion module fuses the extracted multimodal features through a weighted combination method, including a cross-modal gating mechanism. The model building module utilizes the fused features and corresponding... The system constructs predictive models from clinical data, including hierarchical feature interaction networks; a loss function module calculates the loss of the predictive model, including adversarial regularization and temporal smoothing constraints; a model evaluation module evaluates the performance of the predictive model, including time-dependent evaluation and cross-center generalization evaluation units; a model application module applies the predictive model to nuclear medicine imaging data of new patients, providing predictive results and auxiliary diagnostic suggestions, including diagnostic suggestion units, real-time inference acceleration units, and result visualization units; and a federated learning collaboration module supports joint modeling of cross-center nuclear medicine data, employing secure multi-party computation and homomorphic encryption to protect privacy data.
[0022] More specifically, the intelligent quality control unit automatically detects image artifacts using a convolutional neural network and calculates the artifact detection loss using a triplet loss function: ; in, Anchor samples (arbitrary image samples, serving as reference points in the feature space). This is a positive sample without artifacts (an image sample of the same type without artifacts, compared to...). Belonging to the same category (no artifacts) For negative samples with artifacts (image samples with artifacts, compared to...) Belongs to different categories (with artifacts) The margin parameter (used to control the minimum margin between positive and negative samples in the feature space, usually set to 0.5~1.0, and the optimal value is determined through cross-validation). It is a distance metric between feature vectors.
[0023] The intelligent quality control unit also includes a data cleaning subunit, which automatically marks image data that is abnormal due to artifact detection and triggers a manual review process.
[0024] Convolutional Neural Network Architecture: Input nuclear medicine image data (such as PET and SPECT images) is normalized and then input into the CNN; a lightweight convolutional neural network (such as a ResNet-18 variant) is used, which includes multiple 3×3 convolutional layers, batch normalization layers and ReLU activation functions, and finally outputs feature vectors through fully connected layers; through triplet loss function optimization, the network can effectively distinguish between artifact-free samples and artifact-containing samples; Automatic detection and labeling: The CNN performs forward inference on the input image and outputs artifact detection results (such as "normal" or "abnormal"). For image data detected as "abnormal", the data cleaning subunit automatically adds labels (such as recording the "artifact suspicious" label in the metadata).
[0025] Manual review trigger mechanism: The system generates a list of abnormal data and pushes it to human quality control personnel via message queue or visual interface. Human quality control personnel review the images marked as "abnormal" to confirm whether artifacts exist: if artifacts are confirmed, the data is marked as "invalid" and excluded from the training or prediction set; if it is confirmed as a false positive, the data is updated to "normal" and fed back to the model to optimize training.
[0026] More specifically, the temporal registration unit of the data preprocessing module adopts a mutual information-based temporal registration algorithm, which achieves spatiotemporal alignment by maximizing the mutual information entropy between time series. The formula for calculating temporal mutual information is: ; in, , The floating image and reference image sequence at time t. For time frames; The modal alignment unit in the data preprocessing module is used to spatially register PET and SPECT images, achieving intermodal coordinate alignment by minimizing the Euclidean distance loss function. ; in, , For the first The coordinates of the registration points in PET and SPECT images This represents the total number of registration points. The denoising unit of the data preprocessing module uses a statistical noise correction algorithm for PET images and an iterative reconstruction denoising algorithm for SPECT images.
[0027] The temporal registration unit (temporal and spatial alignment based on mutual information) solves the problem of temporal and spatial offset caused by organ movement (such as breathing and heartbeat) in nuclear medicine image time series (such as dynamic PET scans), ensuring that images at different time points are spatially aligned.
[0028] Temporal mutual information: a measure of floating image sequences With reference image sequence The statistical correlation between the two is as follows: The larger the temporal mutual information value, the more similar the spatial structure and gray-level distribution of the two, and the more accurate the registration.
[0029] The formula is: ; in: and Representing single images and The information entropy reflects the uncertainty of pixel grayscale distribution.
[0030] Representing two images and The joint information entropy reflects the common uncertainty of the gray-scale distribution of both.
[0031] Maximizing mutual information: By adjusting spatial transformation parameters (such as translation and rotation), the sum of mutual information of all time frames is maximized, thus achieving spatiotemporal alignment.
[0032] Modal alignment units (PET and SPECT spatial registration) resolve spatial coordinate deviations caused by different imaging principles in multimodal images (such as PET metabolic images and SPECT functional images), ensuring alignment of anatomical structures.
[0033] Euclidean distance loss function: loss function The calculation formula is: ; Specifically, by manually or automatically selecting corresponding anatomical points (such as the pituitary gland and heart valves) in PET and SPECT images, the sum of squares of the differences in registration point coordinates is calculated, and the loss value is minimized by optimizing rotation and translation parameters to achieve intermodal spatial alignment.
[0034] The denoising unit (modal noise suppression) reduces statistical noise in nuclear medicine images (such as Poisson noise in PET and low count noise in SPECT), improving the reliability of subsequent feature extraction.
[0035] Modal strategies: PET images employ statistical noise correction algorithms (such as Maximum A posteriori probability estimation, MAP), which, based on the Poisson distribution characteristics of noise, suppresses noise through iterative optimization while preserving details in high-metabolic regions such as tumors. SPECT images employ iterative reconstruction denoising algorithms (such as Ordered Subset Expectation Maximization, OSEM), which reduce random noise and enhance the visibility of low-contrast structures by fitting projection data through multiple iterations.
[0036] Data preprocessing module workflow: Temporal registration stage: Input a floating image sequence with multiple time frames and reference image sequence (e.g., different respiratory phase images from dynamic PET); initialize spatial transformation parameters (e.g., translation vector, rotation matrix); iteratively adjust parameters using gradient descent or optimization algorithms to maximize... This aligns the floating image with the reference image in space and time; it outputs the space-time aligned image sequence. The spatial positions of each time frame are consistent.
[0037] Modal alignment stage: Input spatiotemporally aligned PET images and original SPECT images (or other modalities). Select at least 10 anatomical registration points with the same name (e.g., automatically extracted using feature detection algorithms); calculate the coordinate differences of the registration points in PET and SPECT, and optimize the rigid transformation parameters (rotation) using the least squares method. Peaceful relocation Minimize Euclidean distance loss The output spatially aligned PET and SPECT composite images have the same coordinates for corresponding anatomical structures.
[0038] Noise reduction stage: Input spatiotemporally and modally aligned PET and SPECT images; Modal processing: PET images are processed using statistical noise correction algorithms, such as iteratively solving for noise-free signals through maximum a posteriori probability estimation to suppress Poisson noise; SPECT images are reconstructed using the OSEM algorithm, inferring tomographic images from projection data to reduce low-count noise. Output high-quality images after noise reduction, with noise levels reduced by 30%-60% and improved clarity of key structures (such as lesions).
[0039] More specifically, the functional connectivity dynamic analysis unit extracts functional connectivity features by calculating the dynamic correlation coefficients of metabolic time series from different brain regions. The formula for the dynamic correlation coefficient is: ; The core dependent variable, i.e. Time of the first The region and the first The dynamic correlation coefficient of each region; directly outputs the dynamic correlation strength between two regions, with a value range of [-1, 1]. Positive values indicate positive correlation (synchronous changes), and negative values indicate negative correlation (inverse changes). The larger the absolute value, the stronger the correlation, which is used to characterize the functional connectivity state of brain regions. For the covariance term, the core calculation unit, measures the first... Lagging in some areas The signal at time and the first each region The degree of linear correlation between signals at any given time reflects the consistency of the changing trends of the two signals and is the basis for calculating the correlation coefficient (eliminating the influence of the absolute magnitude of the signals). For the first A specific imaging region (such as a specific brain region in nuclear medicine PET / SPECT images) in The timing signal value at a given moment; its meaning: Represents the signal intensity of the region (such as metabolic concentration, image gray value). It is the area code. It is the current moment. It is a time lag; its function is to capture the first... The temporal dynamic characteristics of each region are introduced by hysteresis. It can reflect the temporal correlation of signals (such as the temporal transmission of pathological changes); For the first Each image area in The timing signal value at a given moment; its meaning: Is with Different region numbering (which can be understood as another brain region / image region for paired analysis), the rest are the same. ; as a reference signal, and Paired calculation of covariance constitutes the core analytical object for the dynamic correlation between two regions; For the first The standard deviation of the time-series signal in each region; measuring the standard deviation of the first region's time-series signal; The dispersion (signal fluctuation magnitude) of all time-series signals in each region is used to normalize the covariance, eliminate the interference caused by signal amplitude differences, and make the correlation coefficients comparable. For the first The standard deviation of the time series signal in each region; and Consistency is used to normalize covariance and ensure The values are uniformly set in [-1, 1] to avoid deviations in correlation calculation due to different signal amplitudes in different regions; This is the time lag (implied in the formula, a key parameter); it adjusts the lag correlation of temporal signals and can be set according to clinical data (such as the cycle of pathological changes) to capture the dynamic temporal characteristics of functional connectivity between brain regions (such as signal transduction delay). Cross-modal gating mechanisms use gating vectors The fusion weights of different modal features are dynamically adjusted, and the gating vector is calculated as follows: ; in, For the first Modal features, It is the Sigmoid activation function. , These are trainable parameters; Cross-modal gating mechanisms also include attention mechanisms, which dynamically adjust the mutual information gain between modalities by calculating the mutual information gain. The weight value is calculated using the following formula: ; in, For the first Modal features and predicted labels Mutual information value; The feature fusion module also includes a feature normalization unit, which is used to standardize the fused features, employing the Z-score normalization method. ; in, As a feature of fusion, The characteristic mean, The characteristic standard deviation is denoted as .
[0040] Functional connectivity dynamic analysis unit: captures the temporal dependence of metabolic activities in different brain regions and reveals the dynamic characteristics of brain functional networks.
[0041] formula: ; Dynamic correlation coefficient (DCC): measures brain regions and Time-varying correlation of metabolic signals between them; Time delay parameter ( ): To capture causal relationships between brain regions (such as the delayed response of brain region B after activation of brain region A), usually 1-3 time steps are taken; Positive DCC indicates coactivation (such as synchronized activity between the visual cortex and the occipital lobe), while negative DCC indicates functional antagonism (such as mutual inhibition between the default network and the task activation network).
[0042] Cross-modal gating mechanism: dynamically integrates multimodal features (such as PET metabolic features and MRI structural features) to avoid information redundancy or loss caused by fixed weights.
[0043] Gating vector ( ): ; Through the Sigmoid function By mapping features to the [0,1] interval, the information flow of mode m is adaptively controlled.
[0044] Attention weights ( ): ; Weights are dynamically assigned based on the mutual information (MI) between modal features and predicted labels. For example, if PET features are more critical for disease prediction, then... When the value approaches 1, the weights of other modalities decrease.
[0045] Feature normalization unit: eliminates the dimensional differences between features of different modalities and ensures the numerical stability of fused features.
[0046] Z-score formula: ; Standardization process: Integrating features Scaling to a distribution with a mean of 0 and a standard deviation of 1 prevents gradient vanishing / exploding and improves model training efficiency.
[0047] Work process: Functional connectivity feature extraction: Input preprocessed brain region segmentation images (such as PET metabolic maps), with each brain region corresponding to a time series. For each brain region Calculate different time delays The DCC values are used to construct a dynamic correlation matrix; statistical features of the matrix (such as mean DCC, standard deviation, and clustering coefficient) are extracted to quantify the functional network characteristics. The dynamic functional connectivity feature vector is output to reflect the temporal changes of the brain network.
[0048] Cross-modal feature fusion: Input multimodal features (such as functional connectivity features, spatial metabolic features); through trainable parameters Calculate the gate vector Suppress redundant modal information; calculate the mutual information between each modality and the predicted label. Obtain attention weights Weighted fusion, It combines gating and attention mechanisms to output an adaptively fused multimodal feature vector.
[0049] Feature normalization: Input fused features ; Calculate the mean of the features in the training set. and standard deviation Applying Z-score transform: ; Output the standardized feature vectors for subsequent model training.
[0050] Through time delay parameter This approach captures causal relationships between brain regions, outperforming traditional static functional connectivity analysis. A gating mechanism suppresses noisy modes through a data-driven approach, improving feature quality. An attention mechanism automatically assigns modality weights based on mutual information, highlighting the most critical information for prediction (such as tumor metabolic features). Z-score normalization ensures scale consistency across features from different sources, accelerating model convergence and enhancing generalization ability. In Alzheimer's disease diagnosis, this process can integrate PET metabolic features with MRI structural features, discovering early brain network abnormalities through dynamic functional connectivity. Adaptive weighting ensures that metabolic features dominate in the later stages of the disease, while structural features are more important in the early stages, ultimately improving diagnostic accuracy.
[0051] More specifically, the hierarchical feature interaction network includes shallow local feature extraction layers and deep global interaction layers: The shallow layer uses 3×3×3 convolutional kernels to capture local metabolic features and introduces residual connection structures: ; Deep learning models global spatial dependencies through non-local operations, as shown in the formula: ; in, For the first The feature vectors of each spatial location are updated by nonlocal operations and used to fuse the original local features and global spatial dependency information to achieve a collaborative representation of local metabolic features and global association information. For the first The original input feature vector (features of corresponding voxels or pixels in nuclear medicine images) at each spatial location is used as a residual connection term to preserve the local detail features at that location and avoid losing key local metabolic information when aggregating global information. The normalization coefficient is used to stabilize the numerical scale of the global aggregation term, avoid distortion of the aggregation result due to differences in the number of sampling points or feature amplitude, and ensure that the feature values are within a reasonable range that the model can train. The number of sampling points for nonlocal operations, i.e. the total number of reference locations participating in global dependency calculations, is used to efficiently capture global spatial dependencies while controlling computational complexity, and is adapted to large-scale voxel data of nuclear medicine 3D images. This is a similarity function used to calculate the reference position. eigenvectors With the target location eigenvectors The stronger the correlation between them, the higher the corresponding reference position. For the target location The greater the feature update contribution weight; The deep global interaction layer employs a multi-head self-attention mechanism, which computes multiple attention heads in parallel to capture feature dependencies in different subspaces.
[0052] Shallow local feature extraction layers capture local metabolic features in nuclear medicine images (such as high-metabolic regions at the tumor margin and subtle changes in organ boundaries); 3×3×3 convolution kernels, compared to larger convolution kernels (such as 5×5×5), can preserve spatial details more finely and reduce information blurring. 3D convolution is suitable for processing voxel-level image data (such as the three-dimensional volume of PET / SPECT), capturing features in both spatial and channel dimensions.
[0053] Residual connection structure: ; The output feature vector of the shallow local feature extraction layer is used to fuse the local metabolic features extracted by 3D convolution with the original input features, so as to achieve the collaborative representation of local details and convolution features, and provide an accurate local feature basis for deep global interaction. For 3D convolution operations, a 3×3×3 convolution kernel is used to extract local features from the input features. Its core function is to capture local metabolic features in nuclear medicine images (such as lesion edge details and differences in local tissue metabolic intensity), adapting to the feature extraction needs of three-dimensional voxel-level image data. The original input feature vector of the shallow local feature extraction layer corresponds to the preprocessed nuclear medicine image voxel features (such as denoised and registered PET / SPECT image features), and is used as the original feature term of the residual connection to preserve low-level local details in the input data, avoid the loss of key local information caused by convolution operations, and alleviate the gradient vanishing problem in deep network training.
[0054] Skip connections mitigate the vanishing gradient problem, ensuring the network can learn subtle local differences. This enhances feature reuse capabilities by preserving low-level information (such as texture and edges) from the original input. Deep global interaction layer: Modeling long-distance spatial dependencies (such as functional connections between different brain regions and spatial correlations of systemic tumor metastases).
[0055] Non-local operations: ; Computational logic, similarity function Calculate position With all other positions The degree of correlation; Feature transformation : Reduce computational complexity by compressing feature dimensions through linear mapping; Weighted aggregation: The features at all locations are summed in weights based on their relevance, and the locations are updated. Features; Advantages: It can directly capture the dependency relationship between any two points, break through the limitations of the local receptive field of CNN, and is suitable for capturing diffuse lesions (such as brain atrophy in Alzheimer's disease).
[0056] Multi-head self-attention mechanism: ; Each head: ; Function: To simultaneously focus on the dependencies between different subspaces (such as metabolic intensity, spatial distribution, and temporal variations). To enhance the model's ability to express complex patterns, such as simultaneously capturing intratumoral heterogeneity and the relationship between the tumor and surrounding tissues. Through the design of the hierarchical feature interaction network described above, local and global features can be effectively extracted and fused from nuclear medicine images, improving the accuracy and reliability of disease diagnosis.
[0057] More specifically, the adversarial regularization term enhances the model's generalization ability through generative adversarial networks, and the adversarial loss formula is: ; in, For discriminator, For generator, For the true data distribution, Noise distribution; The temporal smoothing constraint term uses first-order difference loss to force the prediction results of adjacent time points to be continuous. The formula is as follows: ; in, The predicted value at time t; The total loss function of the loss function module is a weighted sum of the adversarial loss, temporal smoothing loss, and task loss: ; in, , For weight parameters, A task-specific loss function.
[0058] Adversarial regularization (Generative Adversarial Networks - GANs): Enhances the model's ability to fit the real data distribution through adversarial training, improves generalization ability, and avoids overfitting.
[0059] Generator (G): Input noise (follows distribution) ), generating data distributions that closely approximate reality. "Pseudo-data" ; Discriminator (D): Determines whether the input data is real data. Or generate data Output probability value or ; Adversarial loss function: ; Training objectives: For discriminator D: Maximize That is, to distinguish between real data and pseudo data as accurately as possible; For generator G (implicit optimization): Minimize That is, the generated data should be as close as possible to the real data, so that D cannot distinguish them; Through the game process, the model is forced to learn the potential distribution characteristics of real data, thereby improving its robustness to edge cases such as noise and low-dose images.
[0060] Temporal smoothing constraint (first-order difference loss): ensures that the prediction results of dynamic nuclear medicine images (such as cardiac-gated PET) are continuous in time series and conform to the temporal consistency of physiological processes; formula: ; in, : Predicted values at any given time (e.g., lesion metabolic activity score).
[0061] By penalizing abrupt changes in predicted values at adjacent time points, the prediction results are forced to change smoothly over time (e.g., tumor growth rate should not jump suddenly); this is suitable for dynamic imaging sequences that require analysis of disease progression or treatment response, avoiding prediction result oscillations caused by noise.
[0062] Total loss function (weighted combination) formula: ; illustrate: Task-specific losses (such as cross-entropy loss for classification tasks and mean squared error for regression tasks). Weight parameters (usually determined through cross-validation, such as...) ), used to balance the importance of different loss terms.
[0063] More specifically, the global model update formula for the federated learning collaborative module is: ; in, For the first Local model parameters at the center, This represents the number of local samples. The total number of global samples supports joint modeling of heterogeneous data sources and aligns differences in image acquisition parameters between different centers through domain adaptation technology. The federated learning collaboration module supports the dynamic addition of new central nodes and achieves incremental updates of the global model through a secure aggregation protocol.
[0064] Global model update mechanism: Without sharing the original data, it aggregates local model parameters from multiple medical centers to construct a globally optimal model, thus protecting patient privacy. Formula: ; Parameter description: : No. Local model parameters for each medical center (such as convolutional kernel weights and bias terms); The number of samples at the center (e.g., the number of PET scan cases at a hospital); Total number of samples from all participating centers.
[0065] Weighted aggregation logic: Centers with larger sample sizes contribute more to the global model, ensuring the model is biased towards distributions with larger data volumes and avoiding overfitting to small centers. This mechanism assigns weights by calculating the proportion of samples in each center, thus aggregating to obtain a more robust global model.
[0066] Differences in equipment parameters (such as PET scanner model and radioactive tracer type) between different medical centers may lead to shifts in image feature distribution and reduce model generalization. By employing adversarial training or the maximum mean difference (MMD) method, the distance between different central feature distributions is minimized. These methods aim to reduce feature bias caused by differences in devices or tracers, enabling the model to learn more general feature representations. Suppose center A uses GE equipment and center B uses Siemens equipment. Domain adaptation techniques allow the model to maintain diagnostic capabilities while ignoring feature differences caused by different devices. Specifically, through adversarial training or MMD methods, the model can learn a mapping that maps image features from different devices to a unified feature space, thus enabling the model to maintain good diagnostic performance on data from different devices.
[0067] More specifically, the model evaluation module also includes a feature importance analysis unit, which quantifies the contribution of each modal feature to the prediction results based on the SHAP value; The time-dependent evaluation unit uses a time-dependent ROC curve to evaluate dynamic prediction performance. It calculates the AUC value by integrating the prediction probabilities at different time points, as shown in the following formula: ; in, express The area under the ROC curve at time t; The cross-center generalization evaluation unit uses leave-one-center cross-validation to evaluate the model's predictive performance on data with unknown centers within the federated learning framework.
[0068] Dynamic ROC curve calculation: For time-series imaging data (such as cardiac-gated PET), we calculate according to time frames. Split the dataset. For each time frame. We calculate the ROC curve between the model's predicted probability and the true label, and obtain the corresponding AUC value, denoted as . .
[0069] Time-dependent AUC integration The formula is: ; in, It is the total number of time frames; To smooth out time series fluctuations, a time sliding window can be set (e.g., (Frame), calculate the average AUC within the window.
[0070] Performance degradation analysis: monitoring The trend over time can be used to identify the time point when the model fails (e.g., (Sudden drop exceeding 0.1); Calculate the AUC volatility coefficient: AUC volatility coefficient The calculation formula is: ; in, and These represent the mean and standard deviation, respectively.
[0071] Introduce a time-weighted factor: assign higher weight to recent time points (such as exponential decay weights); for example, a weighting function can be used. ,in It is the attenuation factor. It is the total number of time frames. This is the current time frame. The weighted AUC can be expressed as: ; Develop an abnormal time point detection algorithm: automatic tagging Time frames that significantly deviate from the baseline. This can be achieved by setting a threshold or using statistical tests such as the z-score test. If If the change exceeds a preset threshold (e.g., z-score>3), then the time frame is marked as abnormal.
[0072] In summary, by calculating dynamic ROC curves, integrating time-dependent AUC data, analyzing performance degradation, and calculating the AUC fluctuation coefficient, we can comprehensively evaluate the performance of models on time-series imagery data and identify potential performance degradation or anomalous time points. Furthermore, introducing time-weighted factors and anomalous time point detection algorithms can further improve the accuracy and practicality of the evaluation.
[0073] More specifically, the real-time inference acceleration unit employs knowledge distillation technology to compress complex models into lightweight models. The distillation loss function is: ; Where CE is the cross-entropy loss and MSE is the mean squared error loss. and These are the predicted probabilities for the student model and the teacher model, respectively. and These are the outputs of the feature layer; The real-time inference acceleration unit achieves efficient inference through the fusion of multiple technologies. It employs knowledge distillation technology, utilizing a distillation loss function composed of cross-entropy loss and mean squared error loss to compress complex models into lightweight models. At the same time, it combines model quantization compression technology and a hardware acceleration engine to successfully reduce the prediction latency to less than 200ms while ensuring that the loss of image prediction accuracy is controlled within ≤2%. This effectively balances inference speed and prediction accuracy, significantly improving real-time inference efficiency.
[0074] More specifically, the results visualization unit supports dynamic time-series curve display, which can plot the changing trend of the indicators of interest over time in real time and overlay clinical event markers for reference; The auxiliary diagnostic suggestion unit automatically generates a structured report based on the prediction results. The structured report includes lesion location information, prediction probability, and links to relevant treatment suggestions, and supports seamless integration with hospital information systems.
[0075] The specific implementation is as follows: In this embodiment, for early screening of neurodevelopmental disorders (such as attention deficit hyperactivity disorder and autism spectrum disorder) in children aged 3-12, a multimodal prediction model is constructed based on nuclear medicine imaging data and clinical data from three children's hospitals (Center A, Center B, and Center C). PET (positron emission tomography) metabolic images and SPECT (single-photon emission computed tomography) functional images are used as core data sources, combined with clinical data such as children's age, height, weight, and clinical symptom scores, to achieve disease risk prediction and assisted diagnosis, while ensuring cross-center data privacy and security.
[0076] Each center is equipped with a GPU server (NVIDIA A100 80GB), with 128GB of memory and 10TB of storage; cross-center connections are made through a dedicated encrypted network link, and model parameter transmission is achieved using federated learning secure aggregation nodes.
[0077] The parameters are set as follows:
[0078] Data acquisition and quality control are as follows: Image acquisition: Each center acquired images of children's brains using a PET scanner (GE Discovery MI) and a SPECT scanner (Siemens Symbia T6). The PET scan lasted 15 minutes (dynamic acquisition, 1 frame every 30 seconds, for a total of 30 frames), and the SPECT scan lasted 20 minutes (static acquisition, 128×128 matrix). Clinical data (age, sex, developmental scale scores, etc.) were collected simultaneously.
[0079] Intelligent quality control: Employs a triplet loss function to detect image artifacts, anchor samples. Select artifact-free, standard pediatric brain images, positive samples Artifact-free clinical images from various centers (3000 cases in total), negative samples Images containing motion artifacts and device noise (500 cases in total). The artifact detection loss was calculated using a CNN model. ,right The images were automatically marked as abnormal and sent to the quality control doctors for review. Finally, 4,800 valid data were selected (Center A=1,920 cases, Center B=1,728 cases, Center C=1,152 cases).
[0080] Data preprocessing is as follows: Timing registration: For PET motion images, the 10th frame is used as the reference image sequence. The remaining frames are floating image sequences. By maximizing temporal mutual information Spatiotemporal alignment was achieved. The average mutual information entropy of each frame was calculated to be 0.82, and the spatial offset error of the registered image was ≤1mm.
[0081] Modal alignment: Fifteen registration points were selected from key anatomical points in the brain (such as the pituitary gland, pineal gland, and cerebellar vermis), and the Euclidean distance loss was minimized. Achieve coordinate alignment between PET and SPECT images with a registration error ≤0.8mm.
[0082] Denoising and normalization: PET images are processed using the MAP algorithm to suppress Poisson noise, and SPECT images are processed using the OSEM algorithm to reduce low-count noise; all images are normalized to scale pixel values to the [0,1] range.
[0083] Feature extraction is as follows: Functional connectivity dynamics analysis: Calculating the dynamic correlation coefficients of metabolic time series data for 12 key brain regions (prefrontal lobe, parietal lobe, temporal lobe, etc.). Example: Prefrontal cortex ( ) and striatum ( )of This indicates that their metabolic activities are highly synchronized.
[0084] Spatial feature extraction: Local metabolic features (such as texture of high metabolic regions in the prefrontal cortex) of PET images and functional activity features of SPECT images are extracted through 3D convolution.
[0085] Metabolic trajectory analysis: Extracting the trajectory of metabolic value changes in each brain region over time (e.g., linear trend, peak time).
[0086] Feature fusion is as follows: Cross-modal gating mechanism: Calculating PET features ( =1), SPECT features ( =2) and clinical data characteristics ( Mutual information (=3) , get weight , , ; through gating vectors (φ is the Sigmoid activation function) Dynamically adjust the information flow rate, and finally merge it into... The 512-dimensional feature vector.
[0087] Feature normalization: Z-score normalization is applied to the fused features to obtain... ,in Mean = 0, Standard deviation = 1.
[0088] The training of the hierarchical feature interaction network is as follows: Shallow local feature extraction: A 256-dimensional local feature vector is output by connecting a 3×3×3 convolution kernel to the residual. .
[0089] Deep global interaction: Global spatial dependencies are modeled using nonlocal operations, as shown in the formula. It combines an 8-head self-attention mechanism to capture multi-subspace feature dependencies and outputs a 512-dimensional global feature vector.
[0090] Loss function optimization: Total loss Among them, the fight against losses (D is the discriminator, G is the generator), temporal smoothing loss The training process is iterated for 200 rounds, with a batch size of 32 per round.
[0091] Cross-center federated learning is as follows: Each center trains a local model based on local data. The parameters of the local model at center A are as follows: Center B Center C .
[0092] Update formula based on global model ( , , , The system aggregates parameters and protects the security of parameter transmission through homomorphic encryption technology; it supports dynamic addition of data from center D (adding 1000 new cases) and achieves incremental updates through a secure aggregation protocol.
[0093] The model evaluation is as follows: Time dependency evaluation: Calculate the time dependency of 20 time frames. ,have to , … According to the formula Calculated This indicates that the dynamic prediction performance is stable.
[0094] Cross-center generalization assessment: Using leave-one-center cross-validation, the model's prediction accuracy in center C was 88.5% and the generalization error was 4.2% when data from center C were removed. Based on SHAP value analysis, PET metabolic features contributed 57% to the prediction results, SPECT functional features contributed 33%, and clinical data contributed 10%.
[0095] Inference efficiency evaluation: Through a knowledge distillation compression model (the student model has 1 / 4 the number of parameters as the teacher model), distillation loss... Combining 8-bit quantization and TensorRT acceleration, the image prediction latency is 180ms and the accuracy loss is 1.5%, meeting the needs of real-time clinical diagnosis.
[0096] The model is applied as follows: Real-time inference: After a new patient's PET / SPECT image is input into the system, the disease risk prediction probability is output within 180ms (e.g., "Attention deficit hyperactivity disorder risk: 89%").
[0097] Results visualization: The trend of prefrontal metabolic values over time is displayed through dynamic time-series curves, with "abnormal clinical symptom scores" markers superimposed; a three-dimensional image visualization map is generated, marking the location of high-risk lesions (such as the area of reduced metabolism in the left prefrontal cortex).
[0098] Assisted diagnostic report: Automatically generates structured reports, including lesion location (left frontal lobe, coordinates X=23mm, Y=18mm, Z=31mm), predicted probability (89%), treatment suggestion links (linked to the latest clinical guidelines), and seamlessly integrates with the hospital's HIS system.
[0099] The implementation results are as follows: Predictive performance: early disease screening The accuracy was 89.3%, sensitivity was 87.6%, and specificity was 90.1%, which was significantly better than the traditional model based on adult data (AUC=0.78).
[0100] Cross-center collaboration: Achieve joint modeling of data from three centers without sharing the original data, with a cross-center generalization error of <5%.
[0101] Clinical applicability: Prediction delay of 180ms, accuracy loss of 1.5%, structured report generation time of <3 seconds, supports doctors to quickly obtain key information, and improves diagnostic efficiency by 40%.
[0102] This invention significantly improves the accuracy, efficiency, and cross-center collaboration capabilities of nuclear medicine image prediction through innovative multi-module design. The cross-modal gating mechanism dynamically assigns weights based on the mutual information between modal features and predicted labels through dual regulation of gating vectors and attention mechanisms, thereby increasing the contribution of key modal features and effectively solving the limitations of fixed-weight fusion. The shallow residual convolutional network, combined with deep nonlocal operations and multi-head self-attention mechanisms, achieves hierarchical modeling of local metabolic features and global spatial dependencies. Compared with traditional convolutional networks, it enhances the ability to capture complex pathological features and significantly improves model accuracy.
[0103] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A medical image prediction model construction system adapted to children's physiological characteristics, characterized in that, include: The data acquisition module is used to acquire multimodal nuclear medicine imaging data and clinical data, including an intelligent quality control unit and an image acquisition unit; The data preprocessing module is used to perform image calibration, denoising, normalization and multimodal temporal registration on the nuclear medicine image data, including a temporal registration unit and a modal alignment unit. The feature extraction module is used to extract multimodal features from preprocessed image data, including a functional connectivity dynamic analysis unit, a spatial feature extraction unit, and a metabolic trajectory analysis unit. The feature fusion module is used to fuse extracted multimodal features through weighted combination, including cross-modal gating mechanisms; The model building module is used to build predictive models using fused features and corresponding clinical data, including a hierarchical feature interaction network. The loss function module is used to calculate the loss of the prediction model, including the adversarial regularization term and the temporal smoothing constraint term; The model evaluation module is used to evaluate the performance of the prediction model, including a time-dependent evaluation unit and a cross-center generalization evaluation unit. The model application module is used to apply the prediction model to the nuclear medicine imaging data of new patients, and to provide prediction results and auxiliary diagnostic suggestions, including a diagnostic suggestion unit, a real-time inference acceleration unit and a result visualization unit. The Federated Learning Collaboration Module supports joint modeling of nuclear medicine data across centers, employing secure multi-party computation and homomorphic encryption to protect privacy data.
2. The medical image prediction model construction system adapted to children's physiological characteristics according to claim 1, characterized in that, The intelligent quality control unit automatically detects image artifacts using a convolutional neural network and calculates the artifact detection loss using a triplet loss function. ; in, For anchor samples, For artifact-free positive samples, For negative samples with artifacts, For interval parameters, It is a distance metric between feature vectors; The intelligent quality control unit also includes a data cleaning subunit, which automatically marks image data that is abnormal in artifact detection and triggers a manual review process.
3. The medical image prediction model construction system adapted to children's physiological characteristics according to claim 2, characterized in that, The time-series registration unit of the data preprocessing module adopts a time-series registration algorithm based on mutual information, which achieves spatiotemporal alignment by maximizing the mutual information entropy between time series. The formula for calculating time-series mutual information is as follows: ; in, , for A sequence of floating images and reference images at different times. For time frames; and Representing single images and The information entropy reflects the uncertainty of pixel grayscale distribution; Representing two images and The joint information entropy reflects the common uncertainty of the gray-scale distribution of both. The modal alignment unit of the data preprocessing module is used to spatially register PET and SPECT images, achieving intermodal coordinate alignment by minimizing the Euclidean distance loss function. ; in, , For the first The coordinates of the registration points in PET and SPECT images This represents the total number of registration points. The denoising unit of the data preprocessing module employs a statistical noise correction algorithm for PET images and an iterative reconstruction denoising algorithm for SPECT images.
4. The medical image prediction model construction system adapted to children's physiological characteristics according to claim 3, characterized in that, The functional connectivity dynamic analysis unit extracts functional connectivity features by calculating the dynamic correlation coefficients of metabolic time series from different brain regions. The formula for the dynamic correlation coefficients is: ; The core dependent variable, i.e. Time of the first The region and the first The dynamic correlation coefficient of each region is directly output; the dynamic correlation strength between two regions is directly output, with a value range of [-1, 1]. For the covariance term, the core calculation unit, measures the first... Lagging in some areas The signal at time and the first each region The degree of linear correlation between signals at different times; For the first Each image area in The time-series signal value at time 10:
00. Represents the signal strength of the region. It is the area code. It is the current moment. It is the time lag; For the first Each image area in The time-series signal value at time 10:
00. Is with Different area codes, the rest are the same. ; as a reference signal, and Paired calculation of covariance constitutes the core analytical object for the dynamic correlation between two regions; For the first The standard deviation of the time-series signal in each region; measuring the standard deviation of the first region's time-series signal; The degree of dispersion of all time-series signals in each region is used to normalize the covariance, eliminate the interference caused by differences in signal amplitude, and make the correlation coefficients comparable. For the first The standard deviation of the time series signal in each region; and Consistency is used to normalize covariance and ensure The values are uniformly set in [-1, 1] to avoid deviations in correlation calculation due to different signal amplitudes in different regions; It represents the time lag; it adjusts the lag correlation of timing signals. The cross-modal gating mechanism uses a gating vector. The fusion weights of different modal features are dynamically adjusted, and the gating vector is calculated as follows: ; in, For the first Modal features, It is the Sigmoid activation function. , These are trainable parameters; The cross-modal gating mechanism also includes an attention mechanism, which dynamically adjusts the intermodal mutual information gain by calculating the intermodal gating gain. The weight value is calculated using the following formula: ; in, For the first Modal features and predicted labels Mutual information value; The feature fusion module also includes a feature normalization unit, which is used to standardize the fused features, using the Z-score normalization method. ; in, As a feature of fusion, The characteristic mean, The characteristic standard deviation is denoted as .
5. The medical image prediction model construction system adapted to children's physiological characteristics according to claim 4, characterized in that, The hierarchical feature interaction network includes a shallow local feature extraction layer and a deep global interaction layer: The shallow layer uses 3×3×3 convolutional kernels to capture local metabolic features and introduces residual connection structures: ; Deep learning models global spatial dependencies through non-local operations, as shown in the formula: ; in, For the first The feature vectors of each spatial location are updated by nonlocal operations and used to fuse the original local features and global spatial dependency information to achieve a collaborative representation of local metabolic features and global association information. For the first The original input feature vectors of each spatial location; These are the normalization coefficients; The number of sampling points for non-local operations; This is a similarity function used to calculate the reference position. eigenvectors With the target location eigenvectors The strength of the correlation between them; This is a feature transformation function used to transform the reference position. The original feature vector Perform linear mapping transformation to extract key feature information of the reference position and compress feature dimensions to reduce the computational cost of global aggregation; The deep global interaction layer employs a multi-head self-attention mechanism, which computes multiple attention heads in parallel to capture feature dependencies in different subspaces.
6. The medical image prediction model construction system adapted to children's physiological characteristics according to claim 5, characterized in that, The adversarial regularization term enhances the model's generalization ability through generative adversarial networks. The adversarial loss formula is as follows: ; in, For discriminator, For generator, For the true data distribution, Noise distribution; The temporal smoothing constraint term uses a first-order difference loss to force the prediction results of adjacent time points to be continuous, and the formula is as follows: ; in, The predicted value at time t; The total loss function of the loss function module is a weighted sum of adversarial loss, temporal smoothing loss, and task loss: ; in, , For weight parameters, A task-specific loss function.
7. The medical image prediction model construction system adapted to children's physiological characteristics according to claim 6, characterized in that, The global model update formula for the federated learning collaborative module is: ; in, For the first Local model parameters at the center, This represents the number of local samples. The total number of global samples supports joint modeling of heterogeneous data sources and aligns differences in image acquisition parameters between different centers through domain adaptation technology. The federated learning collaboration module supports the dynamic addition of new central nodes and achieves incremental updates of the global model through a secure aggregation protocol.
8. The medical image prediction model construction system adapted to children's physiological characteristics according to claim 7, characterized in that, The model evaluation module also includes a feature importance analysis unit, which quantifies the contribution of each modal feature to the prediction result based on the SHAP value. The time-dependent evaluation unit uses a time-dependent ROC curve to evaluate dynamic prediction performance. It calculates the AUC value by integrating the prediction probabilities at different time points, as shown in the following formula: ; in, express The area under the ROC curve at time t; The cross-center generalization evaluation unit uses leave-one-center cross-validation to evaluate the model's predictive performance on data with unknown centers within the federated learning framework.
9. A medical image prediction model construction system adapted to children's physiological characteristics according to claim 8, characterized in that, The real-time inference acceleration unit employs knowledge distillation technology to compress complex models into lightweight models. The distillation loss function is: ; Where CE is the cross-entropy loss and MSE is the mean squared error loss. and These are the predicted probabilities for the student model and the teacher model, respectively. and These are the outputs of the feature layer; The real-time inference acceleration unit also employs model quantization compression technology and a hardware acceleration engine to control the image prediction latency to within 200ms while maintaining a prediction accuracy loss of ≤2%.
10. A medical image prediction model construction system adapted to children's physiological characteristics according to claim 9, characterized in that, The results visualization unit supports dynamic time-series curve display, which can plot the changing trend of the indicators of interest over time in real time and overlay clinical event markers for reference. The auxiliary diagnostic suggestion unit automatically generates a structured report based on the prediction results. The structured report includes lesion location information, prediction probability, and related treatment suggestion links, and supports seamless integration with the hospital information system.