Pancreatic cancer risk prediction method based on machine learning and multi-modal data

By using time-series extrapolation and coupling analysis of multimodal data, combined with prior knowledge of pancreatic cancer pathophysiology and personalized learning, the problems of delayed early warning and insufficient interpretability in pancreatic cancer risk prediction have been solved, enabling early, dynamic, and transparent risk assessment and individualized management.

CN121812151APending Publication Date: 2026-04-07GUANGDONG GENERAL HOSPITAL
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies for pancreatic cancer risk prediction suffer from problems such as delayed early warning, lack of dynamic prediction, and insufficient interpretability, and cannot effectively utilize multimodal time-series data for personalized health management.

Method used

By acquiring multimodal data, using prior constraints from pancreatic cancer pathophysiology to perform temporal extrapolation, a virtual molecular time series is generated. The spatiotemporal coupling strength is calculated using a dynamic time warping algorithm, and a risk prediction model is constructed based on a personalized learning strategy. Explainable artificial intelligence technology is then combined to identify key risk drivers.

Benefits of technology

It enables time-series alignment and coupling analysis of multimodal data, provides early, dynamic, and transparent risk warning tools, can quantify the inherent synergistic relationships in disease evolution, and generate interpretable risk score reports to support personalized health management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121812151A_ABST
    Figure CN121812151A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical information, and discloses a pancreatic cancer risk prediction method based on machine learning and multi-modal data, and the method comprises the steps: obtaining the multi-modal data of a target user; performing time sequence deduction on the molecular biological detection data to generate a virtual molecular time sequence; time sequence signals are extracted from the virtual molecule time sequence and the time sequence behavior monitoring data; calculating the dynamic coupling strength between the two time sequence signals to obtain a space-time coupling coefficient; weighted fusion is carried out on the features, and unified multi-modal feature representation is constructed; carrying out multi-modal feature representation training to obtain a special risk prediction model for the target user; and obtaining a risk quantitative score, and identifying a key risk driving factor which contributes to the score most. According to the invention, through multi-modal time sequence fusion and personalized modeling, early-stage, dynamic and explainable and evaluable pancreatic cancer risks are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical information technology, and in particular to a method for predicting pancreatic cancer risk based on machine learning and multimodal data. Background Technology

[0002] Current clinical pancreatic cancer risk prediction primarily relies on single tumor markers such as CA19-9, or static risk factors based on electronic health records. However, single tumor markers lack sufficient sensitivity in the early stages of disease, and static variables cannot reflect the dynamic evolution of an individual's physiological state. The development of pancreatic cancer involves complex multi-system interactions, and its early prodromal symptoms may manifest as subtle fluctuations in microscopic molecular indicators and minor changes in macroscopic behavioral patterns. Existing assessment models relying on single-dimensional data cannot achieve cross-validation and synergistic capture of these multidimensional, asynchronous, and weak early signals, resulting in significant blind spots in the identification of early pancreatic cancer risk.

[0003] Effective early risk assessment requires a dynamic characterization of disease progression. In reality, behavioral data, which can be acquired continuously at high frequencies, is completely misaligned with molecular data, which can only be acquired discretely at low frequencies, on the timeline. Existing technologies either only allow analysis of single time series or simple static stitching of multimodal data. The lack of a method to extrapolate sparse molecular data into continuous sequences based on pathophysiological laws and to perform time-series alignment and coupling analysis with behavioral data prevents models from answering the crucial clinical question of "how will the patient's current abnormal indicators develop," resulting in a severe deficiency in dynamic predictive capabilities.

[0004] Existing machine learning-based predictive models are mostly data-driven "black boxes," capable of outputting only abstract risk scores but failing to reveal the specific driving factors behind individual risk. This makes it difficult for clinicians to develop personalized intervention strategies. While mechanistic models based on pure medical knowledge are interpretable, they struggle to adapt to complex individual differences and data noise. This disconnect between "data" and "mechanism" results in significant shortcomings in the accuracy, reliability, and clinical operability of existing solutions, preventing the completion of the closed loop from risk warning to personalized health management decisions.

[0005] Therefore, there is an urgent need for a pancreatic cancer risk prediction method that can integrate multimodal time-series data, realize dynamic disease evolution modeling, and provide transparent individual risk attribution. This invention proposes a pancreatic cancer risk prediction method based on machine learning and multimodal data. Summary of the Invention

[0006] The purpose of this invention is to address the problems of delayed early warning, lack of dynamic prediction, and insufficient interpretability in existing pancreatic cancer risk prediction technologies by proposing a pancreatic cancer risk prediction method based on machine learning and multimodal data.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: a method for predicting pancreatic cancer risk based on machine learning and multimodal data, comprising the following steps: Acquire multimodal data of the target user, including time-series behavioral monitoring data, discrete time-point molecular biological detection data, and covariate data; Based on the prior constraint of the pathophysiological evolution sequence of pancreatic cancer, time-series extrapolation is performed on molecular biological detection data at discrete time points to generate a virtual molecular time-series sequence that is synchronized with the behavioral monitoring data. Simulated sequences of inflammatory markers were extracted from virtual molecular time-series sequences, and sequences of diurnal activity intensity variation coefficients reflecting regular changes in activity were extracted from time-series behavioral monitoring data. A dynamic time warping algorithm that integrates medical time priors is used to calculate the dynamic coupling strength between two time series signals and obtain the spatiotemporal coupling coefficient. Based on the spatiotemporal coupling coefficient, features from different modalities are weighted and fused to construct a unified multimodal feature representation; A personalized learning strategy based on the similarity of risk contribution patterns is adopted, and a dedicated risk prediction model for target users is obtained by training multimodal feature representations. Run a dedicated risk prediction model to obtain a quantitative risk score, and identify the key risk drivers that contribute the most to the score based on interpretable artificial intelligence technology.

[0008] The beneficial effects of the technical solution provided by this invention include at least the following: This invention constructs a time-series deduction model based on prior constraints of the pathophysiological evolution of pancreatic cancer, transforming discrete molecular detection data into a virtual sequence synchronized with continuous behavioral data in time. This fundamentally solves the problem of temporal mismatch between "high-frequency behavioral data" and "low-frequency molecular data," laying a unified time benchmark for subsequent multimodal deep fusion analysis and realizing a leap from static snapshot assessment to continuous dynamic evolution modeling.

[0009] This invention employs a dynamic time warping algorithm that integrates medical temporal priors and calculates the spatiotemporal coupling coefficient based on specific weight parameters determined by clinical statistics. This enables the quantification of the intrinsic synergistic relationship between macroscopic behavioral patterns and microscopic molecular signals in disease evolution. It achieves knowledge-weighted feature fusion based on the inherent consistency of data and medical rationality, rather than simple feature splicing, significantly improving the discriminative power and robustness of multimodal feature representation.

[0010] This invention employs a personalized learning strategy based on the similarity of risk contribution patterns and utilizes interpretable artificial intelligence technology to identify key risk drivers. While ensuring the accuracy of model prediction, it transforms the output of the "black box" model into an interpretable report for a specific individual, composed of the contribution of specific features. This transforms abstract risk scores into clinically understandable, auditable, and intervention-friendly decision-making criteria, achieving a breakthrough from group prediction to individualized attribution.

[0011] This invention utilizes a closed-loop technical process of time-series extrapolation, dynamic coupling, and personalized interpretation to comprehensively leverage time-series behavioral monitoring data and molecular biological data. This allows for the capture of synergistic abnormal patterns in the early clinical stages of pancreatic cancer that cannot be detected by single-dimensional methods. Consequently, it provides an earlier, more dynamic, more transparent risk warning tool that can directly guide health management actions, and has significant clinical translational value. Attached Figure Description

[0012] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a schematic diagram of the method flow provided in an embodiment of the present invention; Figure 2 A risk assessment flowchart provided for embodiments of the present invention. Detailed Implementation

[0014] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a pancreatic cancer risk prediction method based on machine learning and multimodal data proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0016] The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0017] The following description, in conjunction with the accompanying drawings, details a specific scheme for a pancreatic cancer risk prediction method based on machine learning and multimodal data provided by the present invention.

[0018] Please see Figure 1 and Figure 2 This document illustrates a flowchart and risk assessment flowchart of a pancreatic cancer risk prediction method based on machine learning and multimodal data, according to an embodiment of the present invention, including the following steps: Step S1: Obtain multimodal data of the target user, including time-series behavioral monitoring data, molecular biological detection data at discrete time points, and covariate data; Step S2: Based on the prior constraints of the pathophysiological evolution sequence of pancreatic cancer, the molecular biological detection data at discrete time points are temporally extrapolated to generate a virtual molecular temporal sequence that is synchronized with the behavioral monitoring data. Step S3: Extract the simulated inflammatory marker sequence from the virtual molecular time series sequence, and extract the diurnal activity intensity variation coefficient sequence reflecting the regular changes in activity from the time series behavior monitoring data; Step S4: Using a dynamic time warping algorithm that integrates medical time priors, the dynamic coupling strength between two time series signals is calculated to obtain the spatiotemporal coupling coefficient. Step S5: Based on the spatiotemporal coupling coefficient, the features from different modalities are weighted and fused to construct a unified multimodal feature representation; Step S6: A personalized learning strategy based on the similarity of risk contribution patterns is adopted, and a dedicated risk prediction model for the target user is trained using multimodal feature representation. Step S7: Run a dedicated risk prediction model to obtain a risk quantification score, and identify the key risk drivers that contribute the most to the score based on interpretable artificial intelligence technology.

[0019] It should be noted that time-series behavior monitoring data refers to high-frequency time-series data that reflects a user's physical activity patterns and is continuously collected through wearable devices (such as medical-grade accelerometers). After processing, quantitative features such as the "coefficient of variation of daily activity intensity" can be extracted. This coefficient is obtained by calculating the ratio of the standard deviation to the mean of minute-level activity intensity within a day, and is used to objectively measure the stability of daily activity patterns.

[0020] Discrete time-point molecular biology detection data refers to biomarker values ​​obtained at a few discontinuous time points (such as physical examinations or follow-up examinations) through invasive or semi-invasive methods such as venous blood testing. Its sampling frequency is much lower than that of behavioral data, exhibiting a "sparse, point-like" appearance. Key indicators include, but are not limited to, inflammatory markers (such as C-reactive protein), tumor markers (such as CA19-9), and metabolic markers (such as glycated hemoglobin).

[0021] The prior constraints on the pathophysiological evolution of pancreatic cancer refer to the knowledge-based rules established based on medical consensus regarding the typical chronological order of early biological events in the development of pancreatic cancer. For example, the constraint might require that, in the model's extrapolation results, the simulated elevation of systemic inflammatory markers should temporally precede the simulated elevation of tumor-specific markers. This constraint guides and standardizes the training and inference process of the temporal extrapolation model, ensuring that the generated virtual sequence conforms to the inherent biological logic of disease development.

[0022] A virtual molecular time series is the output of a time series extrapolation model. It is a continuous sequence of molecular index estimates that are perfectly aligned at time points with the acquired time-series behavioral monitoring data. This sequence is "virtual" because it does not originate from actual dense detections, but is generated by the model based on sparse observations and prior constraints, and is used to address the inherent temporal resolution mismatch between multimodal data.

[0023] The dynamic time warping algorithm, which integrates medical temporal priors, is the core algorithm for calculating the spatiotemporal coupling coefficient. It incorporates medical-specific improvements upon the classic dynamic time warping algorithm. Specifically, when calculating the cumulative alignment distance between two sequences, different weights are applied to each matching point along the path based on medical prior knowledge (such as "changes in inflammatory markers precede changes in behavioral markers"). This ensures that the calculation results not only reflect morphological similarity but also reflect temporal causal relationships consistent with specific disease mechanisms.

[0024] Risk contribution pattern similarity refers to the degree of similarity among different individuals in the combination of core driving factors behind their predicted risk (i.e., which features are most important) and the direction in which these factors act (whether they increase or decrease risk). Personalized learning based on this similarity aims to find historical samples with similar etiological mechanisms for the current target user, thereby enabling more targeted model fine-tuning.

[0025] The core of this invention lies in proposing and implementing a new paradigm for pancreatic cancer risk prediction that is "time-series deduction and coupling under knowledge constraints".

[0026] Specifically, this invention constructs a complete technical system through innovative coupling at the following three levels: First, at the data preprocessing level, this invention creatively solves the fundamental temporal alignment problem between high-frequency continuous behavioral data and low-frequency sparse molecular data in medical scenarios by introducing the "pathophysiological evolution sequence of pancreatic cancer" as a temporal inference model that cannot bypass prior constraints. This model does not perform simple mathematical interpolation, but rather forcibly embeds medical knowledge into the data generation process, thereby inferring a virtual continuous physiological curve that conforms to both individual observations and the general laws of disease development. This provides a reliable and biologically reasonable unified time benchmark for subsequent in-depth cross-modal analysis.

[0027] Second, at the feature fusion level, this invention proposes a "dynamic time warping algorithm based on medical temporal priors" to calculate the spatiotemporal dynamic coupling coefficient. This algorithm not only measures the morphological similarity between behavioral sequences and virtual molecular sequences, but more importantly, by introducing temporal weighting rules based on clinical findings (such as assigning lower penalties to matches where inflammation precedes behavioral changes), the coefficient can quantify the synergistic strength between the two signals that conform to a specific pathological mechanism. This coefficient then serves as an interpretable weight to guide the adaptive weighted fusion of multimodal features, achieving a leap from mechanical splicing to knowledge-guided deep fusion.

[0028] Third, at the model decision-making level, this invention designs a "personalized learning strategy based on the similarity of risk contribution patterns" and combines it with interpretable artificial intelligence technology. This strategy enables the model to be fine-tuned according to the similarity between current users and historical users in terms of "risk-driving mechanisms," rather than using all data. Ultimately, the system outputs not only a risk score, but also an attribution report that clearly lists key risk-driving factors and their contributions, thereby transforming "black box" predictions into transparent, clinically understandable, and intervention-friendly decision-making evidence, completing a closed loop from population risk assessment to individualized attribution guidance.

[0029] In one specific implementation, taking a 55-year-old male user with a history of newly diagnosed type 2 diabetes (diagnosed 8 months ago) as an example, whose body mass index is 26.5 and who has no family history of pancreatic cancer, the specific implementation process of the present invention is demonstrated.

[0030] 1. Multimodal data acquisition The user wore a medical-grade wrist accelerometer continuously for 90 days. The system processed the data in calendar days, calculated the coefficient of variation of daytime activity intensity each day, and finally obtained a time series (CV sequence) of length 90.

[0031] The user underwent three venous blood tests within the past year, yielding three sets of discrete data: C-reactive protein levels were 3.2 mg / L, 4.1 mg / L, and 5.8 mg / L; CA19-9 and glycated hemoglobin levels (details omitted). Covariate data included: age 55 years, male, body mass index 26.5, and a history of newly diagnosed type 2 diabetes (8 months prior).

[0032] 2. Time series deduction based on prior constraints The system loads a pre-trained generative neural network model (such as a variational autoencoder), which has been imposed with prior constraints on the pathophysiological evolution sequence of pancreatic cancer during training. The user's three groups of discrete CRP data are input into the model, and an optimal continuous change path is determined by solving the optimal trajectory optimization problem in the latent variable space (the objective function minimizes the reconstruction error, maximizes the path smoothness, and maximizes the satisfaction of prior constraints). The path is calibrated in combination with the user's diabetes history, and finally a sequence of 90 CRP virtual estimated values corresponding one by one to the 90-day behavior data dates is generated.

[0033] 3. Cross-modal Temporal Feature Coupling and Fusion Extract the above CRP virtual sequence as the molecular temporal signal and extract the CV sequence as the behavioral temporal signal. Use the dynamic time warping algorithm that fuses medical temporal priors to calculate the coupling coefficient: on the DTW path, if the CRP time point index i leads the CV time point index j (i < j), the distance weight of this point pair is set to 0.7 (within the range of the claim 0.5 - 0.8); if it lags (i ≥ j), the weight is set to 1.3 (within the range of 1.2 - 1.5). Calculate the weighted cumulative distance and map it through the Sigmoid function to obtain the spatio-temporal dynamic coupling coefficient of 0.78.

[0034] Take this coefficient as the fusion weight, weight and splice the original features extracted from molecular data, behavioral data, and covariates to construct a unified cross-modal temporal fusion feature vector.

[0035] 4. Personalized Risk Calculation and Attribution Use a pre-trained baseline prediction model (such as XGBoost) and SHAP analysis to construct a risk contribution pattern signature library for historical samples. For the current user, quickly estimate his risk contribution pattern signature based on his fusion feature vector. Use the Jaccard similarity weighted by the feature importance direction (the weights of features with the same direction are doubled) to calculate the matching degree between this signature and each signature in the historical signature library.

[0036] Select the subset of historical samples with the highest similarity according to the matching degree, and use this subset to perform supervised fine-tuning on the baseline model to obtain a dedicated risk prediction model. Input the user's fusion feature vector into this model, and output a comprehensive risk score of 0.28.

[0037] Use SHAP to perform attribution analysis on this prediction, and identify the top 3 features with the largest absolute contribution as the key risk driving factors: C-reactive protein level (contribution: +0.12), new diabetes history (contribution: +0.08), reduction in the coefficient of variation of daily activity intensity (contribution: +0.05). The system generates a final report, clearly listing the risk level (medium risk) and the above key driving factors.

[0038] As one embodiment of the present invention, the steps of performing time-series extrapolation on molecular biological detection data at discrete time points specifically include: A generative neural network model was trained using molecular marker detection data from multiple time points prior to clinical diagnosis of historically diagnosed patients. During the training process, the generative neural network model was subject to prior constraints based on the pathophysiological evolution of pancreatic cancer. The discrete detection data of the target user is input into the trained generative model. By solving the optimal trajectory optimization problem in the latent variable space, a continuously changing path matching the input data is determined. The objective function of the optimal trajectory optimization problem simultaneously minimizes the reconstruction error between the path points and the input data, maximizes the path smoothness, and maximizes the consistency with the prior constraints. Based on the determined continuous change path, parameters are calibrated in conjunction with the individual clinical characteristics of the target user, and a virtual continuous numerical sequence covering the behavioral monitoring period and satisfying prior constraints is output.

[0039] It should be noted that generative neural network models refer to a class of deep learning models that can learn the distribution of training data and generate new samples similar to the training data. In this embodiment, its core function is to learn the typical synergistic change patterns of key molecular indicators (such as CA19-9, C-reactive protein, and glycated hemoglobin) in pancreatic cancer patients before diagnosis.

[0040] Prior constraints on the pathophysiological evolution of pancreatic cancer refer to mandatory rules based on medical knowledge introduced during model training, such as loss functions or regularization terms. For example, a key constraint might be that in any virtual time series generated by the model, the simulated rise of inflammatory markers (such as C-reactive protein) must precede the simulated rise of tumor markers (such as CA19-9) in time. This constraint ensures that the output of the generated model has biological plausibility, rather than being a purely mathematical fit.

[0041] In generative models, the latent variable space refers to a low-dimensional, continuous potential representation space. The encoder maps the input time-series data to a point (latent variable) in this space, which contains key information and structural features of the original data; the decoder then maps this point back to the data space to generate or reconstruct time-series data. In the latent variable space, similar disease progression stages correspond to neighboring points.

[0042] The optimal trajectory optimization problem refers to the mathematical optimization task of finding a continuous path (trajectory) in the latent variable space. This path needs to satisfy three objectives: 1) The key points on the path, after being reconstructed by the decoder, should be as close as possible numerically to the discrete data points actually detected by the target user (minimizing reconstruction error); 2) The path itself should be smooth, avoiding unreasonable and drastic fluctuations (maximizing path smoothness); 3) The evolutionary process represented by the path should conform as closely as possible to the aforementioned "pathophysiological evolutionary sequence of pancreatic cancer" (maximizing consistency with prior constraints).

[0043] Dynamic programming with temporal constraints is an efficient algorithm for solving the aforementioned optimal trajectory optimization problem. Dynamic programming efficiently finds the global optimum by decomposing a complex problem into nested subproblems. In this scenario, "temporal constraints" specifically refer to the requirement that the evolution direction or order of the path in the latent variable space satisfy a certain monotonicity when solving for the path; for example, the dimension of the latent variable corresponding to the severity of the disease should increase over time. This further embeds prior medical knowledge into the optimization process.

[0044] The final output of this step is a virtual continuous numerical sequence that satisfies prior constraints. It is a set of simulated data that is continuous in time (e.g., one estimate per day), smoothly changing in value, and whose internal sequence of changes strictly conforms to the "pathophysiological evolution sequence of pancreatic cancer". This sequence is perfectly aligned with the behavioral monitoring data in terms of time range, but its values ​​are not real measurements. Instead, they are a "virtual" sequence generated based on model inference and knowledge constraints to characterize the continuous changes in the user's potential physiological state.

[0045] As one embodiment of the present invention, the generative neural network model is a variational autoencoder structure. Its training loss function includes reconstruction loss, latent variable distribution regularization loss, and a temporal consistency loss to strengthen prior constraints. The objective function of the optimal trajectory optimization problem is expressed as: L_path=a×L_recon+b×L_smooth+c×L_prior, where L_recon is the reconstruction error term, L_smooth is the path smoothness constraint term, L_prior is the prior consistency term, a, b, and c are adjustable positive weight coefficients, and × is a multiplication sign. The optimal trajectory optimization problem is solved using a dynamic programming algorithm with temporal constraints to ensure that the path satisfies the monotonicity constraint.

[0046] It should be noted that the variational autoencoder structure is a specific generative neural network architecture, consisting of an encoder and a decoder connected by a continuous latent variable space. The encoder maps the input high-dimensional time-series data (such as a set of molecular index sequences) to a probability distribution in the latent variable space (usually represented by mean and variance); the decoder samples a point from this latent variable space and reconstructs it back into the original data space.

[0047] Reconstruction loss is a fundamental loss term used to measure the difference between the data reconstructed by the decoder and the original input data. Commonly used reconstruction losses include mean squared error or cross-entropy loss. Minimizing the reconstruction loss ensures that the model can accurately remember and reproduce the training data.

[0048] The latent variable distribution regularization loss is a unique loss term that uses Kullback-Leibler divergence. This loss term forces the latent variable distribution of the encoder output to approximate a simple prior distribution (such as the standard normal distribution). Its function is to normalize the latent variable space, making it continuous and smooth, avoiding overfitting, and ensuring that meaningful new samples are generated when sampling or interpolating in the latent space.

[0049] The temporal order consistency loss used to reinforce prior constraints is the key innovative loss term of this invention. This loss term mathematically embeds medical knowledge and the evolutionary sequence of pancreatic cancer pathophysiology into the model training. Specifically, it penalizes generated samples that do not conform to the preset temporal order relationships. For example, for a generated three-indicator sequence, if the rise point of the tumor marker (CA19-9) in the decoded sequence is earlier than the rise point of the inflammatory marker (CRP), then the sample will be assigned a higher loss value. By minimizing this loss, the model is guided to learn a generation pattern that conforms to medical principles and has correct temporal dependencies between indicators.

[0050] Dynamic programming with temporal constraints is an efficient algorithm for solving the aforementioned multi-objective, constrained optimal trajectory problem. Dynamic programming decomposes the global pathfinding problem into a series of sequential decision-making subproblems and utilizes the property of "optimal substructure" to efficiently find the global optimum. Here, "temporal constraints" specifically refer to the additional rules applied during the algorithm's solution process. This prevents the algorithm from finding unreasonable paths that "oscillate back and forth" in time, further incorporating the knowledge of irreversible pathophysiology into the solution process.

[0051] As one embodiment of the present invention, the step of calculating the dynamic coupling strength between two timing signals includes: The dynamic time warping algorithm is used to calculate the cumulative alignment distance between the first signal sequence and the second signal sequence, where the first signal sequence is a molecular signal sequence and the second signal sequence is a behavioral signal sequence. When constructing the cumulative alignment distance, for each matching point pair (i, j) on the path, where i is the time index of the molecular signal and j is the time index of the behavioral signal, the local distance contribution weight of this point pair is determined according to the following rules: If i < j, a first weight value is assigned; If i ≥ j, a second weight value is assigned; Where, when the molecular signal sequence is an inflammatory index sequence and the behavioral signal sequence is a coefficient of variation of daily activity intensity sequence, the value range of the first weight value is from 0.5 to 0.8, the value range of the second weight value is from 1.2 to 1.5, and the ratio range of the first weight value to the second weight value is determined based on the statistical average time difference between the inflammatory index leading the behavioral index in the pancreatic cancer patient cohort; Use the weighted local distance to calculate the cumulative alignment distance and map it to a scalar coefficient.

[0052] It should be noted that the dynamic time warping algorithm is a classic algorithm for measuring the similarity between two time series with possibly different lengths. Its core lies in allowing the sequences to be non-linearly "bent" and aligned on the time axis to find the corresponding relationship that makes the overall shapes of the two sequences most matched. Compared with the simple point-by-point Euclidean distance, it can effectively handle the possible phase delay, speed difference or local stretching between two signals, thus more accurately reflecting their inherent temporal correlation.

[0053] The cumulative alignment distance is the core output of the algorithm. After finding the optimal warping alignment path, the algorithm accumulates the distances (such as absolute differences or Euclidean distances) between all matching point pairs on this path, and the sum obtained is the cumulative alignment distance. The smaller this value is, the more similar the overall shapes of the two sequences are in the best alignment.

[0054] The local distance contribution weight is the key innovation point of the present invention. When calculating the cumulative alignment distance, different weights are not assigned to all matching point pairs equally, but are dynamically allocated according to the time index relationship (i, j) of the matching point pairs. This design aims to embed medical prior knowledge into the similarity measurement.

[0055] The two weight values, the first weight value and the second weight value, reflect the "medical temporal prior". Rule setting: When the time point of the molecular signal leads the time point of the behavioral signal (i < j), the first weight value (a value less than 1, such as 0.6) is assigned; when the time point of the molecular signal is synchronous or lags behind the time point of the behavioral signal (i ≥ j), the second weight value (a value greater than 1, such as 1.4) is assigned. The medical principle is that during the progression of pancreatic cancer, the exacerbation of systemic inflammation (molecular signal) usually leads in time to the decline in the regularity of patient activities (behavioral signal). Therefore, a match that conforms to this prior (i < j) is considered "more reasonable" and should be punished less severely (weight < 1); while a match that violates this prior (i ≥ j) is considered "unreasonable" and should be punished more severely (weight > 1).

[0056] The specific value range of the weight values (0.5 to 0.8, 1.2 to 1.5) This range is not arbitrarily set, but is an empirical parameter obtained based on the statistical analysis of a historical cohort of pancreatic cancer patients. By analyzing the continuous monitoring data of a large number of confirmed patients before diagnosis, the average time difference between the rising point of the inflammation index (such as CRP) and the falling point of the activity regularity index (such as the coefficient of variation of daily activity intensity) can be calculated. This time difference statistical information is used to calibrate the specific values and their ratio range of the weight values, making the weight setting supported by clinical data and enhancing the objectivity and reliability of the method.

[0057] As an implementation manner of the present invention, the steps of adopting a personalized learning strategy based on the similarity of risk contribution patterns include: Using a pre-trained baseline prediction model, analyze the risk decision basis of each sample in the historical training set, and extract the risk contribution pattern signature composed of core features and their influence directions; Analyze the multi-modal feature representation of the target user and estimate its risk contribution pattern; Adopt the Jaccard similarity weighted by the direction of feature importance, calculate the matching degree between the estimated pattern and the patterns of each historical sample, where the overlapping features with the same direction are given 2 times the weight in the similarity calculation; According to the level of the matching degree, screen out the subset of historical samples similar to the target user's pattern; Use the subset of historical samples screened out to perform supervised fine-tuning on the baseline prediction model to obtain a dedicated risk prediction model.

[0058] It should be noted that the pre-trained baseline prediction model refers to a machine learning model pre-trained using a complete training set containing a large number of historical patients. This model learns the general mapping relationship from multi-modal features to pancreatic cancer risk and serves as the starting point and knowledge basis for personalized learning.

[0059] Risk decision-making criteria refer to identifying which input features, and how, influenced the final prediction for each prediction from the baseline prediction model. This is revealed through interpretable artificial intelligence techniques, which quantify the contribution of each feature to the prediction result for a specific sample.

[0060] Risk contribution pattern signature is a concise and structured representation of the core driving mechanism behind an individual's risk.

[0061] Jaccard similarity based on feature importance orientation is an improved similarity metric used to calculate the degree of matching between two risk contribution pattern signatures. Standard Jaccard similarity only calculates the ratio of the intersection size to the union size of two feature sets.

[0062] The historical sample subset refers to the portion of samples selected from all historical training samples based on their matching degree with the target user's predicted signature. These samples are considered to be most comparable to the current user in terms of "cause of illness".

[0063] Supervised fine-tuning is a transfer learning technique. Starting with a pre-trained baseline model, its parameters are trained (i.e., fine-tuned) on a selected subset of historical samples similar to the target user, using the true labels (whether the user has the disease) of these samples as supervision. The fine-tuning process involves making small adjustments to the model parameters based on the data distribution of this subset, making the model more adaptable to learning the specific risk pattern represented by the current user, thereby achieving personalization.

[0064] The dedicated risk prediction model is the final model obtained after the aforementioned personalized fine-tuning. This model inherits the general knowledge learned from massive amounts of data by the baseline model, and is optimized for the unique risk-driven patterns of the current target users. Therefore, it can provide them with a more accurate and personalized risk assessment than the general model.

[0065] As one embodiment of the present invention, the step of analyzing the multimodal feature representation of a target user and predicting its risk contribution pattern includes: The multimodal feature representation of the target user is input into a pre-trained baseline prediction model to obtain a preliminary risk score; The same interpretability analysis method as that used to extract risk contribution pattern signatures from historical samples was employed to analyze the decision-making process of the baseline prediction model for this initial score. The contribution value and direction of influence of each feature in the quantification of multimodal feature representation to the initial score; Based on the absolute value of the contribution, the top K features and their corresponding influence directions are selected to form the estimated risk contribution pattern signature of the target user.

[0066] It should be noted that interpretability analysis methods specifically refer to a class of post-hoc analysis techniques capable of revealing the internal decision-making logic of complex machine learning models. In this invention, the Shapley sum interpretation method is preferred. This method originates from cooperative game theory and calculates a SHAP value for each input feature in the model's prediction. This value represents the marginal contribution of that feature to the current specific prediction result relative to the average contribution of all features. The SHAP value can fairly and consistently quantify the impact of each feature.

[0067] Sort by the absolute value of contribution values. This involves arranging all features in descending order of their absolute values ​​after obtaining their contribution values. The purpose of this step is to select the few features with the greatest influence from all features that may affect the prediction, focusing on the core risk drivers.

[0068] The top K features refer to the first K features in the above ranking. K is a preset integer parameter (e.g., K=5) used to control the number of features included in the estimated signature, ensuring that it both summarizes the main sources of risk and maintains simplicity. These K features are considered to be the core driving factors constituting the current user risk.

[0069] The estimated risk contribution pattern signature of the target user is the final output of this step. It is a structured metadata containing two parts: 1) the identity identifiers of K core features; and 2) the direction of influence (positive or negative) of each of these K core features. This signature is derived from the preliminary analysis of the target user based on the baseline model. It depicts the "risk profile" of the target user from the perspective of the general model and is the core basis for subsequent pattern matching with historical samples and fine-tuning of the personalized model. Because it is generated before the final personalized model is determined, it is called the "estimated" signature.

[0070] As one embodiment of the present invention, the same interpretability analysis method is Shapley and the interpretation method; the rules constituting the estimated risk contribution pattern signature of the target user also include: when the absolute difference between the contribution values ​​of two features is less than a preset threshold, they are sorted according to the data modality priority of the feature source, wherein features from virtual molecular time series sequences or original behavior monitoring time series sequences have higher priority.

[0071] It should be noted that the Shapley sum explanation method is an attribution method derived from cooperative game theory, used to explain model-independent, consistent, and theoretically sound predictions made by any machine learning model. This method calculates a numerical value called a SHAP value for each feature, which accurately quantifies the magnitude and direction of that feature's contribution to a specific prediction result relative to the average contribution of all features. In this invention, the SHAP method ensures the consistency and comparability of historical sample signatures with the target user's estimated signature extraction process.

[0072] The preset threshold is a small, positive number (e.g., 0.01 or 0.005). This threshold is used to determine whether the absolute values ​​of the contributions of two features are statistically or substantially "close enough".

[0073] The rules for constructing the predicted risk contribution pattern signature of the target user improve the signature generation logic. Based on the standard "selecting the top K features in descending order of absolute contribution value," a tie-breaker rule is added. This rule ensures that when feature contribution is difficult to quantify, the signature can more stably and specifically reflect the risk signals derived from dynamic multimodal data, which are the focus of this invention's technical system. This enhances the robustness and clinical relevance of subsequent pattern matching and personalized modeling. This reflects the invention's emphasis on and protection of its core innovative dimension (dynamic temporal fusion) while pursuing interpretability.

[0074] As one embodiment of the present invention, the step of calculating the matching degree between the predicted pattern and the historical sample pattern includes: Compare the core feature sets that the two models focus on, and calculate the number of overlapping features. For each overlapping feature, determine whether its role in the two modes is consistent. By combining the number of overlapping features and the proportion of features with the same direction of action, a quantified pattern matching score is calculated.

[0075] It should be noted that the two modes refer to the two risk contribution pattern signatures to be compared. In the context of personalized learning strategies, they usually refer to: 1) the predicted mode, which is the risk contribution pattern signature predicted based on the multimodal feature representation of the target user; and 2) the historical sample mode, which is the risk contribution pattern signature extracted from a specific sample in the historical training set.

[0076] The core feature set refers to the set of all feature identifiers contained in a risk contribution pattern signature. Typically, a pattern signature consists of the top K core features in terms of contribution, so the size of its core feature set is K (e.g., K=5). This set represents the top factors that are of most concern to the individual (or the prediction) and influence their risk score.

[0077] The quantified pattern matching score is the final output of the above calculation process and is a specific numerical value (usually between 0 and 1). This score quantifies the overall similarity between two individuals in terms of risk-driven mechanisms: A score close to 1 indicates that the two patterns are highly similar; they not only focus on most of the same core risk characteristics, but these characteristics also affect their respective risks in exactly the same direction. This means that the two individuals are likely to have very similar "causes of illness" or risk profiles.

[0078] A score close to 0 indicates that the two patterns are very different, either focusing on almost different core features, or even if there are a few overlapping features, their effects are opposite.

[0079] As one embodiment of the present invention, the step of identifying the key risk driver factors that contribute the most to the score includes: An interpretability analysis was performed on the decision logic of the risk quantification score generated by the dedicated risk prediction model, and the contribution value of each feature constituting the multimodal feature representation to the score was calculated. All features are sorted according to the absolute value of their contribution. From the ranking results, select the top N features with the largest absolute value of contribution to form the initial set of key factors; From the initial set of key factors, features whose eigenvalues ​​originate from virtual molecular time series or original behavioral monitoring time series are identified and determined as key risk drivers.

[0080] It should be noted that the decision-making logic refers to the calculation rules, weight allocation, and judgment criteria executed internally by the dedicated risk prediction model when processing the fused feature vectors of target users. This is the intrinsic reason for the final generation of the risk score. This logic is usually complex and opaque.

[0081] The contribution value refers to the marginal contribution of each input feature to the current specific risk score (for the target user), quantified using the aforementioned interpretability analysis techniques (such as SHAP). This value can be positive or negative: a positive value indicates that the feature pushes the score towards higher risk (risk factor), while a negative value indicates that it pushes the score towards lower risk (protection factor). The absolute value directly reflects the magnitude of the feature's influence.

[0082] The initial key factor set refers to the set of N features (N is a preset integer, such as 3 or 5) with the highest absolute value of contribution selected from the above ranking results. This set contains the most important factors for this risk assessment from a pure contribution perspective and is the key basis for subsequent screening.

[0083] Key risk drivers are the final output and the core insights presented to users or clinicians. It is not simply an initial set of key factors, but rather a further filtered selection of features whose characteristics originate from virtual molecular time-series sequences or original behavioral monitoring time-series sequences. This selection rule is based on the following core considerations: As one embodiment of the present invention, the step of identifying features whose characteristic values ​​originate from virtual molecular time series sequences or original behavior monitoring time series sequences includes: For each feature in the multimodal feature representation, maintain a metadata label that identifies the type of its original data source. The source types include virtual molecular time series, original behavior monitoring time series, and static covariate data. In the initial set of key factors, read the metadata tag for each feature; Features tagged with virtual molecular time series or original behavior monitoring time series in metadata were selected and identified as key risk drivers.

[0084] It should be noted that, in the context of this invention, metadata tags refer to identifiers attached to each specific feature that describe its metadata or attributes. It is not a numerical value of the feature itself, but rather descriptive data about the feature's "origin" or "category." In computer systems, this is typically achieved by adding a specific attribute field to each feature in the data structure.

[0085] Identifying the original data source type means that the core function of metadata tags is to explicitly record the most original data category on which the feature was generated or extracted. This design ensures that the lineage or origin of each feature is traceable throughout the entire data processing and model inference chain.

[0086] The source type refers to the specific value of the metadata tag, which is divided into three categories based on the feature's generation path: Virtual molecular time series: Features of this source type are derived directly or indirectly from virtual continuous sequences generated by time series extrapolation models. For example, features such as "mean CRP" and "CRP rise slope" extracted from a 90-day CRP virtual sequence should be tagged as "virtual molecular time series" in metadata. These features are a new source of information reflecting potential physiological dynamics, created by this invention by injecting medical knowledge into the data generation process.

[0087] Raw behavioral monitoring time series: Features of this type are derived directly from the processing and calculation of raw behavioral signals collected by wearable devices. For example, features such as "coefficient of variation of daytime activity intensity" and "average daily total energy consumption" are calculated directly from acceleration and heart rate data, and their metadata tags should be marked as "raw behavioral monitoring time series". These features are the product of objective, continuous behavioral monitoring.

[0088] Static covariate data: Features of this type have values ​​directly taken from the user's basic background or medical history information, without involving complex time-series processing or inference. For example, features such as "age", "gender", and "whether there is a history of new-onset diabetes" should be labeled as "static covariate data" in their metadata.

[0089] The key risk drivers identified are the feature set that was ultimately retained after screening from the aforementioned sources. These factors not only contribute the most to the final risk score (ranked by contribution), but more importantly, they all originate from the dynamic, time-series data modalities that the technical system of this invention focuses on. Therefore, these "key risk drivers" are presented in the final report. This enhances the novelty, specificity, and operability of the clinical insights output by the entire technical solution.

[0090] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for predicting pancreatic cancer risk based on machine learning and multimodal data, characterized in that, Includes the following steps: Acquire multimodal data of the target user, including time-series behavioral monitoring data, discrete time-point molecular biological detection data, and covariate data; Based on the prior constraint of the pathophysiological evolution sequence of pancreatic cancer, time-series extrapolation is performed on molecular biological detection data at discrete time points to generate a virtual molecular time-series sequence that is synchronized with the behavioral monitoring data. Simulated sequences of inflammatory markers were extracted from virtual molecular time-series sequences, and sequences of diurnal activity intensity variation coefficients reflecting regular changes in activity were extracted from time-series behavioral monitoring data. A dynamic time warping algorithm that integrates medical time priors is used to calculate the dynamic coupling strength between two time series signals and obtain the spatiotemporal coupling coefficient. Based on the spatiotemporal coupling coefficient, features from different modalities are weighted and fused to construct a unified multimodal feature representation; A personalized learning strategy based on the similarity of risk contribution patterns is adopted, and a dedicated risk prediction model for target users is obtained by training multimodal feature representations. Run a dedicated risk prediction model to obtain a quantitative risk score, and identify the key risk drivers that contribute the most to the score based on interpretable artificial intelligence technology.

2. The pancreatic cancer risk prediction method based on machine learning and multimodal data according to claim 1, characterized in that, The steps for performing time-series extrapolation of molecular biological detection data at discrete time points specifically include: A generative neural network model was trained using molecular marker detection data from multiple time points prior to clinical diagnosis of historically diagnosed patients. During the training process, the generative neural network model was subject to prior constraints on the pathophysiological evolution sequence of pancreatic cancer. The discrete detection data of the target user is input into the trained generative model. By solving the optimal trajectory optimization problem in the latent variable space, a continuously changing path matching the input data is determined. The objective function of the optimal trajectory optimization problem simultaneously minimizes the reconstruction error between the path points and the input data, maximizes the path smoothness, and maximizes the consistency with the prior constraints. Based on the determined continuous change path, parameters are calibrated in conjunction with the individual clinical characteristics of the target user, and a virtual continuous numerical sequence covering the behavioral monitoring period and satisfying prior constraints is output.

3. The pancreatic cancer risk prediction method based on machine learning and multimodal data according to claim 2, characterized in that, The generative neural network model is a variational autoencoder structure. Its training loss function includes reconstruction loss, latent variable distribution regularization loss, and a temporal consistency loss to strengthen prior constraints. The objective function of the optimal trajectory optimization problem is expressed as: L_path = a × L_recon + b × L_smooth + c × L_prior, where L_recon is the reconstruction error term, L_smooth is the path smoothness constraint term, L_prior is the prior consistency term, a, b, and c are adjustable positive weight coefficients, and × is a multiplication sign. The optimal trajectory optimization problem is solved using a dynamic programming algorithm with temporal constraints to ensure that the path satisfies the monotonicity constraint.

4. The pancreatic cancer risk prediction method based on machine learning and multimodal data according to claim 1, characterized in that, The steps for calculating the dynamic coupling strength between two time-series signals include: The cumulative alignment distance between the first signal sequence and the second signal sequence is calculated using a dynamic time warping algorithm. The first signal sequence is a molecular signal sequence, and the second signal sequence is a behavioral signal sequence. When constructing the cumulative alignment distance, for each matching point pair (i, j) on the path, where i is the time index of the molecular signal and j is the time index of the behavioral signal, the local distance contribution weight of this point pair is determined according to the following rules: If i < j, a first weight value is assigned; If i ≥ j, a second weight value is assigned; Among them, when the molecular signal sequence is an inflammatory index sequence and the behavioral signal sequence is a coefficient of variation sequence of daily activity intensity, the value range of the first weight value is 0.5 to 0.8, and the value range of the second weight value is 1.2 to 1.

5. The ratio range of the first weight value to the second weight value is determined based on the statistical average time difference of the inflammatory index leading the behavioral index in the pancreatic cancer patient cohort; Use the weighted local distance to calculate the cumulative alignment distance and map it to a scalar coefficient.

5. The method for predicting the risk of pancreatic cancer based on machine learning and multimodal data according to claim 1, the steps of adopting a personalized learning strategy based on the similarity of risk contribution patterns include: Using a pre-trained baseline prediction model, analyze the risk decision basis of each sample in the historical training set, and extract the risk contribution pattern signature composed of core features and their influence directions; Analyze the multimodal feature representation of the target user and estimate its risk contribution pattern; Adopt the Jaccard similarity weighted by the direction of feature importance, and calculate the matching degree between the estimated pattern and the patterns of each historical sample. Among them, the overlapping features with the same direction are given 2 times the weight in the similarity calculation; According to the level of the matching degree, screen out the subset of historical samples similar to the target user's pattern; Use the screened subset of historical samples to perform supervised fine-tuning on the baseline prediction model to obtain a dedicated risk prediction model.

6. The pancreatic cancer risk prediction method based on machine learning and multimodal data according to claim 5, characterized in that, The steps of analyzing the multimodal feature representation of the target user and estimating its risk contribution pattern include: Input the multimodal feature representation of the target user into the pre-trained baseline prediction model to obtain a preliminary risk score; Adopt the same interpretability analysis method as extracting the risk contribution pattern signature of historical samples to analyze the decision-making process of the baseline prediction model for this preliminary score; Quantify the contribution degree value and influence direction of each feature in the multimodal feature representation to the preliminary score; Sort according to the absolute value of the contribution degree value, and select the top K features and their corresponding influence directions to form the estimated risk contribution pattern signature of the target user.

7. The pancreatic cancer risk prediction method based on machine learning and multimodal data according to claim 6, characterized in that, The same interpretability analysis method is the Shapley additive explanation method; The rules for forming the estimated risk contribution pattern signature of the target user also include: when the absolute value difference of the contribution degree values of two features is less than a preset threshold, sort according to the data modality priority of the feature sources, and the features from the virtual molecular time series or the original behavior monitoring time series have higher priority.

8. The pancreatic cancer risk prediction method based on machine learning and multimodal data according to claim 5, characterized in that, The steps of calculating the matching degree between the estimated pattern and the historical sample pattern include: Compare the core feature sets concerned by the two patterns and calculate the number of their overlapping features; For each overlapping feature, judge whether the direction of its role in the two patterns is the same; By combining the number of overlapping features and the proportion of features with the same direction of action, a quantified pattern matching score is calculated.

9. The pancreatic cancer risk prediction method based on machine learning and multimodal data according to claim 1, wherein the step of identifying the key risk drivers that contribute most to the score includes: An interpretability analysis was performed on the decision logic of the risk quantification score generated by the dedicated risk prediction model, and the contribution value of each feature constituting the multimodal feature representation to the score was calculated. All features are sorted according to the absolute value of their contribution. From the ranking results, select the top N features with the largest absolute value of contribution to form the initial set of key factors; From the initial set of key factors, features whose eigenvalues ​​originate from virtual molecular time series or original behavioral monitoring time series are identified and determined as key risk drivers.

10. The pancreatic cancer risk prediction method based on machine learning and multimodal data according to claim 9, characterized in that, The step of identifying features whose characteristic values ​​originate from virtual molecular time series sequences or original behavior monitoring time series sequences includes: For each feature in the multimodal feature representation, maintain a metadata tag that identifies its original data source type, which includes virtual molecular time series, original behavior monitoring time series, and static covariate data; In the initial set of key factors, read the metadata tag for each feature; Features tagged with virtual molecular time series or original behavior monitoring time series in metadata were selected and identified as key risk drivers.

Citation Information

Cited By

  • Oral precancerous lesion intelligent follow-up system and method

    CN122245695A

  • Oral precancerous lesion intelligent follow-up system and method

    CN122245695B