SLC31A1-based castration-resistant prostate cancer prognosis prediction model

By constructing a predictive model that integrates SLC31A1 expression levels and methylation levels, the problem of integrating dynamic changes in existing technologies has been solved, enabling dynamic assessment and treatment strategy recommendations for castration-resistant prostate cancer, and improving the accuracy and stability of prediction.

CN121983308APending Publication Date: 2026-05-05何嘉炜
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
何嘉炜
Filing Date
2026-01-19
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies struggle to integrate multidimensional regulatory information of SLC31A1, failing to reflect its dynamic changes in castration-resistant prostate cancer, resulting in insufficient predictive accuracy and stability, and an inability to provide temporal insights.

Method used

By constructing a prognostic prediction model for castration-resistant prostate cancer based on SLC31A1, the expression level of messenger ribonucleic acid of SLC31A1 and the promoter methylation level are integrated to calculate the functional activity integration value. Dynamic morphological features are extracted through time series analysis and matched with a feature library of poor prognostic patterns to output a quantitative prognostic risk level and treatment strategy.

Benefits of technology

It enables dynamic assessment of castration-resistant prostate cancer, providing more stable biological status indicators and clinically interpretable risk assessment results, enhancing the model's usability and relevance to treatment recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983308A_ABST
    Figure CN121983308A_ABST
Patent Text Reader

Abstract

The invention discloses an SLC31A1-based castration-resistant prostate cancer prognosis prediction model, and belongs to the technical field of bioinformatics and tumor precision medical treatment. The model is realized through the following data processing flow: acquiring cross-time-point molecular and clinical data of a patient, calculating a functional activity integration value of each time point and constructing a time sequence curve by fusing a transcription expression quantity of SLC31A1 and a methylation level of a specific regulation and control region; extracting dynamic morphological features of the curve, and performing matching analysis on the dynamic morphological features and a bad prognosis mode feature library pre-stored based on historical data; and inputting the obtained matching degree, the current activity integration value and the clinical features into a decision function, outputting a quantitative disease progress risk level, and further providing a tendency suggestion of a treatment strategy. According to the method, the dynamic, quantitative and explainable accurate prediction of the disease progress risk of the castration-resistant prostate cancer patient is realized, and a basis is provided for individualized treatment decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of bioinformatics and precision oncology, and more specifically, to a prognostic prediction model for castration-resistant prostate cancer based on SLC31A1. Background Technology

[0002] Tumor prognosis prediction is a crucial basis for clinical decision-making. Currently, prognostic assessments based on biomarkers mostly employ static analysis models, which involve detecting the expression levels of specific genes at a single time point (usually at diagnosis) and constructing predictive models by combining them with clinicopathological parameters. This approach has the following limitations: First, it cannot reflect the dynamic evolution of tumor molecular characteristics during disease progression, especially under treatment intervention, a process directly related to treatment response and disease progression. Second, it typically relies on single-dimensional information (such as messenger RNA expression abundance) and fails to integrate other key regulatory information affecting gene functional status (such as DNA methylation modification), potentially affecting the stability and universality of the biomarkers.

[0003] In the prior art, Chinese patent CN116859048A discloses the use of the transmembrane protein SLC31A1 as a prognostic biomarker for tumors, indicating that its expression level is associated with the prognosis of various tumors, including prostate cancer. This patent highlights the prognostic value of SLC31A1. However, its technical solution still falls under the aforementioned static analysis paradigm, and its claims mainly involve detecting the expression level of SLC31A1 (protein or nucleic acid) at a single time point. This approach does not involve analysis of the epigenetic regulatory aspects of SLC31A1 function (such as promoter methylation), nor does it utilize time-series detection data from the same patient at multiple treatment or follow-up points.

[0004] Therefore, when dealing with castration-resistant prostate cancer, a disease characterized by high heterogeneity and dynamic evolution, there is room for improvement in the accuracy, stability, and temporal insights provided by existing technologies. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a prognostic prediction model for castration-resistant prostate cancer based on SLC31A1, in order to solve the problems mentioned in the background art. Specifically, the prior art has difficulty integrating the multi-dimensional regulatory information of SLC31A1, analyzing the dynamic changes of it over time, and outputting risk assessment results that are more relevant to clinical decision-making.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a prognostic prediction model for castration-resistant prostate cancer based on SLC31A1, the internal data processing logic of which includes the following steps executed sequentially: S1. Input the time-point molecular and clinical data of the target patient; S2. Based on molecular data, calculate the SLC31A1 functional activity integration value at each time point. The calculation of this value integrates the SLC31A1 messenger ribonucleic acid expression level at the corresponding time point with the methylation level of a specific CpG site located in the promoter region of the SLC31A1 gene. The methylation level participates in the integration with negative weights to comprehensively characterize the epigenetically regulated SLC31A1 functional activity. S3. The SLC31A1 functional activity integration values ​​at each time point are constructed into an activity time-series curve according to time sequence, and the dynamic morphological features of the curve are extracted. The dynamic morphological features include the key inflection points of the curve, the slope values ​​of each curve segment, and the cumulative deviation value of the curve relative to a preset baseline, so as to quantify the dynamic pattern of activity change. S4. The extracted dynamic morphological features are matched with a pre-stored adverse prognostic pattern feature library. The adverse prognostic pattern feature library contains a set of morphological features extracted from the activity time-series curves of patients with poor prognosis in the historical cohort. The matching is used to identify whether the current patient presents a known adverse prognostic evolution trajectory. S5. Based on the matching degree and the SLC31A1 functional activity integration value at the current time point, output the quantitative prognostic risk level for the target patient through a predefined risk stratification rule.

[0007] Furthermore, in step S2, the calculation of the SLC31A1 functional activity integration value also introduces a third variable, which is the expression level of a specific microRNA that has a known regulatory relationship with SLC31A1, in order to introduce post-transcriptional regulatory information and further refine the functional activity integration value.

[0008] Furthermore, in step S3, the extraction of dynamic morphological features specifically includes: S3.1 Identify the points in the activity time-series curve where the first derivative is zero or the second derivative is an extreme value, as the key turning points, to locate the moment when the trend changes. S3.2 Calculate the average slope of the curve segment defined by the adjacent key inflection points, and use it as the segment slope value to quantify the rate and direction of change in each time period; S3.3 Calculate the area enclosed by the active time-series curve and the straight line connecting the start and end points of the curve, and use it as the cumulative deviation value to measure the degree of deviation of the overall change path from the linear assumption.

[0009] Furthermore, the pre-stored adverse prognostic pattern feature library mentioned in step S4 includes: First mode feature: The slope value of the activity time-series curve between two consecutive detection time points is greater than a first preset slope threshold, which is used to identify the mode of rapid increase in activity; The second mode feature is that the activity time-series curve contains no fewer than three monotonically increasing steps, and the serum prostate-specific antigen level recorded at the first clinical testing time point after each step increases by more than a preset percentage threshold compared to the recorded value at the previous clinical testing time point. This is used to identify the step-like growth pattern of activity accompanied by a rebound in clinical indicators.

[0010] Furthermore, the matching degree calculation in step S4 is a weighted calculation, wherein, within the time span of the active time series curve, the dynamic morphological features extracted in the latter 50% of the time period have a higher calculation weight than the features extracted in the first 50% of the time period, so that the recent change pattern has a greater impact on the prognostic judgment.

[0011] Furthermore, the predefined risk stratification rule in step S5 is a decision function. The input variables of this decision function include the matching degree, the integrated value of SLC31A1 functional activity at the current time point, and the Gleason score in the clinical data. Risk quantification is achieved through comprehensive calculation of these variables.

[0012] Furthermore, the model also includes step S6: S6. If the quantitative prognostic risk level output in step S5 is a preset high-risk level, then the treatment strategy tendency inference step is executed; this step outputs a corresponding preferred treatment strategy category based on the specific adverse prognostic pattern matched by the dynamic morphological features, in order to provide decision support.

[0013] Furthermore, the association rule used in the treatment strategy preference inference step S6 is as follows: If the dynamic morphological features match the first pattern features, then the output preferred treatment strategy category is copper metabolism intervention therapy; If the dynamic morphological features match the second pattern features, then the output preferred treatment strategy category is DNA damage repair targeted therapy.

[0014] Furthermore, the specific CpG site located in the promoter region of the SLC31A1 gene mentioned in step S2 refers to a CpG dinucleotide site located in a region of 200 to 50 base pairs upstream of the transcription start site of the SLC31A1 gene. This region is a common functional region that regulates gene transcription.

[0015] Furthermore, in the data processing logic within the model, the negative weights in step S2, the adverse prognostic pattern feature library and matching threshold in step S4, and the risk stratification rule parameters in step S5 are all obtained by training a machine learning model using a time-series training dataset of castration-resistant prostate cancer patients containing known disease progression time and overall survival data, thereby optimizing the model parameters based on real-world data.

[0016] The technical effects and advantages of this invention are as follows: First, by fusing multi-dimensional molecular information to calculate an integrated functional activity value, a more stable indicator of biological state than a single expression level is provided. Existing techniques typically detect the messenger RNA expression level of SLC31A1 separately. This invention fuses this expression level with the methylation level of its specific regulatory regions to obtain a comprehensive integrated functional activity value. Since DNA methylation is a key epigenetic mechanism regulating gene expression, this integration process simultaneously incorporates transcriptional signals and epigenetic repression signals at the computational level. The resulting integrated value reflects the potential functional activity of SLC31A1 in tumor cells more accurately than single expression level data, reducing assessment bias caused by fluctuations in a single indicator or changes in regulatory state.

[0017] Second, by analyzing the temporal patterns of functional activity changes, a dynamic assessment of disease progression risk is achieved. This invention is not limited to single-time-point analysis, but rather connects the integrated functional activity values ​​from various time points into a time-series curve, extracting quantitative features such as slope changes and cumulative bias. These features describe the trajectory of activity evolution over time. By matching and comparing this trajectory with typical change patterns summarized from historical data of patients with poor prognoses, the individual patient's disease progression tendency can be assessed. This method allows the assessment to encompass dynamic information during disease evolution, rather than just the initial state.

[0018] Third, by establishing a logical connection from feature matching to risk grading and then to treatment recommendations, the clinical interpretability and practicality of the model results are enhanced. This invention combines the matching results of the aforementioned dynamic features with the patient's current functional activity state and baseline clinical characteristics, generating a clear risk level through a trainable decision function. Furthermore, the model associates different feature matching patterns with specific underlying biological mechanisms, thereby outputting propensity-based treatment strategy recommendations. For example, a rapidly rising pattern is associated with copper metabolism intervention. This design ensures that the model's output is no longer an abstract risk score, but a structured information containing risk level and potential intervention direction, helping clinicians understand the basis for prediction and consider subsequent treatment options. Attached Figure Description

[0019] Figure 1This is a schematic diagram of the core processing flow of the prognostic prediction model of the present invention.

[0020] Figure 2 This is a schematic diagram of the dynamic feature extraction and pattern matching process of the present invention.

[0021] Figure 3 This is a schematic diagram of the poor prognosis pattern matching logic structure of the present invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] As attached Figures 1 to 3 The specific implementation details of the SLC31A1-based prognostic prediction model for castration-resistant prostate cancer are as follows: Example 1: Model Construction and Internal Logic Implementation This embodiment illustrates the construction and execution process of the model's internal data processing logic.

[0024] I. Data Preparation and Input This model requires the input of the target patient at at least two different time points (which can be labeled as time point 1, time point 2, ..., time point 3). ,in Molecular detection data (integers greater than 1) and corresponding clinical follow-up data.

[0025] The time points are associated with specific clinical events or follow-up points. For example, time point 1 may correspond to the date of diagnosis of castration-resistant prostate cancer. Time point 2 corresponds to the first efficacy assessment after the start of the initial systemic treatment; Subsequent timelines can be determined based on routine clinical follow-up plans or as needed for disease assessment.

[0026] At each time point, tumor tissue samples (e.g., obtained via biopsy) and peripheral venous blood samples must be collected from the target patient. The tumor tissue samples are used for the following molecular tests, and the peripheral venous blood samples are used to obtain clinical serum biochemical markers.

[0027] The specific steps for acquiring molecular detection data are as follows: 1. RNA Sequencing and Expression Quantification: Total RNA was extracted from tumor tissue samples. Whole transcriptome sequencing was performed using a high-throughput sequencing platform (such as Illumina's NovaSeq series sequencers) to obtain raw sequencing data (reads, usually in FASTQ format).

[0028] Subsequently, the raw data were processed using a bioinformatics analysis workflow, which typically includes: aligning reads to the human reference genome GRCh38 / hg38 using sequence alignment software (such as HISAT2); calculating the expression abundance of each gene using transcript assembly and quantification software (such as StringTie); and outputting a tab-delimited text file containing the transcripts per million (TPM) value for each gene. The TPM value corresponding to the SLC31A1 gene was read from this file and denoted as... .

[0029] To eliminate the influence of extreme values ​​and conform to the normal distribution assumption, a logarithmic transformation is performed on this expression level. For example, a base-10 logarithmic transformation is used: ,in For the first The expression level after logarithmic transformation at each time point.

[0030] 2. Deoxyribonucleic acid (DNA) methylation sequencing and site quantification: Genomic DNA was extracted from the same tumor tissue sample. Probes were designed targeting the transcription start site and upstream promoter region of the SLC31A1 gene for hybridization capture, and the captured DNA was subjected to bisulfite treatment and sequencing.

[0031] Sequencing data is processed using methylation analysis software (such as Bismark), and the steps include: aligning the sequencing reads to a reference genome that has undergone bisulfite conversion; The methylation level of each CpG dinucleotide site was calculated, defined as the percentage of methylated cytosine reads at that site relative to the total number of covered reads (the sum of methylated and unmethylated reads). Based on historical studies or independent training data, a specific CpG site within this region that showed a statistically negative correlation with SLC31A1 messenger RNA expression levels was selected (the specific site can be determined by consulting annotation information in public databases such as UCSCGenomeBrowser, NCBIGene, or MethHC under the GRCh38 / hg38 version). This site was selected at a specific time point. The methylation level is denoted as Its value is Real numbers between 100 and 100.

[0032] Clinical Data Acquisition: During each clinical visit at each time point, serum samples were obtained through venous blood collection. Serum prostate-specific antigen (PSA) concentrations were measured using standard clinical laboratory methods (such as the Roche Cobase series electrochemiluminescence immunoassay analyzer and its reagents), typically expressed in nanograms per milliliter (ng / mL). Simultaneously, the Gleason score (GS) at the time of diagnosis was obtained from the patient's medical record. Integers or half-integers between 10 and 10.

[0033] II. Data Standardization and Alignment To ensure effective integration of data from different sources and dimensions, the raw data is standardized and the time series is aligned.

[0034] 1. Molecular data standardization: For data obtained from multiple time points... and The values ​​are standardized based on the distribution of a reference dataset.

[0035] During the model training phase, this reference dataset is the training dataset itself. The mean of each metric across all samples in the training dataset is calculated. ) and standard deviation ( ).

[0036] Then, for each time point of the training set, validation set, and future new samples, the original values ​​are uniformly calculated using the values ​​obtained from the training set. and Standardize using Z-scores: ; ; Standardized and The mean is Standard deviation is The real number. If a third variable is introduced (such as the expression level of a specific microRNA), the same standardization process is required to obtain... .

[0037] 2. Clinical Data Time Alignment: The time points for molecular testing and clinical PSA testing may not be completely consistent. To link the two, a time mapping relationship needs to be established. One implementation method is: for each molecular testing time point... Find the clinical PSA test record closest to the molecular time point and use its PSA value as the associated PSA value for that molecular time point, denoted as . If PSA records are available before and after the molecular detection time point, their average value can be taken, or a linear interpolation calculation can be performed. This method ensures that each molecular time point... All are associated with a clinical PSA value .

[0038] III. Calculation of SLC31A1 Functional Activity Integration Value This step integrates the standardized molecular information from each time point into a single quantitative index, called the SLC31A1 functional activity integration value (denoted as SLC31A1). ).

[0039] 1. Calculation formula: SLC31A1 functional activity integration value Calculated using the following linear combination formula: ; in: This represents the standardized expression level of SLC31A1 messenger RNA.

[0040] This represents the standardized methylation level at a specific CpG site. Since DNA methylation typically suppresses gene expression, it is represented as a negative term in the formula (…). This is used in the calculation to reflect its negative regulatory effect. It is a positive number.

[0041] The standardized expression level of a specific microRNA (if enabled). The sign (positive or negative) depends on the known regulatory direction (inhibition or promotion) of the miRNA on SLC31A1, and its specific value is determined by training.

[0042] , , These are the weight coefficients of the model, which are real numbers. When the value is zero, it is equivalent to not using a third variable.

[0043] Weighting coefficient , , The specific value will be determined through the subsequent model training process.

[0044] 2. Example parameters: Before model training begins, the weight coefficients can be initialized to specific values, such as... , , The training process will optimize these values.

[0045] IV. Extraction of Dynamic Morphological Features This step applies to different time points. Mathematical analysis is performed on the sequence of values ​​to extract morphological features that describe their patterns of change over time.

[0046] 1. Construction of activity time-series curves: The horizontal axis is plotted in chronological order (time unit: month), and... Using the vertical axis as the plot, plot the points... Connect the line segments sequentially to form a piecewise linear activity time series curve.

[0047] 2. Inflection Point Identification: An inflection point is the location on the activity time series curve where the trend changes. For discrete data point sequences... Identification is performed using the numerical difference method. The first-order difference between adjacent data points is calculated. For interior points ( from arrive ),if and The signs are opposite (i.e.) ), then the time node It was identified as a turning point.

[0048] 3. Segmented Slope Calculation: After identifying the inflection point, the entire activity time series curve is divided into several monotonic intervals. For each monotonic interval... (All within the interval) Using the same signs, calculate the average slope as the piecewise slope value for that interval. .

[0049] The calculation formula is: ,in and These are the start and end points of the interval, respectively. value, and This corresponds to the absolute time (in months). The unit is "standardized activity unit / month".

[0050] 4. Cumulative Deviation Calculation: The cumulative deviation value is used to quantify the degree to which the overall activity time series curve deviates from the linear trend. First, a baseline is defined, which is the first point connecting the sequences. and the last point straight line Its equation is .

[0051] Then, the area (absolute value) of the region enclosed between the original activity time-series curve and the baseline is calculated as the cumulative bias value. Since the data is discrete, the area can be approximated using the trapezoidal rule: .

[0052] The unit is "standardized activity unit month". The larger the value, the stronger the fluctuation in the activity change process.

[0053] V. Matching of adverse prognostic patterns This step compares the extracted dynamic morphological features with a predefined pattern feature library that represents poor prognosis and calculates the matching degree.

[0054] 1. Construction of a feature library for poor prognostic patterns: This feature library is derived from the analysis of historical patient cohort data. Patients in the historical cohorts have known prognostic outcomes (such as time to disease progression).

[0055] By analyzing the activity time-series curves of patients with poor prognosis (e.g., those with disease progression time shorter than a certain clinical consensus threshold), common morphological features are identified and quantified into matchable patterns. As a specific implementation method, the pattern feature library includes the following two patterns: Pattern features (Rapidly Rising Type): This pattern is defined as the existence of a monotonic interval in the activity time series curve, with its segmented slope values... Exceeding a set positive threshold . The distribution of the slope of the curve for patients with poor historical prognosis can be determined by statistical analysis, for example, by taking the higher percentile.

[0056] Pattern features (Step-by-step increasing type): This mode is defined as having at least three consecutive time points in the activity time series curve. The value is strictly monotonically increasing (i.e.) Furthermore, this increase is associated with a rebound in clinical serum PSA levels.

[0057] The correlation is defined as the PSA value recorded at the next clinical visit time point after each incremental step is completed. PSA value compared to the previous clinical visit The increase exceeds a certain percentage (For example, ), This can be determined through historical data analysis.

[0058] threshold and These are part of the pattern feature library, which is determined during the model training phase.

[0059] 2. Matching Degree Calculation: The dynamic morphological features extracted from the target patient are compared with each item in the pattern feature library. Matching Degree Initialize to The matching rules are as follows: If there exists any ,but , For matching pattern features The score obtained.

[0060] If there exists a pattern feature If three consecutive time points are defined, then , For matching pattern features The score obtained.

[0061] and These are the trainable parameters of the model, and their initial value can be set to 1.0.

[0062] 3. Time Weighting Adjustment: As a preferred implementation method, to reflect the clinical prior knowledge that recent biological changes may have a greater impact on current prognosis, a time weight is introduced into the matching degree calculation. Therefore, a time weighting function is defined. This function maps time to a weighting factor. For example, a linear weighting function could be used: applying the entire observation timeline. Normalize to the interval [0, 1], let .definition ,in It is greater than And the parameters are trainable.

[0063] This makes the later ( near Features are assigned higher weights, and for each matched feature, its normalized time is calculated based on the time when the feature mainly occurs. And obtain the weights The final weighted matching degree The calculation is as follows: ; The summation applies to all matched features.

[0064] VI. Risk Level Output This step combines a weighted matching degree with the patient's current status and outputs the final risk level through a predefined decision rule.

[0065] 1. Decision function input preparation: Calculate the weighted matching degree... The integrated functional activity value of SLC31A1 at the current (most recent) time point. (Right now ), and baseline clinical characteristics, namely the Gleason score. , which serves as the input to the decision function.

[0066] To ensure that different variables are of consistent magnitude, they can be standardized or normalized before being input into the decision function. For example, Z-score standardization can be performed using the mean and standard deviation of each variable in the training set.

[0067] 2. Decision function form: The decision function maps the above inputs to a risk score. This function can be linear or non-linear. As a specific implementation, a linear function is used: ; in, , , , These are the regression coefficients of the model, obtained through training. Before training, these coefficients can be initialized to... .

[0068] 3. Risk Level Classification: Based on the calculated risk score This is done by dividing the risk into discrete levels. This is typically achieved by setting a threshold. For example, setting two thresholds. and ( ): like If so, it is judged as low risk.

[0069] like If so, it is judged as medium risk.

[0070] like If so, it is judged as high risk.

[0071] threshold and This can be determined by optimizing classification performance on the training set (e.g., maximizing discriminative power).

[0072] VII. Inference of Treatment Strategy Tendency This step is optional and is usually triggered only when the risk level is determined to be high, and is intended to provide a reference for treatment direction.

[0073] 1. Basis for inference: The inference is based on which adverse prognostic pattern feature contributed the main score in the matching degree calculation.

[0074] 2. Inference Rules: As a preferred implementation method, the following mapping relationship is established: If the matching degree mainly comes from pattern features For (rapidly rising type), the preferred treatment strategy is "copper metabolism intervention therapy".

[0075] If the matching degree mainly comes from pattern features (Stepwise incremental type) In this case, the preferred treatment strategy category is "DNA damage repair targeted therapy".

[0076] 3. Output: This step outputs a text-based treatment strategy recommendation as supplementary information to the prognostic report.

[0077] Example 2: Training and Determination of Model Parameters This embodiment illustrates how to use machine learning methods to determine all the parameters that need to be learned in Embodiment 1 using historical data.

[0078] 1. Training Dataset Preparation: Collect historical cohort data of castration-resistant prostate cancer patients with complete follow-up information. For each patient in the cohort, the following are required: time-series raw data of tumor molecular testing (for computation). , (etc.), time-series clinical serum PSA test data, baseline Gleason score, etc. And clearly recorded time to disease progression (TTP, in months) and / or overall survival (OS, in months) data.

[0079] 2. Parameter set: The parameters that need to be determined during training include (but are not limited to): Reference mean required for standardization , , , .

[0080] Weighting coefficients in the calculation of functional activity integration value , , .

[0081] Threshold parameters in the feature library of poor prognostic patterns , .

[0082] The score in the matching degree calculation , and time weighting function parameters .

[0083] Regression coefficients in the decision function , , , and risk level classification thresholds , .

[0084] 3. Training process: a. Constructing a trainable computational graph: Inputting the raw data as described in Example 1 into the risk score The entire output process is implemented as a differentiable computational graph. This requires transforming discontinuous or discrete decision operations in the process into differentiable approximations. Specifically: For threshold comparisons (e.g.) (This is followed by a seemingly unrelated sentence about using the sigmoid function for a smooth approximation.) ,in It is the sigmoid function. It is a large positive number used to control the steepness of the approximation.

[0085] For logical judgments in pattern matching, similarly differentiable logical operations are used for approximation.

[0086] This transformation allows the entire process to be represented by differentiable operations, thus enabling gradient backpropagation. This allows the model to automatically learn parameters in the mapping process from molecular data to activity index and then to morphological feature matching, including thresholds that would otherwise be required in traditional methods, achieving end-to-end optimization with a focus on prognostic prediction.

[0087] b. Define the loss function: The loss function measures the difference between the model's prediction and the actual prognosis. Since the prognostic endpoint is time-to-tribulation (TTP) data, a loss function based on survival analysis is used, such as the negative logarithmic form of the partial likelihood loss function for the Cox proportional hazards model. This loss function causes the model to output a higher risk score R for patients with shorter TTPs and a lower R for patients with longer TTPs.

[0088] c. Optimization algorithm: Use stochastic gradient descent or its improved algorithms (such as the Adam optimizer) to iteratively optimize all parameters of the entire computation graph on the training dataset to minimize the loss function.

[0089] d. Validation and Early Stopping: Use a separate validation dataset to monitor the model's performance on unseen data during training (e.g., calculate the C-index). Stop training when the validation set performance no longer improves to prevent overfitting.

[0090] e. Parameter Fixation: After training, a set of optimal parameters will be obtained. These parameters are fixed and combined with fixed data processing logic (steps one through seven) to obtain a complete, deployable prognostic prediction model. This model is used to process input data from new patients and generate prognostic prediction reports.

[0091] 4. Explanation of process relationships: Step 1 (data preparation) is the input interface for the model; Step 2 (standardization and alignment) is the data preprocessing step; Steps 3 through 6 constitute the computational part of the model, transforming the input into features and scores; Step 7 is the application based on the scores. This training process uses historical data to back-optimize all adjustable parameters in the computational part, enabling the model to generate predictions for new input data.

[0092] The above embodiments demonstrate the complete implementation process of the technical solution protected by the claims of this invention. Those skilled in the art can adjust the specific parameters and calculation details within the scope defined by the claims based on the above description, and such adjustments should all be considered within the protection scope of this invention.

Claims

1. A prognostic prediction model for castration-resistant prostate cancer based on SLC31A1, characterized in that, Its internal data processing logic includes the following steps executed sequentially: S1. Input the time-point molecular and clinical data of the target patient; S2. Based on the molecular data, calculate the SLC31A1 functional activity integration value at each time point. The calculation of this value integrates the SLC31A1 messenger ribonucleic acid expression level at the corresponding time point with the methylation level of a specific CpG site located in the promoter region of the SLC31A1 gene, wherein the methylation level participates in the integration with a negative weight. S3. Construct the SLC31A1 functional activity integration value at each time point into an activity time-series curve according to time sequence, and extract the dynamic morphological features of the curve. The dynamic morphological features include the key inflection points of the curve, the slope values ​​of each curve segment, and the cumulative deviation value of the curve relative to a preset baseline. S4. The extracted dynamic morphological features are matched with a pre-stored poor prognosis pattern feature library, which contains a set of morphological features extracted from the activity time-series curves of patients with poor prognosis in the historical cohort. S5. Based on the matching degree and the SLC31A1 functional activity integration value at the current time point, output the quantitative prognostic risk level for the target patient through a predefined risk stratification rule.

2. The prognostic prediction model for castration-resistant prostate cancer based on SLC31A1 according to claim 1, characterized in that, In step S2, the calculation of the SLC31A1 functional activity integration value also introduces a third variable, which is the expression level of a specific microRNA that has a known regulatory relationship with SLC31A1.

3. The prognostic prediction model for castration-resistant prostate cancer based on SLC31A1 according to claim 1, characterized in that, In step S3, the extraction of dynamic morphological features specifically includes: S3.1 Identify the points in the activity timeline curve where the first derivative is zero or the second derivative is an extreme value, and use them as the key inflection points; S3.2 Calculate the average slope of the curve segment defined by adjacent key inflection points, as the segment slope value; S3.3 Calculate the area enclosed by the active time-series curve and the straight line connecting the start and end points of the curve, and use it as the cumulative deviation value.

4. The prognostic prediction model for castration-resistant prostate cancer based on SLC31A1 according to claim 1, characterized in that, The pre-stored adverse prognostic pattern feature library mentioned in step S4 includes: First mode feature: The slope value of the activity time-series curve between two consecutive detection time points is greater than a first preset slope threshold; The second mode feature is that the activity time-series curve contains no fewer than three monotonically increasing steps, and the serum prostate-specific antigen level recorded at the first clinical testing time point after each step increases by more than a preset percentage threshold compared to the recorded value at the previous clinical testing time point.

5. The prognostic prediction model for castration-resistant prostate cancer based on SLC31A1 according to claim 4, characterized in that, The matching degree calculation in step S4 is a weighted calculation, wherein, within the time span of the active time series curve, the dynamic morphological features extracted in the latter 50% of the time period have a higher calculation weight than the features extracted in the first 50% of the time period.

6. The prognostic prediction model for castration-resistant prostate cancer based on SLC31A1 according to claim 1, characterized in that, The predefined risk stratification rule in step S5 is a decision function whose input variables include the matching degree, the integrated value of SLC31A1 functional activity at the current time point, and the Gleason score in the clinical data.

7. The prognostic prediction model for castration-resistant prostate cancer based on SLC31A1 according to claim 1, characterized in that, The model further includes step S6: S6. If the quantitative prognostic risk level output in step S5 is a preset high-risk level, then the treatment strategy tendency inference step is executed; this step outputs a corresponding preferred treatment strategy category based on the specific adverse prognostic pattern matched by the dynamic morphological features.

8. The prognostic prediction model for castration-resistant prostate cancer based on SLC31A1 according to claim 7, characterized in that, The association rule used in the treatment strategy bias inference step S6 is as follows: If the dynamic morphological features match the first pattern features, then the output preferred treatment strategy category is copper metabolism intervention therapy; If the dynamic morphological features match the second pattern features, then the output preferred treatment strategy category is DNA damage repair targeted therapy.

9. The prognostic prediction model for castration-resistant prostate cancer based on SLC31A1 according to claim 1, characterized in that, The specific CpG site in the promoter region of the SLC31A1 gene mentioned in step S2 refers to a CpG dinucleotide site located in the region of 200 to 50 base pairs upstream of the transcription start site of the SLC31A1 gene.

10. The prognostic prediction model for castration-resistant prostate cancer based on SLC31A1 according to claim 1, characterized in that, In the data processing logic within the model, the negative weights in step S2, the adverse prognostic pattern feature library and matching threshold in step S4, and the risk stratification rule parameters in step S5 are all obtained by training a machine learning model using a time-series training dataset of castration-resistant prostate cancer patients containing known disease progression time and overall survival data.

Citation Information

Patent Citations

  • Tumor marker transmembrane protein SLC31A1 and application thereof

    CN116859048A