Model hypothesis-independent survival analysis marker discovery method and device

By constructing a model-free biomarker selection optimization model in the regenerating nuclear Hilbert space, quantifying dependencies and eliminating redundant features, the reliability problem caused by the reliance on assumptions in existing methods is solved, and robust biomarker discovery and prediction are achieved.

CN121963846APending Publication Date: 2026-05-01TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2025-12-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing methods for discovering survival biomarkers rely on linear and proportional hazards assumptions, which make it difficult to capture the complexity of real-world survival dynamics. This leads to model misspecification and reduced reliability, affecting the generalization ability of prediction tasks and clinical applications.

Method used

By mapping the features of biomarkers to the survival outcome kernel into the regeneration kernel Hilbert space, a model-free survival analysis biomarker selection optimization model is constructed. The dependency relationship is quantified using the nucleated survival outcome dependency module, and redundant features are eliminated by the nucleated minimum redundancy module to identify variables with significant survival associations.

Benefits of technology

Without requiring parameter assumptions, this method effectively screens biomarker combinations with strong predictive power and low redundancy, improving the discriminative power and interpretability of biomarker combinations and ensuring the robustness and reliability of predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963846A_ABST
    Figure CN121963846A_ABST
Patent Text Reader

Abstract

The invention provides a model hypothesis-independent survival analysis marker discovery method and device, and the method comprises the steps: collecting the follow-up survival data of a plurality of patients, and building a survival data set, samples of the data set comprise multi-dimensional features of biomarkers reflecting prognosis conditions of patients, actual observation time and labels indicating whether interested events are observed or not; constructing a model-free survival analysis marker selection optimization model, wherein the optimization model comprises a nucleation survival outcome dependency module for quantifying a dependency relationship between each feature and a survival outcome and a nucleation minimum redundancy module for eliminating redundant features; and solving the optimization model by utilizing the survival data set to obtain a screening result of the biomarker for predicting the prognosis condition of the patient. According to the method, the biomarker combination with high predictability and low redundancy can be effectively screened out under the condition that a preset model form is not needed, so that the discrimination capability and interpretability of the biomarker combination are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent life science technology, and specifically relates to a method and apparatus for discovering survival analysis biomarkers that are independent of model assumptions. Background Technology

[0002] In survival analysis, survival-related biomarkers serve as biological indicators for predicting prognosis or early disease risk. Their core value lies in identifying combinations of biomarkers with strong predictive power and the fewest possible number, which is crucial for reliable decision-making and practical clinical applications.

[0003] Existing methods for discovering survival biomarkers often rely on assumptions about the data generation process, failing to capture the complexities of real-world survival dynamics. This leads to significant model misspecification problems, affecting their generalization ability and reliability. In clinical practice, the large-scale application of unreliable survival biomarkers will seriously impact public health and waste medical resources.

[0004] Recent high-impact studies on early disease detection have employed the well-known regularized Cox proportional hazard (Cox-PH) model to identify protein biomarkers strongly associated with various diseases. The effectiveness of such regularized Cox-PH models heavily relies on two fundamental assumptions in the data generation process: the linearity assumption and the proportional hazard (PH) assumption. The PH assumption requires that the hazard ratio between patient groups remains constant over time, while the linearity assumption presupposes a linear relationship between the logarithm of the hazard ratio and the covariate. Due to the complex and dynamic nature of living systems, these assumptions are easily violated in real biological processes. Traditional biomarker discovery methods, relying on fragile statistical assumptions, often identify variables with spurious survival-related factors. These spurious associations caused by model misspecification, due to changes in covariates driven by population heterogeneity, often cannot be reproduced in independent cohorts. This poses a serious generalization challenge to predictive tasks and carries significant risks for clinical applications.

[0005] For example, recent research increasingly reveals unique age-dependent epigenetic changes at different stages of life, suggesting that epigenetic markers based on nonlinear or segmented relationships may more accurately characterize human lifespan than linear age-related markers. More research indicates strong nonlinear biological associations between plasma proteomic markers or DNA methylation markers and aging outcomes. Furthermore, the impact of certain genes on survival is dynamic and may change over time, violating the PH hypothesis. For instance, the expression levels of immune-related genes such as interleukin-6 (IL-6) or genes involved in cell cycle regulation may be associated with early tumor progression, but their significance may disappear or change over time. In several clinical cohorts predominantly lymph node-positive, female patients with ER-positive (ER+, encoded by ESR1 and ESR2 genes) tumors had better recurrence-free survival in the first two years after diagnosis than ER-negative (ER-) patients, but this effect reversed or diminished over time. This strong evidence suggests that the underlying biological processes in reality are often highly complex.

[0006] The most common approach to predictive biomarker identification currently is to impose sparsity regularization penalties on the Cox PH model, including methods such as Lasso and elastic networks. While these methods have been successful in promoting coefficient sparsity, their effectiveness still depends on whether real-world data conforms to the assumptions of the Cox PH model. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of existing technologies and propose a model-assumption-independent method and apparatus for discovering survival biomarkers. This invention maps the features of biomarkers to a regeneration kernel Hilbert space, enabling the identification of variables with significant survival associations without parametric assumptions. When the kernel-quantified survival dependence of biomarkers is redundant, this invention can reliably select effective biomarkers.

[0008] A first aspect of this invention proposes a method for discovering survival analysis biomarkers independent of model assumptions, comprising:

[0009] Collect follow-up survival data from multiple patients to establish a survival dataset; the samples in the survival dataset include: multidimensional features of biomarkers reflecting patient prognosis, actual observation time, and labels indicating whether an event of interest was observed;

[0010] An optimization model for selecting biomarkers in model-free survival analysis is constructed. The optimization model includes: a kernelized survival outcome dependency module for quantifying the dependency between each feature and the survival outcome, and a kernelized minimum redundancy module for eliminating redundant features.

[0011] Using the survival dataset, the optimization model is solved to obtain the screening results of biomarkers for predicting patient prognosis.

[0012] In one specific embodiment of the present invention, it further includes:

[0013] The survival dataset is denoted as ;

[0014] in, Let be the covariate vector of the multidimensional features of biomarkers for the i-th patient, denoted by dimension d. ; The actual observation time for the i-th patient. This represents the survival time of the i-th patient. Let be the censoring time for the i-th patient; This indicates whether an event of interest was observed in the i-th patient; if ,but This indicates that the event of interest to the i-th patient has been observed; otherwise, , indicating that no event of interest to the i-th patient was observed; n is the number of samples.

[0015] In one specific embodiment of the present invention, it further includes:

[0016] The optimized model for biomarker selection in model-free survival analysis is expressed as follows:

[0017]

[0018] in, It is a set of feature importance for biomarkers. The importance of the i-th feature; To adjust the hyperparameters of sparsity; The k-th feature in the multidimensional feature covariates of a biomarker Survival Outcome The nuclear independence measure, where T represents the actual observation time. Labels indicating whether an event of interest was observed; HSIC For use in measuring features and A measure of kernel independence that determines the nonlinear dependence between them; and HSIC All are non-negative values.

[0019] In one specific embodiment of the present invention, it further includes:

[0020]

[0021] Among them, matrix , and The expressions are as follows:

[0022]

[0023] in, ()and () represent kernel calculation methods; Indicator symbol;

[0024] The k-th feature of the biomarkers for the i-th patient;

[0025]

[0026] in, It is a vector of all 1s. for The Gram matrix.

[0027] In one specific embodiment of the present invention, it further includes:

[0028] use Solve the optimization model;

[0029] During the solution process, the solution is iterated over d features, where the index of the current feature is denoted as j.

[0030] Calculate redundant terms ;

[0031] Get dependencies ;

[0032] Soft threshold update ;

[0033] At the end of each iteration, a set of feature importance values ​​is obtained. ;

[0034] When the iteration reaches the preset termination condition, the solution is complete, and the final set of feature importance is output. ;

[0035] After the solution is completed, all features are arranged in descending order of importance, and the biomarkers corresponding to the top m features with the highest importance are selected as the screening results for biomarkers used to predict patient prognosis.

[0036] A second aspect of the present invention provides a model assumption-independent survival analysis biomarker discovery device, comprising:

[0037] The survival data acquisition module is used to collect follow-up survival data from multiple patients and establish a survival dataset. The samples in the survival dataset include: multidimensional features of biomarkers reflecting patient prognosis, actual observation time, and labels indicating whether an event of interest was observed.

[0038] An optimization model building module is used to build an optimization model for selecting biomarkers in model-free survival analysis. The optimization model includes: a kernelized survival outcome dependency module for quantifying the dependency between each feature and the survival outcome, and a kernelized minimum redundancy module for eliminating redundant features.

[0039] The biomarker screening module is used to solve the optimization model using the survival dataset to obtain the screening results of biomarkers for predicting patient prognosis.

[0040] In one specific embodiment of the present invention, it further includes:

[0041] The survival dataset is denoted as ;

[0042] in, Let be the covariate vector of the multidimensional features of biomarkers for the i-th patient, denoted by dimension d. ; The actual observation time for the i-th patient. This represents the survival time of the i-th patient. Let be the censoring time for the i-th patient; This indicates whether an event of interest was observed in the i-th patient; if ,but This indicates that the event of interest to the i-th patient has been observed; otherwise, , indicating that no event of interest to the i-th patient was observed; n is the number of samples.

[0043] In one specific embodiment of the present invention, it further includes:

[0044] The optimized model for biomarker selection in model-free survival analysis is expressed as follows:

[0045]

[0046] in, It is a set of feature importance for biomarkers. The importance of the i-th feature; To adjust the hyperparameters of sparsity; The k-th feature in the multidimensional feature covariates of a biomarker Survival Outcome The nuclear independence measure, where T represents the actual observation time. Labels indicating whether an event of interest was observed; HSIC For use in measuring features and A measure of kernel independence that determines the nonlinear dependence between them; and HSIC All are non-negative values.

[0047] In one specific embodiment of the present invention, it further includes:

[0048]

[0049] Among them, matrix , and The expressions are as follows:

[0050]

[0051] in, ()and () represent the kernel calculation method kernel; Indicator symbol;

[0052] The k-th feature of the biomarkers for the i-th patient;

[0053]

[0054] in, It is a vector of all 1s. for The Gram matrix.

[0055] In one specific embodiment of the present invention, it further includes:

[0056] use Solve the optimization model;

[0057] During the solution process, the solution is iterated over d features, where the index of the current feature is denoted as j.

[0058] Calculate redundant terms ;

[0059] Get dependencies ;

[0060] Soft threshold update ;

[0061] At the end of each iteration, a set of feature importance values ​​is obtained. ;

[0062] When the iteration reaches the preset termination condition, the solution is complete, and the final set of feature importance is output. ;

[0063] After the solution is completed, all features are arranged in descending order of importance, and the biomarkers corresponding to the top m features with the highest importance are selected as the screening results for biomarkers used to predict patient prognosis.

[0064] A third aspect of the present invention provides an electronic device comprising:

[0065] At least one processor; and a memory communicatively connected to said at least one processor;

[0066] The memory stores instructions executable by the at least one processor, the instructions being configured to perform the aforementioned model hypothesis-independent survival analysis biomarker discovery method.

[0067] A fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions for causing the computer to perform the above-described model hypothesis-independent survival analysis biomarker discovery method.

[0068] Features and beneficial effects of the present invention:

[0069] Existing methods are typically based on linear Cox proportional hazards models and rely on strong assumptions such as linearity and proportional hazards. These assumptions are often difficult to satisfy in the complex mechanisms of human survival, leading to model misspecification and reduced reliability. This invention fundamentally solves the problem of predictive reliability caused by the invalidity of model assumptions by avoiding reliance on easily violated assumptions, thus ensuring the robustness of biomarker discovery.

[0070] The core feature of this invention lies in the joint optimization of two nucleation modules: a nucleation survival outcome dependency module, used to quantify the dependency relationship between each biomarker feature and the survival outcome; and a nucleation minimum redundancy module, used to eliminate redundant features. This joint optimization mechanism can effectively screen biomarker combinations with strong predictive power and low redundancy without requiring a pre-defined model, thereby improving the discriminative power and interpretability of biomarker combinations and demonstrating stable predictive performance across multiple cancer datasets.

[0071] This invention can be widely applied to survival data analysis in transcriptomics and proteomics, and is suitable for prognostic prediction and early disease detection of various cancer types, such as in real-world data environments like cohort studies (e.g., the Lifelines Cohort Study) and biobanks. Its value lies in its ability to robustly identify key biomarkers across cancer types, effectively distinguishing patient subtypes, thereby providing a reliable basis for clinical decision-making and contributing to the development of precision medicine and personalized treatment strategies. Attached Figure Description

[0072] Figure 1 This is an overall flowchart of a survival analysis biomarker discovery method that is independent of model assumptions, according to an embodiment of the present invention. Detailed Implementation

[0073] This invention proposes a method and apparatus for discovering survival analysis biomarkers that are independent of model assumptions, which will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0074] A first aspect of this invention proposes a method for discovering survival analysis biomarkers independent of model assumptions, comprising:

[0075] Collect follow-up survival data from multiple patients to establish a survival dataset; the samples in the survival dataset include: multidimensional features of biomarkers reflecting patient prognosis, actual observation time, and labels indicating whether an event of interest was observed;

[0076] An optimization model for selecting biomarkers in model-free survival analysis is constructed. The optimization model includes: a kernelized survival outcome dependency module for quantifying the dependency between each feature and the survival outcome, and a kernelized minimum redundancy module for eliminating redundant features.

[0077] Using the survival dataset, the optimization model is solved to obtain the screening results of biomarkers for predicting patient prognosis.

[0078] In a specific embodiment of the present invention, the overall process of the method for discovering survival analysis biomarkers independent of model assumptions is as follows: Figure 1 As shown, it includes the following steps:

[0079] 1) Collect follow-up survival data from multiple patients to establish a survival dataset; the samples in the survival dataset include: multidimensional features of biomarkers reflecting patient prognosis, actual observation time, and labels indicating whether an event of interest was observed.

[0080] In this embodiment, the follow-up survival data comes from the same cohort of patients, and the obtained survival dataset is denoted as... .

[0081] in, Let d be the covariate vector of the multidimensional features of the biomarkers of the i-th patient. In this embodiment, the dimension is denoted as d. In one specific embodiment of the present invention, the multidimensional features of the biomarker are the patient's genomic and proteomic features, where d is the dimension of the omics features. The actual observation time for the i-th patient. This represents the survival time of the i-th patient. Let be the censoring time for the i-th patient. This indicates whether an event of interest was observed in the i-th patient; if ,but This indicates that an event of interest to the i-th patient has been observed, such as the patient's death; otherwise, , indicating that the event of interest to the i-th patient has not been observed, such as no patient death has been observed so far. n is the number of samples.

[0082] 2) Construct a model-free survival analysis biomarker selection optimization model, with the following expression:

[0083]

[0084] in, It is a set of feature importance for biomarkers. Let represent the importance of the i-th feature. To adjust the hyperparameters of sparsity, The k-th feature in the multidimensional feature covariates of a biomarker Survival Outcome The kernel independence measure (called the kernel log-rank statistic KLRS), where T represents the actual observation time. Labels indicating whether an event of interest was observed; HSIC It is a kernel-based independence measure (Hilbert-Schmidt Independence Criterion HSIC) used to measure features and The non-linear dependency between them. KLRS and HSIC always take non-negative values. For either KLRS or HSIC, the value of the indicator is zero if and only if the two input variables corresponding to the indicator are independent, i.e., HSIC =0 or =0. If the feature With output A strong correlation leads to an increase in the KLRS value, resulting in Increase accordingly to optimize the objective function; conversely, if and Independent (i.e., KLRS=0) Will be Regular terms Compress to zero ( (Controlling sparsity intensity). Furthermore, if and Strong correlation (i.e. redundant features), HSIC ( , A larger value will force or Approaching zero means that redundant features will be automatically eliminated by the objective function. All The non-negative constraint ensures that its value reflects the importance of the feature for marker selection.

[0085]

[0086] Among them, matrix , and These are all intermediate variables, constructed from the kernel function and survival data respectively, and their expressions are as follows:

[0087]

[0088] in, ()and () represent kernel calculation methods, typically Gaussian kernels. For indication purposes.

[0089] Let k be the biomarker feature of the i-th patient.

[0090]

[0091] in, It is a vector of all 1s. for The Gram matrix.

[0092] 3) Use the survival dataset obtained in step 1) to solve the optimization model established in step 2) to obtain the screening results of biomarkers for predicting patient prognosis.

[0093] In this embodiment, using Solve the optimization model established in step 2).

[0094] During the solution process, the solution is iterated over d features, where the index of the current feature is denoted as j.

[0095] Calculate redundant terms ;

[0096] Get dependencies ;

[0097] Soft threshold update ;

[0098] At the end of each iteration, a set of feature importance values ​​is obtained. ;

[0099] The solution is complete when the iteration reaches the preset termination condition (i.e., iteration reaches convergence or the maximum number of iterations is reached), and the final set of feature importance is output. .

[0100] After the solution is completed, all features are arranged in descending order of importance. In this embodiment, the biomarkers corresponding to the top m features with the highest importance are selected as the screening results for biomarkers used to predict patient prognosis.

[0101] To achieve the above embodiments, a second aspect of the present invention provides a model assumption-independent survival analysis biomarker discovery device, comprising:

[0102] The survival data acquisition module is used to collect follow-up survival data from multiple patients and establish a survival dataset. The samples in the survival dataset include: multidimensional features of biomarkers reflecting patient prognosis, actual observation time, and labels indicating whether an event of interest was observed.

[0103] An optimization model building module is used to build an optimization model for selecting biomarkers in model-free survival analysis. The optimization model includes: a kernelized survival outcome dependency module for quantifying the dependency between each feature and the survival outcome, and a kernelized minimum redundancy module for eliminating redundant features.

[0104] The biomarker screening module is used to solve the optimization model using the survival dataset to obtain the screening results of biomarkers for predicting patient prognosis.

[0105] In one specific embodiment of the present invention, it further includes:

[0106] The survival dataset is denoted as ;

[0107] in, Let d be the covariate vector of the multidimensional features of biomarkers for the i-th patient, with dimension denoted as d. ; The actual observation time for the i-th patient. This represents the survival time of the i-th patient. Let be the censoring time for the i-th patient; This indicates whether an event of interest was observed in the i-th patient; if ,but This indicates that the event of interest to the i-th patient has been observed; otherwise, , indicating that no event of interest to the i-th patient was observed; n is the number of samples.

[0108] In one specific embodiment of the present invention, it further includes:

[0109] The optimized model for biomarker selection in model-free survival analysis is expressed as follows:

[0110]

[0111] in, It is a set of feature importance for biomarkers. The importance of the i-th feature; To adjust the hyperparameters of sparsity; The k-th feature in the multidimensional feature covariates of a biomarker Survival Outcome The nuclear independence measure, where T represents the actual observation time. Labels indicating whether an event of interest was observed; HSIC For use in measuring features and A measure of kernel independence that determines the nonlinear dependence between them; and HSIC All are non-negative values.

[0112] In one specific embodiment of the present invention, it further includes:

[0113]

[0114] Among them, matrix , and The expressions are as follows:

[0115]

[0116] in, ()and () represent kernel calculation methods; Indicator symbol;

[0117] The k-th feature of the biomarkers for the i-th patient;

[0118]

[0119] in, It is a vector of all 1s. for The Gram matrix.

[0120] In one specific embodiment of the present invention, it further includes:

[0121] use Solve the optimization model;

[0122] During the solution process, the solution is iterated over d features, where the index of the current feature is denoted as j.

[0123] Calculate redundant terms ;

[0124] Get dependencies ;

[0125] Soft threshold update ;

[0126] At the end of each iteration, a set of feature importance values ​​is obtained. ;

[0127] When the iteration reaches the preset termination condition, the solution is complete, and the final set of feature importance is output. ;

[0128] After the solution is completed, all features are arranged in descending order of importance, and the biomarkers corresponding to the top m features with the highest importance are selected as the screening results for biomarkers used to predict patient prognosis.

[0129] To implement the above embodiments, a third aspect of the present invention provides an electronic device, comprising:

[0130] At least one processor; and a memory communicatively connected to said at least one processor;

[0131] The memory stores instructions executable by the at least one processor, the instructions being configured to perform the aforementioned model hypothesis-independent survival analysis biomarker discovery method.

[0132] To implement the above embodiments, a fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions for causing the computer to execute the above-described model hypothesis-independent survival analysis biomarker discovery method.

[0133] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0134] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to perform a model assumption-independent survival analysis biomarker discovery method according to the above embodiments.

[0135] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0136] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0137] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0138] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this application pertain.

[0139] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0140] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0141] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0142] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0143] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for discovering survival analysis biomarkers independent of model assumptions, characterized in that, include: Collect follow-up survival data from multiple patients and establish a survival dataset; The survival dataset includes samples of: multidimensional features of biomarkers reflecting patient prognosis, actual observation time, and labels indicating whether an event of interest was observed; An optimization model for selecting biomarkers in model-free survival analysis is constructed. The optimization model includes: a kernelized survival outcome dependency module for quantifying the dependency between each feature and the survival outcome, and a kernelized minimum redundancy module for eliminating redundant features. Using the survival dataset, the optimization model is solved to obtain the screening results of biomarkers for predicting patient prognosis.

2. The method according to claim 1, characterized in that, Also includes: The survival dataset is denoted as ; in, Let d be the covariate vector of the multidimensional features of biomarkers for the i-th patient, with dimension denoted as d. ; The actual observation time for the i-th patient. This represents the survival time of the i-th patient. Let be the censoring time for the i-th patient; This indicates whether an event of interest was observed in the i-th patient; if ,but This indicates that the event of interest to the i-th patient has been observed; otherwise, , indicating that no event of interest to the i-th patient was observed; n is the number of samples.

3. The method according to claim 2, characterized in that, Also includes: The optimized model for biomarker selection in model-free survival analysis is expressed as follows: in, It is a set of feature importance for biomarkers. The importance of the i-th feature; To adjust the hyperparameters of sparsity; The k-th feature in the multidimensional feature covariates of a biomarker Survival Outcome The nuclear independence measure, where T represents the actual observation time. Labels indicating whether an event of interest was observed; HSIC For use in measuring features and A measure of kernel independence that determines the nonlinear dependence between them; and HSIC All are non-negative values.

4. The method according to claim 3, characterized in that, Also includes: Among them, matrix , and The expressions are as follows: in, ()and () represent kernel calculation methods; Indicator symbol; The k-th feature of the biomarkers for the i-th patient; in, It is a vector of all 1s. for The Gram matrix.

5. The method according to claim 4, characterized in that, Also includes: use Solve the optimization model; During the solution process, the solution is iterated over d features, where the index of the current feature is denoted as j. Calculate redundant terms ; Get dependencies ; Soft threshold update ; At the end of each iteration, a set of feature importance values ​​is obtained. ; When the iteration reaches the preset termination condition, the solution is complete, and the final set of feature importance is output. ; After the solution is completed, all features are arranged in descending order of importance, and the biomarkers corresponding to the top m features with the highest importance are selected as the screening results for biomarkers used to predict patient prognosis.

6. A device for discovering survival analysis biomarkers independent of model assumptions, characterized in that, include: The survival data acquisition module is used to collect follow-up survival data from multiple patients and establish a survival dataset. The survival dataset includes samples of: multidimensional features of biomarkers reflecting patient prognosis, actual observation time, and labels indicating whether an event of interest was observed; An optimization model building module is used to build an optimization model for selecting biomarkers in model-free survival analysis. The optimization model includes: a kernelized survival outcome dependency module for quantifying the dependency between each feature and the survival outcome, and a kernelized minimum redundancy module for eliminating redundant features. The biomarker screening module is used to solve the optimization model using the survival dataset to obtain the screening results of biomarkers for predicting patient prognosis.

7. The apparatus according to claim 6, characterized in that, Also includes: The survival dataset is denoted as ; in, Let d be the covariate vector of the multidimensional features of biomarkers for the i-th patient, with dimension denoted as d. ; The actual observation time for the i-th patient. This represents the survival time of the i-th patient. Let be the censoring time for the i-th patient; This indicates whether an event of interest was observed in the i-th patient; if ,but This indicates that the event of interest to the i-th patient has been observed; otherwise, , indicating that no event of interest to the i-th patient was observed; n is the number of samples.

8. The apparatus according to claim 7, characterized in that, Also includes: The optimized model for biomarker selection in model-free survival analysis is expressed as follows: in, It is a set of feature importance for biomarkers. The importance of the i-th feature; To adjust the hyperparameters of sparsity; The k-th feature in the multidimensional feature covariates of a biomarker Survival Outcome The nuclear independence measure, where T represents the actual observation time. Labels indicating whether an event of interest was observed; HSIC For use in measuring features and A measure of kernel independence that determines the nonlinear dependence between them; and HSIC All are non-negative values.

9. The apparatus according to claim 8, characterized in that, Also includes: Among them, matrix , and The expressions are as follows: in, ()and () represent kernel calculation methods; Indicator symbol; The k-th feature of the biomarkers for the i-th patient; in, It is a vector of all 1s. for The Gram matrix.

10. The apparatus according to claim 9, characterized in that, Also includes: use Solve the optimization model; During the solution process, the solution is iterated over d features, where the index of the current feature is denoted as j. Calculate redundant terms ; Get dependencies ; Soft threshold update ; At the end of each iteration, a set of feature importance values ​​is obtained. ; When the iteration reaches the preset termination condition, the solution is complete, and the final set of feature importance is output. ; After the solution is completed, all features are arranged in descending order of importance, and the biomarkers corresponding to the top m features with the highest importance are selected as the screening results for biomarkers used to predict patient prognosis.