Dynamic risk stratification method and device for acute myelogenous leukemia

By fusion and feature screening of multi-source data, combined with random survival forest model and dynamic clustering technology, the problems of insufficient utilization of genomic data and insufficient analysis of feature interaction effects in the existing technology are solved, and a more accurate and personalized risk assessment of acute myeloid leukemia is achieved.

CN120015326AInactive Publication Date: 2025-05-16ZHEJIANG LAB
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510467530.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing prognostic evaluation system for acute myeloid leukemia is difficult to make full use of complex genomic data generated by high-throughput sequencing, and cannot capture the complex interactions between gene features.

Method used

By fusion of multi-source data, a normalized feature set was constructed, and feature screening was performed using univariate Cox regression analysis, LASSO-punished Cox regression analysis, and multivariate Cox regression analysis to extract feature sets with independent prognostic value. Then, a random survival forest model is constructed based on these characteristics, an individualized risk score is generated, and the patients are divided into low, medium and high risk levels through dynamic clustering.

Benefits of technology

It improves the accuracy and personalization of prognostic assessment, overcomes the limitations of the linear assumption of traditional Cox risk models, enhances the predictive ability of risk scores, and provides clinicians with a more accurate basis for patient risk stratification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015326A_ABST
    Figure CN120015326A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic risk stratification method and device for acute myelogenous leukemia (AML). The method comprises the following steps: fusing multi-source data, and constructing a standardized feature set; performing feature screening through univariate Cox regression analysis, LASSO punishment Cox regression analysis and multivariate Cox regression analysis, and extracting a feature set with an independent prognosis value from the standardized feature set; and based on the feature set with the independent prognosis value, constructing a random survival forest model, and generating an individualized risk score, and carrying out dynamic clustering on the individualized risk scores so as to divide the patients into low, medium and high risk levels. The method and the device provided by the invention provide intelligent support for AML precise diagnosis and treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of medical artificial intelligence technology, and in particular to a dynamic risk stratification method and device for acute myeloid leukemia. Background Art

[0002] Acute myeloid leukemia (AML) is a highly heterogeneous blood tumor with significant differences in patient survival time. The current prognostic evaluation system based on the ELN 2022 guidelines mainly relies on traditional cytogenetic analysis technology and limited molecular marker detection (such as FLT3-ITD, NPM1, etc.), which makes it difficult to fully utilize the complex genomic data generated by high-throughput sequencing. The existing Cox proportional-hazards model assumes a linear risk relationship and cannot capture the complex interactions between genetic features. Therefore, the development of an intelligent prognostic system that can integrate multidimensional biomarkers, capture nonlinear association features, and has biological interpretability has become a key scientific issue in optimizing the precise diagnosis and treatment of AML.

[0003] Therefore, there is an urgent need to develop an intelligent system that can improve the accuracy and personalization of prognostic assessment by integrating multidimensional biomarkers and machine learning technology. Summary of the invention

[0004] The purpose of the present invention is to provide a dynamic risk stratification method and device for acute myeloid leukemia, aiming to solve the problems of insufficient utilization of genomic data and insufficient analysis of interaction effects between features in the existing prognostic evaluation system.

[0005] This manual adopts the following technical solutions: This manual provides a dynamic risk stratification method for acute myeloid leukemia, including: Fuse multi-source data and construct a standardized feature set; Performing feature screening by univariate Cox regression analysis, LASSO-penalized Cox regression analysis, and multivariate Cox regression analysis to extract a feature set with independent prognostic value from the standardized feature set; and Based on the feature set with independent prognostic value, a random survival forest model is constructed to generate an individualized risk score; the individualized risk score is dynamically clustered to divide the patients into low, medium and high risk levels.

[0006] Optionally, the feature screening is performed by univariate Cox regression analysis, LASSO-penalized Cox regression analysis, and multivariate Cox regression analysis, and a feature set with independent prognostic value is extracted from the standardized feature set, including: Perform the univariate Cox proportional hazard regression analysis on each feature in the normalized feature set, screen out features with a p value < 0.2, and obtain a preliminary feature set; The preliminary feature set is input into the Cox model with LASSO penalty, the regularization parameter is determined by S-fold cross validation, and the sparse feature set is screened out; Multivariate Cox stepwise regression optimization is performed on the sparse feature set to obtain the feature set with independent prognostic value.

[0007] Optionally, fusing multi-source data to construct a normalized feature set includes: Performing binary logic coding on discrete features in the multi-source data, and retaining original value ranges for continuous features in the multi-source data; Count the feature missing rate of each sample and delete samples with feature missing rate ≥ 20%; For discrete features in the retained samples, the majority imputation method is used to handle missing values; for continuous features, the median interpolation method is used to handle missing values.

[0008] Optionally, the individualized risk score is calculated based on a weighted cumulative risk function:

[0009] in, For the CHF estimation of each tree, output risk score .

[0010] Optionally, dynamically clustering the individualized risk scores to divide patients into low, medium, and high risk levels includes: dynamically clustering the individualized risk scores using a Gaussian mixture model and Bayesian decision technology to divide patients into low, medium, and high risk levels.

[0011] Optionally, the objective function of the LASSO-penalized Cox model for sparse processing is:

[0012] in, is the partial likelihood function, and the regularization parameter Determined by S-fold cross validation, is the total number of features.

[0013] This specification provides a dynamic risk stratification device for acute myeloid leukemia, comprising: Data fusion processing module, used to fuse multi-source data and construct a standardized feature set; A progressive feature screening module, for performing feature screening through univariate Cox regression analysis, LASSO-penalized Cox regression analysis, and multivariate Cox regression analysis, to extract a feature set with independent prognostic value from the standardized feature set; The risk scoring and stratification module is used to construct a random survival forest model based on the feature set with independent prognostic value to generate an individualized risk score; and dynamically cluster the individualized risk score to divide the patient into low, medium and high risk levels.

[0014] Optionally, the device also includes a verification and interpretation module for evaluating model performance and providing multi-granularity feature interpretation.

[0015] The present specification provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the dynamic risk stratification method for acute myeloid leukemia is implemented.

[0016] The present specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned dynamic risk stratification method for acute myeloid leukemia when executing the program.

[0017] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects: It can be seen from the above methods that the accuracy and personalization level of prognostic evaluation are improved through multi-source data fusion processing and three-level progressive feature screening; the limitations of the linear assumptions of the traditional Cox risk model are overcome by adopting the random survival forest algorithm, thereby improving the predictive ability of the risk score; through the Gaussian mixture model and dynamic clustering technology, patients are divided into low-risk, medium-risk and high-risk groups according to individualized risk scores, providing clinicians with a more accurate stratification basis. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation on this specification. In the drawings: Figure 1 A schematic flow chart of a dynamic risk stratification method for acute myeloid leukemia provided by an embodiment of the present invention is shown; Figure 2 A schematic diagram of a process flow of multi-source data fusion processing steps provided by an embodiment of the present invention is shown; Figure 3 A schematic diagram of a process of three-level progressive feature screening steps provided by an embodiment of the present invention is shown; Figure 4A system block diagram of a dynamic risk stratification device for acute myeloid leukemia provided by an embodiment of the present invention is shown; Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.

[0020] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.

[0021] Figure 1 A schematic diagram of a process flow of a dynamic risk stratification method for acute myeloid leukemia provided by one embodiment of the present invention. Figure 1 As shown, a dynamic risk stratification method for acute myeloid leukemia in an embodiment of the present invention may include the following steps: Step 100: Fuse multi-source data and construct a normalized feature set.

[0022] In this specification, multi-source data is fused to construct a standardized feature set, such as Figure 2 As shown, including: Step 101: Perform binary logic coding on discrete features in multi-source data, and retain the original value range for continuous features in multi-source data; Multi-source data include multi-dimensional data such as genomics, clinical indicators, treatment response and imaging. Multi-source data can be obtained from medical information systems. The acquired data include molecular biological characteristics, clinical baseline characteristics, and survival outcome data. Molecular biological characteristics include gene mutation spectrum (for example, FLT3-ITD, NPM1, CEBPA and other mutations, or mutation frequency), chromosomal abnormality events (del(5q), t(8;21) etc.), fusion gene expression (PML-RARA, RUNX1-RUNX1T1 etc.). Clinical baseline characteristics include demographic baseline (such as gender, age, race) and disease burden indicators (for example, initial white blood cell count). Survival outcome data include overall survival (OS_time) and survival status (OS).

[0023] Discrete features refer to variables with finite categories or non-continuous values. They cannot be infinitely subdivided and are usually used to describe the category, state or grouping of things, such as gene mutation, gender, pathological classification, etc. Continuous features refer to variables with infinite possible values ​​that can be accurately measured. They can be infinitely subdivided and are usually used to describe the quantitative attributes of things, such as age, white blood cell count, etc. Perform binary logic encoding on discrete features. Specifically, the complex karyotype of chromosomes, gene mutation status, the presence of fusion genes, gender and other discrete features are uniformly binary coded. The specific coding rules are as follows: the complex karyotype of chromosomes is based on the International Cytogenetic Nomenclature (ISCN) standard, and is recorded as 1 (positive) when ≥3 abnormal karyotypes are detected, and 0 (negative) when the number of abnormal karyotypes is <3 or the test is normal; the gene mutation status is based on the second-generation sequencing test results, and the mutant type and wild type are assigned values ​​of 1 and 0 respectively; the fusion gene feature is verified by fluorescence in situ hybridization (FISH) or RT-PCR technology, and is coded as 1 when the target gene fusion event exists, and coded as 0 when no specific fusion is detected; the gender feature is directly mapped to male 1 and female 0 according to the biological sex. For continuous features such as age and peripheral blood leukocyte count, the original value range is retained.

[0024] Step 102: Count the feature missing rate of each sample, and delete samples with feature missing rate ≥ 20%; Step 103: For discrete features in the retained samples, the majority imputation method is used to process missing values, and for continuous features, the median interpolation method is used to process missing values.

[0025] The fusion of multi-source data can improve the prediction accuracy and generalization ability of the model and enhance the robustness and integrity of the data.

[0026] Step 200: Perform feature screening (three-level progressive feature screening) through univariate Cox regression analysis, LASSO-penalized Cox regression analysis, and multivariate Cox regression analysis to extract a feature set with independent prognostic value from the standardized feature set.

[0027] In this specification, feature screening is performed by univariate Cox regression analysis, LASSO-penalized Cox regression analysis, and multivariate Cox regression analysis to extract a feature set with independent prognostic value from the standardized feature set, such as Figure 3 As shown, including: Step 201: Perform the univariate Cox proportional hazard regression analysis on each feature in the normalized feature set, screen out features with p-values ​​< 0.2, and obtain a preliminary feature set.

[0028] Univariate Cox regression analysis: Preliminary screening of features significantly associated with survival outcomes, elimination of variables with no statistical significance, and reduction of noise in subsequent analysis. Specifically, univariate Cox regression analysis was performed independently for each feature (such as gene mutation status, age), and features with a hazard ratio (HR) p value < 0.2 were screened out, retaining a set of features with potential prognostic trends The Cox model is as follows:

[0029] in, is the baseline hazard function, is the regression coefficient, is the feature vector.

[0030] Step 202: Input the preliminary feature set into the Cox model with LASSO penalty, determine the regularization parameter through S-fold cross validation, and screen out the sparse feature set.

[0031] Cox regression analysis with LASSO penalty: It performs sparse processing on high-dimensional data (such as proteomics and genomics), eliminates collinear features, and retains non-zero coefficient variables. Specifically, the feature set The Cox model with LASSO (least absolute shrinkage and selection operator) penalty is input, and LASSO regression is used to sparse the high-dimensional features. The objective function is:

[0032] in, is the partial likelihood function, and the regularization parameter Determined through 10-fold cross validation, is the total number of features, filtering sparse feature sets .

[0033] Step 203: Perform multivariate Cox stepwise regression optimization on the sparse feature set to obtain a feature set with independent prognostic value.

[0034] Multivariate Cox regression analysis: Determine the key features that are independently related to prognosis and output a feature set with independent prognostic value. Specifically, based on the Akaike Information Criterion (AIC), the sparse feature set Multivariate Cox stepwise regression optimization was performed to finally obtain a set of features with independent prognostic value:

[0035] in, is the number of model parameters, is the likelihood function value, and the final output is a feature set with independent prognostic value .

[0036] The three-level progressive feature screening can balance efficiency and accuracy through phased optimization, and enhance the model's interpretability and clinical practicality. Through multi-source data fusion processing to integrate multi-dimensional clinical and genomic data, and using three-level progressive feature screening to extract key prognostic factors, it makes full use of the complex genomic data generated by high-throughput sequencing, which can improve the accuracy of AML prognosis assessment and provide patients with personalized risk prediction.

[0037] Step 300: Based on a set of features with independent prognostic value, a random survival forest model is constructed to generate an individualized risk score; the individualized risk score is dynamically clustered to divide the patients into low, medium, and high risk levels.

[0038] In this specification, the Random Survival Forest (RSF) model is configured with 800 decision trees, a maximum tree depth of 10 layers, and a node splitting criterion that maximizes the log-rank statistic:

[0039] in, is the time point of the event. For the The number of events at a time point, is the number of risk set samples.

[0040] When building the random survival forest model, a dynamic training mechanism is added. After each tree is trained, the out-of-bag (OOB) error of the current model is calculated, the OOB error is monitored in real time, the OOB error sequence of the most recent 10 iterations is recorded, and its fluctuation range is calculated. When the error fluctuation range for 10 consecutive iterations is ≤±2%, early stopping is triggered.

[0041] At the model terminal, risk quantification is performed. The individualized risk score is calculated based on the weighted cumulative hazard function (CHF) of the terminal node:

[0042] in, For the CHF estimation of each tree, output risk score .

[0043] By adopting the random survival forest algorithm, the nonlinear interactions between genetic features can be captured, overcoming the limitations of the linear assumption of the traditional Cox proportional hazard model, thereby improving the predictive ability of the risk score.

[0044] In this specification, dynamically clustering the individualized risk scores to divide patients into low, medium, and high risk levels includes: dynamically clustering the individualized risk scores using a Gaussian mixture model and Bayesian decision technology to divide patients into low, medium, and high risk levels.

[0045] More specifically, the risk score was modeled using a three-component Gaussian mixture model to determine the probability density of risk distribution. The EM algorithm was used to optimize the model parameters, and the optimal number of clusters was determined based on the Bayesian Information Criterion (BIC). The critical value of risk stratification was determined by the intersection of the probability density function, and the patients were divided into low-risk, medium-risk, and high-risk groups.

[0046] A risk distribution probability density model is constructed based on a three-component Gaussian mixture model (GMM), and the EM algorithm is used for iterative parameter optimization:

[0047] in , , Respectively The mean, variance, and weight of a Gaussian distribution.

[0048] The optimal number of clusters is verified according to the Bayesian Information Criterion (BIC), and the stratification critical value is finally determined by the intersection of the probability density function: Low risk group: (corresponding to the 95% confidence upper limit of the normal distribution); Medium risk group: ; High-risk group: (corresponding to the 95% confidence lower limit of the normal distribution).

[0049] in, is the mean of the three Gaussian distributions.

[0050] In order to objectively evaluate the constructed model, this manual also provides model validation and interpretation. The model performance is evaluated by temporal consistency index, dynamic AUC and Brier score, and multi-granularity feature interpretation is provided by using SHAP value and force-directed graph.

[0051] The time-dependent C-index and dynamic AUC (evaluation time window t = 12 / 24 / 36 months) of the test set were calculated.

[0052] The deviation of the predicted survival rate from the actual Kaplan-Meier curve was quantified by the Brier score:

[0053] The clinical utility was assessed by evaluating the survival differences between risk groups using the log-rank test.

[0054] Multi-granularity feature explanation includes group explanation and individual explanation.

[0055] The group explanation quantifies the global contribution of the feature through the SHAP (shapley additive explanations) value, generating a feature importance ranking matrix and an interaction effect dependency graph:

[0056] in, For the complete set of features, is a feature subset, is the model prediction function; Individual explanations are based on force plots to intuitively display the driving factors of single-sample risk scores.

[0057] Through verification and explanation, the model prediction results are made more biologically interpretable, which helps clinicians understand and trust the prediction results.

[0058] The above is a method for implementing the dynamic risk stratification of acute myeloid leukemia in this specification. Based on the same idea, this specification also provides a corresponding dynamic risk stratification device for acute myeloid leukemia, such as Figure 4 The dynamic risk stratification device for acute myeloid leukemia in the embodiment of the present invention may include: The data fusion processing module 10 is used to fuse multi-source data and construct a standardized feature set; A progressive feature screening module 20, for performing feature screening through univariate Cox regression analysis, LASSO-penalized Cox regression analysis, and multivariate Cox regression analysis, to extract a feature set with independent prognostic value from the standardized feature set; The risk scoring and stratification module 30 is used to construct a random survival forest model based on the feature set with independent prognostic value to generate an individualized risk score; and dynamically cluster the individualized risk score to divide the patient into low, medium and high risk levels.

[0059] In this specification, the dynamic risk stratification device for acute myeloid leukemia in the embodiment of the present invention may also include: a verification and interpretation module for evaluating model performance and providing multi-granularity feature interpretation.

[0060] The data fusion processing module 10 may include a data acquisition component, which extracts multi-dimensional data from the hospital's HIS (clinical data) and LIS (laboratory and genomic data) systems by being compatible with internationally accepted medical information exchange standards, and provides complete input information for subsequent risk stratification; a data preprocessing component that implements an automated feature encoding and missing value processing pipeline.

[0061] The progressive feature screening module 20 may include a univariate Cox regression analysis component, which performs the univariate Cox proportional risk regression analysis on each feature in the normalized feature set to obtain a preliminary screening feature set; a LASSO-penalized Cox component, which is configured with a LASSO optimizer with 10-fold cross-validation to screen out a sparse feature set; and a multivariate Cox regression analysis component, which performs multivariate Cox stepwise regression optimization to ultimately obtain a feature set with independent prognostic value.

[0062] The risk scoring and stratification module 30 may include a training component that deploys a Spark-based distributed random survival forest algorithm based on a multi-tree ensemble strategy and a dynamic training strategy; a risk quantification component that generates risk scores and confidence intervals in real time based on a cumulative risk function and a weighting strategy; a GMM modeling component that dynamically adjusts Gaussian grouping based on a three-GMM distribution probability model and an EM algorithm; and a Bayesian decision component that automatically calculates stratification critical values ​​based on the BIC criterion.

[0063] The verification and interpretation module can include a performance dashboard that visualizes indicators such as C-index, dynamic AUC, grouping curves, and inter-group differences; a dual interpreter that generates a SHAP dependency graph and an individualized feature attribution-oriented graph.

[0064] This specification also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 Provided is a dynamic risk stratification approach for acute myeloid leukemia.

[0065] This manual also provides Figure 5 The one shown corresponds to Figure 1 A schematic diagram of the electronic device. Figure 5 As shown, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The described dynamic risk stratification method for acute myeloid leukemia.

[0066] In the 1990s, it was very clear whether the improvement of a technology was hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with hardware entity modules. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0067] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.

[0068] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0069] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0070] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0071] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0072] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0073] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0074] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0075] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0076] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0077] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0078] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0079] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0080] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0081] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.

Claims

1. A dynamic risk stratification method for acute myeloid leukemia, characterized in that: The method comprises the following steps: Fuse multi-source data and construct a standardized feature set; Performing feature screening by univariate Cox regression analysis, LASSO-penalized Cox regression analysis, and multivariate Cox regression analysis to extract a feature set with independent prognostic value from the standardized feature set; and Based on the feature set with independent prognostic value, a random survival forest model is constructed to generate an individualized risk score; the individualized risk score is dynamically clustered to divide the patients into low, medium and high risk levels.

2. The method according to claim 1, characterized in that The feature screening is performed by univariate Cox regression analysis, LASSO-penalized Cox regression analysis, and multivariate Cox regression analysis, and the feature set with independent prognostic value extracted from the standardized feature set includes: Perform the univariate Cox proportional hazard regression analysis on each feature in the normalized feature set, screen out features with a p value < 0.2, and obtain a preliminary feature set; The preliminary feature set is input into the Cox model with LASSO penalty, the regularization parameter is determined by S-fold cross validation, and the sparse feature set is screened out; Multivariate Cox stepwise regression optimization is performed on the sparse feature set to obtain the feature set with independent prognostic value.

3. The method according to claim 1, characterized in that The fusing of multi-source data to construct a standardized feature set includes: Performing binary logic coding on discrete features in the multi-source data, and retaining original value ranges for continuous features in the multi-source data; Count the feature missing rate of each sample and delete samples with feature missing rate ≥ 20%; For discrete features in the retained samples, the majority imputation method is used to handle missing values; for continuous features, the median interpolation method is used to handle missing values.

4. The method according to claim 1, characterized in that: The individualized risk score is calculated based on the weighted cumulative risk function: in, For the CHF estimation of each tree, output risk score .

5. The method according to claim 1, characterized in that The dynamically clustering the individualized risk scores to classify the patients into low, medium, and high risk levels includes: dynamically clustering the individualized risk scores using a Gaussian mixture model and a Bayesian decision technique to classify the patients into low, medium, and high risk levels.

6. The method according to claim 3, characterized in that The objective function of the LASSO-penalized Cox model for sparse processing is: in, is the partial likelihood function, and the regularization parameter Determined by S-fold cross validation, is the total number of features.

7. A dynamic risk stratification device for acute myeloid leukemia, characterized in that: The device comprises: Data fusion processing module, used to fuse multi-source data and construct a standardized feature set; A progressive feature screening module, for performing feature screening through univariate Cox regression analysis, LASSO-penalized Cox regression analysis, and multivariate Cox regression analysis, to extract a feature set with independent prognostic value from the standardized feature set; The risk scoring and stratification module is used to construct a random survival forest model based on the feature set with independent prognostic value to generate an individualized risk score; and dynamically cluster the individualized risk score to divide the patient into low, medium and high risk levels.

8. The device according to claim 7, characterized in that The device also includes a verification and explanation module for evaluating model performance and providing multi-granularity feature explanations.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Statistical analysis method and system based on multi-omics and clinical data and storage medium

    CN111913999A

  • Tumor prognosis prediction method, system and equipment based on novel image clustering features and storage medium

    CN116189890A

  • Method for constructing prognosis model of adult acute myelogenous leukemia patient based on telomere-related gene

    CN117352171A

  • Diagnostic marker of acute myelogenous leukemia prognosis model as well as application and prediction method thereof

    CN119193843A

  • System and method for predicting risk of recurrence of hepatocellular carcinoma in patient after surgery

    CN119811654A