Sparse explanations of feature selection from high-dimensional health data for biomedical events

The method addresses the lack of explainability in survival data analysis by using a machine learning approach to generate sparse explanations of feature selection from high-dimensional health data, enhancing understanding and trust in predictive models for medical applications.

WO2025104506A1PCT designated stage expired Publication Date: 2025-05-22NEC LAB EURO GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2024/051423
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-14
Filing Date
2024-02-15
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing technologies for analyzing survival data in high-dimensional health data lack the ability to provide explanations for feature selection, making it difficult to understand and trust the predictions, especially in critical fields like medicine.

Method used

A computer-implemented machine learning method that generates sparse explanations of feature selection by using an embedding machine learning model to encode and decode training data, compute sparsified feature selections, and update the model based on reconstruction and partial likelihood losses.

Benefits of technology

The method provides human-readable, explainable feature selections, enabling better understanding and trust in survival predictions, and facilitating more targeted treatments in medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024051423_22052025_PF_FP_ABST
    Figure IB2024051423_22052025_PF_FP_ABST
Patent Text Reader

Abstract

A machine-learning method generates sparse explanations of feature selection from high¬ dimensional data. A training phase includes: encoding training data into latent space to generate first training embeddings, which are decoded to generate first reconstructed training data; computing sparsified feature selections of the first reconstructed training data; calculating reconstruction losses; calculating negative partial likelihood losses; encoding the sparsified feature selections of the first reconstructed training data into the latent space to generate second training embeddings; calculating negative double-pass partial likelihood losses based on the second training embeddings; and updating an embedding machine learning model used for the encoding and decoding based on the reconstruction losses, negative partial likelihood losses and negative double-pass partial likelihood losses. The method has applications including, but not limited to, use cases in medicine / healthcare, treatment, diagnosis or biomedical events, to optimize predictions or support decision making.
Need to check novelty before this filing date? Find Prior Art

Description

SPARSE EXPLANATIONS OF FEATURE SELECTION FROM HIGH-DIMENSIONAL HEALTH DATA FOR BIOMEDICAL EVENTSCROSS-REFERENCE TO RELATED APPLICATION

[0001] Priority is claimed to U.S. Provisional Patent Application No. 63 / 548,414, filed on November 14, 2023, the entire disclosure of which is hereby incorporated by reference herein. FIELD

[0002] The present disclosure relates to artificial intelligence (Al) and machine learning (ML), and in particular to a method, system, data structure, computer program product and computer-readable medium for providing sparse explanations of feature selection, having applications in medicine / healthcare such as being applied to high-dimensional health data for biomedical events.BACKGROUND

[0003] Many complex systems are influenced by high-dimensional data, which cannot be readily understood by humans. This understanding -barrier makes effective and predictable human interaction with these complex systems extremely difficult..

[0004] Biomedical events are an important category of complex systems that can be influenced by high-dimensional data. For example, a human patient’s survival from cancer may be influenced by complex molecular data — such as gene expression and mutational patterns — that define the level of aggressiveness of a tumor, which can lead to a worse prognosis for one patient compared to another who has the same tumor type. Because the ultimate goal of cancer treatment is to elongate a patient’s survival — ideally also in a way that ensures quality of life — it is crucial to understand underlying biological factors. This is because understanding such biological factors helps to select treatment options that will improve survival for one particular patient, and are not just created to cover a broad majority of patients with the same cancer type. Additionally, a better human understanding can lead to the discovery of new biomarkers or novel targets for targeted treatments like immunotherapies. Similarly, electronic health records (EHRs) can also be high-dimensional datasets that are difficult to navigate, yet having a sufficient understanding could be essential to effective treatment of the patient.

[0005] Accordingly, the present inventors have recognized that a methodology is needed to narrow down the factors that influence complex systems to a point where it is easier for a human to understand. Of particular importance to the present inventors are biological events, for which having a methodology that helps narrow down the factors that actually influence a patient’s individual survival, and therefore, allow more targeted treatment, would prove to be a significant advancement in patient care. The present inventors have also determined that it is important tohave a methodology that identifies these factors to be explainable by design to be transparent because there otherwise would not be any trust in the methodology by patients and doctors, and therefore the method would not be of practical use.

[0006] Nevertheless, existing technology for analyzing survival data does not address the need to produce any explanations tailored for high-dimensional data when predicting the patient's risk to gain a better understanding of individual combined risk factors. “Highdimensional data” as used the machine learning context of the present disclosure means that the minimum number of features is in the order of thousands.

[0007] For example, existing methods of analyzing survival in cancer cohorts include tools like Kaplan-Meier curves (see Kaplan, E. L.; and Meier, P. 1958. Nonparametric estimation from incomplete observations. Journal of the American statistical association, 53(282): 457-481, which is hereby incorporated by reference herein), where the survival rates of patient groups are visualized and gaps, in particular lower survival in one group compared to the other, are identified. Methods such as DeepSurv (see Katzman, J. L.; Shaham, U.; Cloninger, A.; Bates, J.; Jiang, T.; and Kluger, Y. 2018. DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC medical research methodology, 18(1): 1- 12, which is hereby incorporated by reference herein) promise a high accuracy of predicted risks in terms of c-index using deep neural networks. However, such deep architecture can be seen as black boxes that lack any means of interpretability or understanding, which renders them unsafe to use especially in high-risk areas such as the medical field.

[0008] As another example of an existing technology, there is the so-called support vector regression for censored data (see Khan, F. M.; and Zubek, V. B. 2008. Support vector regression for censored data (SVRc): a novel tool for survival analysis. In 2008 Eighth IEEE International Conference on Data Mining, 863-868. IEEE, which is hereby incorporated by reference herein), which is an old alternative to deep learning that produces support vectors that could give some insights into the problem in high-dimensional data, such as gene expression data (mRNA), micro-RNA data (miRNA), or extensive EHRs. There also exists a methodology called classification-by-components (see Saralajew, S.; Holdijk, L.; Rees, M.; Asan, E.; and Vihmann, T. 2019. Classification-by-components: Probabilistic modeling of reasoning over a set of components. Advances in Neural Information Processing Systems 32, which is hereby incorporated by reference herein), which is a prototype-based learning method that works best for image data, but lacks the ability to deal with censored data. The present inventors have found that these technologies also fall short of improving human interpretation. This is because, for example, high-dimensional data can be particularly difficult to navigate even if explainability methods (like feature importance) are applied.SUMMARY

[0009] An aspect of the present disclosure provides a computer-implemented machinelearning method that generates sparse explanations of feature selection from high-dimensional data. The method includes a training phase that includes: (a) using an embedding machine learning model, encoding training data into latent space to generate first training embeddings; (b) using the embedding machine learning model, decoding the first training embeddings to generate first reconstructed training data; (c) computing sparsified feature selections of the first reconstructed training data; (d) calculating reconstruction losses based on the first reconstructed training data and the training data; (e) calculating negative partial likelihood losses based on the first training embeddings; (f) using the embedding machine learning model, encoding the sparsified feature selections of the first reconstructed training data into the latent space to generate second training embeddings; (g) calculating negative double-pass partial likelihood losses based on the second training embeddings; and (h) updating the embedding machine learning model based on the reconstruction losses, the negative partial likelihood losses, and the negative double-pass partial likelihood losses. The method has applications including, but not limited to, use cases in medicine / healthcare, treatment, diagnosis or biomedical events, to optimize predictions or support decision makingBRIEF DESCRIPTION OF THE DRAWINGS

[0010] Subject matter of the present disclosure will be described in even greater detail below based on the exemplary figures. All features described and / or illustrated herein can be used alone or combined in different combinations. The features and advantages of various embodiments will become apparent by reading the following detailed description with reference to the attached drawings, which illustrate the following:

[0011] FIG. 1 schematically illustrates a method and system according to an embodiment of the present disclosure, which includes a training phase and a production phase;

[0012] FIG. 2 illustrates an autoencoder for generating a double-pass partial likelihood according to an embodiment of the present disclosure; and

[0013] FIG. 3 is a block diagram of an exemplary processing system, which can be configured to perform any and all operations disclosed herein.DETAILED DESCRIPTION

[0014] Aspects of the present disclosure are directed to explaining complex systems in terms of learned sparse representations. While the present disclosure is provide largely with respect to explaining biological systems (specifically in terms of human healthcare) — a particularly important category of such tasks — a person of ordinary skill in the art would readily understand that the present disclosure is not so limited. For example, according to the particularimplementation, an event could also be in other domains such as the event of ailing electrical parts (e.g., as in predictive maintenance).

[0015] Decision making in the healthcare field is of critical importance because a patient’s quality of life, or even survival, may be irreversibly impacted by such decisions. For example, many treatment decisions may be of one path, meaning that, after the decision is made, any change becomes almost impossible. Accordingly, aspects of the present disclosure not only process data to make predictions about a patient’s condition and treatment, but also provide explanations of the decisions and recommendations. Such explanations can only be of practical use if they are human-readable. Thus, embodiments of the present disclosure provide human- readable explanations by learning and outputting sparse representations. In this way, the learning of sparse representations is equivalent to explainable feature selection.

[0016] According to one aspect of the present disclosure, a patient's survival risks are explained in terms of learned sparse representations. For instance, a method is provided for the scenario where the patient's data is high-dimensional, such that representations in the original space are difficult for doctors and patients to interpret and understand. The method processes the data to make predictions and outputs sparse representations. With this, clinicians can identify a marker for a better survival prognosis, but also a marker for a worse prognosis that can potentially be exploited as drug targets. These sparse representations represent an explicit interpretable feature selection.

[0017] According to one aspect, the present disclosure provides a method for generating sparse representations. The method includes a training phase, and a production phase. In the training phase, the method includes: embedding patient training data through an encoded / decoder structure, i.e., an embedding model; computing a sparsified feature selection (e.g., using sparsification model), which is passed back to embedding model; re-passing information through the sparsification model to compute double-pass partial likelihood (DPPL) in order to guarantee faithful and sparse feature selection; computing a loss functions and gradients, and updating the embedding model; and generating the sparse representation as an output. In the production phase, the method includes: embedding the patient test data through the encoded / decoder structure, i.e., the embedding model; and generating sparse representation as outputs. Depending on the implementation, one or both of the training and production phase may be instantiated.

[0018] In a first exemplary aspect, the present disclosure provides a computer-implemented machine -learning method for generating sparse explanations of feature selection from highdimensional data The method includes a training phase, which includes:a. using an embedding machine learning model, encoding training data into latent space to generate first training embeddings; b. using the embedding machine learning model, decoding the first training embeddings to generate first reconstructed training data; c. computing sparsified feature selections of the first reconstructed training data; d. calculating reconstruction losses based on the first reconstructed training data and the training data; e. calculating negative partial likelihood losses based on the first training embeddings; f. using the embedding machine learning model, encoding the sparsified feature selections of the first reconstructed training data into the latent space to generate second training embeddings; g. calculating negative double-pass partial likelihood losses based on the second training embeddings; and h. updating the embedding machine learning model based on the reconstruction losses, the negative partial likelihood losses, and the negative double-pass partial likelihood losses.

[0019] According to a first implementation of the first aspect, the training phase of the method may further include iteratively executing operations a through h to further update the embedding machine learning model until satisfying a predetermined optimization criteria.

[0020] According to a second implementation, in any implementation of the first aspect, the training phase of the method may further include: using the updated embedding machine learning model to encode the training data into the latent space to generate final training embeddings; using the updated embedding machine learning model to decode the final training embeddings to generate final reconstructed training data; computing final sparsified feature selections of the final reconstructed training data; and saving at least the final sparsified feature selections of the final reconstructed training data as sparse prototypes.

[0021] According to a third implementation, in any implementation of the first aspect, the high-dimensional data may be patient data.

[0022] In a fourth implementation, further to the third implementation of the first aspect, the patient data may include: gene expression data, micro-RNA data, electronic health records, or blood test results.

[0023] In a fifth implementation, further to the fourth implementation of the first aspect, the patient data may include vectors respectively corresponding to patients, each vector having a plurality of coefficients. Here, the embedding machine learning model may include a variational autoencoder architecture having an encoder and a decoder. The operation of using the embedding machine learning model to encode the training data into the latent space to generatethe first training embeddings may then include: obtaining a sample from a fixed distribution; for each of the vectors, using the encoder, determining a respective mean and a respective standard deviation; and for each of the vectors, transforming the sample into a respective one of the first training embeddings using a parameterization technique based on the sample, the respective mean, and the respective standard deviation.

[0024] In a sixth implementation, further to the fifth implementation of the first aspect, the operation of using the embedding machine learning model to decode the first training embeddings to generate first reconstructed training data may include using the decoder, generating the first reconstructed training data as a reconstructed vector of coefficients based on the embeddings. Here, the operation of calculating reconstruction losses based on the first reconstructed training data and the training data includes using the formula: ll - x\\ where x is the vector of coefficients from the high-dimensional data and x is the reconstructed vector of coefficients.

[0025] In a seventh implementation, further to the fifth or sixth implementation of the first aspect, the operation of computing sparsified feature selections of the first reconstructed training data may include: for each reconstructed vector, reduce a number of non-zero entries by using Li norm and a masked thresholding function, to generate a reconstructed sparse sample of the reconstructed vector.

[0026] In an eighth implementation, further to the seventh implementation of the first aspect, the masked thresholding function may include setting a smallest predefined percentage of features in the reconstructed vector to zero.

[0027] In an ninth implementation, further to any of the fifth through eighth implementations of the first aspect, using the embedding machine learning model to encode the sparsified feature selections of the first reconstructed training data into the latent space to generate the second training embeddings includes: providing as input to the encoder the reconstructed sparse sample of each of the reconstructed vectors as sparse vectors; obtaining a sample from a fixed distribution; for each of the sparse vectors, using the encoder, determining a respective mean and a respective standard deviation; and for each of the sparse vectors, transforming the sample into a respective one of the second training embeddings using the parameterization technique based on the sample, the respective mean, and the respective standard deviation.

[0028] In a tenth implementation, further the ninth implementation of the first aspect, the operation of calculating the negative double-pass partial likelihood losses based on the second training embeddings may include maximizing the following double-pass partial likelihood: mDPPL = ^[ln(r(x0))) + ln(r(x0))] —0=1where an instance is given in a form of (x0, t0, 50), t0is either as an observation time of the instance or a censoring time for right-censored individuals. 80is an event indicator where 80= 0 for censored cases and 80= 1 otherwise, x0is the vector for a patient, x0is the reconstructed vector for a patient , r(x0) isarisk set.

[0029] In an eleventh implementation, further to any implementation of the first aspect, the method further may further include a production phase executed after completion of the training phase. The production phase may include: using the embedding machine learning model, encoding test data into the latent space to generate test embeddings; using the embedding machine learning model, decoding the test embeddings to generate reconstructed test data; and computing sparsified feature selections of the reconstructed test data.

[0030] In an twelfth implementation, further to the eleventh implementation of the first aspect, the production phase may further include comparing the sparsified feature selections to the sparse prototypes and making a prediction regarding the test data based on comparison.

[0031] According to a further aspect, the present disclosure provides a computer system comprising one or more hardware processors which, alone or in combination, are configured to execute a machine-learning method for generating sparse explanations of feature selection from high-dimensional data according to one, or a combination, of the aspects described herein.

[0032] According to a further aspect, the present disclosure provides a tangible, non- transitory computer-readable medium having instructions thereon which, upon being executed by one or more hardware processors, alone or in combination, provide for execution a machinelearning method for generating sparse explanations of feature selection from high-dimensional data according to one, or a combination, of the aspects described herein.

[0033] Aspects of the present disclosure provide for at least the following improvements and technical advantages over existing technology:

[0034] Aspects of the present disclosure enable the creation of faithful sparsification of patient data in addition to predicting the patient’s risk. Faithfulness is a direct result of the double-pass likelihood function introduced in the present disclosure. This is an importantimprovement over existing technology because understanding, explaining, and trusting the learned models of survival and hazard functions are important for making decisions on recommending the best treatment for patients.

[0035] Aspects of the present disclosure enhance computer functionality, security, trustworthiness, reliability, explainability and accuracy of machine learning systems by the introduction and implementation of a novel objective function, double-pass partial likelihood (DPPL), that guarantees faithful and sparse feature selection.

[0036] Aspects of the present disclosure advantageously introduce a masked thresholding procedure on the reconstructed data for the creation of sparse feature selection. Embodiments further introducing a novel matching approach.

[0037] Aspects of the present disclosure provide human-understandable and explainable risk functions used in survival analysis.

[0038] In contrast to existing methods, embodiments implemented according to aspects of the present disclosure do not need to rely on manual grouping of patients (e.g., a priori). While at the same time, such aspects do enable the identification of new features among large cohorts.

[0039] What follows is a discussion of exemplary embodiments implementing aspects of the present disclosure, which is given by way of reference to one or more figures of the present disclosure. As a person of ordinary skill in the art would readily comprehend, the exemplary embodiments are merely illustrative and do not represent the limits of the present disclosure. Rather, a person of ordinary skill in the art would consider the disclosure as a whole in understanding the aspects and advantages provided herein. What’s more, a person of ordinary skill in the art would understand that the exemplary features described herein may be variously combined and modified all without departing from the scope of the present disclosure.

[0040] FIG. 1 illustrates an exemplary method and system 100 for the creation of sparse representations according to an embodiment of the present disclosure. The method and system 100 has two instantiations: 1) the training phase 101, and 2) the production phase 102. The method and system 100, as illustrated in FIG. 1, has four main conceptual modules that differ in their functionalities: A) embedding module 103; B) optimization module 104; C) sparsification module 105; and D) output module 106.

[0041] The embedding module 103 is implemented with an embedding machine learning model, which takes data as an input and embeds the data into latent space and can transform data from the latent space back into the input feature space. In the embodiment shown, the embedding module 103 takes the patient(s) data (i.e., training patient data 107 and test patient data 108) as input and embeds it in a latent space. By way of illustration, the patient data 107 could be gene expression (RNA) data 109, blood test results 110, or any EHR data 111. Theoutput of the embedding module 103 is forwarded to the optimization module 104 (in the training phase) or to the output module 106 in the validation or production phases.

[0042] The embedding module 103 may implement the embedding machine learning model with any encoding / decoding architecture (such as variational autoencoders / autodecoders) that transforms the features of the input into the encoded space (i.e., encoded as embeddings), without departing from the scope of the present disclosure.

[0043] In machine learning, an "embedding" refers to a mapping or transformation of input data from a high-dimensional feature space into a lower-dimensional latent space. This transformation is designed to capture meaningful relationships, patterns, or representations inherent in the original data while reducing its dimensionality.

[0044] The process of embedding involves encoding the input data, such as categorical variables or high-dimensional features, into a lower-dimensional space known as the latent space. This encoding is achieved through mathematical techniques that preserve essential information while discarding less crucial details. By doing so, embeddings enable the model to effectively learn and operate on a more compact and informative representation of the input data, facilitating tasks like clustering, classification, or recommendation systems by capturing relevant similarities or structures within the data.

[0045] The optimization module 104 implements machine learning algorithms that compute loss functions and gradients, and update the embedding machine learning model based on the computed loss functions and gradients. Exemplary loss functions optimized with the optimization module 104 are discussed in further detail below. The optimization model may be repeatedly executed to incrementally update the embedding machine learning model.Optimization may end when a predetermined number of iterations is reached or when the loss is reduced below a predetermined threshold.

[0046] The sparsification module 105 includes a sparsification machine learning model that sparsifies the output of the embedding module 103 and propagates the result to the output module 106. The sparsification module 105 also repasses the sparse output to the embedding module.

[0047] In the field of machine learning, “sparsification” refers to a methodology aimed at reducing the complexity of a model by selectively removing or reducing certain parameters, features, or connections within the model while retaining its essential functionality and performance. This process often involves identifying and eliminating redundancies, irrelevant data, or less influential components, resulting in a more streamlined and efficient model that requires fewer resources for storage, computation, or inference without significantly compromising its predictive capability or accuracy. Embodiments implementing aspects of thepresent disclosure may utilize various sparsification machine learning models (e.g., using a thresholding function or LI Regularization (Lasso)) without departing from the scope of the present disclosure.

[0048] The output module 106 creates the sparse representations to be output as sparse prototypes 112 / 113.

[0049] Exemplary embodiments of the training phase 101 and production phase 102 instantiations of the method are described as follows:

[0050] In the training phase 101, the parameters of the embedding model of the embedding module 103 are trained using a gradient-based optimizer implemented in the optimization module 104.

[0051] In particular, in the training phase 101, the training patient data 107 (e.g., RNA data 109, blood test results 110, and / or EHR data 111) is passed as input to the embedding module 103. The embedding module 103, using an encoder of the embedding machine learning model, encodes the test data into the latent space, and then propagates the encoded test data to the optimization module 104 and the sparsification module 105.

[0052] In the training phase, the sparsification module 105 applies the sparsification procedures to reduce the number of non-zero entries, e.g., via executing a LI norm and a masked thresholding function. The sparse output is enforced to be faithful via repassing it through the embedding module. As mentioned herein, the embedding module can execute any embedding function that is capable of mapping the data to an embedding space, and where reconstruction is possible.

[0053] According to an embodiment, the output module may be implemented with a fitted decoder. In this embodiment, the fitted decoder receives the embedded sample, and reconstructs the sample. The sparsification aspects of the output module are executed via a sparsification module, which may be external to the output module itself. In such a case, the sparsification module would pass the sparsified sample to the output module.

[0054] The optimization module 104 computes a plurality of different losses. In the embodiment of FIG. 1, these losses are the construction loss Lossl, the negative partial likelihood Loss2, and the negative double-pass partial likelihood Loss3.

[0055] Upon completion of the training phase 101, the final output of the embedding module 103 and sparsification module 105 are received as input to the output model 106. The output of the output model 106 are the learned representations and the sparsified representations. These may be saved as sparse porotypes 112 for use in the production phase 102.

[0056] In the production phase 102, the test patient data 108 are embedded using the embedding module 103 after completing the training phase at least once. In the production phase102, the output module 106 computes the decoded and sparsified representations. In the production phase, the learned network (i.e., encoder-decoder) and sparsification procedure are fixed. Thus, the whole architecture is applied to new samples (e.g., of patient data).

[0057] FIG. 2 shows an exemplary embodiment implementing aspects of the present disclosure that employs a variational autoencoder architecture. As with FIG 1, FIG. 2 illustrates a machine learning system and method where patient data are used as input, and the generated sparse representations are provided as output.

[0058] In FIG. 2, data associated representing each patient is represented as a vector x of coefficients. This vector is passed to the encoder-decoder system 200.

[0059] For each vector x, an encoder 201 calculates the mean pxand the standard deviation <JX. Using the parametrization trick, z~N(jxx, <7X) is obtained. The encoder may be implemented with a variational autoencoder, however any encoding function capable of mapping data to an embedding space, where the mapped data is capable of reconstruction, may be used without departing from the scope of the present disclosure.

[0060] In an exemplary embodiment implementing variational autoencoders (VAEs), for example, the parameterization trick involves the representation of the latent variable z using a parameterization that facilitates the training process through backpropagation.

[0061] In VAEs, the latent variable z represents the encoded, continuous latent space where the input data is mapped — that is, z is an embedding representing the high-dimensional input data in lower-dimensional space. To enable the training of VAEs using gradient-based optimization techniques, the parameterization trick is used. Instead of directly sampling the embedding z from the learned distribution, which involves randomness and makes it challenging to compute gradients through the sampling operation, the trick involves introducing a deterministic transformation that involves sampling from a simple, fixed distribution (like a standard normal distribution) and then transforming these samples to obtain the embedding z.

[0062] One practice is to sample e from a simple distribution (like a standard normal distribution, N(0,l)) and then parameterize the embedding z as: z=p+oOeHere: p and o respectively are the mean and standard deviation (or log -variance, depending on the implementation) obtained from the encoder. Also, e is a sample from a fixed distribution.

[0063] This parameterization enables the model to generate the embedding z in a way that is differentiable, allowing gradients to flow through the sampling process during backpropagation. This makes it possible to optimize the VAE by minimizing the loss function using stochastic gradient descent or other optimization techniques.

[0064] Next, the decoder 202 receives the embedding z, and reconstructs the sample x as reconstructed as x. The encoder-decoder 200 is then optimized using several loss functions. The decoder is implemented with corollary architecture to the encoder such that the mapped data is faithfully reconstructed. In the figure, P(x|z) represents the probability distribution of the output x given the embedding z.

[0065] First, the encoder-decoder system 200 utilizes the evidence lower bound (ELBO). The ELBO is a concept in variational inference, which is a technique used in probabilistic modeling and machine learning for approximating complex probability distributions. In particular, the ELBO serves as an element in the training of variational autoencoders (VAEs).

[0066] In the context of variational inference, the ELBO represents a lower bound on the log marginal likelihood of the data. It may calculated using the Kullback-Leibler (KL) divergence, a measure of how one probability distribution diverges from a second, usually a simpler, distribution.

[0067] For a given probability distribution, the ELBO is expressed as the difference between the logarithm of the likelihood of the observed data and the KL divergence between the approximate (variational) distribution and a simpler, chosen prior distribution. Mathematically, it can be formulated as:ELBO = E[log p(x|z)] - KL[q(z|x) || p(z)]Here, E[log p(x|z)] is the expected log-likelihood of the data given the latent variables; and KL[q(z|x) || p(z)] is the KL divergence between the variational distribution (q) of the latent variables given the observed data (x) and the prior distribution (p) of the latent variables.

[0068] In variational inference, maximizing the ELBO is equivalent to minimizing the KL divergence, leading to a more efficient approximation of the true posterior distribution.

[0069] Maximizing the ELBO helps in training models to generate data that closely resembles the original data distribution while ensuring that the learned representations are meaningful and can be effectively manipulated for various tasks in machine learning.

[0070] Additionally, the encoder-decoder system 200 computes the reconstruction loss Ik - %||

[0071] Further, in order to enforce a correct estimation of the hazard function, the partial likelihood of the correct ranking of the samples in the set is maximized. A hazard function is a term of art, and may be implemented using a log-linear function by using a linear layer in the hazard head. An instance takes the form (x0, t0, 80~), where t0is either the observation time of the event or the censoring time for right-censored individuals. Censoring, here, means that the target event for that individual was not observed before the termination of the study (or leaving the study), thus, participating with the partial information of surviving at least until t0. 80is theevent indicator, i.e., 80= 0 for censored cases and 80= 1 otherwise. The probability for this instance to have an event at the time t0is maximized, conditioned on the risk set R(x0) this is termed the partial likelihood. The log partial likelihood takes the form:where R (t0) is the risk set for the sample x0. By way of illustration, in an implementation the risk set R(t0) is the risk set for the sample x0may be the set of patients that are still alive when x0gets the event. The risk head in the model by be implemented using a linear layer.

[0072] Finally, each reconstructed sample goes through sparsification procedures that reduce the number of non-zero entries by using the L1norm and a masked thresholding function. This results in a newly reconstructed sparse sample x.

[0073] One possibility for the thresholding function is to set the smallest 95% of the features to zero. In order to enforce faithful creation of sparse representations, x is repassed through the network and a new loss function is applied on its embedding to guarantee that the new sample causes a ranking similar to the ranking on the original data (this may only be done in the training phase). Hence, the following likelihood is maximized, referred to herein as the double-pass partial likelihood: mDPPL = ^[ln( o=l

[0074] At least the reconstruction loss, negative partial likelihood, and negative double-pass partial likelihood are used to update the machine learning model, including updating the gradients. Optimization may be performed with the help of gradient-based methods, such as ADAM.

[0075] Aspects of the present disclosure thus provide for general improvements to computers in machine learning systems, in particular providing enhanced computer functionality for faithful sparsification and handling of censored data, as well as providing for self-explaining learned models for providing explanations for feature selection from high-dimensional data, thereby increasing security, trust, reliability and accuracy of the machine learning systems. Moreover, embodiments according to the present disclosure can be practically applied to use cases to effect further improvements in a number of technical fields including, but not limited to, medical (e.g., digital medicine, personalized healthcare, Al-assisted drug or vaccine development, survivability, disease prediction, diagnosis and treatment, etc.) and predictive maintenance and smart cities (e.g., automated traffic or vehicle control, smart districts, smart buildings, smart industrial plants, smart agriculture, energy management, etc.)

[0076] In an exemplary embodiment, the present disclosure can be practically applied to diseases with well-defined events / survival and underlying genetic factors. For a disease like cancer, survival is one of the main measurements to identify vulnerable subgroups. Additionally, over-expression of genes and the presence of certain mutations belong to the main features investigated for this disease. Gene expression, mutational patterns or other kinds of molecular information can also play a role for other types of diseases that are characterized survival or other critical events, like cardiovascular diseases with heart attacks or neurological diseases with seizure events. The input to the system includes genetic / molecular data, for example gene expression or mutation information in the form of tables for every patient. The inputs can be taken from measurements from laboratory testing equipment from patient samples, or electronic health records or databases that record this information. The system according to an embodiment of the present invention takes molecular information to create sparsified representations. These representations contain only the important, i.e., meaningful over- or under-expressed genes, mutations etc., that define the representation, for example with an increased likelihood of early events like decreased survival.

[0077] In an exemplary embodiment, the present disclosure can be practically applied for identification of fitness in terms of survivability of single cells or clones. Single cells or clones, which is a group of identical cells, possess individual genetic features like expressed genes or mutational patterns that distinguish them from other cells or clones. These features can decide whether a cell or clone is able to survive in a given condition, e.g., clones within a tumor, which is referred to as the clone’s fitness. Identifying features of clones that are more likely to survive can help in designing treatments particularly targeting these clones. The input to the system includes genetic / molecular data, for example gene expression or mutation information in the form of tables for every cell or clone. The system according to an embodiment of the present invention takes molecular information to create sparsified representations. These representations contain only the important, i.e., meaningful over- or under-expressed genes, mutations etc., that define a representation, for example a tumor clone with an increased fitness. The outputs can be used to generate diagnoses, prescriptions and / or treatments in an automated manner, while providing human-understandable explanations for the decisions.

[0078] In an exemplary embodiment, the present disclosure can be practically applied to electronic health record (EHR) data for diseases with well-defined events / survival. Diseases that are characterized survival or other critical events can also be defined by information from EHRs. EHRs might also contain a lot of irrelevant information that needs to be filtered in order to identify critical measurements that can be predictive for survival or events like heart attacks and seizures. The input to the system includes measurements from medical or laboratory equipmentsuch as diagnostic results, blood tests, etc., taken from patient samples along with other patient EHR data or patient data stored in databases. The system according to an embodiment of the present invention the takes EHR information and creates sparsified representations. These representations contain only the important, i.e., meaningful measurements, increased blood values, etc., that define a representation, for example with an increased likelihood of early events like decreased survival. The outputs can be used to generate diagnoses, prescriptions and / or treatments in an automated manner, while providing human-understandable explanations for the decisions.

[0079] Referring to FIG. 3, a processing system 300 can include one or more processors 302, memory 304, one or more input / output devices 306, one or more sensors 308, one or more user interfaces 310, and one or more actuators 312. Processing system 300 can be representative of each computing system disclosed herein.

[0080] Processors 302 can include one or more distinct processors, each having one or more cores. Each of the distinct processors can have the same or different structure. Processors 302 can include one or more central processing units (CPUs), one or more graphics processing units (GPUs), circuitry (e.g., application specific integrated circuits (ASICs)), digital signal processors (DSPs), and the like. Processors 302 can be mounted to a common substrate or to multiple different substrates.

[0081] Processors 302 are configured to perform a certain function, method, or operation (e.g., are configured to provide for performance of a function, method, or operation) at least when one of the one or more of the distinct processors is capable of performing operations embodying the function, method, or operation. Processors 302 can perform operations embodying the function, method, or operation by, for example, executing code (e.g., interpreting scripts) stored on memory 304 and / or trafficking data through one or more ASICs. Processors 302, and thus processing system 300, can be configured to perform, automatically, any and all functions, methods, and operations disclosed herein. Therefore, processing system 300 can be configured to implement any of (e.g., all of) the protocols, devices, mechanisms, systems, and methods described herein.

[0082] For example, when the present disclosure states that a method or device performs task “X” (or that task “X” is performed), such a statement should be understood to disclose that processing system 300 can be configured to perform task “X”. Processing system 300 is configured to perform a function, method, or operation at least when processors 302 are configured to do the same.

[0083] Memory 304 can include volatile memory, non-volatile memory, and any other medium capable of storing data. Each of the volatile memory, non-volatile memory, and anyother type of memory can include multiple different memory devices, located at multiple distinct locations and each having a different structure. Memory 304 can include remotely hosted (e.g., cloud) storage.

[0084] Examples of memory 304 include a non-transitory computer-readable media such as RAM, ROM, flash memory, EEPROM, any kind of optical storage disk such as a DVD, a Blu- Ray® disc, magnetic storage, holographic storage, a HDD, a SSD, any medium that can be used to store program code in the form of instructions or data structures, and the like. Any and all of the methods, functions, and operations described herein can be fully embodied in the form of tangible and / or non-transitory machine-readable code (e.g., interpretable scripts) saved in memory 304.

[0085] Input-output devices 306 can include any component for trafficking data such as ports, antennas (i.e., transceivers), printed conductive paths, and the like. Input-output devices 306 can enable wired communication via USB®, DisplayPort®, HDMI®, Ethernet, and the like. Input-output devices 306 can enable electronic, optical, magnetic, and holographic, communication with suitable memory 306. Input-output devices 306 can enable wireless communication via WiFi®, Bluetooth®, cellular (e.g., LTE®, CDMA®, GSM®, WiMax®, NFC®), GPS, and the like. Input-output devices 306 can include wired and / or wireless communication pathways.

[0086] Sensors 308 can capture physical measurements of environment and report the same to processors 302. User interface 310 can include displays, physical buttons, speakers, microphones, keyboards, and the like. Actuators 312 can enable processors 302 to control mechanical forces.

[0087] Processing system 300 can be distributed. For example, some components of processing system 300 can reside in a remote hosted network service (e.g., a cloud computing environment) while other components of processing system 300 can reside in a local computing system. Processing system 300 can have a modular design where certain modules include a plurality of the features / functions shown in FIG. 3. For example, I / O modules can include volatile memory and one or more processors. As another example, individual processor modules can include read-only-memory and / or local caches.

[0088] While subject matter of the present disclosure has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive. Any statement made herein characterizing the invention is also to be considered illustrative or exemplary and not restrictive as the invention is defined by the claims. It will be understood that changes and modificationsmay be made, by those of ordinary skill in the art, within the scope of the following claims, which may include any combination of features from different embodiments described above.

[0089] The terms used in the claims should be construed to have the broadest reasonable interpretation consistent with the foregoing description. For example, the use of the article “a” or “the” in introducing an element should not be interpreted as being exclusive of a plurality of elements. Likewise, the recitation of “or” should be interpreted as being inclusive, such that the recitation of “A or B” is not exclusive of “A and B,” unless it is clear from the context or the foregoing description that only one of A and B is intended. Further, the recitation of “at least one of A, B and C” should be interpreted as one or more of a group of elements consisting of A, B and C, and should not be interpreted as requiring at least one of each of the listed elements A, B and C, regardless of whether A, B and C are related as categories or otherwise. Moreover, the recitation of “A, B and / or C” or “at least one of A, B or C” should be interpreted as including any singular entity from the listed elements, e.g., A, any subset from the listed elements, e.g., A and B, or the entire list of elements A, B and C.

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented machine-learning method for generating sparse explanations of feature selection from high-dimensional data, the method comprising a training phase comprising: using an embedding machine learning model, encoding training data into latent space to generate first training embeddings; using the embedding machine learning model, decoding the first training embeddings to generate first reconstructed training data; computing sparsified feature selections of the first reconstructed training data; calculating reconstruction losses based on the first reconstructed training data and the training data; calculating negative partial likelihood losses based on the first training embeddings; using the embedding machine learning model, encoding the sparsified feature selections of the first reconstructed training data into the latent space to generate second training embeddings; calculating negative double-pass partial likelihood losses based on the second training embeddings; and updating the embedding machine learning model based on the reconstruction losses, the negative partial likelihood losses, and the negative double-pass partial likelihood losses.

2. The computer-implemented method according the claim 1, wherein the training phase further comprises an optimization process performed by iteratively executing operations of claim 1 to further update the embedding machine learning model until satisfying a predetermined optimization criteria, and wherein the optimization process minimizes negatives of the partial likelihood losses or maximizes the negatives of the partial likelihood losses when not negated.

3. The computer-implemented method according to any of claims 1-2, wherein the training phase of the method further comprises: using the updated embedding machine learning model to encode the training data into the latent space to generate final training embeddings; using the updated embedding machine learning model to decode the final training embeddings to generate final reconstructed training data;computing final sparsified feature selections of the final reconstructed training data; and saving at least the final sparsified feature selections of the final reconstructed training data as sparse prototypes.

4. The computer-implemented method according to any of claims 1-3, wherein the highdimensional data is patient data.

5. The computer-implemented method according to claim 4, wherein the patient data comprises: gene expression data, micro-RNA data, electronic health records, or blood test results.

6. The computer-implemented method according to claim 4, wherein the patient data comprises a plurality of vectors respectively corresponding to a plurality of patients, each vector comprising a plurality of coefficients, wherein the embedding machine learning model comprises a variational autoencoder architecture comprising an encoder and a decoder, wherein the operation of using the embedding machine learning model to encode the training data into the latent space to generate the first training embeddings comprises: obtaining a sample from a fixed distribution; for each of the vectors, using the encoder, determining a respective mean and a respective standard deviation; and for each of the vectors, transforming the sample into a respective one of the first training embeddings using a parameterization technique based on the sample, the respective mean, and the respective standard deviation.

7. The computer-implemented method according to claim 6, wherein the operation of using the embedding machine learning model to decode the first training embeddings to generate first reconstructed training data comprises using the decoder, generating the first reconstructed training data as a reconstructed vector of coefficients based on the embeddings, and wherein the operation of calculating reconstruction losses based on the first reconstructed training data and the training data comprises using the formula:II*- *112 where x is the vector of coefficients from the high-dimensional data and x is the reconstructed vector of coefficients.

8. The computer-implemented method according to claims 6 or 7, wherein the operation of computing sparsified feature selections of the first reconstructed training data comprises: for each reconstructed vector, reduce a number of non-zero entries by using Li norm and a masked thresholding function, to generate a reconstructed sparse sample of the reconstructed vector.

9. The computer implemented method according to claim 8, wherein the masked thresholding function may comprise setting a smallest predefined percentage of features in the reconstructed vector to zero.

10. The computer implemented method according to any of claims 6-9, wherein using the embedding machine learning model to encode the sparsified feature selections of the first reconstructed training data into the latent space to generate the second training embeddings comprises: providing as input to the encoder the reconstructed sparse sample of each of the reconstructed vectors as sparse vectors; obtaining a sample from a fixed distribution; for each of the sparse vectors, using the encoder, determining a respective mean and a respective standard deviation; and for each of the sparse vectors, transforming the sample into a respective one of the second training embeddings using the parameterization technique based on the sample, the respective mean, and the respective standard deviation.

11. The computer-implemented method according to claim 10, wherein the operation of calculating the negative double-pass partial likelihood losses based on the second training embeddings comprises maximizing the following double-pass partial likelihood: mDPPL = ^[ln(r(x0))) + ln(r(x0))] — o=lwhere an instance is given in a form of (x0, t0, 50), t0is either as an observation time of the instance or a censoring time for right-censored individuals. 80is an event indicator where 80= 0 for censored cases and 80= 1 otherwise, x0is the vector for a patient, x0is the reconstructed vector for a patient , r(x0) is a risk set.

12. The computer-implement method according to any of claims 1-11, the method further comprising a production phase executed after completion of the training phase, the production phase comprising: using the embedding machine learning model, encoding test data into the latent space to generate test embeddings;using the embedding machine learning model, decoding the test embeddings to generate reconstructed test data; and computing sparsified feature selections of the reconstructed test data.

13. The computer-implemented method according to claim 12, the production phase further comprising comparing the sparsified feature selections to the sparse prototypes and making a prediction regarding the test data based on comparison.

14. A computer system comprising one or more hardware processors which, alone or in combination, are configured to execute a machine-learning method for generating sparse explanations of feature selection from high-dimensional data, the method comprising a training phase comprising: using an embedding machine learning model, encoding training data into latent space to generate first training embeddings; using the embedding machine learning model, decoding the first training embeddings to generate first reconstructed training data; computing sparsified feature selections of the first reconstructed training data; calculating reconstruction losses based on the first reconstructed training data and the training data; calculating negative partial likelihood losses based on the first training embeddings; using the embedding machine learning model, encoding the sparsified feature selections of the first reconstructed training embeddings into the latent space to generate second training embeddings; calculating negative double-pass partial likelihood losses based on the second training embeddings; and updating the embedding machine learning model based on the reconstruction losses, the negative partial likelihood losses, and the negative double-pass partial likelihood losses.

15. A tangible, non-transitory computer-readable medium having instructions thereon which, upon being executed by one or more hardware processors, alone or in combination, provide for execution a machine -learning method for generating sparse explanations of feature selection from high-dimensional data, the method comprising a training phase comprising: using an embedding machine learning model, encoding training data into latent space to generate first training embeddings; using the embedding machine learning model, decoding the first training embeddings to generate first reconstructed training data; computing sparsified feature selections of the first reconstructed training data;calculating reconstruction losses based on the first reconstructed training data and the training data; calculating negative partial likelihood losses based on the first training embeddings; using the embedding machine learning model, encoding the sparsified feature selections of the first reconstructed training embeddings into the latent space to generate second training embeddings; calculating negative double-pass partial likelihood losses based on the second training embeddings; and updating the embedding machine learning model based on the reconstruction losses, the negative partial likelihood losses, and the negative double-pass partial likelihood losses.