Computing systems and methods for event prediction using a machine learning model

CA3263074A1Pending Publication Date: 2026-09-21THE TORONTO DOMINION BANK
0 Cites 0 Cited by

Patent Information

Application Number
CA3263074
Authority / Receiving Office
CA · CA
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2026-09-21
Patent Text Reader

Abstract

Methods and systems for performing event prediction using a machine learning model. The methods include training, using a diffusion model and a plurality of synthetic tabular datasets representing a prior, a machine learning model to generate a tabular dataset that has a distribution that is consistent with the prior; conditioning the trained machine learning model to: receive (i) an input tabular dataset comprising one or more datapoints, each datapoint comprising one or more elements, and (ii) information identifying one or more elements of the input tabular dataset that are to be predicted, and predict the identified one or more elements of the input tabular dataset based on the other elements of the input tabular dataset; and using the conditioned trained machine learning model to perform event prediction.
Need to check novelty before this filing date? Find Prior Art

Description

COMPUTING SYSTEMS AND METHODS FOR EVENT PREDICTION USING A MACHINE LEARNING MODEL TECHNICAL FIELD 5

[0001] The disclosed example embodiments relate to computer-implemented methods and systems for event prediction using a machine learning model and more specifically, predicting the likelihood of an event occurring using a variant of a priordata fitted network. BACKGROUND 10

[0002] Machine learning models can be used to predict the likelihood of an event occurring by training machine learning models to identify patterns that indicate the likelihood of the event. Such models can be used for event prediction in a variety of industries including, but not limited to, healthcare, education, manufacturing, energy and utilities, technology and cybersecurity, real estate and construction, transportation 15 and logistics, education, and hospitality and travel.

[0003] Traditionally machine learning models are trained to perform a specific task. Specifically, traditionally machine learning models are trained on a specific labelled dataset to learn and replicate the relationship between the input(s) and the output(s) in the labelled dataset. For example, as shown in FIG. 1, a machine learning 20 model may be trained to predict an output 𝑦 based on an input 𝑥 from a training dataset comprising (𝑥, 𝑦) pairs. At inference time the trained model receives a new 𝑥 and generates a prediction for 𝑦.

[0004] However, when a model is trained on a specific dataset, it is difficult to use that model for a different dataset. In other words, a model trained to make a 25 prediction based on one dataset may be suboptimal at making a prediction based on another dataset. For example, a machine learning model trained to predict the likelihood of a flight being delayed in response to a set of features related to a flight based on a training dataset comprising (set of features related to a flight, whether the flight was delayed) pairs may be suboptimal at predicting the likelihood of an individual 30 defaulting on a credit card in response to a set of features related to the individual and their credit card in accordance with a dataset that comprises (set of features related to an individual and their credit card, whether the individual defaulted on the credit card) pairs. To obtain a prediction based on the other dataset (e.g., the dataset that– 2 – comprises (set of features related to the individual and their credit card, whether the individual defaulted on the credit card) pairs, it may be necessary to train another model on the other dataset.

[0005] Generating and maintaining a separate model for each event for which 5 predictions may be desired is both labour and time intensive and requires significant computing resources to store each model. SUMMARY

[0006] The following summary is intended to introduce the reader to various aspects of the detailed description, but not to define or delimit any invention. 10

[0007] A first aspect provides a system for performing event prediction, the system comprising: a memory, a communication interface, and at least one processor operatively coupled to the memory and the communication interface; the at least one processor configured to: train, using a diffusion model and a plurality of synthetic tabular datasets representing a prior, a machine learning model to generate a tabular 15 dataset that has a distribution that is consistent with the prior; condition the trained machine learning model to: receive (i) an input tabular dataset comprising one or more datapoints, each datapoint comprising one or more elements, and (ii) information identifying one or more elements of the input tabular dataset that are to be predicted, and predict the identified one or more elements of the input tabular dataset based on 20 the other elements of the input tabular dataset; and use the conditioned trained machine learning model to perform event prediction.

[0008] The at least one processor may be configured to use the conditioned trained machine learning model to perform event prediction by providing the conditioned trained machine learning model with (i) an event input tabular dataset 25 comprising a plurality of datapoints and (ii) event information identifying one or more elements of the event input tabular dataset to be predicted, and the plurality of datapoints may comprise one or more training datapoints and an input datapoint, and the event information may identify an element of the input datapoint to be predicted.

[0009] Each training datapoint may comprise a set of elements that represent 30 a set of features of a record and an element that represents an occurrence of the event for the record, and the input datapoint may comprise a set of elements that represent the set of features for a new record.– 3 –

[0010] The input datapoint may comprises an element that represents an occurrence of the event for the new record and the element that represents the occurrence of the event for the new record may not comprise valid data; and the event information may identify the element of the input datapoint that represents the 5 occurrence of the event for the new record as an element to be predicted such that the conditioned trained machine learning model predicts the occurrence of the event for the new record.

[0011] The event information may comprise a mask.

[0012] The machine learning model may comprise a transformer architecture. 10

[0013] The machine learning model may be configured to tokenize each element of the event input tabular dataset.

[0014] Using the conditioned trained machine learning model to perform event prediction may comprise using the conditioned trained machine learning model to predict a likelihood of an entity requesting a financial product within a predetermined 15 period of time.

[0015] The financial product may be a credit card, unsecured line of credit, unsecured loan, real estate secured lending, investing account, chequing account, personal investment, savings account, term life policy, overdraft protection, trade account, travel medical insurance, or balance protection insurance. 20

[0016] Using the conditioned trained machine learning model to perform event prediction may comprise using the conditioned trained machine learning model to predict fraudulent activity related to a financial product.

[0017] Predicting the fraudulent activity related to the financial product may comprises predicting an account take over risk, predicting mule account risk, 25 predicting money laundering account risk, predicting application fraud, or predicting breach of an account.

[0018] Using the conditioned trained machine learning model to perform event prediction may comprise using the conditioned trained machine learning model to predict delinquency related to a financial product. 30

[0019] The financial product may comprise a credit card, an unsecured line of credit, an unsecured loan, or a real estate secured lending.– 4 –

[0020] The at least one processor may be further configured to generate one or more synthetic tabular datasets of the plurality of synthetic tabular datasets by sampling a Bayesian neural network.

[0021] The at least one processor may be further configured to generate one or 5 more synthetic tabular datasets of the plurality of synthetic tabular datasets by sampling a structural causal model.

[0022] The at least one processor may be configured to train the machine learning model using the diffusion model and the plurality of synthetic tabular datasets by diffusing each of the plurality of synthetic tabular datasets in accordance with a 10 diffusion process to generate a plurality of diffused datasets and adjusting parameters of the machine learning model so that the machine learning model reverses the diffusion process on the plurality of diffused datasets.

[0023] Adjusting the parameters of the machine learning model so that the machine learning model reverses the diffusion process may comprise adjusting the 15 parameters of the machine learning model so that the machine learning model generates, in response to receiving a diffused dataset, the corresponding synthetic tabular dataset.

[0024] A second aspect provides a method for performing event prediction, the method executed in a computing environment comprising at least one processor, a 20 communication interface, and memory, and the method comprising: training, using a diffusion model and a plurality of synthetic tabular datasets representing a prior, a machine learning model to generate a tabular dataset that has a distribution that is consistent with the prior; conditioning the trained machine learning model to: receive (i) an input tabular dataset comprising one or more datapoints, each datapoint 25 comprising one or more elements, and (ii) information identifying one or more elements of the input tabular dataset that are to be predicted, and predict the identified one or more elements of the input tabular dataset based on the other elements of the input tabular dataset; and using the conditioned trained machine learning model to perform event prediction. 30

[0025] Using the conditioned trained machine learning model to perform event prediction may comprise providing the conditioned trained machine learning model with an event input tabular dataset comprising a plurality of datapoints and event– 5 – information identifying one or more elements of the event input tabular dataset to be predicted, wherein the plurality of datapoints comprise one or more training datapoints and an input datapoint, and the event information identifies an element of the input datapoint to be predicted. 5

[0026] According to some aspects, the present disclosure provides a nontransitory computer-readable medium storing computer-executable instructions. The computer-executable instructions, when executed, configure a processor to perform any of the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS 10

[0027] The drawings included herewith are for illustrating various examples of articles, methods, and systems of the present specification and are not intended to limit the scope of what is taught in any way. In the drawings: FIG. 1 is a schematic diagram illustrating inputs and outputs of a traditional trained machine learning model; 15 FIG. 2 is a schematic diagram illustrating inputs and outputs of a prior-data fitted network (PFN); FIG. 3 is a schematic diagram illustrating inputs and outputs of a diffusion PFN; FIG. 4 is a block diagram of an example system for generating a conditioned diffusion PFN and using the conditioned diffusion PFN to perform event prediction; 20 FIG. 5 is a block diagram of an example implementation of the cloud-based computing cluster of FIG. 4 configured to generate a conditioned diffusion PFN and use the conditioned diffusion PFN to perform event prediction; FIG. 6A is a schematic diagram of an example Bayesian neural network (BNN); FIG. 6B is a schematic diagram of an example Structural Causal Model (SCM); 25 FIG. 6C is a schematic diagram of example SCMs sampled from a prior; FIG. 7A is a schematic diagram illustrating a first example of using a conditioned diffusion PFN to generate tabular data; FIG. 7B is a schematic diagram illustrating a second example of using a conditioned diffusion PFN to generate tabular data; 30 FIG. 8 is a block diagram of an example computer; and– 6 – FIG. 9 is a flow diagram of an example method for event prediction using a machine learning model. DETAILED DESCRIPTION

[0028] Described herein are methods and systems for performing event 5 prediction using a variant of a prior-data fitted network (PFN) trained to process tabular data. As described below, once trained, the described PFN variant can be used to predict a variety of different events, without retraining the model.

[0029] Mueller et al. in “Transformers Can Do Bayesian Inference”, International Conference of Learning Representations, 2021, introduce the concept of 10 a PFN which is a large machine learning model, such as a Transformer encoder, which is trained offline once, to approximate Bayesian inference on synthetic datasets from a prior. Specifically, as shown in FIG. 2, a PFN is designed to receive an input 𝑥, and a test or reference dataset, 𝐷, comprising an arbitrary number 𝑛 of datapoints 𝑑(𝑖) = (𝑥(𝑖), 𝑦(𝑖)) = (𝑑1(𝑖), … , 𝑑𝑘(𝑖)) where 𝑑𝑘(𝑖) = 𝑦 and 𝑦(𝑖) is the output for 𝑥(𝑖) (i.e., 𝐷 = 15 {𝑑(1), … , 𝑑(𝑛)} = {(𝑥(𝑖), 𝑦(𝑖))}𝑖 𝑛=1); and predict an output 𝑦 based on 𝑥 and 𝐷. More specifically the PFN is configured to generate the posterior predictive distribution (PPD) 𝑝(𝑦|𝑥, 𝐷) – i.e., the distribution over possible values of 𝑦 for the input 𝑥 – in a single forward pass.

[0030] A PFN is pre-trained on a plurality of distinct test datasets which are 20 synthetically generated from a prior. This pre-training may be referred to as the prior fitting stage or phase, or the offline phase. Specifically, a plurality of synthetic datasets are generated by sampling a prior. In particular, each synthetic dataset comprise a plurality of samples / datapoints 𝑑(𝑖) = (𝑥(𝑖), 𝑦(𝑖)). For each of the plurality of synthetic datasets a test dataset 𝐷 is generated which comprises a subset of the samples 𝑑(𝑖) = 25 (𝑥(𝑖), 𝑦(𝑖)) in the synthetic dataset. The number of samples that form a test dataset are randomly selected from the synthetic dataset. The synthetic datasets and test datasets have an arbitrary size such that different synthetic datasets and different test datasets may have a different number of samples. The samples from a synthetic dataset which do not form part of the corresponding test dataset are referred to as the 30 holdouts. The PFN is trained to predict or generate the holdouts for each synthetic dataset from the corresponding test dataset.– 7 –

[0031] For example, if a synthetic dataset comprises samples / datapoints 𝑑(1) = (𝑥(1), 𝑦(1)), 𝑑(2) = (𝑥(2), 𝑦(2)), 𝑑(3) = (𝑥(3), 𝑦(3)), 𝑑(4) = (𝑥(4), 𝑦(4)) and 𝑑(5) = (𝑥(5), 𝑦(5)), and the test dataset 𝐷 comprises only samples / datapoints 𝑑(1) = (𝑥(1), 𝑦(1)), 𝑑(2) = (𝑥(2), 𝑦(2)), and 𝑑(5) = (𝑥(5), 𝑦(5)), then the PFN is trained to (i) predict 𝑦(3) 5 from 𝑥(3) and 𝐷; and (ii) predict 𝑦(4) from 𝑥(4) and 𝐷. Specifically, Muller et al. propose updating the parameters of the machine learning model via gradient descent to minimizing the log-likelihood.

[0032] The PFN learns properties to generalize over the plurality of synthetic datasets. In this manner the PFN is trained to learn the posterior predictive distribution 10 (PPD) 𝑝(𝑦|𝑥, 𝐷). In some cases, there may be millions of synthetic datasets.

[0033] Mueller et al. indicate that many neural architectures can be used as the machine learning model. However, they explored architectures based on the Transformer encoder and a novel regression head for regression problems. Specifically, Mueller et al. describe an architecture based on a Transformer encoder 15 without positional encodings, which makes it invariant to permutations in the dataset. The Transformer encodes each datapoint (e.g., each feature vector (𝑥) and label (𝑦) combination) as a token, allowing token representations to attend to each other. For example, training samples / datapoints 𝑑(1) = (𝑥(1), 𝑦(1)), 𝑑(2) = (𝑥(2), 𝑦(2)), 𝑑(3) = (𝑥(3), 𝑦(3)), are transformed to 3 tokens, which attend to each other and test samples 20 𝑥(4), and 𝑥(5) attend only to the training samples.

[0034] Once the PFN has been pre-trained, the PFN can receive a new test or reference dataset 𝐷 and an input or query 𝑥 and predict 𝑦 (i.e., generate 𝑝(𝑦|𝑥, 𝐷)). This may be referred to as the prediction or online stage or phase. The test or reference dataset can change from one inference to another such that the PFN 25 performs in-context learning. Thus, the weights of a pre-trained PFN do not need to be updated to be able to generate a prediction for a new dataset. In contrast, the PFN has this knowledge built in.

[0035] Mueller et al. demonstrated that PFNs can be used to perform tabular classification problems. However, their work was limited to 30 training samples, 30 balanced binary classification and 60 features.

[0036] Hollman et al. in “TabPFN: A Transformer That Solves Tabular Classification Problems In a Second”, The Eleventh International Conference on– 8 – Learning Representations, 2022, built on the PFN concept to develop TabPFN, a trained Transformer that can do supervised classification for small tabular datasets. Specifically, Hollman et al. designed a prior based on Bayesian Neural Networks and Structural Causal Models (SCMs) to model complex feature dependencies and 5 potential causal mechanisms underlying tabular data and trained a PFN on samples from the prior. In particular, the prior has a large space of structural causal model with preference for simple structure. After training, the trained TabPFN model accepts training samples (𝑥𝑡𝑟𝑎𝑖𝑛, 𝑦𝑡𝑟𝑎𝑖𝑛) and test features 𝑥𝑡𝑒𝑠𝑡, and yields predictions 𝑦𝑡𝑒𝑠𝑡 for the entire test set in a single forward pass. 10

[0037] By extending the prior (vs PFN), TabPFN can handle imbalanced classes and multi-classification problems. Furthermore, while PFNs use an encoder layer that accepts fixed dimensional inputs, Tab PFN accepts datasets with different numbers of dimensions. Hollman et al. proposed a 12 layer Transformer, embeddings size 512, hidden size 1024 in feed forward layers and 4-head attention. 15

[0038] It has been shown that a TabPFN can generalize to virtually any tabular dataset (a dataset comprising columns and rows) through in context learning.

[0039] Since PFNs and TabPFNs are configured to predict the label(s) 𝑦, PFNs and TabPFNs are suitable for classification and regression problems. Specifically, classification and regression are both supervised machine learning techniques that 20 use labeled data (e.g. (𝑥(𝑖), 𝑦(𝑖)) pairs) to find patterns and predict outcomes. They differ in the type of output. Specifically, classification predicts categorical output, such as a label or class from a predefined list. In contrast, regression predicts continuous numerical output, such as a real-valued number that can vary within a range. However, it would be desirable to have a foundation model that can perform generative 25 tasks (e.g., generate data) on / for tabular data.

[0040] Ma et al. in “TabPFGen = Tabular Data Generation with TabPFN”, NeurIPS 2023 Second Table Representation Workshop, 2023, devised a technique to turn TabPFN into an energy-based generative model, which is referred to as TabPFGen. TabPFGen leverages the strong in-context performance of TabPFN to 30 devise a class-conditional generative model. In particular, TabPFGen generates 𝑝(𝑥|𝑦) using TabPFN. Specifically, given a trained TabPFN model 𝑓, TabPFGen obtains 𝑓(𝑥𝑠𝑦𝑛𝑡ℎ) using (𝑥𝑡𝑟𝑎𝑖𝑛, 𝑦𝑡𝑟𝑎𝑖𝑛) as training data. Then the class conditioned– 9 – energy 𝐸(𝑥|𝑦) ≔ −𝑓(𝑥)[𝑦] is computed using 𝐸(𝑥𝑠𝑦𝑛𝑡ℎ|𝑦𝑠𝑦𝑛𝑡ℎ) ≔ −𝑓(𝑥𝑠𝑦𝑛𝑡ℎ)[𝑦𝑠𝑦𝑛𝑡ℎ]. The stochastic gradient Langevin dynamics (SGLD) method is then used to sample from this energy-based model to generate a batch of 𝑥𝑠𝑦𝑛𝑡ℎ. TabPFN harnesses a pre-trained TabPFN to generate an energy based model for tabular data generation 5 without additional training.

[0041] Accordingly, TabPFN provides a way to sample from the in contextdistribution 𝑝(𝑑|𝐷). In other words, it gives the distribution of a datapoint given a training dataset. However, it would be more beneficial to be able to generate the whole unconditional distribution 𝑝(𝐷). Specifically, once a generative model has learned the 10 full unconditional distribution 𝑝(𝐷), then in principle the generative model can then generate from any conditional 𝑝((𝑑𝑗(1𝑖1), … , 𝑑𝑗(𝑙𝑖𝑙))|𝐷𝑡𝑒𝑠𝑡\(𝑑𝑗(1𝑖1), … , 𝑑𝑗(𝑙𝑖𝑙))) – i.e., it can then impute any missing entries in a dataset given the rest of a dataset. Thus, any combination of individual entries or elements 𝑑 ( 𝑗 𝑖) of a dataset can be imputed.

[0042] While it is possible to argue that in principle learning 𝑝(𝑑|𝐷) is sufficient 15 for imputing individual entries or elements 𝑑𝑗(𝑖)of a datapoint, learning 𝑝(𝐷) allows the model to generate entries that are correlated across datapoints. Imputations that are correlated across datapoint are highly desirable since it is desirable for imputations to be consistent across datapoints. Furthermore, learning 𝑝(𝐷) is conceptually simpler and easier to implement in terms of time and computing resources that learning 20 𝑝(𝑑|𝐷).

[0043] Accordingly described herein are methods and systems for performing event prediction using a machine learning model by generating a generative foundation model for tabular data and using that generative foundation model to predict the event. Specifically, in the examples described herein first, a variant of a 25 PFN, which is referred to herein as a diffusion PFN, is generated by training, using a diffusion model and a plurality of different synthetic datasets from a prior, a large machine learning model (e.g., a transformer model) to generate an entire synthetic dataset which has a distribution that is consistent with the prior. The plurality of synthetic datasets may be generated, for example, in a similar manner to how the 30 synthetic datasets are generated for a PFN or TabPFN (e.g., by randomly sampling a prior). Accordingly, where PFNs and TabPFNs generate 𝑝(𝑦|𝑥, 𝐷), a diffusion PFN– 10 – generates 𝑝(𝐷) directly. Thus, as shown in FIG. 3, a diffusion PFN can generate an entire synthetic dataset (e.g., a set of datapoints (𝑥,𝑦)) from scratch. A diffusion PFN is therefore a generative foundation model.

[0044] Once the machine learning model (i.e., diffusion PFN) has been trained 5 to generate an entire synthetic dataset, the trained machine learning model is conditioned to perform imputation (i.e., estimating or inferring missing values in a dataset using algorithms and other datapoints). A conditioned generative model takes additional inputs as conditions to control the generation process. In some cases, the trained machine learning model may be conditioned to impute any missing entries in 10 an input dataset given the rest of a dataset – i.e., the trained machine learning model can be conditioned to generate 𝑝((𝑑𝑗(1𝑖1), … , 𝑑𝑗(𝑙𝑖𝑙))|𝐷𝑡𝑒𝑠𝑡\(𝑑𝑗(1𝑖1), … , 𝑑𝑗(𝑙𝑖𝑙))). As described in more detail below, this may be implemented, for example, by conditioning the trained machine learning model down to receive (i) an input dataset and (ii) information (e.g., a mask) indicating which datapoints or elements of the input dataset are to be 15 predicted or generated; and generate the identified elements of the input dataset based on the other elements of the input dataset.

[0045] Once the trained machine learning model has been conditioned to impute any missing entries in a dataset given the rest of the dataset, the conditioned and trained machine learning model (which may be referred to herein as the 20 conditioned diffusion PFN) can be used to perform event prediction. Specifically, the conditioned diffusion PFN can be used to predict the likelihood of the event based on features related to the event by providing the conditioned diffusion PFN with an input tabular dataset that comprises a plurality of datapoints, wherein each datapoint comprises a plurality of elements. The plurality of datapoints comprise one or more 25 test or training datapoints and an input datapoint. Each training datapoint comprises a set of elements that represent a set of features related to an example (e.g., historical) record and an element indicating whether the event occurred for that record. The input datapoint (which may, in some cases, be the last datapoint in the input tabular dataset) comprises a set of elements that represent the set of features for a new record but 30 does not comprise a valid element that represents whether the event occurred. The conditioned diffusion PFN is then instructed (e.g., via a mask) to generate or predict the missing element of the input tabular dataset – i.e., the element representing the event occurrence for the new record – based on the other elements of the input– 11 – dataset. In other words, the conditioned diffusion PFN generates a prediction of whether the event occurred based on the set of features for the new record and the training datapoints.

[0046] For example, if the objective is to predict whether an individual will 5 default on a credit card (e.g. whether the individual will miss multiple required payments over a period of time such that the credit card issuer writes the debt off as a loss) then an input tabular dataset with 𝑛 datapoints may be provided to the conditioned diffusion PFN wherein: each of datapoints 1 to 𝑛-1 are training datapoints that comprise (i) elements that represent a set of features related to an example 10 individual that are relevant to credit card default such as, but not limited to, age, credit score, salary etc., (ii) and an element that represents whether the individual defaulted on the credit card; and datapoint 𝑛 is the input datapoint that comprises elements that represent to the set of features for a new individual who has applied for a credit card. The input data point (i.e., datapoint 𝑛) does not comprise a valid element that 15 represents whether the individual defaulted on their credit card. The conditioned diffusion PFN is then instructed (e.g., via a mask) to predict the missing element based on the other elements of the input tabular dataset. In other words, the mask instructs the conditioned diffusion PFN to predict whether the individual will default on the credit card based on the set of features for the individual and the training datapoints. 20

[0047] Reference is now made to FIG. 4, which illustrates a block diagram of an example computing system 400 for performing event prediction using a machine learning model. Computing system 400 comprises a source database system 402, an enterprise data provisioning platform (EDPP) 404 operatively coupled to the source database system 402, and a cloud-based computing cluster 406 that is operatively 25 coupled to the EDPP 404.

[0048] Source database system 402 has one or more databases, of which three are shown for illustrative purposes: database 408a, database 408b and database 408c. One or more of the databases of the source database system 402 may contain confidential information that is subject to restrictions on export. One or more export 30 modules 410a, 410b, 410c may periodically (e.g., daily, weekly, monthly, etc.) export data from the databases 408a, 408b, 408c to the EDPP 404. In some instances, the data is exported on an ad hoc basis.– 12 –

[0049] EDPP 404 receives source data exported by the export modules 410a, 410b, 410c of source database system 402, processes it and exports the processed data to an application database within the cloud-based computing cluster 406. For example, a parsing module 412 of EDPP 404 may perform extract, transform and load 5 (ETL) operations on the received source data.

[0050] In many environments, access to the EDPP may be restricted to relatively few users, such as administrative users. However, with appropriate access permissions, data relevant to a document or group of documents (e.g., a client document) may be exported via reporting and analysis module 414 or an export 10 module 416a, 416b, 416c. In particular, parsed data can then be processed and transmitted to the cloud-based computing cluster 406 by a reporting and analysis module 414. Alternatively, one or more export modules 416a, 416b, 416c can export the parsed data to the cloud-based computing cluster 406.

[0051] In some cases, there may be confidentiality and privacy restrictions 15 imposed by governmental, regulatory, or other entities on the use or distribution of the source data. These restrictions may prohibit confidential data from being transmitted to computing systems that are not “on-premises” or within the exclusive control of an organization, for example, or that are shared among multiple organizations, as is common in a cloud-based environment. In particular, such privacy restrictions may 20 prohibit the confidential data from being transmitted to distributed or cloud-based computing systems, where it can be processed by machine learning systems, without appropriate anonymization or obfuscation of personal identifiable information (PII) in the confidential data. Moreover, such “on-premises” systems typically are designed with access controls to limit access to the data, and thus may not be resourced or 25 otherwise suitable for use in broader dissemination of the data. In some cases, to comply with such restrictions, one or more module of EDPP 404 may “de-risk” data tables that contain confidential data prior to transmission to cloud-based computing cluster 406. In some cases, this de-risking process may obfuscate or mask elements of confidential data, or may exclude certain elements, depending on the specific 30 restrictions applicable to the confidential data. The specific type of obfuscation, masking or other processing is referred to as a “data treatment.”

[0052] The cloud-based computing cluster 406 is configured to generate a diffusion PFN, condition the generated diffusion PFN to generate a conditioned– 13 – diffusion PFN, and use the conditioned diffusion PFN and data received from the EDPP 404 to perform event prediction. The cloud-based computing cluster 406 includes an interface 418, which facilitates data communication with one or more user devices 420. 5

[0053] In some environments, the EDPP may be omitted. In such cases the cloud-based computing cluster 406 may receive the data, for use with the conditioned diffusion PFN, directly from the source database system 402.

[0054] Reference is now made to FIG. 5, which illustrates an example implementation of the cloud-based computing cluster 406 of FIG. 4. As described 10 above, the cloud-based computing cluster 406 is configured to perform event prediction using a machine learning model by generating a generative foundation model for tabular data and using that generative foundation model to perform the event prediction. In the example shown in FIG. 5 the cloud-based computing cluster 406 comprises a first system 502 which is configured to generate a generative foundation 15 model, which is referred to herein as a diffusion PFN 504, which can generate a tabular dataset; a second system 506 which is configured to modify the diffusion PFN 504 to generate a conditioned diffusion PFN 508 which can perform imputation on an input tabular dataset; and a third system 510 which is configured to use the conditioned diffusion PFN 508 to perform event prediction. 20

[0055] In some cases, one or more components of the cloud-based computing cluster 406 may be implemented by one or more computers within the cloud-based computing cluster, such as, but not limited to, computer 800 described below with respect to FIG. 8. In some cases, one or more components of the cloud-based computing cluster 406 may be implemented as virtual machines within the cloud- 25 based computing cluster 406.

[0056] The first system 502 is configured to generate the diffusion PFN 504 by training a machine learning model 512 to generate a plurality of synthetic tabular datasets 514 (i.e., to generate 𝑝(𝐷)) using a diffusion model (e.g., using diffusion modelling). In the example shown in FIG. 5, the first system 502 comprises the 30 machine learning model 512, a diffusion model 516 and an evaluation module 518.

[0057] The plurality of synthetic tabular datasets 514 are synthetic datasets 𝐷 generated by sampling a prior. Each synthetic dataset, 𝐷, comprises an arbitrary– 14 – number 𝑛 of datapoints 𝑑(𝑖) = (𝑥(𝑖), 𝑦(𝑖)) = (𝑑1(𝑖), … , 𝑑𝑘(𝑖)) where 𝑑𝑘(𝑖) = 𝑦 and 𝑦(𝑖) is the output for 𝑥(𝑖) (i.e., 𝐷 = {𝑑(1), … , 𝑑(𝑛)} = {(𝑥(𝑖), 𝑦(𝑖))}𝑖 𝑛=1). Each 𝑑𝑗(𝑖) of a datapoint 𝑖 is referred to as an entry or an element of the datapoint. An entry that forms part of 𝑥(𝑖) is referred to as feature, and an entry that forms part of 𝑦(𝑖) is referred to as a target 5 or a label. 𝑛 and 𝑘 can vary between datasets such that the number of datapoints per dataset and the number of entries per datapoint may vary between datasets.

[0058] In some cases, the plurality of synthetic tabular datasets 514 are generated from a prior that represents tabular datasets in the real world. In some cases, the prior is based on Structural Causal Models (SCMs) and / or Bayesian Neural 10 Networks to model complex feature dependencies and potential causal mechanisms underlying tabular data as described by Hollman et al.

[0059] Specifically, since tabular data often exhibits causal relationships between columns, in some cases, the prior for the diffusion PFN may be based on SCMs that model causal relationships. An SCM is a collection 𝑍 ≔ ({𝑧1, … , 𝑧𝑘}) of 15 structural assignments (called mechanisms): 𝑧𝑖 = 𝑓𝑖 (𝑧𝑃𝐴𝐺(𝑖), 𝜖𝑖), where 𝑃𝐴𝐺(𝑖)is the set of parents of the node 𝑖 (its direct causes) in an underlying DAG 𝐺 (the causal graph), 𝑓𝑖 is a (potentially nonlinear) deterministic function and 𝜖𝑖 is a noise variable. Causal relationships in 𝐺 are represented by directed edges pointing from causes to effects and each mechanism 𝑧𝑖 is assigned to a node in 𝐺. An example SCM is shown in FIG. 20 6B.

[0060] To create a prior based on SCMs a plurality of datasets may be generated wherein each dataset is based on one randomly-sampled SCM (including the DAG structure and deterministic functions 𝑓𝑖). Given an SCM, a set of nodes 𝑧𝑋 in the causal graph 𝐺 are sampled (one for each feature) along with one node 𝑧𝑦 from 𝐺. 25 For each SCM and set of nodes (𝑧𝑋, 𝑧𝑦) a plurality (e.g., 𝑛) samples are generated by sampling all noise variables in the SCM 𝑛 times. These are then propagated through the graph and the values for all the identified nodes (𝑧𝑋, 𝑧𝑦) are retrieved for all 𝑛 samples. FIG. 6C shows example SCMs sampled from the prior.

[0061] In some cases, the prior may also, or alternatively, be based on 30 Bayesian Neural Networks (BNNs). A BNN is a type of neural network that incorporates uncertainty into its predictions. Specifically, when a regular neural– 15 – network makes a prediction, it gives a simple value based on the inputs. However, BNNs not only make a prediction but qualify the uncertainly. This is achieved by treating the weights of the BNN as distributions instead of fixed numbers. An example BNN is shown in FIG. 6A. In some cases, datasets for the prior may be generated by 5 (1) sampling a BNN architecture; (2) sampling model weights for the selected architecture; (3) for each datapoint, an input is sampled for each feature; (4) for each datapoint, the selected features are fed through the BNN with sampled noise variables and the output 𝑦 is used as the target for that set of features.

[0062] In some cases, the datasets may be sampled from either one or the other 10 prior (SCN prior or BNN prior) with equal probability.

[0063] Where the SCN or the BNN returns a scalar label, in order to generate synthetic classification labels for imbalanced multi-class datasets the scalar labels may be transformed to discrete class labels. In some cases, this may be implemented by splitting the values of the scalar labels into intervals that map to class labels by: (1) 15 sampling the number of classes 𝑁𝑐~𝑝(𝑁𝑐); (2) sampling 𝑁𝑐-1 class bounds 𝐵𝑖 randomly from the set of continuous targets 𝑦; and (3) mapping each scalar label 𝑦 to the index of the unique interval that contains it.

[0064] Generating synthetic datasets from a combination of SCMs and BNNs has proven a powerful inductive basis for applying large models, such as machine 20 learning models with a transformer backbone, to small datasets.

[0065] In some cases, the synthetic tabular datasets 514 may be stored in a repository 520 of the cloud-based computing cluster 406. The repository 520 may be any mechanism or device, such as memory, that can store digital information. In some cases, the cloud-based computing cluster 406 may receive the plurality of synthetic 25 tabular datasets 514 from, for example, the source database system 402 or the EDPP 404. However, in other cases, the cloud-based computing cluster 406 may comprise a synthetic dataset generator 522 which is configured to generate the plurality of synthetic tabular datasets 514 in accordance with any of the above methods.

[0066] The machine learning model 512 is a large and / or powerful machine 30 learning model, such as, but not limited to, a deep learning model with a transformer architecture. In some cases, the machine learning model 512 is a deep learning model with a transformer architecture, however, instead of having one token per datapoint– 16 – (i.e., per 𝑑(𝑖) = (𝑥(𝑖), 𝑦(𝑖)) = (𝑑1(𝑖), … , 𝑑𝑦(𝑖))) like TabPFN, there is a token for each 𝑑𝑗(𝑖) in a dataset – i.e., for each entry in the dataset. This allows any entry (𝑑𝑗(𝑖)) in a dataset 𝐷 to be imputed by its outputs. An example tokenizer which may be used to implement this is described in Zhang et al., “Mixed-Type Tabular Data Synthesis With Score- 5 Based Diffusion in Latent Space”, arXiv preprint arXiv:2310.0956, 2023. Specifically, in the example tokenizer the dataset 𝐷 is represented as a matrix (𝑑𝑗(𝑖))𝑖,𝑗 and each column is converted into an 𝑛-dimensional vector. First one-hot encoding is used to process categorical features. Each datapoint is represented as a vector. Then, a linear transformation is applied for numerical columns which creates an embedding 10 lookup table for columns, where each category is assigned a learnable 𝑛-dimensional vector. Now, each dataset is expressed as the stack of the embeddings of all columns. Having a token for each entry in a dataset, versus having a token for each datapoint, makes the context much larger.

[0067] Regarding positional encoding, where the dataset is represented by a 15 matrix, the machine learning model 512 uses both column and row indices. However, it is desirable that the order of the columns and rows does not matter. Accordingly, in some cases, this may be addressed by using a random noise vector for each row 𝑧 𝑟~𝑁(0, 𝐼) and a random noise vector for each column 𝑧𝑐~𝑁(0, 𝐼). In some cases, either or both of these random noise vectors may be added to the encoding of an entry 20 in the dataset (e.g., (𝑑𝑗(𝑖))𝑖,𝑗).

[0068] The diffusion model 516 and the evaluation module 518 are used to train the machine learning model 512 to generate entire datasets (i.e., 𝑝(𝐷) ) using diffusion modelling. Fundamentally, diffusion modelling works by corrupting training data, and then teaching a model to recover the original data by reversing this corruption process. 25 Specifically, the model is trained to iteratively undo a forward corruption process 𝑞 that corrupts clean data 𝑐 and defines latent variables 𝑧𝑡 for 𝑡 ∈ [0,1] that represent progressively noisy versions of 𝑐. After training, the trained machine learning model can be used to generate data that mimics the training data by simply passing random sampled noise to the trained model. Training a model using diffusion modelling has 30 shown to be an effective method to train a model to generate continuous data but it also has been shown to handle the generation of discrete data.– 17 –

[0069] The process of corrupting an input is called forward diffusion. In the example of FIG. 5, the forward diffusion is performed by the diffusion model 516. Specifically, the diffusion model 516 is configured to receive a synthetic tabular dataset 514 and corrupt the received synthetic tabular dataset to generate a corrupted or 5 diffused dataset 524. The diffusion model 516 may gradually corrupt the received synthetic dataset over a series of steps. In some cases, the diffusion model 516 may be a patent variable model which maps to the latent space using a Markov chain. Specifically, this chain may gradually add Gaussian noise to the received synthetic dataset over a number of steps. In these cases, each step of the forward diffusion 10 process may be defined by equation (1) where 𝜀 = 𝑁(0, 𝐼) and (𝛼𝑡)𝑡∈[0,1] is a noise schedule, monotonically decreasing in 𝑡. This has been shown to work well for continuous data. 𝑧𝑡 = √𝛼𝑡𝑐 + √1 − 𝛼𝑡𝜀 (1)

[0070] In other cases, the diffusion model 516 may implement masked or 15 absorbing state diffusion. In masked diffusion or absorbing state diffusion, during the forward diffusion step the data is gradually transformed or “absorbed” into a specific state. This state could be a form of noise or even an absorbing state (a state that once reached, doesn’t evolve further). This is done using a mask. Specifically, in each step a subset of the elements in the input are masked out. This gradually destroys the 20 information in the input. If enough steps are performed you end up with a fully masked output. This has been shown to work well for discrete data. Any suitable masked or absorbing diffusion method may be used.

[0071] One example method of masked diffusion which may be implemented by the diffusion model 516 is described in Sahoo et al., “Simple and Effective Masked 25 Diffusion Language Models”, arXiv e-prints, pages arXiv-2406, 2024. In this example, scalar discrete random variables with 𝐾 categories are denoted as “one-hot” column vectors and 𝑉 ∈ {𝑐 ∈ {0,1}𝐾: ∑𝐾 𝑖=1 𝑐𝑖 = 1} is defined as the set of all vectors. 𝐶𝑎𝑡(. ; 𝜋) is defined at the categorical distribution over 𝐾 classes with probabilities given by 𝜋 ∈ ∆𝐾, where ∆𝐾 denotes the K-simplex. They start with a forward process that 30 interpolates between clean data 𝑐 and a target distribution 𝐶𝑎𝑡(. ; 𝜋) forming a direct extension of the Gaussian diffusion described above. If 𝑞 defines a sequence of increasingly noisy latent variables 𝑧𝑡 where the time step 𝑡 runs from 𝑡 = 0 (least noisy)– 18 – to 𝑡 = 1 (most noise), then the marginal of 𝑧𝑡 conditional on 𝑐 at time 𝑡 is shown in equation (2) where 𝛼𝑡 ∈ [0,1] is a strictly decreasing function in 𝑡, with 𝛼0 ≈ 1 and 𝛼1 ≈ 0. 𝑞(𝑧𝑡|𝑐) = 𝐶𝑎𝑡(𝑧𝑡; 𝛼𝑡𝑐 + (1 − 𝛼𝑡)𝜋) (2) 5

[0072] This implies transition probabilities 𝑞(𝑧𝑡|𝑧𝑠) = 𝐶𝑎𝑡(𝑧𝑡; 𝛼𝑡|𝑠𝑧𝑠 + (1 − 𝛼𝑡|𝑠)𝜋) where 𝛼𝑡|𝑠 = 𝛼𝑡 / 𝛼𝑠. This indicates that during each diffusion step from 𝑠 → 𝑡, a fraction of the probability mass is transferred to the prior distribution 𝜋. The reverse posterior is given in equation (3). 𝑞(𝑧𝑠|𝑧𝑡, 𝑐) = 𝐶𝑎𝑡 (𝑧𝑠; [𝛼𝑡|𝑠𝑧𝑡+(1𝛼−𝑡𝛼𝑧𝑡𝑡|𝑇𝑠)𝑐1+𝜋(Τ1𝑧−𝑡𝛼]⊙𝑡)[ 𝑧𝛼𝑡𝑇𝑠𝑐𝜋+(1−𝛼𝑠)𝜋]) (3) 10

[0073] In masked diffusion, 𝜋 is set to 𝑚. At each noising step 𝑡, the input 𝑐 transitions to a ‘masked’ state 𝑚 with some probability. If an input transitions to 𝑚 at any time 𝑡’, it will remain in this state for all 𝑡 > 𝑡’: 𝑞(𝑧𝑡|𝑧𝑡′ = 𝑚) = 𝐶𝑎𝑡(𝑧𝑡; 𝑚). At time 𝑇, all inputs are masked with probability 𝐼. The marginal of the forward process is given by equation (4). 15 𝑞(𝑧𝑡|𝑐) = 𝐶𝑎𝑡(𝑧𝑡; 𝛼𝑡𝑐 + (1 − 𝛼𝑡)𝜋) (4)

[0074] From properties of the masking process the posterior 𝑞(𝑧𝑠|𝑧𝑡, 𝑐) simplifies to equation (5). 𝑞(𝑧𝑠|𝑧𝑡, 𝑐) = {𝐶𝑎𝑡 (𝑧𝑠; (1−𝛼𝑠)𝑚1𝐶𝑎𝑡 −+𝛼(𝛼𝑡 (𝑠−𝑧𝛼𝑠;𝑡)𝑧𝑐𝑡)),, 𝑧 𝑧𝑡 𝑡 ≠ = 𝑚 𝑚 (5)

[0075] During training, synthetic tabular datasets 514 of the plurality of synthetic 20 datasets are provided to the diffusion model 516 where they are diffused or corrupted via a diffusion process to generate corresponding diffused datasets 524.

[0076] The machine learning model 512 is then trained to perform the reverse diffusion process on the diffused datasets 524. Specifically, during training, each of the diffused datasets 524 output by the diffusion model 516 is provided to the machine 25 learning model 512 where the machine learning model 512 generates an output therefore (e.g., a predicted dataset 526). The goal is to have the machine learning model 512 generate, in response to processing a diffused dataset 524, the corresponding synthetic tabular dataset 514. This is accomplished by the evaluation module 518 comparing the output of the machine learning model 512 in response to– 19 – a diffused dataset 524 (i.e., the predicted dataset 526) to the corresponding synthetic tabular dataset 514, generating a loss function therefrom and adjusting the parameters 528 of the machine learning model 512 to minimize the loss function. The loss function that is minimized may be any suitable loss function and may be based on the 5 diffusion process used by the diffusion model 516 to corrupt the synthetic datasets. In some cases, training may comprise adjusting the parameters of the machine learning model 512 to minimize the variational upper or lower bound on the negative log likelihood.

[0077] For example, where the diffusion model 516 implements the diffusion 10 process represented by equations (2) and (3), the parameters of the machine learning model 512 (wherein the machine learning model is denoted 𝑝𝜃) may be adjusted so as to maximize the variational lower bound on log-likelihood (ELBO). Specifically, given a number of discretization steps 𝑇, defining 𝑠(𝑖) = (𝑖 − 1) / 𝑇 and 𝑡(𝑖) = 𝑖 / 𝑇 and using 𝐷𝐾𝐿[. ] to denote the Kullback-Leibler divergence, the Negative ELBO can be 15 expressed by equation (6) where ∑𝑇 𝑖=1 𝐷𝐾𝐿[𝑞(𝑧𝑠(𝑖)|𝑧𝑡(𝑖), 𝑐||𝑝𝜃(𝑧𝑠(𝑖)|𝑧𝑡(𝑖))] is the diffusionrelated loss (ℒ𝑑𝑖𝑓𝑓𝑢𝑠𝑖𝑜𝑛). Ε 𝑞[−log 𝑝𝜃 (𝑐|𝑧𝑡(0))] + ∑𝑇 𝑖=1 𝐷𝐾𝐿[𝑞(𝑧𝑠(𝑖)|𝑧𝑡(𝑖), 𝑐||𝑝𝜃(𝑧𝑠(𝑖)|𝑧𝑡(𝑖))] + 𝐷𝐾𝐿[𝑞(𝑧𝑡(𝑇)|𝑐||𝑝𝜃(𝑧𝑡(𝑇))] (6)

[0078] In another example, where the diffusion model 516 implements the 20 diffusion process represented by equations (4) and (5), as per Sahoo, the parameters of the machine learning model 512 may be adjusted to so as to maximize the variational lower bound on log-likelihood (ELBO) which is simplified, with respect to equation (6) to equation (7) where ∑𝑇 𝑖=1 [𝛼𝑡1(−𝑖)− 𝛼𝑡𝛼(𝑖𝑠)(𝑖) 𝑙𝑜𝑔〈𝑐𝜃(𝑧𝑡(𝑖)), 𝑐〉] is the diffusionrelated loss (ℒ𝑑𝑖𝑓𝑓𝑢𝑠𝑖𝑜𝑛). Ε 25 𝑞[−log 𝑝𝜃 (𝑐|𝑧𝑡(0))] + ∑𝑇 𝑖=1 [𝛼𝑡1(−𝑖)𝛼−𝑡𝛼(𝑖𝑠)(𝑖) 𝑙𝑜𝑔〈𝑐𝜃(𝑧𝑡(𝑖)), 𝑐〉] + 𝐷𝐾𝐿[𝑞(𝑧𝑡(𝑇)|𝑐||𝑝𝜃(𝑧𝑡(𝑇))] (7)

[0079] Once the machine learning model 512 has been trained to generate the synthetic datasets (i.e., it has learned how to reverse the diffusion process) then the resulting trained machine learning model, which may be referred to as a diffusion PFN 504, can be used to generate a synthetic dataset from scratch. Specifically, the 30 trained machine learning model (i.e., diffusion PFN 504) may be provided with a– 20 – diffused dataset (e.g., a random dataset) of the desired shape and size. The trained machine learning model (i.e., diffusion PFN 504) then removes the noise in the diffused dataset step by step to create a new dataset of the desired shape and size. Since the diffusion PFN 504 has learned to reverse the process from synthetic 5 datasets from a prior, the new dataset it generates has a distribution that is consistent with the prior it was trained on.

[0080] However, once the machine learning model has been trained to generate datasets in accordance with the prior (i.e., to generate 𝑝(𝐷)) (e.g., once the diffusion PFN 504 has been generated), the trained machine learning model (i.e., the diffusion 10 PFN 504) is modified or conditioned, by the second system 506 (i.e., the conditioning module 530 thereof), to perform imputation. The original trained machine learned model (e.g., diffusion PFN 504) is referred to as unconditioned or an unconditional model and the modified model (e.g., conditioned diffusion PFN 508) may be referred to as a conditioned or conditional model. An unconditional diffusion-trained model 15 models a data distribution 𝑝(𝐴) whereas a conditional diffusion model models a conditional distribution 𝑝(𝐴|𝐵), where the output of the model is guided by additional information 𝐵. In other words, a conditional or conditioned diffusion-trained model takes additional inputs as conditions to control the generation process.

[0081] In the image context, in-painting is something that is commonly 20 performed with a machine learning model that has been trained, using diffusion modelling, to generate images. In-painting is the process of inferring missing parts in an image based on available regions specified by a binary mask (which may be referred to as a segmentation mask). Specifically, a machine learning model that has been trained using a diffusion model to generate images may be able to generate all 25 of the pixels of an image together in response to receiving a set of random inputs that match the desired shape and size of the image. However, such a trained model can also be modified to condition on available image regions to produce high quality inferences (e.g., missing image regions). Specifically, a modified model can be generated that leverages the trained model’s ability to generate images while 30 conditioning on the observed part(s) of an image so that the modified model generates missing parts of the image based on the observed part(s) of the image.

[0082] In some cases, this may be implemented by generating a mask (e.g., a binary mask) which indicates which pixels of the image are to be generated (e.g., the– 21 – pixels that are missing) and generating a modified model to perform conditioned processing – i.e., to use the provided context (the pixels of the image that are not to be generated or modified) and the mask to generate the identified pixels using the reverse diffusion process. 5

[0083] The model may be modified to implement conditional processing in a number of different ways. In some cases, this may be implemented via a conditional denoising autoencoder 𝜖𝜃(𝑧𝑡, 𝑡, 𝑟) which allows the synthesis (i.e., image generation) process to be controlled through inputs 𝑟, such as, but not limited to, the map and the input image. See, for example, Rombach et al., “High-Resolution Image Synthesis 10 with Latent Diffusion Models”, Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 10684-10695, 2022, for a description of conditioning a diffusion model for inpainting.

[0084] Inpainting can be seen as one example or form of imputation (i.e., estimating or inferring missing values in a dataset using algorithms and other 15 datapoints). Thus, methods similar to those used to condition a model trained to generate an image to perform inpainting can be used to condition the trained machine learning model (e.g., diffusion PFN 504) to perform imputation – i.e., generate missing data from a dataset. Specifically, the trained machine learning model may be conditioned down, for example, to receive (i) an input dataset and (ii) a mask indicating 20 which datapoints or elements of the input dataset are to be generated or predicted; and generate or predict the identified elements based on the other elements of the input dataset. Once the trained machine learning model (e.g., diffusion PFN 504) has been conditioned to generate missing data from a dataset (e.g., the conditioned diffusion PFN 508 has been generated) it can be used to perform a variety of tasks. 25

[0085] For example, as shown in FIG. 7A, a conditioned diffusion PFN may be used to perform in-context classification or regression by providing the conditioned diffusion PFN with an input dataset that comprises a plurality of example datapoints (𝑥1, 𝑦1; 𝑥2, 𝑦2, ... 𝑥𝑛−1, 𝑦𝑛−1) and a partial datapoint, with only, for example, a set of features (e.g., 𝑥𝑛), but without the corresponding target (e.g., 𝑦𝑛) – and causing the 30 conditioned diffusion PFN to generate the missing target (e.g., 𝑦𝑛). In this way the conditioned diffusion PFN acts like a PFN (or TabPFN) described above in that it can generate a prediction 𝑦, based on a test dataset (𝑥1, 𝑦1; 𝑥2, 𝑦2, ... 𝑥𝑛−1, 𝑦𝑛−1) and an input 𝑥𝑛. Thus, it can perform in-context learning.– 22 –

[0086] In another example, as shown in FIG. 7B, a conditioned diffusion PFN may be used to perform in-context generation of datapoints by providing the conditioned diffusion PFN with an input dataset that comprises a plurality of datapoints (𝑥1, 𝑦1; 𝑥2, 𝑦2, ... 𝑥𝑛−1, 𝑦𝑛−1) and causing the conditioned diffusion PFN to generate 5 an additional datapoint (e.g., 𝑥𝑛, 𝑦𝑛) that is consistent with the provided datapoints.

[0087] In another example, where a mask is used to identify the elements of an input tabular dataset to generate, then the conditioned diffusion PFN may be used to generate an entire dataset by providing the conditioned diffusion PFN with an input tabular dataset of the desired shape and size with random or noise datapoints, and 10 generating a mask that indicates that all of the datapoints of the dataset are to be generated.

[0088] Furthermore, since a diffusion PFN learns the distribution 𝑝(𝐷), when a conditioned version of the diffusion PFN is used to perform classification or regression, the conditioned diffusion PFN can not only output a final prediction (a number) but it 15 can predict a distribution over those numbers. So, it can perform uncertainty quantification. This is the difference between saying the result is 800 versus saying with 90% confidence the result is between 790 and 810. Accordingly, by learning the entire distribution 𝑝(𝐷), a diffusion PFN can generate density estimates to get uncertainty estimates. 20

[0089] Accordingly, a conditioned diffusion PFN can be used to perform all of the tasks that a PFN or TabPFN can perform (e.g., classification and regression), but it can also be used to perform a number of other tasks such as, but not limited to, generation of datapoints consistent with a dataset and complete dataset generation, making it a much more valuable and flexible model. 25

[0090] In the examples described herein, once the conditioned diffusion PFN 508 has been generated, the conditioned diffusion PFN 508 is used in the third system 510 to perform event prediction.

[0091] In some cases, the conditioned diffusion PFN 508 is used to perform event prediction by providing the conditioned diffusion PFN 508 with: (1) an input 30 tabular dataset 532 comprising a plurality of datapoints, each datapoint comprising one or more elements, and (2) information (e.g., a mask 534) identifying one or more elements of the input tabular dataset 532 that are to be generated or predicted. The– 23 – plurality of datapoints in the input tabular dataset 532 comprise one or more training datapoints and an input datapoint.

[0092] Each training datapoint represents a set of features related to an example (e.g., historical) record and whether the event occurred for that record. In 5 particular, each training datapoint comprises a set of elements that represent a set of features related to the record and one or more elements that represents whether the event occurred. The input datapoint represents the set of features for a new record, but the one or more elements representing whether the event occurred is missing or does not have valid data. The information (e.g., mask 534) then identifies the missing 10 one or more elements (i.e., the one or more elements representing whether the event occurred for the new record) as the element(s) to be generated or predicted.

[0093] In response to receiving the input tabular dataset 532 and the information (e.g., mask 534) identifying the elements of the input tabular dataset 532 to be generated or predicted, the conditioned diffusion PFN 508 generates or predicts 15 the identified elements of the input tabular dataset 532 (e.g., prediction 536) based on the other elements of the input tabular dataset 532 (i.e., the set of features for the input datapoint and the training datapoints). In some cases, the output of the conditioned diffusion PFN 508 may be in the form of an output tabular dataset that comprises (1) all of the elements of the input tabular dataset 532 except the identified elements, and 20 (2) the prediction for the identified elements (e.g., the prediction 536). The prediction 536 can then be extracted from the output tabular dataset.

[0094] What set of features are in the datapoints will vary based on the event. For example, if the event is an individual defaulting on their credit card the set of features may comprise features such as, but not limited to, the age of the individual, 25 the credit score of the individual, the salary or income of the individual, etc. Therefore, in this example, the training datapoints demonstrate the relationship between the set of features and default of a credit card so that the conditioned diffusion PFN 508 can predict whether an individual will default on the credit card based on the set of features for that individual. 30

[0095] In some cases, all or a portion of the data 538 used to generate the input tabular dataset 532 may be received from the source database system 402 or the EDPP 404 via, for example, a data ingestor 540 and stored in the repository 520. In– 24 – some cases, the training datapoints of the input tabular dataset 532 may be received from the source database system 402 of the EDPP 404 and the input datapoint may be received from a user via a user device 420 that is connected over a data communication link 542 to the user interface 418 of the cloud-based computing cluster 5 406. For example, the user may input the elements of the input datapoint via a web browser 546 or some other application that operates on the user device 420.

[0096] In some cases, the prediction 536 generated by the conditioned diffusion PFN 508 is provided to the user device 420. For example, the prediction 536 may be provided to the user via the web browser 546 or some other application that operates 10 on the user device 420.

[0097] The conditioned diffusion PFN 508 may be used to predict the same event for multiple different new records using the same training datapoints. For example, the same training datapoints maybe used to predict (a) whether a first individual will default on a credit card, and (b) whether a second, different, individual 15 will default on a credit card. In these cases, it may be that only the input datapoint differs between predicting (a) and (b).

[0098] In the same way that a PFN can operate on different datasets without having to retrain the PFN, the conditioned diffusion PFN 508 may be used to predict different events, or even the same event using different features, by providing the 20 conditioned diffusion PFN 508 with different training datapoints. For example, one set of training datapoints may be used to predict a first event and different set of training datapoints may be used to predict, a second event; or one set of training datapoints may be used to predict a first event using a first set of features and a different set of training datapoints may be used to predict the same first event using a second set of 25 features. Example Computer

[0099] Reference is now made to FIG. 8 which illustrates a simplified block diagram of an example computer 800. Computer 800 is an example implementation of a computer which may implement the source database system 402, EDPP 404, one 30 or more components of the cloud-based computing cluster 406 of FIGS. 4 and 5 and / or the method 900 of FIG. 9. Computer 800 has at least one processor 802 operatively coupled to at least one memory 804, at least one communications interface 806 (also– 25 – referred to herein as a network interface), and at least one input / output (I / O) device 808.

[00100] The at least one memory 804 includes a volatile memory that stores instructions executed or executable by the processor 802, and input and output data 5 used or generated during execution of the instructions. The memory 804 may also include non-volatile memory used to store input and / or output data – e.g., within a database – along with program code containing executable instructions.

[00101] The processor 802 may transmit or receive data via the communications interface 806 and may also transmit or receive data via any additional input / output 10 device 808 as appropriate.

[00102] In some cases, the processor 802 includes a system of central processing units (CPUs) 810. In other cases, the processor 802 includes a system of one or more CPUs 810 and one or more Graphical Processing Units (GPUs) 812 that are coupled together. For example, any combination of the machine learning models 15 512, diffusion PFNs 504, and conditioned diffusion PFNs 508 described herein may execute neural network computations on CPU and GPU hardware, such as the system of CPUs 810 and GPUs 812 of FIG. 8. Method

[00103] Reference is now made to FIG. 9 which illustrates an example method 20 900 for performing event prediction using a machine learning model. The method 900 beings at block 902 where a machine learning model (e.g., a machine learning model with a transformer architecture) is trained, using a diffusion model and a plurality of synthetic tabular datasets representing a prior, to generate a tabular dataset that has a distribution that is consistent with the prior. In other words, the machine learning 25 model is trained to generate 𝑝(𝐷). As described above, in some cases, the prior may be based on Bayesian Neural Networks (BNNs) and / or Structural Causal Models (SCMs) to model complex feature dependencies and potential causal mechanisms underlying tabular data and the plurality of synthetic tabular datasets may be generated by sampling one or more BNNs and / or one or more SCMs. 30

[00104] As described above, training the machine learning model using a diffusion model and the plurality of synthetic tabular datasets may comprises diffusing each of the plurality of synthetic tabular datasets, passing each of the plurality of– 26 – diffused datasets through the machine learning model to generate a predicted dataset, and adjusting the parameters of the machine learning model based on a comparison of each of the predicted datasets and the corresponding synthetic tabular dataset so that the machine learning model performs the reverse of the diffusion process. In 5 some cases, the adjusting of the parameters of the machine learning model may be iterative. For example, in some cases (1) diffused datasets may be generated for a set of synthetic tabular datasets, (2) the set of synthetic tabular datasets may be processed by the machine learning model to generate corresponding predicted datasets, and (3) the parameters of the machine learning model may be adjusted or 10 updated based on a comparison of the predicted datasets and their corresponding synthetic tabular datasets. Steps (1), (2) and (3) may then be repeated for a set of synthetic tabular datasets (which may be the same set or a different set from the one used in the previous iteration) and so on. This may be repeated until an error or loss metric based on a comparison of the predicted datasets and their corresponding 15 synthetic tabular datasets reaches a certain level. Once the machine learning model has been trained, the method 900 proceeds to block 904.

[00105] At block 904, the trained machine learning model (e.g., the diffusion PFN) is conditioned to perform the generation of all or a portion of a dataset based on additional information. In some cases, the additional information comprises (i) an 20 input tabular dataset comprising one or more datapoints, each datapoint comprising one or more elements, and (ii) information (e.g., a mask) identifying one or more elements of the input tabular dataset that are to be generated, and the conditioned trained machine learning model (e.g., the conditioned diffusion PFN) is configured to generate the identified one or more elements of the input tabular dataset based on the 25 other elements of the input tabular dataset. The identified one or more elements of the input tabular dataset that are to be generated may be any combination of the elements of the input tabular dataset. For example, the identified elements may comprise all the elements of one or more datapoints of the input tabular dataset, or the identified elements may comprise only a subset of the elements of one or more 30 datapoints of the input tabular dataset.

[00106] As described above, once the trained model has been conditioned in this way to generate identified elements of an input tabular dataset based on the other elements of the input tabular dataset it can be used to perform a number of tasks in– 27 – context on a new, unseen dataset, including, but not limited to, classification, regression, data generation, data imputation, density estimation and qualifying model uncertainly and / or detecting anomalous inputs.

[00107] Once the trained machine learning model has been conditioned, the 5 method 900 proceeds to block 906.

[00108] At block 906, the conditioned and trained machine learning model (e.g., conditioned diffusion PFN) is used to perform event prediction. In some cases, the conditioned diffusion PFN is used to perform event prediction by providing the conditioned diffusion PFN with: (1) an input tabular dataset comprising a plurality of 10 datapoints, each datapoint comprising one or more elements, and (2) information (e.g., a mask) identifying one or more elements of the input tabular dataset that are to be generated or predicted. The plurality of datapoints in the input tabular dataset comprise one or more training datapoints and an input datapoint.

[00109] Each training datapoint represents a set of features related to an 15 example (e.g., historical) record and whether the event occurred for that record. In particular, each training datapoint comprises a set of elements that represent a set of features related to the record and one or more elements that represent whether the event occurred. The input datapoint represents the set of features for a new record, but the one or more elements representing whether the event occurred is missing or 20 does not have valid data. The information (e.g., mask) then identifies the missing one or more elements (i.e., the one or more elements representing whether the event occurred for the new record) as the element(s) to be generated or predicted.

[00110] In response to receiving the input dataset and the information (e.g., mask 534) identifying the elements of the input tabular dataset to be generated or predicted, 25 the conditioned diffusion PFN generates or predicts the identified elements of the input tabular dataset (e.g., the prediction) based on the other elements of the input tabular dataset (i.e., the set of features for the input datapoint and the training datapoints). This causes the conditioned diffusion PFN to perform event prediction. Example Use Cases 30

[00111] The conditioned diffusion PFN 508 described above can be used to perform event prediction in a variety of industries including, but not limited to, healthcare, education, manufacturing, energy and utilities, technology and– 28 – cybersecurity, real estate and construction, transportation and logistics, education, hospitality and travel, and financial.

[00112] Specifically, the conditioned diffusion PFN 508 described above may be able to predict whether an entity (e.g., an individual or a business / enterprise) will apply 5 for, or request, a financial product (or a change to a financial product) within a predetermined time in the future (e.g., within the next three months). The prediction generated by the conditioned diffusion PFN 508 may then be used, for example, to determine whether to target the entity (e.g., individual or business / enterprise) for marketing of that financial product. Examples of financial products for which 10 acquisition thereof within a predetermined period in the future may be predicted include, but are not limited to, a credit card (CC), unsecured line of credit (ULOC), unsecured loan (ULOAN), real estate secured lending (RESL), balance transfer, direct investing, wealth, chequing account, personal investment account, savings account, term life policy, overdraft protection, trade account, travel medical insurance policy, 15 balance protection insurance, and a credit limit increase to a financial product such as a ULOC, HELOC and CC. The features that are used to predict whether an entity (e.g., individual or business / enterprise) will obtain or request a financial product may vary based on the financial product.

[00113] The conditioned diffusion PFN 508 described above may also, or 20 alternatively, be used to predict financial product attrition. For example, the conditioned diffusion PFN 508 described above may be used to predict whether a credit card holder will cancel their credit card. Such a prediction generated by the conditioned diffusion PFN 508 may be used by the financial product provider to proactively contact the client to persuade them to keep, or continue with, the product. 25

[00114] The conditioned diffusion PFN 508 described above may also, or alternatively, be used to predict features or events related to a financial product. For example, the conditioned diffusion PFN 508 may be used to predict ULOC utilization, HELOC utilization, a credit limit decrease and / or cash flow management (e.g., estimated money in / money out) for an account. 30

[00115] The conditioned diffusion PFN 508 described above may also, or alternatively, be used to predict whether, an entity (e.g., individual or business / enterprise) will become delinquent with respect to a financial product. The– 29 – prediction generated by the conditioned diffusion PFN 508 can then be used to determine whether the financial product provider is to provide the financial product to the entity. For example, the conditioned diffusion PFN 508 may used to predict whether an entity will be delinquent with respect to one or more of a credit card, ULOC, 5 ULON, and RESL.

[00116] The conditioned diffusion PFN 508 described above may also, or alternatively, be used to predict fraudulent activity with respect to a financial product. In these cases, the prediction generated by the conditioned diffusion PFN 508 may be used to take proactive action such as, for example, re-issuing a new credit card for an 10 CC account which has been identified as being at risk for fraudulent use. For example, the conditioned diffusion PFN 508 may be used to detect the risk or likelihood of an account being fraudulently taken over, the risk or likelihood of mule fraud for an account (i.e., the risk that an entity moves or transfers ill-gotten funds via the account), or the risk or likelihood of an account being used for money laundering. In another 15 example, the conditioned diffusion PFN 508 may also be able to predict that an application for a financial product (e.g., CC) is fraudulent, and / or there has been an account breach (e.g., a fraudulent transaction) on a CC or a debit account.

[00117] Various systems or processes have been described to provide examples of embodiments of the claimed subject matter. No such example embodiment 20 described limits any claim and any claim may cover processes or systems that differ from those described. The claims are not limited to systems or processes having all the features of any one system or process described above or to features common to multiple or all the systems or processes described above. It is possible that a system or process described above is not an embodiment of any exclusive right granted by 25 issuance of this patent application. Any subject matter described above and for which an exclusive right is not granted by issuance of this patent application may be the subject matter of another protective instrument, for example, a continuing patent application, and the applicants, inventors or owners do not intend to abandon, disclaim or dedicate to the public any such subject matter by its disclosure in this document. 30

[00118] For simplicity and clarity of illustration, reference numerals may be repeated among the figures to indicate corresponding or analogous elements. In addition, numerous specific details are set forth to provide a thorough understanding of the subject matter described herein. However, it will be understood by those of– 30 – ordinary skill in the art that the subject matter described herein may be practiced without these specific details. In other instances, well-known methods, procedures, and components have not been described in detail so as not to obscure the subject matter described herein. 5

[00119] The terms “coupled” or “coupling” as used herein can have several different meanings depending in the context in which these terms are used. For example, the terms coupled or coupling can have a mechanical, electrical or communicative connotation. For example, as used herein, the terms coupled or coupling can indicate that two elements or devices are directly connected to one 10 another or connected to one another through one or more intermediate elements or devices via an electrical element, electrical signal, or a mechanical element depending on the particular context. Furthermore, the term “operatively coupled” may be used to indicate that an element or device can electrically, optically, or wirelessly send data to another element or device as well as receive data from another element or device. 15

[00120] As used herein, the wording “and / or” is intended to represent an inclusive-or. That is, “X and / or Y” is intended to mean X or Y or both, for example. As a further example, “X, Y, and / or Z” is intended to mean X or Y or Z or any combination thereof.

[00121] Terms of degree such as "substantially", "about", and "approximately" 20 as used herein mean a reasonable amount of deviation of the modified term such that the result is not significantly changed. These terms of degree may also be construed as including a deviation of the modified term if this deviation would not negate the meaning of the term it modifies.

[00122] Any recitation of numerical ranges by endpoints herein includes all 25 numbers and fractions subsumed within that range (e.g., 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.90, 4, and 5). It is also to be understood that all numbers and fractions thereof are presumed to be modified by the term "about" which means a variation of up to a certain amount of the number to which reference is being made if the result is not significantly changed. 30

[00123] Some elements herein may be identified by a part number, which is composed of a base number followed by an alphabetical or subscript-numerical suffix– 31 – (e.g., 112a, or 112b). All elements with a common base number may be referred to collectively or generically using the base number without a suffix (e.g., 112).

[00124] The systems and methods described herein may be implemented as a combination of hardware or software. In some cases, the systems and methods 5 described herein may be implemented, at least in part, by using one or more computer programs, executing on one or more programmable devices including at least one processing element, and a data storage element (including volatile and non-volatile memory and / or storage elements). These systems may also have at least one input device (e.g., a pushbutton keyboard, mouse, a touchscreen, and the like), and at least 10 one output device (e.g., a display screen, a printer, a wireless radio, and the like) depending on the nature of the device. Further, in some examples, one or more of the systems and methods described herein may be implemented in or as part of a distributed or cloud-based computing system having multiple computing components distributed across a computing network. For example, the distributed or cloud-based 15 computing system may correspond to a private distributed or cloud-based computing cluster that is associated with an organization. Additionally, or alternatively, the distributed or cloud-based computing system be a publicly accessible, distributed or cloud-based computing cluster, such as a computing cluster maintained by Microsoft Azure™, Amazon Web Services™, Google Cloud™, or another third-party provider. 20 In some instances, the distributed computing components of the distributed or cloudbased computing system may be configured to implement one or more parallelized, fault-tolerant distributed computing and analytical processes, such as processes provisioned by an Apache Spark™ distributed, cluster-computing framework or a Databricks™ analytical platform. Further, and in addition to the CPUs described 25 herein, the distributed computing components may also include one or more graphics processing units (GPUs) capable of processing thousands of operations (e.g., vector operations) in a single clock cycle, and additionally, or alternatively, one or more tensor processing units (TPUs) capable of processing hundreds of thousands of operations (e.g., matrix operations) in a single clock cycle. 30

[00125] Some elements that are used to implement at least part of the systems, methods, and devices described herein may be implemented via software that is written in a high-level procedural language such as object-oriented programming language. Accordingly, the program code may be written in any suitable programming– 32 – language such as Python or Java, for example. Alternatively, or in addition thereto, some of these elements implemented via software may be written in assembly language, machine language or firmware as needed. In either case, the language may be a compiled or interpreted language. 5

[00126] At least some of these software programs may be stored on a storage media (e.g., a computer readable medium such as, but not limited to, read-only memory, magnetic disk, optical disc) or a device that is readable by a general or special purpose programmable device. The software program code, when read by the programmable device, configures the programmable device to operate in a new, 10 specific, and predefined manner to perform at least one of the methods described herein.

[00127] Furthermore, at least some of the programs associated with the systems and methods described herein may be capable of being distributed in a computer program product including a computer readable medium that bears computer usable 15 instructions for one or more processors. The medium may be provided in various forms, including non-transitory forms such as, but not limited to, one or more diskettes, compact disks, tapes, chips, and magnetic and electronic storage. Alternatively, the medium may be transitory in nature such as, but not limited to, wire-line transmissions, satellite transmissions, internet transmissions (e.g., downloads), media, digital and 20 analog signals, and the like. The computer usable instructions may also be in various formats, including compiled and non-compiled code.

[00128] While the above description provides examples of one or more processes or systems, it will be appreciated that other processes or systems may be within the scope of the accompanying claims. 25

[00129] To the extent any amendments, characterizations, or other assertions previously made (in this or in any related patent applications or patents, including any parent, sibling, or child) with respect to any art, prior or otherwise, could be construed as a disclaimer of any subject matter supported by the present disclosure of this application, Applicant hereby rescinds and retracts such disclaimer. Applicant also 30 respectfully submits that any prior art previously considered in any related patent applications or patents, including any parent, sibling, or child, may need to be revisited.

Claims

What is claimed is:

1. A system for performing event prediction, the system comprising: a memory, a communication interface, and at least one processor operatively 5 coupled to the memory and the communication interface; the at least one processor configured to: train, using a diffusion model and a plurality of synthetic tabular datasets representing a prior, a machine learning model to generate a tabular dataset 10 that has a distribution that is consistent with the prior; condition the trained machine learning model to: receive (i) an input tabular dataset comprising one or more datapoints, each datapoint comprising one or more elements, and (ii) information identifying one or more elements of the 15 input tabular dataset that are to be predicted, and predict the identified one or more elements of the input tabular dataset based on the other elements of the input tabular dataset; and use the conditioned trained machine learning model to perform event 20 prediction.

2. The system of claim 1, wherein the at least one processor is configured to use the conditioned trained machine learning model to perform event prediction by providing the conditioned trained machine learning model with (i) an event 25 input tabular dataset comprising a plurality of datapoints and (ii) event information identifying one or more elements of the event input tabular dataset to be predicted, wherein the plurality of datapoints comprise one or more training datapoints and an input datapoint, and the event information identifies an element of the input datapoint to be predicted. 30 3. The system of claim 2, wherein each training datapoint comprises a set of elements that represent a set of features of a record and an element that represents an occurrence of the event for the record, and the input datapoint– 34 – comprises a set of elements that represent the set of features for a new record.

4. The system of claim 3, wherein: 5 the input datapoint comprises an element that represents an occurrence of the event for the new record and the element that represents the occurrence of the event for the new record does not comprise valid data; and 10 the event information identifies the element of the input datapoint that represents the occurrence of the event for the new record as an element to be predicted such that the conditioned trained machine learning model predicts the occurrence of the event for the new record. 15 5. The system of claim 1, wherein the event information comprises a mask.

6. The system of claim 1, wherein the machine learning model comprises a transformer architecture. 20 7. The system of claim 6, wherein the machine learning model is configured to tokenize each element of the event input tabular dataset.

8. The system of claim 1, wherein using the conditioned trained machine learning model to perform event prediction comprises using the conditioned 25 trained machine learning model to predict a likelihood of an entity requesting a financial product within a predetermined period of time.

9. The system of claim 8, wherein the financial product is a credit card, unsecured line of credit, unsecured loan, real estate secured lending, 30 investing account, chequing account, personal investment, savings account, term life policy, overdraft protection, trade account, travel medical insurance, or balance protection insurance.– 35 – 10.The system of claim 1, wherein using the conditioned trained machine learning model to perform event prediction comprises using the conditioned trained machine learning model to predict fraudulent activity related to a financial product. 5 11.The system of claim 10, wherein predicting the fraudulent activity related to the financial product comprises predicting an account take over risk, predicting mule account risk, predicting money laundering account risk, predicting application fraud, or predicting breach of an account. 10 12.The system of claim 1, wherein using the conditioned trained machine learning model to perform event prediction comprises using the conditioned trained machine learning model to predict delinquency related to a financial product. 15 13.The system of claim 12, wherein the financial product comprises a credit card, an unsecured line of credit, an unsecured loan, or a real estate secured lending. 20 14.The system of claim 1, wherein the at least one processor is further configured to generate one or more synthetic tabular datasets of the plurality of synthetic tabular datasets by sampling a Bayesian neural network. 15.The system of claim 1, wherein the at least one processor is further 25 configured to generate one or more synthetic tabular datasets of the plurality of synthetic tabular datasets by sampling a structural causal model. 16.The system of claim 1, wherein the at least one processor is configured to train the machine learning model using the diffusion model and the plurality of 30 synthetic tabular datasets by diffusing each of the plurality of synthetic tabular datasets in accordance with a diffusion process to generate a plurality of diffused datasets and adjusting parameters of the machine learning model so that the machine learning model reverses the diffusion process on the plurality of diffused datasets.– 36 – 17.The system of claim 16, wherein adjusting the parameters of the machine learning model so that the machine learning model reverses the diffusion process comprises adjusting the parameters of the machine learning model 5 so that the machine learning model generates, in response to receiving a diffused dataset, the corresponding synthetic tabular dataset. 18.A method for performing event prediction, the method executed in a computing environment comprising at least one processor, a communication 10 interface, and memory, and the method comprising: training, using a diffusion model and a plurality of synthetic tabular datasets representing a prior, a machine learning model to generate a tabular dataset that has a distribution that is consistent with the prior; 15 conditioning the trained machine learning model to: receive (i) an input tabular dataset comprising one or more datapoints, each datapoint comprising one or more elements, and (ii) information identifying one or more elements of the input tabular dataset that are to be predicted, and predict the identified one or 20 more elements of the input tabular dataset based on the other elements of the input tabular dataset; and using the conditioned trained machine learning model to perform event prediction. 25 19.The method of claim 18, wherein using the conditioned trained machine learning model to perform event prediction comprises providing the conditioned trained machine learning model with an event input tabular dataset comprising a plurality of datapoints and event information identifying 30 one or more elements of the event input tabular dataset to be predicted, wherein the plurality of datapoints comprise one or more training datapoints and an input datapoint, and the event information identifies an element of the input datapoint to be predicted.– 37 – 20.A non-transitory computer readable medium storing computer executable instructions which, when executed by at least one computer processor, cause the at least one computer processor to carry out a method for performing event prediction, the method comprising: 5 training, using a diffusion model and a plurality of synthetic tabular datasets representing a prior, a machine learning model to generate a tabular dataset that has a distribution that is consistent with the prior; 10 conditioning the trained machine learning model to: receive (i) an input tabular dataset comprising one or more datapoints, each datapoint comprising one or more elements, and (ii) information identifying one or more elements of the input tabular dataset that are to be predicted, and predict the identified one or more elements of the input tabular dataset based on the other elements of the 15 input tabular dataset; and using the conditioned trained machine learning model to perform event prediction.