A machine learning model for calculating a feature vector that encodes a peripheral distribution

The system encodes input distributions as feature vectors and applies machine learning to efficiently simulate complex stochastic processes, addressing computational complexity and enabling real-time decision-making.

JP2025524518APending Publication Date: 2025-07-30TOWERS WATSON SOFTWARE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024577030
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-30
Filing Date
2023-06-30
Publication Date
2025-07-30

AI Technical Summary

Technical Problem

Existing models of correlated stochastic processes are computationally complex and time-consuming, making them impractical for large datasets or situations with strict time constraints.

Method used

A system that uses processors to encode input marginal distributions as feature vectors, compute dependencies, and apply a trained machine learning model to generate output distribution feature vectors, thereby simplifying the computation process.

Benefits of technology

Enables faster and more efficient simulation of complex stochastic processes using fewer computational resources, allowing for real-time decision-making in applications like energy management and insurance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025524518000001_ABST
    Figure 2025524518000001_ABST
Patent Text Reader

Abstract

An arithmetic system 10 is provided that includes one or more processors 12 configured to receive a plurality of input marginal distributions 20 during a runtime phase. The one or more processors may further be configured to receive one or more dependencies 22 among the plurality of input marginal distributions. The one or more processors may further be configured to calculate respective plural input distribution feature vectors 30 that encode the plurality of input marginal distributions. Based at least in part on the plural input distribution feature vectors and the one or more dependencies, the one or more processors may further be configured to calculate respective one or more first output distribution feature vectors 50 that encode one or more first output marginal distributions in a first trained machine learning model. The one or more processors may further be configured to output the one or more first output distribution feature vectors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to modeling of stochastic processes. In particular, although not particularly limited, the present invention relates to modeling of correlation distributions that are executed more quickly and efficiently on processor hardware than conventional methods of peripheral distribution simulation modeling.

Background Art

[0002] Models of correlated stochastic processes are used in fields such as weather forecasting, power grid management, supply chain management, and insurance. The inputs to these models can be statistical distributions for various quantities, and the outputs of these models can be statistical distributions of output variables that can be used to determine actions to be taken and / or to control one or more devices or systems. For example, in energy-related applications, scheduling of other power generation resources is performed to meet demand based on estimates of renewable energy output such as the availability of wind, solar, and hydroelectric power generation resources. The inputs to the model may be sufficiently correlated, and in order to obtain an accurate output from the model, it is necessary to represent the dependency relationship between the input distributions.

[0003] In order to accurately model the dependency relationship between the input distributions and its influence on the output distribution, many existing models encode non-linear behavior. These non-linear behaviors can model complex interactions in the system that the model simulates. However, models of such correlated stochastic processes can be highly computationally complex and may require a long time to execute. Due to the highly complex computations involved in such models, their use in large datasets or situations with strict time constraints may become impractical. Aspects and embodiments of the present invention have been devised in consideration of the above.

Summary of the Invention

Means for Solving the Problems

[0004] According to one aspect of the present disclosure, during a runtime phase, an arithmetic system is provided that includes one or more processors configured to receive a plurality of input marginal distributions. The one or more processors may further be configured to receive one or more dependencies between two or more of the input marginal distributions. The one or more processors may further be configured to compute respective plural input distribution feature vectors that encode the plurality of input marginal distributions. Based at least in part on the plural input distribution feature vectors and the one or more dependencies, the one or more processors may further be configured to compute one or more first output distribution feature vectors that each encode one or more first output marginal distributions in a first trained machine learning model. The one or more processors may further be configured to output the one or more first output distribution feature vectors.

[0005] This summary of the invention is provided to introduce a simplified set of concepts that are further described in the following detailed description for implementing the invention. This summary of the invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Further, the claimed subject matter is not limited to embodiments that solve some or all of the disadvantages described in any part of this disclosure.

Brief Description of the Drawings

[0006]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 6

Figure 7A

Figure 7B

Figure 7C

Figure 8A

Figure 8B

Figure 9

[0007] FIG. 1 schematically shows an exemplary computing environment 1 in which a correlation probability process may be modeled to address the above problems. The computing environment 1 of FIG. 1 may include a server system 2, a client computing device 3, and a controlled device 4. In the client computing device 3, a plurality of input marginal distributions 20 and one or more copulas 26 may be computed from an input data set 5 and transmitted to the server system 2. In the server system 2, data set preprocessing 6 may be performed to compute one or more dependencies 22 from the copulas 26. The data set preprocessing 6 may further include encoding the input marginal distributions as respective plurality of input distribution feature vectors 30. The dependencies 22 and the input distribution feature vectors 30 may be input to a first trained machine learning model 40 in which a plurality of output distribution feature vectors 50 are computed. As described in more detail below, the first trained machine learning model 40 may be trained to simulate a target model.

[0008] The server system may transmit the output distribution feature vector 50 to the client computing device 3. Based at least in part on the output distribution feature vector 50, the client computing device 3 may generate one or more commands 7 for the controlled device 4 in the control program 58. The controlled device 4 may be, for example, a device included in an energy grid and configured to supply power. As another example, the controlled device 4 may be a financial transaction computing device that may program financial transactions. As a third example, the controlled device 4 may be an inventory management computing device configured to set safety inventory levels. The controlled device 4 may be configured to programmatically execute one or more commands 7 received from the client computing device 3.

[0009] FIG. 2 schematically shows an arithmetic system 10 that may be included in the exemplary arithmetic environment 1 of FIG. 1. For example, the arithmetic system 10 may instantiate the server system 2. The arithmetic system 10 may include one or more processors 12 configured to execute instructions for performing arithmetic processes. The arithmetic system 10 may further include one or more memory devices 14 communicatively connected to the one or more processors 12. The one or more memory devices 14 may include, for example, one or more volatile memory devices and / or one or more non-volatile memory devices.

[0010] The arithmetic system 10 may be instantiated in a single physical computing device or in a plurality of communicatively connected physical computing devices. For example, at least a part of the arithmetic system 10 may be provided as a server computing device installed in a data center. In such an example, the arithmetic system 10 may further include one or more client computing devices configured to communicate with one or more server computing devices via a network.

[0011] Figure 2 shows one or more processors 12 of the computing system 10 during the runtime phase. During the runtime phase, one or more processors 12 can be configured to receive a plurality of input marginal distributions 20. The plurality of input marginal distributions 20 can be probability distributions or frequency distributions. Each of the plurality of input marginal distributions 20 can represent a corresponding plurality of probabilities or frequencies associated with different values of the dependent variable for a plurality of values of the respective independent variable. Each of the plurality of input marginal distributions 20 can have a different independent variable - dependent variable combination. In some examples, the plurality of input marginal distributions 20 can be the distribution of one dependent variable (e.g., total energy demand) across a plurality of different independent variables. Alternatively, the plurality of input marginal distributions 20 can include the distribution of a plurality of different dependent variables (e.g., available energy supplies from different energy sources) across a plurality of independent variables.

[0012] During the runtime phase, one or more processors 12 can further be configured to receive one or more dependencies 22 between two or more of the input marginal distributions 20. In some examples, one or more dependencies 22 can be one or more correlation coefficients between pairs of the input marginal distributions 20. The correlation coefficients can be shown in a correlation matrix 24 in such examples. The correlation matrix 24 can be an N×N matrix, where N is the number of input marginal distributions 20. Alternatively, the dependency 22 can be specified in some other type of data structure.

[0013] In some examples, one or more processors 12 may be configured to compute at least in part one or more dependencies 22 by computing a copula 26 over two or more respective variables associated with two or more of the input marginal distributions 20. The copula 26 may be used as an input on which one or more dependencies 22 are computed. For example, the copula 26 may be a Gaussian copula. In such an example, one or more dependencies 22 may be the correlation coefficients of the copula 26. In other examples, one or more copulas 26 may be some other type of copulas such as a Clayton copula, a Gumbel copula, or a t-copula. In an example where the dependencies 22 are shown in the correlation matrix 24, one or more processors 12 may be further configured to utilize the copula 26 when computing the correlation matrix 24.

[0014] In some examples, one or more processors 12 may be configured to receive one or more copulas 26 as empirical copulas that include empirical data. In such an example, the copula 26 may further include portions that are synthetically computed by one or more processors 12.

[0015] After receiving the plurality of input marginal distributions 20, one or more processors 12 may be further configured to compute respective plural input distribution feature vectors that encode the plurality of input marginal distributions 20. In the example of FIG. 2, the input distribution feature vectors are plural first input distribution feature vectors 30. As described in more detail below, one or more additional sets of input distribution feature vectors may also be computed for other sets of input marginal distributions. By encoding the plurality of input marginal distributions 20 as the first input distribution feature vectors 30, the plurality of input marginal distributions 20 may be compressed so that they can be post-processed using fewer computing resources.

[0016] As shown in the example of FIG. 2, each of the plurality of first input distribution feature vectors 30 may include a corresponding plurality of first input quantile values 32 for the corresponding input marginal distribution 20. In one example, the first input quantile values 32 for the input marginal distribution 20 may be quantile values located at significance levels of 0.500, 0.900, 0.950, 0.990, and 0.995 in the input marginal distribution 20. In other examples, some other set of quantile values may be included in the first input distribution feature vector 30.

[0017] The plurality of first input distribution feature vectors 30 may further include a corresponding plurality of first input moments 34 for each of the plurality of input marginal distributions 20. The plurality of first input moments 34 included in the first input distribution feature vector 30 generated for the input marginal distribution 20 may include, for example, the mean, variance, skewness, kurtosis, and hyperkurtosis of the input marginal distribution 20. The plurality of first input moments 34 may include higher-order moments in some examples, or one or more of the above moments may be excluded. The plurality of first input moments 34 may be represented in the form of raw moments or central moments in the first input distribution feature vector 30.

[0018] By representing the first input distribution feature vector 30 with the plurality of first input quantile values 32 and the plurality of first input moments 34, the shape of the corresponding input marginal distribution 20 can be represented in a form that is more efficiently processed from the perspectives of processor utilization, memory utilization, and time, as compared to using the input marginal distribution 20 in an uncompressed form. As described in more detail below, since the plurality of first input distribution feature vectors 30 are encoded in vector form, they are more suitable for processing by a machine learning model than the raw input marginal distribution 20.

[0019] In some examples, as an alternative to the plurality of first input quantile values 32 and the plurality of first input moments 34, the first input distribution feature vector 30 may instead include a plurality of coordinates of a spline 35 or a plurality of coefficients of a mixture model 36.

[0020] After encoding the plurality of input marginal distributions 20 by the plurality of first input distribution feature vectors 30, one or more processors 12 may be further configured to input the plurality of first input distribution feature vectors 30 and one or more dependencies 22 into a first trained machine learning model 40. The first trained machine learning model 40 may be a deep neural network configured to receive the plurality of first input distribution feature vectors 30 and one or more dependencies 22 at an input layer. In the first trained machine learning model 40, one or more processors 12 may be further configured to calculate one or more first output distribution feature vectors 50 based at least in part on the plurality of input distribution feature vectors 30 and one or more dependencies 22.

[0021] At an output layer of the first trained machine learning model 40, one or more processors 12 may be further configured to output one or more first output distribution feature vectors 50. The one or more first output distribution feature vectors 50 may be an output for one or more additional arithmetic processes 58.

[0022] One or more first output distribution feature vectors 50 may each encode one or more first output marginal distributions 56. In some examples, one or more processors 12 may further be configured to compute one or more estimates of one or more first output marginal distributions 56 from one or more first output distribution feature vectors 50. In such examples, an additional computational process 58 in which one or more processors 12 output one or more first output distribution feature vectors 50 may be a graphical user interface (GUI) generation module in which a visual representation of one or more first output marginal distributions 56 is generated for display to a user on a display device. As another alternative, one or more additional computational processes may be downstream programs configured for a particular use case scenario. For example, in the use case scenarios described below, an energy grid resource allocation program may be configured to receive a predictive distribution of the availability of power sources, or an insurance risk assessment program may be configured to receive a prediction of an overall risk for use in decision-making according to a given program logic. As another example, the additional computational process 58 may be an inventory management system configured to receive a projection of a predicted demand for use in inventory pooling and safety stock setting.

[0023] Each of the one or more first output distribution feature vectors 50 may include a plurality of first output quantile values 52 and a plurality of first output moments 54 for a corresponding first output marginal distribution 56 of the one or more first output marginal distributions 56. In some examples, similar to the plurality of first input distribution feature vectors 30, each of the one or more first output distribution feature vectors 50 may include quantile values located at significance levels of 0.500, 0.900, 0.950, 0.990, and 0.995 in the first output marginal distribution 56. Each of the one or more first output distribution feature vectors 50 may further include the mean, variance, skewness, kurtosis, and hyperkurtosis of the first output marginal distribution 56. In other examples, other sets of first output quantile values 52 and / or first output moments 54 may be included in each first output distribution feature vector 50.

[0024] As described in more detail below, the first trained machine learning model 40 can be a proxy for a more complex prediction model with a higher computational cost. Therefore, in the first trained machine learning model 40, one or more processors 12 can be configured to reproduce the output of the complex model in the form of one or more first output distribution feature vectors 50. Similar to the first input distribution feature vector 30, each of the one or more first output distribution feature vectors 50 may be compressed as compared to the first marginal output distribution 56 to be encoded, thereby enabling more efficient computation of additional outputs related to the first output marginal distribution 56.

[0025] In some examples, as schematically shown in FIG. 3, one or more processors 12 can be further configured to execute a second trained machine learning model 60 configured to receive, as input, one or more first output distribution feature vectors 50. The second trained machine learning model 60 is configured to receive a plurality of distribution feature vectors as input, including a plurality of first output distribution feature vectors 50 and / or one or more distribution feature vectors from several other input sources.

[0026] In the second trained machine learning model 60, based at least in part on the one or more first output distribution feature vectors 50, one or more processors 12 can be further configured to compute one or more second output distribution feature vectors 70. One or more processors 12 can be further configured to output the one or more second output distribution feature vectors 70.

[0027] One or more second output distribution feature vectors 70 may each encode one or more second output marginal distributions 76. As shown in the example of FIG. 3, each of the one or more second output distribution feature vectors 70 may include a plurality of second output quantile values 72 and a plurality of second output moments 74 of the corresponding second output marginal distribution 76. In some examples, the one or more second output distribution feature vectors 70 may have the same format as the plurality of first input distribution feature vectors 30 and the one or more first output distribution feature vectors 50.

[0028] In some examples, as schematically shown in FIG. 4, one or more processors 12 may be configured to execute three or more trained machine learning models configured to receive a distribution feature vector as an input and / or generate a distribution feature vector as an output. In the example of FIG. 4, in a third trained machine learning model 80, one or more processors 12 are further configured to compute one or more third output distribution feature vectors 100 based at least in part on a plurality of third input distribution feature vectors 90. Each of the plurality of third input distribution feature vectors 90 may include a plurality of third input quantile values 92 and a plurality of third input moments 94. Additionally, each of the one or more third output distribution feature vectors 100 may include a plurality of third output quantile values 102 and a plurality of third output moments 104. Each of the one or more third output distribution feature vectors 100 may encode one or more third output marginal distributions 106.

[0029] In the example of FIG. 4, a second trained machine learning model 60 is further configured to receive one or more third output distribution feature vectors 100 as an input when one or more second output distribution feature vectors 70 are computed. Thus, the second trained machine learning model 60 is configured to receive distribution feature vectors from both the first trained machine learning model 40 and the third trained machine learning model 80.

[0030] The inputs and outputs of the machine learning model executed by one or more processors 12 may be configured according to a structure different from that shown in FIGS. 3 and 4. In some examples, the second trained machine learning model 60 may be configured to receive distribution feature vectors from three or more trained machine learning models. The input to the second machine learning model 60 may additionally or alternatively include one or more additional input distribution feature vectors that have not been received from a machine learning model. Additionally or alternatively, the second trained machine learning model 60 may be configured to output a second output distribution feature vector 70 to one or more additional machine learning models. The plurality of machine learning models may generally be arranged in any directed acyclic graph.

[0031] One or more processors 12 may further be configured to receive one or more additional dependency relationships 122 among a plurality of marginal distributions, encoded by one or more first output distribution feature vectors 50 and one or more third output distribution feature vectors 100, as shown in the example of FIG. 4. In some examples, one or more processors 122 may be configured to receive one or more additional copulas 126 on which one or more additional dependency relationships 122 may be computed. For example, one or more additional copulas 126 may include one or more Gaussian copulas. In such examples, one or more additional dependency relationships 122 may be correlation coefficients among the plurality of marginal distributions and may also be included in an additional correlation matrix 124. The second trained machine learning model 60 may further be configured to receive one or more additional dependency relationships 122 as inputs when one or more second output distribution feature vectors 70 are computed. Therefore, the second trained machine learning model 60 may be configured to consider the dependency relationship between one or more first output distribution feature vectors 50 and one or more third output distribution feature vectors 100.

[0032] Figure 5A schematically shows a target model 210 that is trained to be reproduced by a first trained machine learning model 40, as described in more detail below. The target model 210 can be a probabilistic model or a rule-based operation model. The target model 210 is shown during the generation of training data, which is the training data by which the first trained machine learning model 40 is configured to be trained. As shown in Figure 5A, the target model 210 can be configured to receive a plurality of training input marginal distributions 220 and one or more training dependencies 222 between the training input marginal distributions 220. The one or more training dependencies 222 can be a correlation coefficient or an indicator of some other type of dependency. The target model 210 can be configured to output a plurality of training output marginal distributions 256 that can be used to provide ground truth for the first trained machine learning model 40 during training. When the training data is generated, one or more processors 12 can be configured to generate a plurality of sets of training input marginal distributions 220, training dependencies 222, and corresponding training output marginal distributions 256.

[0033] In some examples, at least a portion of the plurality of training input marginal distributions 220 may include empirical data. Additionally or alternatively, at least a portion of the plurality of training input marginal distributions 220 can be synthetically generated. In an example where at least a portion of the plurality of training input marginal distributions 220 is synthetically generated, the plurality of training input marginal distributions 220 can be, for example, a beta distribution, a gamma distribution, a lognormal distribution, a normal distribution, a Pareto distribution, a uniform distribution, a Weibull distribution, a Bernoulli distribution, a binomial distribution, a Poisson distribution, a negative binomial distribution, or some other type of probability distribution.

[0034] According to the example of FIG. 5A, a training data set 212 from which a first trained machine learning model 40 can be trained is schematically shown in FIG. 5B. As shown in the example of FIG. 5B, one or more processors 12 can be configured to synthetically generate a plurality of training input distribution feature vectors 230 and training output distribution feature vectors 250 in a Monte Carlo sample generation module 300. The exemplary Monte Carlo sample generation module shown in FIG. 5B includes an initial sample generation module 310 and a quantum-inspired algorithm 320. Examples of suitable quantum-inspired algorithms include simulated annealing, simulated quantum annealing, population annealing, and parallel tempering. Each of these utilizes a temperature parameter to control the extent to which a Monte Carlo algorithm can accept a random proposed value that has a higher cost than the previous time step, according to an acceptance / rejection policy. In the initial sample generation module 310, one or more processors 12 can be configured to perform rank matching according to a plurality of training input marginal distributions 220 and one or more training dependencies 222 to generate a set of initial distribution samples 312 and a set of initial copula samples 314 that are output to the quantum-inspired algorithm 320.

[0035] In the quantum-inspired algorithm 320, one or more processors 12 may further be configured to iteratively replace the initial distribution samples and the initial copula samples included in the set generated by the initial sample generation module 310 such that the value of the disagreement function 322 of the quantum-inspired algorithm 320 is reduced. By reducing the value of the disagreement function 322, the training output distribution feature vector 250 may be able to more accurately represent the training output marginal distribution 256. The quantum-inspired algorithm 320 may be configured to output a plurality of training input distribution feature vectors 230 and one or more training output distribution feature vectors 250. As shown in the example of FIG. 6, each of the plurality of training input distribution feature vectors 230 may include a plurality of training input quantile values 232 and a plurality of training input moments 234, and each of the plurality of training output distribution feature vectors 250 may include a plurality of training output quantile values 252 and a plurality of training output moments 254.

[0036] Returning to the example of FIG. 5B, when the training data set 212 is generated, one or more processors 12 may further be configured to execute a correlation matrix generator 330 to generate a training dependency 222 in some examples. In the correlation matrix generator 330, one or more processors 12 may be configured to execute a convex optimization solver 332 that generates a training correlation matrix 224 for the training dependency 222 corresponding to the training input distribution feature vector 230.

[0037] FIG. 6 schematically shows a first trained machine learning model 40 during a training phase that occurs before the runtime phase. During the training phase, one or more processors 12 may further be configured to train the first trained machine learning model 40 using a training dataset 212 to reproduce the output of a target model 210. The training phase may include a plurality of training iterations. During each training iteration, the first trained machine learning model 40 may be configured to receive a training input distribution feature vector 230 and a corresponding training dependency 222. In some examples, the training dependency 222 may be encoded in a training correlation matrix 224. The first trained machine learning model 40 may further be configured to compute a candidate output distribution feature vector 260 based at least in part on the training input distribution feature vector 230 and the training dependency 222. The candidate output distribution feature vector 260 may include a plurality of candidate output quantile values 262 and a plurality of candidate output moments 264.

[0038] During each training iteration included in the training phase, one or more processors 12 may be further configured to compute a value of a loss function 270 based at least in part on a training output distribution feature vector 250 and a candidate output distribution feature vector 260. The loss function 270 may be a measure of the distance (e.g., L1 or L2 norm) between the training output distribution feature vector 250 and the candidate output distribution feature vector 260. One or more processors 12 may further compute a loss gradient 272 from the value of the loss function 270 and perform gradient descent in the first trained machine learning model 40 to update one or more parameters of the first trained machine learning model 40 based at least in part on the loss gradient 272. Thus, one or more processors 12 may be configured to train the first trained machine learning model 40 to output a candidate output distribution feature vector 260 that approximately matches the training output distribution feature vector 250, thereby simulating the target model 210.

[0039] Returning now to FIG. 7A, a flowchart of an exemplary method 400 used in the computing system is provided. Method 400 includes steps performed by the computing system during the runtime phase. In step 402, method 400 may include receiving a plurality of input marginal distributions. Each of the plurality of input marginal distributions may be a probability distribution or a frequency distribution of a dependent variable as a function of an independent variable.

[0040] In step 404, method 400 may further include receiving one or more dependencies between two or more of the input marginal distributions. In some examples, one or more dependencies between two or more input marginal distributions may be correlation coefficients between pairs of input marginal distributions. These correlation coefficients may be shown in a correlation matrix.

[0041] When step 404 is executed, in some examples, steps 404A and 404B may be computed. In step 404A, method 400 may further include computing a copula over two or more respective variables. The copula may be, for example, a Gaussian copula, a Clayton copula, a Gumbel copula, a T copula, or some other type of copula. In some examples, at least a portion of the data included in the copula may be empirical data. In step 404B, method 400 may further include computing one or more dependencies based at least in part on the copula. Therefore, in examples where steps 404A and 404B are executed, processing may be performed on the input copula to obtain dependencies between the input marginal distributions.

[0042] In step 406, method 400 may further include computing respective input distribution feature vectors that encode a plurality of input marginal distributions. Each input distribution feature vector may include a plurality of input quantile values and a plurality of input moments for the corresponding input marginal distribution. The input distribution feature vector may include, for example, five input quantile values and five input moments for the input marginal distribution. Alternatively, each input distribution feature vector may include a plurality of coordinates of a spline or a plurality of coefficients of a mixture model. Thus, the input marginal distributions may be compressed so that they can be more efficiently processed by a machine learning model.

[0043] In step 408, method 400 may further include calculating, in a first trained machine learning model, one or more first output distribution feature vectors that respectively encode one or more first output marginal distributions. The one or more first output distribution feature vectors may be calculated based at least in part on a plurality of input distribution feature vectors and one or more dependencies. Each output distribution feature vector may include a plurality of output quantile values and a plurality of output moments for a corresponding first output marginal distribution. For example, the first output distribution feature vector may have the same format as the input distribution feature vector. In step 410, method 400 may further include outputting the one or more first output distribution feature vectors. The one or more first output distribution feature vectors may be an output for additional calculation processes such as a GUI generation module.

[0044] Figures 7B - 7C show additional steps of method 400 that may be performed in an example where the inputs and outputs of a plurality of trained machine learning models are chained together. Figure 7B shows additional steps performed when a second trained machine learning model is executed in a computing system. In step 412, method 400 may further include calculating, in a second trained machine learning model, one or more second output distribution feature vectors based at least in part on the one or more first output distribution feature vectors. The one or more second output distribution feature vectors may respectively encode one or more second output marginal distributions. In step 414, method 400 may further include outputting the one or more second output distribution feature vectors.

[0045] Each of the one or more second output distribution feature vectors may include respective pluralities of output quantile values and respective pluralities of output moments. In some examples, each of the one or more second output distribution feature vectors may have the same format as the plurality of input distribution feature vectors and / or the one or more first output distribution feature vectors.

[0046] FIG. 7C shows additional steps of a method 400 that can be performed in an example where a third trained machine learning model is executed in an arithmetic system according to the example of FIG. 7B. In step 416, the method 400 may further include one or more third output distribution feature vectors in the third trained machine learning model. The one or more third output distribution feature vectors may each encode one or more third output marginal distributions. The one or more third output distribution feature vectors may be calculated based at least in part on a plurality of third input distribution feature vectors.

[0047] Similar to the one or more first output distribution feature vectors, each of the one or more third output distribution feature vectors may include a plurality of output quantile values and a plurality of output moments. Each of the one or more third output distribution feature vectors may have the same format as the one or more first output distribution feature vectors.

[0048] In step 418, the method 400 may further include calculating one or more additional dependencies among a plurality of marginal distributions including one or more first output marginal distributions and one or more third output marginal distributions encoded by the one or more first output distribution feature vectors and the one or more third output distribution feature vectors. The plurality of additional dependencies may, in some examples, be expressed as additional correlation coefficients included in an additional correlation matrix. In some examples, calculating the plurality of additional dependencies may include calculating an additional copula across two or more dependent variables. In such examples, the additional correlation matrix may be generated based at least in part on the additional copula.

[0049] Method 400 may further include being performed when one or more second output distribution feature vectors are computed in step 412. In step 412A, method 400 may further include receiving, in a second trained machine learning model, one or more third output distribution feature vectors as inputs when one or more second output distribution feature vectors are computed. Additionally, method 400 may further include, in step 412B, receiving, in a second trained machine learning model, one or more additional dependencies as inputs when one or more second output distribution feature vectors are computed. Thus, the inputs to the second trained machine learning model may include one or more first output distribution feature vectors, one or more third output distribution feature vectors, and one or more additional dependencies. Distribution feature vectors received from one or more additional machine learning models or from computational processes other than machine learning models may alternatively be received as inputs in the second trained machine learning model. The stage at which the distribution feature vectors are processed may be arranged in any directed acyclic graph.

[0050] FIG. 8A shows a flowchart of an exemplary method 500 that a first trained machine learning model may be trained in a computing system during a training phase that occurs prior to a runtime phase. In step 502, method 500 may include computing a plurality of training output marginal distributions in a target model. The plurality of training output distributions are computed based at least on a plurality of training input marginal distributions and one or more training dependencies among the plurality of training input marginal distributions. For example, one or more training dependencies may be one or more training correlation coefficients included in a training correlation matrix. The target model may be a machine learning model or, alternatively, some other type of computational model that does not utilize machine learning.

[0051] In step 504, method 500 may further include generating a training data set for the first trained machine learning model. The training data set may include a plurality of training input distribution feature vectors that encode a plurality of training input marginal distributions. In some examples, the plurality of training input distribution feature vectors may be computed by compressing the plurality of training input marginal distributions. In other examples, the plurality of training input marginal distributions used as inputs to the target model may be computed from the training input distribution feature vectors. The training data set may further include one or more training dependencies among the training input marginal distributions. Additionally, the training data set may further include a plurality of training output distribution feature vectors that encode a plurality of training output marginal distributions. The plurality of training output distribution feature vectors may be computed from the training output marginal distributions output by the target model.

[0052] In step 506, method 500 may further include training the first trained machine learning model using the training data set to reproduce the output of the target model. Step 506 may include a plurality of training iterations in which the first trained machine learning model generates one or more candidate output distribution feature vectors based at least in part on the plurality of training input distribution feature vectors and the one or more training dependencies. The one or more candidate output distribution feature vectors and the one or more corresponding training output distribution feature vectors may be used as inputs for computing a value of a loss function that indicates a distance between the one or more candidate output distribution feature vectors and the one or more training output distribution feature vectors.

[0053] Figure 8B shows, in some examples, additional steps of method 500 of Figure 8A that may be performed when generating a training dataset for a first trained machine learning model. The steps of Figure 8B may be performed prior to the steps of Figure 8A. In step 508, method 500 may further include synthetically generating a plurality of training input marginal distributions in a Monte Carlo sample generation module. The plurality of training input marginal distributions may be, for example, beta distributions, gamma distributions, lognormal distributions, normal distributions, Pareto distributions, uniform distributions, Weibull distributions, Bernoulli distributions, binomial distributions, Poisson distributions, negative binomial distributions, or other types of probability distributions. In an example where step 508 is performed, method 500 may further include, in step 510, synthetically generating a training copula in the Monte Carlo sample generation module. The Monte Carlo sample generation module may include, for example, an initial sample generation module and a quantum inspired algorithm. In step 512, method 500 may further include computing a training correlation matrix including one or more training dependencies, at least partially based on the training copula. Thus, inputs for the target model may be generated.

[0054] In some examples, alternatively or additionally to the synthetically generated training input marginal distributions generated in the steps shown in Figure 7B, at least a portion of the plurality of training input marginal distributions may include empirical data.

[0055] According to one exemplary use case scenario, the target model can be a model that predicts power supply. In this example, the input marginal distribution can be the probability distribution of weather conditions. For example, the input marginal distribution can show different ranges of probabilities for temperature, rainfall, cloud cover, and wind speed. Since weather conditions are the result of individual weather phenomena, the input marginal distributions may be correlated. The output marginal distribution can be the probability distribution of the amount of power available from different power sources such as solar power, wind power, and hydroelectric power. The different power sources shown in the output marginal distribution may correspond to different power generation methods or to specific physical facilities capable of generating power. In the target model, one or more processors can be configured to generate a predicted distribution of the amount available from different power sources as a function of the input distribution of weather conditions. The one or more processors can further be configured to train a first trained machine learning model for simulating the target model. Thus, in the first trained machine learning model, one or more processors can be configured to generate a predicted distribution of power source availability more quickly and with fewer computational resources compared to the target model.

[0056] In another exemplary use case scenario, the target model can be configured to output the distribution of financial risks in an insurance setting. The target model can be configured to output, for example, the distribution of the total losses of an insurance company. The input marginal distribution in this example can be the probability distribution or frequency distribution of different types of insurance claims. Since large-scale events such as wildfires and floods can lead to a large number of insurance claims, the input marginal distributions can be correlated. Additionally, the target model can exhibit non-linear behavior when modeling systems such as reinsurance. By simulating the target model with the first trained machine learning model, an updated prediction of the overall risk can be generated more quickly, thereby enabling the user to adapt decision-making in real-time or near real-time.

[0057] In a third exemplary use case scenario, the target model may be configured to output a predicted demand distribution for products in a supply chain. The input marginal distribution may be a demand prediction for stock-keeping units (“SKUs”) across different customer regions. Natural factors such as weather phenomena or traditional preferences and anthropogenic factors may result in a correlated distribution of neighboring regions. The target model may be set to output the distribution of the total demand across the region. By simulating the target model with a first trained machine learning model, a prediction of regional product demand can be generated more quickly, thereby enabling decision-making regarding inventory pooling and safety stock setting for large-scale inventory management.

[0058] By using the above-described system and method, the marginal distributions used as inputs to, or generated as outputs of, models of stochastic processes can be compressed by encoding them as distribution feature vectors. This encoding may enable machine learning models that simulate these models to be trained more quickly and with fewer computational resources. Thus, distribution feature vector encoding may enable the training of machine learning models that efficiently simulate complex models of stochastic processes.

[0059] In some embodiments, the methods and processes described herein may be associated with the computing systems of one or more computing devices. In particular, such methods and processes may be implemented as a computer-application program or service, an application-program interface (API), a library, and / or other computer-program products.

[0060] FIG. 9 schematically shows a non - limiting embodiment of an arithmetic system 600 capable of implementing one or more of the above - described methods and processes. The arithmetic system 600 is shown in a simplified form. The arithmetic system 600 may embody the above - described arithmetic system 10 illustrated in FIG. 2. The components of the arithmetic system 600 may be included in one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, game devices, mobile computing devices, mobile communication devices (e.g., smartphones) and / or other arithmetic devices, as well as wearable computing devices such as smart wristwatches and head - mounted augmented reality devices.

[0061] The arithmetic system 600 includes a logical processor 602, a volatile memory 604, and a non - volatile storage device 606. The arithmetic system 600 may optionally include a display subsystem 608, an input subsystem 610, a communication subsystem 612, and / or other components not shown in FIG. 9.

[0062] The logical processor 602 includes one or more physical devices configured to execute instructions. For example, the logical processor may be configured to execute instructions that are part of one or more of an application, program, routine, library, object, component, data structure, or other logical constructs. Such instructions may be implemented to perform a task, implement a data type, transform the state of one or more components, achieve a technical effect, or reach a desired result in other ways.

[0063] A logical processor may include one or more physical processors (hardware) configured to execute software instructions. Additionally, or alternatively, a logical processor may include one or more hardware logic circuits or firmware devices configured to execute logic or firmware instructions implemented in hardware. The processor of the logical processor 602 may be single-core or multi-core, and the instructions executed thereby may be configured for sequential processing, parallel processing, and / or distributed processing. The individual components of the logical processor may optionally be remotely located and / or distributed among two or more separate devices configured for cooperative processing. Aspects of the logical processor may be virtualized and may be executed by remotely accessible network-connected computing devices configured in a cloud-computing configuration. In such cases, it will be understood that these virtualized aspects may be executed on different physical logical processors of different machines.

[0064] The non-volatile memory device 606 includes one or more physical devices configured to hold instructions executable by a logical processor for implementing the methods and processes described herein. When such methods and processes are implemented, the state of the non-volatile memory device 606 may change, for example, to hold different data.

[0065] The non-volatile memory device 606 can include a physical device that is removable and / or embedded. The non-volatile memory device 606 can include optical memory (e.g., CD, DVD, HD-DVD, Blu-Ray Disc, etc.), semiconductor memory (e.g., ROM, EPROM, EEPROM, FLASH memory, etc.), and / or magnetic memory (e.g., hard disk drive, floppy disk drive, tape drive, MRAM, etc.), or other mass storage device technologies. The non-volatile memory device 606 can include devices that are non-volatile, dynamic, static, read-write, read-only, sequential access, location-addressable, file-addressable, and / or content-addressable. It will be appreciated that the non-volatile memory device 606 can be configured to hold instructions even when the power to the non-volatile memory device 606 is turned off.

[0066] The volatile memory 604 can include a physical device that includes random access memory. The volatile memory 604 is typically utilized by the logic processor 602 to temporarily store information during the processing of software instructions. It will be appreciated that the volatile memory 604 typically does not continue to store instructions when the power to the volatile memory 604 is turned off.

[0067] Aspects of the logic processor 602, the volatile memory 604, and the non-volatile memory device 606 may be integrated into one or more hardware-logic components. Such hardware-logic components can include, for example, a field-programmable gate array (FPGA), a program- and application-specific integrated circuit (PASIC / ASIC), a program- and application-specific standard product (PSSP / ASSP), a system-on-chip (SOC), and a complex programmable logic device (CPLD).

[0068] The terms "module", "program", and "engine" may be used to describe aspects of an arithmetic system 600 typically implemented in software by a processor to execute a particular function using a portion of volatile memory, the function including conversion processing that specially configures the processor to execute this function. Thus, a module, program, or engine may be instantiated using a portion of volatile memory 604 via a logic processor 602 that executes instructions held by a non-volatile storage device 606. It will be understood that different modules, programs, and / or engines may be instantiated from the same application, service, code block, object, library, routine, API, function, etc. Similarly, the same module, program, and / or engine may be instantiated by different applications, services, code blocks, objects, routines, APIs, functions, etc. The terms "module", "program", and "engine" may include individual or multiple ones of executable files, data files, libraries, drivers, scripts, database records, etc.

[0069] When included, a display subsystem 608 may be used to represent a visual representation of data held by a non-volatile storage device 606. The visual representation may be in the form of a graphical user interface (GUI). The methods and processes described herein change data held by a non-volatile storage device and thus transform the state of the non-volatile storage device, and accordingly, the state of the display subsystem 608 may also be transformed to visually represent the underlying change in data. The display subsystem 608 may include one or more display devices that utilize virtually any type of technology. Such display devices may be combined with the logic processor 602, volatile memory 604, and / or non-volatile storage device 606 in a shared housing, or such display devices may be peripheral display devices.

[0070] When included, the input subsystem 610 may include, or interface with, one or more user input devices such as a keyboard, mouse, touch screen, or game controller. In some embodiments, the input subsystem may include, or interface with, selected natural user input (NUI) components. Such components may be built-in or peripheral, and the conversion and / or processing of input actions may be performed on-board or off-board. Exemplary NUI components may include a microphone for speech and / or voice recognition; an infrared camera, color camera, stereo camera, and / or depth camera for machine vision and / or gesture recognition; a head tracker, eye tracker, accelerometer, and / or gyroscope for motion detection and / or intent recognition; and / or an electric field sensing component for evaluating brain activity; and / or any other suitable sensor.

[0071] When included, the communication subsystem 612 may be configured to communicatively connect the various computing devices described herein to each other and to other devices. The communication subsystem 612 may include wired and / or wireless communication devices that are compatible with one or more different communication protocols. By way of non-limiting example, the communication subsystem may be configured for communication via a wireless telephone network or a wired or wireless local or wide area network such as an HDMI (registered trademark) connection via Wi-Fi. In some embodiments, the computing system 600 may be capable of transmitting and / or receiving messages to and from other devices via a network such as the Internet through the communication subsystem.

[0072] In the following paragraphs, a plurality of aspects according to the present disclosure will be described. According to one aspect of the present disclosure, during a runtime phase, an arithmetic system is provided that includes one or more processors configured to receive a plurality of input marginal distributions. The one or more processors may further be configured to receive one or more dependencies among two or more of the input marginal distributions. The one or more processors may further be configured to compute respective plural input distribution feature vectors that encode the plurality of input marginal distributions. Based at least in part on the plural input distribution feature vectors and the one or more dependencies, the one or more processors may further be configured to compute one or more first output distribution feature vectors that each encode one or more first output marginal distributions in a first trained machine learning model. The one or more processors may further be configured to output the one or more first output distribution feature vectors. A potential technical advantage of such a configuration is that the first output marginal distribution can be estimated more quickly and efficiently in processor hardware than some other types of marginal distribution simulation models.

[0073] According to this aspect, each input distribution feature vector of the plural input distribution feature vectors may include plural input quantile values and plural input moments, plural coordinates of a spline, or plural coefficients of a mixture model for a corresponding input marginal distribution of the plural input marginal distributions. A potential technical advantage of such a configuration is that different types of features of the input marginal distribution can be represented in the input to the first trained machine learning model.

[0074] According to this aspect, each output distribution feature vector of the one or more first output distribution feature vectors may include plural output quantile values and plural output moments for a corresponding first output marginal distribution of the one or more first output marginal distributions. A potential technical advantage of such a configuration is that the first output distribution feature vector can parameterize the first output marginal distribution.

[0075] According to this aspect, one or more processors may be further configured to respectively compute, in a second trained machine learning model, one or more second output distribution feature vectors that encode one or more second output marginal distributions. The one or more second output distribution feature vectors may be computed based at least in part on the one or more first output distribution feature vectors. One or more processors may be further configured to output the one or more second output distribution feature vectors. A potential technical advantage of such a configuration is that a toolchain of multiple simulation models with high computation costs can be efficiently simulated with multiple trained machine learning models linked together.

[0076] According to this aspect, in a third trained machine learning model, one or more processors may be further configured to respectively compute one or more third output distribution feature vectors that encode one or more third output marginal distributions. The second trained machine learning model may be further configured to receive, as inputs, the one or more third output distribution feature vectors when the one or more second output distribution feature vectors are computed. A potential technical advantage of such a configuration is that a toolchain of multiple simulation models with high computation costs can be efficiently simulated with multiple trained machine learning models arranged in a directed acyclic graph.

[0077] According to this aspect, one or more processors may be further configured to compute one or more additional dependencies among a plurality of marginal distributions including one or more first output marginal distributions and one or more third output marginal distributions encoded by one or more first output distribution feature vectors and one or more third output distribution feature vectors. The second trained machine learning model is further configured to receive one or more additional dependencies as inputs when one or more second output distribution feature vectors are computed. A potential technical advantage of such a configuration is that additional correlations among the distribution feature vectors received as inputs in the trained machine learning model can be reflected in the output of the trained machine learning model.

[0078] According to this aspect, one or more dependencies among two or more input marginal distributions may be the correlation coefficients between pairs of input marginal distributions shown in the correlation matrix. A potential technical advantage of such a configuration is that the dependencies can be expressed in a form in which they can be used as inputs in the first trained machine learning model.

[0079] According to this aspect, one or more processors may be further configured to at least partially compute one or more dependencies among two or more input marginal distributions by computing a copula over two or more respective dependent variables. A potential technical advantage of such a configuration is that the dependencies can be estimated from the input marginal distributions.

[0080] According to this aspect, during the training phase that occurs before the runtime phase, one or more processors are further configured to at least partially train a first trained machine learning model by computing a plurality of training output marginal distributions in a target model based at least on a plurality of training input marginal distributions and one or more training dependencies among the plurality of training input marginal distributions. The training of the first training machine learning model may further include generating a training dataset that includes a plurality of training input distribution feature vectors that encode the plurality of training input marginal distributions. The training dataset may further include one or more training dependencies. The training dataset may further include a plurality of training output distribution feature vectors that encode the plurality of training output marginal distributions. The training of the first training machine learning model may further include using the training dataset to train the first trained machine learning model to reproduce the output of the target model. A potential technical advantage of such a configuration is that the first trained machine learning model can be trained to simulate the behavior of the target model, which may be slower and less efficient than the first trained machine learning model.

[0081] According to this aspect, at least a portion of the plurality of training input marginal distributions may be empirical data. A potential technical advantage of such a configuration is that the first trained machine learning model can be trained to simulate an empirical process.

[0082] According to this aspect, one or more processors may be configured to synthetically generate at least a portion of a plurality of training input marginal distributions in a Monte Carlo sample generation module. A potential technical advantage associated with such a configuration is that the first trained machine learning model can be trained to accurately represent the correlations in the first output distribution feature vector without the user having to collect a large amount of empirical correlation data for use as training data.

[0083] According to another aspect of the present disclosure, a method for use in an arithmetic system is provided. This method may be implemented on a computer. The method may include receiving a plurality of input marginal distributions during a runtime phase. The method may further include receiving one or more dependencies between two or more of the input marginal distributions. The method may further include calculating respective plurality of input distribution feature vectors that encode the plurality of input marginal distributions. Based at least in part on the plurality of input distribution feature vectors and the one or more dependencies, the method may further include calculating, in a first trained machine learning model, one or more first output distribution feature vectors that each encode one or more first output marginal distributions. The method may further include outputting the one or more first output distribution feature vectors. A potential technical advantage associated with such a configuration is that the first output marginal distribution can be estimated more quickly and efficiently than some other types of marginal distribution simulation models.

[0084] According to this aspect, each input distribution feature vector of the plurality of input distribution feature vectors may include a plurality of input quantiles and a plurality of input moments, a plurality of coordinates of a spline, or a plurality of coefficients of a mixture model, for the corresponding input marginal distribution of the plurality of input marginal distributions. A potential technical advantage associated with such a configuration is that different types of feature quantities of the input marginal distribution can be represented in the input to the first trained machine learning model.

[0085] According to this aspect, each output distribution feature vector of one or more first output distribution feature vectors may include a plurality of output quantile values and a plurality of output moments with respect to the corresponding first output marginal distribution of one or more first output marginal distributions. A potential technical advantage of such a configuration is that the first output distribution feature vector can parameterize the first output marginal distribution.

[0086] According to this aspect, the method may further include, in a second trained machine learning model, respectively calculating one or more second output distribution feature vectors that encode one or more second output marginal distributions. The one or more second output distribution feature vectors may be calculated based at least in part on the one or more first output distribution feature vectors. The method may further include outputting the one or more second output distribution feature vectors. A potential technical advantage of such a configuration is that a toolchain of multiple simulation models with high computational costs can be efficiently simulated with multiple trained machine learning models linked together.

[0087] According to this aspect, the method may further include, in a third trained machine learning model, respectively calculating one or more third output distribution feature vectors that encode one or more third output marginal distributions. The method may further include calculating one or more additional dependency relationships among a plurality of marginal distributions including one or more first output marginal distributions and one or more third output marginal distributions encoded by the one or more first output distribution feature vectors and the one or more third output distribution feature vectors. When the one or more second output distribution feature vectors are calculated in the second trained machine learning model, the method may further include receiving the one or more third output distribution feature vectors and the one or more additional dependency relationships as inputs. A potential technical advantage of such a configuration is that a toolchain of multiple simulation models with high computational costs can be efficiently simulated with multiple trained machine learning models arranged in a directed acyclic graph.

[0088] According to this aspect, one or more dependencies between two or more input marginal distributions can be the correlation coefficients between the input marginal distributions that are pairs shown in the correlation matrix. A potential technical advantage of such a configuration is that the dependencies can be expressed in a form in which they can be used as inputs in the first trained machine learning model.

[0089] According to this aspect, the method may further include at least partially calculating one or more dependencies between two or more input marginal distributions by calculating a copula over two or more respective dependent variables. A potential technical advantage of such a configuration is that the dependencies can be estimated from the input marginal distributions.

[0090] According to this aspect, the method may further include training a first trained machine learning model during a training phase that occurs before the runtime phase. Training the first trained machine learning model may include computing a plurality of training output marginal distributions based at least on a plurality of training input marginal distributions and one or more training dependencies among the plurality of training input marginal distributions in a target model. Training the first trained machine learning model may further include generating a training dataset that includes a plurality of training input distribution feature vectors that encode the plurality of training input marginal distributions. The training dataset may further include one or more training dependencies. The training dataset may further include a plurality of training output distribution feature vectors that encode the plurality of training output marginal distributions. Training the first trained machine learning model may further include using the training dataset to train the first trained machine learning model to reproduce the output of the target model. A potential technical advantage of such a configuration is that the first trained machine learning model can be trained to simulate the behavior of a target model that may be slower and less efficient than the first trained machine learning model.

[0091] According to another aspect of the present disclosure, during a runtime phase, an arithmetic system is provided that includes one or more processors configured to receive a plurality of input marginal distributions. The one or more processors may further be configured to receive a correlation matrix indicating one or more correlation coefficients for each pair of the input marginal distributions. The one or more processors may further be configured to compute respective plural input distribution feature vectors that encode the plurality of input marginal distributions. Each of the plural input distribution feature vectors may include respective plural input quantile values and respective plural input moments. Based at least in part on the plural input distribution feature vectors and the plural correlation coefficients, the one or more processors may further be configured to compute, in a trained machine learning model, respective one or more first output distribution feature vectors that encode one or more output marginal distributions. Each of the one or more output distribution feature vectors may include respective plural output quantile values and respective plural output moments. The one or more processors may further be configured to output the one or more first output distribution feature vectors. A potential technical advantage of such a configuration is that the first output marginal distribution can be estimated more quickly and efficiently than some other types of marginal distribution simulation models.

[0092] Features described in the context of separate aspects and embodiments of the present invention may be used together and / or interchangeably. Similarly, where features are described in the context of a single embodiment for the sake of brevity, these may also be provided individually or in any suitable sub-combination. Features described in relation to a system may have corresponding features definable in relation to a method, and vice versa, and these embodiments are specifically envisaged.

[0093] As used herein, "and / or" is defined as an inclusive logical disjunction or (∨) as defined by the following truth table.

[0094]

Table 1

[0095] The configurations and / or approaches described in this specification are of an exemplary nature, and it will be understood that these specific embodiments or examples should not be construed in a limiting sense because numerous variations are possible. The specific routines or methods described in this specification may represent one or more of any number of processing strategies. Accordingly, the various acts illustrated and / or described may be performed in the order illustrated and / or described, in other orders, in parallel, or may be omitted. Similarly, the order of the above processes may be changed.

[0096] The subject matter of the present disclosure includes all novel and non-obvious combinations and sub-combinations of any and all of the various processes, systems and configurations, and other features, functions, acts and / or characteristics disclosed herein, as well as their equivalents.

Claims

**Claim 1** An arithmetic system comprising: During the runtime phase: Receiving a plurality of input marginal distributions; Receiving one or more dependency relationships among two or more of the input marginal distributions; Calculating respective plurality of input distribution feature vectors for encoding the plurality of input marginal distributions; Based at least in part on the plurality of input distribution feature vectors and the one or more dependency relationships, calculating in a first trained machine learning model, one or more first output distribution feature vectors for encoding one or more first output marginal distributions respectively; and Outputting the one or more first output distribution feature vectors An arithmetic system comprising one or more processors configured as such. **Claim 2** Each input distribution feature vector of the plurality of input distribution feature vectors is for a corresponding input marginal distribution of the plurality of input marginal distributions: A plurality of input quantile values and a plurality of input moments; A plurality of coordinates of a spline; or A plurality of coefficients of a mixture model The arithmetic system according to claim 1, comprising the above. **Claim 3** Each output distribution feature vector of the one or more first output distribution feature vectors includes a plurality of output quantile values and a plurality of output moments for a corresponding first output marginal distribution of the one or more first output marginal distributions. The arithmetic system according to claim 2. **Claim 4** The one or more processors are further configured to: Based at least in part on the one or more first output distribution feature vectors, calculating in a second trained machine learning model, one or more second output distribution feature vectors for encoding one or more second output marginal distributions respectively; and Outputting the one or more second output distribution feature vectors The arithmetic system according to claim 1, configured as such. **Claim 5** In a third trained machine learning model, the one or more processors are further configured to calculate respectively one or more third output distribution feature vectors for encoding one or more third output marginal distributions; and The second trained machine learning model is further configured to receive the one or more third output distribution feature vectors as an input when the one or more second output distribution feature vectors are calculated. The arithmetic system according to claim 4. **Claim 6** The one or more processors are further configured to compute one or more additional dependencies among a plurality of marginal distributions including the one or more first output marginal distributions and the one or more third output marginal distributions encoded by the one or more first output distribution feature vectors and the one or more third output distribution feature vectors; The computing system according to claim 5, wherein the second trained machine learning model is further configured to receive the one or more additional dependencies as inputs when the one or more second output distribution feature vectors are computed. **Claim 7** The computing system according to claim 1, wherein the one or more dependencies among the two or more input marginal distributions are correlation coefficients among the input marginal distributions that are pairs shown in a correlation matrix. **Claim 8** The computing system according to claim 1, wherein the one or more processors are further configured to at least partially compute the one or more dependencies among the two or more input marginal distributions by computing a copula over two or more respective dependent variables. **Claim 9** During a training phase performed before the runtime phase, the one or more processors further: compute, in a target model, a plurality of training output marginal distributions based at least on a plurality of training input marginal distributions and one or more training dependencies among the plurality of training input marginal distributions; generate a training dataset comprising: a plurality of training input distribution feature vectors encoding the plurality of training input marginal distributions; the one or more training dependencies; and a plurality of training output distribution feature vectors encoding the plurality of training output marginal distributions ; and train the first trained machine learning model using the training dataset to reproduce the output of the target model so as to at least partially train the first trained machine learning model, the computing system according to claim 1. **Claim 10** The computing system according to claim 9, wherein at least a portion of the plurality of training input marginal distributions includes empirical data. **Claim 11** 10. The computing system of claim 9, wherein the one or more processors are configured to synthetically generate at least a portion of the plurality of training input marginal distributions in a Monte Carlo sample generation module.

12. 1. A method for use in a computing system, comprising, during a runtime phase: Accepting multiple input marginal distributions; receiving one or more dependencies between two or more of the input marginal distributions; computing respective plurality of input distribution feature vectors encoding the plurality of input marginal distributions; computing one or more first output distribution feature vectors, each encoding one or more first output marginal distributions, in a first trained machine learning model based at least in part on the plurality of input distribution feature vectors and the one or more dependencies; and outputting the one or more first output distribution feature vectors; A method comprising:

13. Each input distribution feature vector of the plurality of input marginal distributions is a function of: Multiple input quantiles and multiple input moments; multiple coordinates of the spline; or Multiple coefficients in mixed models 13. The method of claim 12, comprising:

14. 14. The method of claim 13, wherein each output distribution feature vector of the one or more first output distribution feature vectors comprises a plurality of output quantiles and a plurality of output moments for a corresponding first output marginal distribution of the one or more first output marginal distributions.

15. Computing one or more second output distribution feature vectors encoding one or more second output marginal distributions, respectively, in a second trained machine learning model based at least in part on the one or more first output distribution feature vectors; and outputting the one or more second output distribution feature vectors. The method of claim 12 further comprising:

16. Computing, in the third trained machine learning model, one or more third output distribution feature vectors encoding the one or more third output marginal distributions, respectively; computing one or more additional dependencies between a plurality of marginal distributions, including the one or more first output marginal distributions and the one or more third output marginal distributions, encoded by the one or more first output distribution feature vectors and the one or more third output distribution feature vectors; and In the second trained machine learning model, when the one or more second output distribution feature vectors are calculated, receiving the one or more third output distribution feature vectors and the one or more additional dependencies as inputs The method according to claim 15, further comprising: **Claim 17** The method according to claim 12, wherein the one or more dependencies between the two or more input marginal distributions are correlation coefficients between the input marginal distributions that are pairs shown in a correlation matrix. **Claim 18** The method according to claim 12, further comprising at least partially calculating the one or more dependencies between the two or more input marginal distributions by calculating a copula over two or more respective dependent variables. **Claim 19** During a training phase performed before the runtime phase: In a target model, calculating a plurality of training output marginal distributions based at least on a plurality of training input marginal distributions and one or more training dependencies between the plurality of training input marginal distributions; A training dataset comprising: A plurality of training input distribution feature vectors encoding the plurality of training input marginal distributions; The one or more training dependencies; and A plurality of training output distribution feature vectors encoding the plurality of training output marginal distributions Generating a training dataset; and Using the training dataset to train the first trained machine learning model to reproduce the output of the target model The method according to claim 12, further comprising at least partially training the first trained machine learning model thereby. **Claim 20** An arithmetic system comprising: During a runtime phase: Receiving a plurality of input marginal distributions; Receiving a correlation matrix indicating one or more correlation coefficients for each pair of the input marginal distributions; Calculating respective plurality of input distribution feature vectors encoding the plurality of input marginal distributions, wherein each of the plurality of input distribution feature vectors includes respective plurality of input quantile values and respective plurality of input moments; Based at least in part on the plurality of input distribution feature vectors and the plurality of correlation coefficients, in a trained machine learning model, respectively calculate one or more first output distribution feature vectors that encode one or more output marginal distributions, where each of the one or more output distribution feature vectors includes a plurality of respective output quantile values and a plurality of respective output moments; and output the one or more first output distribution feature vectors One or more processors configured to include an arithmetic system.