Resampling Simulation Results of Correlation Events

The arithmetic system optimizes Monte Carlo simulations by selecting a representative sequence using a quantum-based routine, addressing computational inefficiencies and maintaining accuracy with fewer simulations.

JP2025523578APending Publication Date: 2025-07-23TOWERS WATSON SOFTWARE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024577057
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-30
Filing Date
2023-06-30
Publication Date
2025-07-23

AI Technical Summary

Technical Problem

Existing Monte Carlo-based models for correlated stochastic processes require significant computational resources and time due to the need for a large number of simulations, while reducing the number of simulations compromises accuracy.

Method used

An arithmetic system and method that uses a quantum-based optimization routine to select a representative sequence of simulations, minimizing the disagreement metric within an optimization threshold, thereby reducing the number of simulations required while maintaining accuracy.

Benefits of technology

Achieves the same accuracy as conventional methods using 1/10 to 1/1000 of the simulations, enhancing predictive ability and reducing processing time and memory usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025523578000001_ABST
    Figure 2025523578000001_ABST
Patent Text Reader

Abstract

An arithmetic system 702 is provided that includes a processor configured to receive a simulation sample 710 including a plurality of simulations 711 for a plurality of correlation probability variables. The processor may at least partially generate a surrogate cumulative distribution model 714 by estimating a plurality of surrogate model parameters 718. Based at least in part on the surrogate cumulative distribution model, the processor may select one or more subsets of the plurality of simulations. In each of one or more resampling iterations 748, the processor may calculate a disagreement score until it is determined that the sum of the disagreement scores for each of the subsets meets an optimization threshold. Based at least in part on this sum, the processor may sample one or more resampled simulations 730. The processor may replace one or more simulations included in one or more subsets with one or more resampled simulations. The processor may output one or more subsets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to modeling of stochastic processes. In particular, without limitation, the present invention relates to modeling of correlation distributions. The present invention can be used to generate one or more compressed subsets of simulation samples, whereby when one or more subsets are used as input, the processor hardware can execute the Monte Carlo algorithm more efficiently.

Background Art

[0002] Models of correlated stochastic processes are used in fields such as weather forecasting, power grid management, finance, supply chain management, and insurance. In some cases, these models are based on Monte Carlo simulations. For example, in energy-related applications, scheduling of other power generation resources is performed to meet demand based on estimated values of renewable energy output such as the availability of wind, solar, and hydroelectric power generation resources. However, the probability variables representing the availability of renewable energy output may be correlated with meteorological events. Therefore, the total power generation can be estimated using the Monte Carlo method.

[0003] As another example, in the estimation of insurance claims over time, decision-making by insurance companies is performed to maintain capital. Claims from different claimants may be correlated due to underlying events (e.g., severe weather). The Monte Carlo method can be used to estimate the total losses due to such events.

[0004] As yet another example, in retail applications, companies set the inventory levels of various products at regional distribution centers to meet the estimated customer demand. However, customer demand is correlated between neighboring regions due to social influences. For example, when the Chicago baseball team wins the championship game, there may be a demand for baseball memorabilia above the average in other cities and towns near Chicago. The Monte Carlo method can be used to estimate the approximate sales volume of baseball memorabilia.

[0005] Rather than focusing on so-called tail statistics of relatively infrequent events (e.g., financial losses due to a flood occurring multiple times in a single season once every 500 years), many models aim to predict the entire distribution of a variable. Such models can quantify the components of a risk profile that are the main factors of a risk indicator, the focus of capital allocation, or the selected results near the central region of the distribution as well as near its tail. For example, in insurance-related applications, this distribution can be used to predict the average weather-related financial losses modeled near the central region of the distribution. The distribution can also be used to allocate risk capital based on the likelihood of financial losses exceeding the average, which can be modeled at the tail end of the distribution.

[0006] However, there are technical challenges in developing such complex models. For example, running a sufficient number of Monte Carlo simulations to develop an accurate model can require a large amount of processor time and memory usage. Reducing the number of simulations can accelerate model generation, but there is also a possibility that the accuracy will be sacrificed compared to a model based on more simulations. The aspects and embodiments of the present invention have been devised in consideration of the above. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0007] To address these problems, an arithmetic system and method for outputting a plurality of representative simulation results based on stratification are provided. In one exemplary aspect, for a plurality of correlation variables, each simulation includes a first predetermined number of simulations (from Monte Carlo simulation samples) that include a plurality of initial simulation results for a plurality of variables, one or more target statistics, a second predetermined number of layers for each variable and each target statistic, and a processor configured to receive a cumulative distribution function for each variable and each target statistic. For each of each variable and one or more target statistics, the unit interval of the cumulative distribution function is divided into the second predetermined number of layers, and the base of the cumulative distribution function is divided into a plurality of bins such that each bin of the plurality of bins corresponds to one layer. An initial disagreement score is determined based on the amount of values in each bin, the first predetermined number of simulations, and the second predetermined number of layers for the variable. An initial sum of the initial disagreement scores is determined. At least one of the plurality of initial simulations is removed based on a determination that the initial sum of the initial disagreement scores is not within an optimization threshold. At least one other simulation is added to the remaining one or more initial simulations. For each variable and for each of the one or more target statistics, the amount of values in one or more bins corresponding to at least one of the plurality of initial simulation results and the amount of values in one or more bins corresponding to at least one other simulation result are used to generate an updated disagreement score. The arithmetic system further determines an updated sum of the updated disagreement scores and is configured to output a plurality of representative simulations representing the cumulative distribution function across the layers based on the updated sum of the updated disagreement scores.

[0008] According to another aspect of the present disclosure, there is provided an arithmetic system comprising a processor configured to receive a simulation sample including a plurality of simulations for a plurality of correlation probability variables. Each simulation may include a plurality of simulation results. The processor may further be configured to at least partially generate a surrogate cumulative distribution model by estimating a plurality of surrogate model parameters based at least in part on the plurality of simulation results. Based at least in part on the surrogate cumulative distribution model having the surrogate model parameters, the processor may further be configured to select one or more subsets of the plurality of simulations. In each of one or more resampling iterations, the processor may further be configured to calculate the sum of one or more disagreement scores for one or more of the subsets until it is determined that the sum of the one or more disagreement scores satisfies an optimization threshold. Based at least in part on the sum of the one or more disagreement scores, the processor may further be configured to sample one or more resampled simulations for the plurality of correlation probability variables from a plurality of simulations included in the simulation sample and not yet included in one or more of the subsets. The processor may further be configured to replace one or more simulations included in one or more of the subsets with one or more of the resampled simulations. The processor may further be configured to output the simulations included in one or more of the subsets after performing one or more resampling iterations.

[0009] In another exemplary aspect, an arithmetic system is provided that includes a processor configured to receive a plurality of simulations, discrete distribution functions, and one or more cumulative distribution models. Each simulation includes a plurality of simulation results. One or more conditional cumulative distribution models are generated based at least in part on the discrete distribution functions and the one or more cumulative distribution models. The ranges of the one or more cumulative distribution models and the one or more conditional cumulative distribution models are stratified into a plurality of layers. A sum of disagreement scores is computed for the plurality of simulations based at least in part on the cumulative distribution models and the conditional cumulative distribution models. One or more resampling iterations are performed until it is determined that the sum of each of the one or more disagreement scores meets an optimization threshold. In each of the one or more resampling iterations, one or more resampled simulations are generated based at least in part on the one or more cumulative distribution models. An updated sum of disagreement scores is generated for the plurality of simulations in which one or more simulations are replaced with the one or more resampled simulations. One or more simulations are replaced with the one or more resampled simulations based on a policy. The plurality of simulations are output after performing the one or more resampling iterations.

[0010] This Summary of the Invention is provided to introduce a simplified set of concepts that are further described in the Detailed Description of the Invention below. This Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Further, the claimed subject matter is not limited to embodiments that solve any or all of the disadvantages noted in any part of this disclosure.

Brief Description of the Drawings

[0011]

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 6C

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11A

Figure 11B

Figure 11C

Figure 11D

Figure 12

Figure 13

Figure 14

Figure 15A

Figure 15B

Figure 16

[0012] Section 1. Introduction

[0013] As described above, there are technical challenges in the development and use of Monte Carlo-based models related to the correlation probability process. Generally, by performing a larger number of Monte Carlo simulations, a more uniform sampling can be obtained across the entire distribution compared to what can be achieved with a smaller number of simulations. This may result in a more accurate model. However, performing a larger number of Monte Carlo simulations requires both time and processing resources for execution and can increase the computational cost. On the other hand, reducing the number of simulations may lead to a decrease in the accuracy of the model. There is an opportunity to address these conflicting technical challenges and improve accuracy while reducing the processing time for performance.

[0014] To address these issues, a system and method for modeling correlation variables using a quantum-based optimization routine to achieve accurate and efficient performance are disclosed herein. By using the output of the quantum-based optimization routine, it is possible to construct a model based on a relatively small number of Monte Carlo simulations while maintaining the same accuracy as a model based on a larger number of simulations. With the system and method disclosed herein, it may be possible to achieve the same accuracy using only 1 / 10 to 1 / 1000 of the number of simulations compared to conventional approaches. Briefly, the system and method disclosed herein minimize the disagreement metric in the sequence of simulations within an optimization threshold. This reduces the sampling error and enhances the predictive ability of the simulations. As a result, it is possible to obtain a complete distribution with a significantly smaller number of simulations, up to several orders of magnitude less.

[0015] Briefly described, FIG. 1A shows an arithmetic system 102 configured to implement a modified Iman Conover algorithm to accurately and efficiently simulate multivariate data having a known distribution. The arithmetic system 102 receives, as input, a plurality of marginal distributions representing input data, correlates the marginal distributions using a copula, generates a user-specified number of simulation results, and generates a low-discrepancy sequence of dependent variables using a technique for estimating the aggregate distribution of the dependent variables to generate low-discrepancy samples. The reasons for this efficiency are generally related to three different techniques, as shown as (1)-(3) in FIG. 1A: (1) stratification of an event-driven model for generating input simulations, (2) generation of a stratification-based discrepancy metric for selecting simulation data with reduced discrepancies, and (3) generation of a low-discrepancy sequence while capturing the cumulative distribution function of a dependent variable, such as an aggregate value, a minimum value, a maximum value, or other target statistic for which an operation using the dependent variable is desired. Each of these contributions is described in detail in the following sections.

[0016] FIG. 1A shows an example of an arithmetic system 102 for modeling correlated variables. In some examples, the arithmetic system 102 comprises a server arithmetic system (e.g., a cloud-based server or a plurality of distributed cloud servers). In other examples, the arithmetic system 102 may include other suitable types of arithmetic devices, such as a desktop computer, a laptop computer, etc. Additional aspects of the arithmetic system 102 are described in more detail below with reference to FIG. 16.

[0017] The computing system 102 is configured to receive data representing a plurality of correlation variables as an input peripheral distribution. Some examples of suitable variables include, but are not limited to, the sale of one or more predefined stock keeping units (SKUs) in a geographic region, insurance losses in one or more business fields, and the electricity generated at one or more power plants. These three specific examples are shown to aid in the understanding of the present invention; however, it will be recognized that the data types are not limited thereto. In the example shown in FIG. 1A, the computing system 102 is configured to receive a first input data 104, a second input data 106, and a third input data 108. In some examples, the first input data 104 includes empirical data 110, the second input data 106 includes a first lognormal distribution 112, and the third input data 108 includes an aggregate value 114 of an event-driven model 125. The first input data 104, the second input data 106, and the third input data 108 represent independent correlation variables used by the computing system 102 to model the output data 116.

[0018] In some examples, the event-driven model 125 includes a frequency-severity model 126. The event-driven model 125 can include uncertainty in the number of events that occur and uncertainty in the values associated with each event. Dependent variables such as the sum of the values associated with each event (e.g., an aggregated risk model) are often of interest. The values associated with each event can additionally or alternatively be considered in the development of models such as insurance claim amount models, levee breach models, power surge models, stock price over-time models, operational risk models, and the like. The frequency-severity model 126 associates the frequency with which an event occurs with the severity of the value associated with the event. As an illustrative example, the frequency-severity model can be for weather events. For example, the frequency-severity model 126 can model a first event 128, such as an empirical measurement of the total output of wind power over a period (e.g., a calendar year), based on a discrete number of second events 130, such as storm events that occur during that year. As described in more detail below, in some examples, the systems and methods disclosed herein are configured to apply policies to stratify an event-driven model 125, such as a frequency-severity model 126, into a plurality of layers that are mutually exclusive and collectively exhaustive subgroups of data that form non-overlapping partitions of the input data. Stratifying the input data in this manner reduces sampling error compared to using the same number of Monte Carlo simulation random samples that are not fully stratified. It will be appreciated that the conditional cumulative distribution of each of the plurality of dependent variables can be conditioned on a predetermined amount of correlated events occurring. For example, in the wind power example above, Monte Carlo simulations that produce inadequately stratified simulation results are considered because, for example, if fewer than three storms occur in a year, they will include little or no power generation examples over several years.A model generated based on such simulation results may relatively accurately reflect the correlations in other layers in the frequency-severity model 126 or the correlations related to the aggregated values of all events (for example, the average power generated by all typhoons occurring in any year), but this model is likely to inadequately reflect the correlations related to one or more target statistics on the condition that less than three storms occur. In addition, the computing system 102 can achieve a more uniform sampling across the whole distribution with fewer simulations. As a result, a large number of randomly selected simulations may be required to achieve the same accuracy, thus reducing the usage of the processor and memory. The layering of the event-driven model will be described in more detail below with reference to FIGS. 12 to 15.

[0019] The computing system 102 is further configured to receive a copula 118. The copula 118 describes the correlations among the first input data 104, the second input data 106, and the third input data 108. The copula 118 may have various forms of distributions 120. For example, the copula may be a Gaussian copula having a Gaussian distribution, or an Archimedean copula such as a Clayton copula or a Gumbel copula. In addition, the copula 118 may have a related correlation matrix 122.

[0020] Output data 116 represents a dependent variable that is a function of the first input data 104, the second input data 106, and the third input data 108. In some examples, output data 116 represents an aggregated value of the first input data 104, the second input data 106, and the third input data 108. Some examples of suitable aggregated values include, but are not limited to, aggregated sales, aggregated insurance losses, or aggregated power generation. For example, the first input data 104 may represent the amount of power generated by a solar array, the second input data 106 may represent the amount of power generated by a wind power plant with relatively stable low-intensity winds, and the third input data 108 may represent the amount of power generated by higher-intensity wind events that occur less frequently than those reflected in the second input data 106. In these examples, output data 116 represents the aggregated power generation as a function of the first input data 104, the second input data 106, and the third input data 108. In other examples, output data 116 may include any other suitable output generated based on the first input data 104, the second input data 106, and the third input data 108. For example, output data 116 may include non-aggregated data.

[0021] Output data 116 is generated by a correlator 124 using, for example, the Iman-Conover method to generate correlation samples. The following paragraphs describe examples of systems and methods for reducing discrepancies in simulation samples within an optimization threshold of a distribution. As described in more detail below with reference to FIGS. 2-6, an arithmetic system 102 can stratify a cumulative distribution function (CDF) based on a discrepancy metric, as shown at 132. In some examples, the CDF is an aggregated distribution function for a first input data 104, a second input data 106, and a third input data 108. The discrepancy metric is calculated based on the CDF for a plurality of initial Monte Carlo simulation results. The arithmetic system 102 further determines whether the discrepancy metric is reduced by replacing one or more samples. In these aspects of the systems and methods described below, by reducing the sampling error for the initial Monte Carlo simulation results, the predictive power of a model based on a relatively small representative pool of simulations is improved while maintaining the same accuracy as a model based on a larger number of simulations.

[0022] The arithmetic system 102 is also configured to model an aggregated distribution, as shown at 134, and / or compress a Monte Carlo simulation, as shown at 136. Thereby, the arithmetic system 102 can generate, at an output stage, a representative sequence of simulations that accurately reflects the CDF of the aggregated distribution without requiring additional simulations. These aspects are described in more detail below with reference to FIGS. 7-11.

[0023] In some examples, referring now to FIG. 1B, the output data 116 is the first output data that is input to at least the second correlator 138. In this way, it is possible to combine the correlation and stratified properties of the first output data 116 with other models. For example, the second correlator 138 is further configured to receive as input the third output data 140 from the third correlator 142. The second correlator 138 generates second output data 144 based on the first output data 116, the third output data 140, and the second copula 146. In some examples, the first output data 116 represents the aggregated power generation in the Mid-Atlantic region of the United States, and the third output data 140 represents the aggregated power generation in the southeastern United States. The second output data 144 output by the second correlator 138 represents the aggregated power generation in the eastern United States.

[0024] Section 2. Stratification Based on the Mismatch Score Referring first to FIG. 2, examples regarding the outputs of a plurality of representative simulations based on stratification are disclosed. As described above, in some examples, a sequence of Monte Carlo simulations

Number

[0025]

Number

[0026] Here,

Number

Number

Math

Math

Math

[0027] In some cases, the initial Monte Carlo simulation data 206 is affected by simulation or sampling errors. The Koksma-Hlawka inequality suppresses the error in the estimation of the statistic with respect to the total variation of the statistic and the * -discrepancy of the sample sequence.

Math

Math

[0028]

Math

[0029] Here, μ is

Math

Math

Math

Number

[0030] In other cases, * - A simple form of inconsistency is used. For example,

Number

[0031]

Number

[0032] In equation (3),

Number

Number

Number

Number

Number

Number

Number

[0033] To address these problems, an example regarding the selection of a representative sequence of stratification-based simulations is disclosed. Briefly described, the initial mismatch score is determined based on the amount of values in each bin related to a layer of factors or variables, a first predetermined number of simulations (e.g., the amount of simulations), and a second predetermined number (e.g., an amount). As described in more detail below, this formula for the mismatch score is computationally tractable. A representative sequence of simulations is selected in an iterative optimization process in which at least one simulation result value is removed from a set of initial simulation result values and at least one other simulation result value is added to this set. The updated mismatch score is determined using the amount of values in one or more bins corresponding to the removed simulation result value and the added simulation result value. As described in more detail below, computing the updated mismatch score in this way has a lower computational cost than recomputing the initial mismatch score for the entire set. Thereby, the computing system can select and output a plurality of representative simulation result values that can be used in downstream processing (e.g., Iman-Conover-based analysis of correlation variables).

[0034] FIG. 2 shows an example of a computing system 202 for selecting a representative sequence of simulations based on stratification. In some examples, the computing system 202 embodies the computing system 102 of FIG. 1A. In other examples, the computing system 202 is a separate computing system.

[0035] The calculation system 202 is configured to receive, for a plurality of correlation variables, a first predetermined number 212 of simulations 204 from Monte Carlo simulation samples. Each simulation includes a plurality of initial simulation results 206 for a plurality of variables (e.g., the initial simulation results include Monte Carlo simulation data for at least one dimension j). For simplicity, the initial simulation results 206 are described herein as one-dimensional. In other examples, the initial simulation results 206 include two or more dimensions. In still other examples, the initial simulation results 206 include ten or more dimensions.

[0036] FIG. 3 shows a plot of the initial simulation results 206 in the form of Monte Carlo simulation data for an insurance loss model. FIG. 3 also shows the CDF 208 for the initial Monte Carlo simulation data. In some examples, the CDF 208 is provided in the form of a function or includes synthetic data generated based on a function. It will also be understood that in other examples, the CDF can take the form of a discrete table including an empirical distribution. In some examples, the initial simulation results 206 are derived from 50,000 to 250,000 simulations. In other examples, as described in more detail below, the initial simulation results 206 are derived from a smaller number of simulations (e.g., less than 50,000 simulations). For random sampling, with more simulation results (e.g., 50,000 to 250,000), a closer approximation of the CDF 208 is obtained than with a smaller number of simulation results. However, as described above, performing such a large number of simulations can be computationally expensive. In the simplified example shown in FIG. 3, the plot includes ten simulations. With ten sample points, an initial approximate distribution 211 offset from the CDF 208 is formed due to the small sample size.

[0037] The calculation system 202 is further configured to receive a CDF 208 for each variable and a second predetermined number 210 of layers for the variable. The unit intervals of the CDF 208 are divided into the second predetermined number of layers, and the base of the CDF 208 (e.g., the elements of the domain of the CDF 208 that are not mapped to zero) is divided into a plurality of bins such that each bin corresponds to one of the layers. In this way, the number of layers 210 controls the accuracy of the estimated distribution obtained from one or more representative simulation result values 218 output by the calculation system. The CDF 208 is also stratified for each of the one or more target statistics 213. For example, the target statistic 213 may include one or more of the following: an aggregation result or a dependent variable value, a minimum value, a maximum value, or any other target statistic for which an operation is desired based on the dependent variable. Similarly, the unit intervals of the CDF 208 for each target statistic are divided according to the second predetermined number 210 of layers.

[0038] The number of layers 210 can be greater than or equal to a first predetermined number 212 of sample points 204. The calculation system 202 may attempt to place at least one simulation in each bin. Therefore, the greater the number of layers 210, the more the CDF 208 is divided into a larger number of bins 216, and as a result, the representative simulation results 218 can be more evenly distributed across the unit interval 214 than when the number of layers 210 is small. In some examples, the number of layers 210 is between 10,000 and 50,000 layers. In other examples, the number of layers 210 is greater than 50,000 layers (e.g., 250,000 layers). In yet other examples, the number of layers 210 is less than 50,000 layers (e.g., 1,000 layers).

[0039] For example, the CDF 208 in FIG. 3 has a unit interval 214 of [0,1], which represents the probability that the insurance loss plotted in FIG. 3 is evaluated to be below the selected dollar value on the x-axis. The unit interval is divided into ten bins 216A - 216J corresponding to deciles (0.1, 0.2, …, 1.0), each representing a range of insurance losses (in millions of dollars). For example, the first bin 216A representing the probability value in the range of 0 - 0.1 corresponds to a loss of $9 million or less. The tenth bin 216J representing the probability value in the range of 0.9 - 1.0 corresponds to a loss exceeding $33 million.

[0040] The computing system 202 is further configured to place the value x for each initial simulation 204 into one of the plurality of bins 216A - 216J. To place each value x, the computing system 202

Number

Number

Number

[0041]

Number

[0042] In Equation (4),

Number

Number

Number

Number

[0043] Referring to FIG. 3 again, the first bin 216A contains one simulation. The second bin 216B contains two simulations. The fifth bin 216E contains three simulations. The sixth bin 216F contains one simulation. The eighth bin 216H contains two simulations. The ninth bin 216I contains one simulation. The third bin 216C, the fourth bin 216D, the seventh bin 216G, and the tenth bin 216J do not contain any of the initial simulations 204.

[0044] The initial mismatch score 220 is determined based on the amount of values in each bin 216, the first predetermined number 212 of simulations 204, and the second predetermined number 210 of layers, with respect to a variable or a target statistic. As described in more detail below, the initial mismatch score measures the deviation between the initial Monte Carlo simulation data and the CDF.

[0045] In some examples, when determining the initial mismatch score 220, the determination of the initial bin-by-bin mismatch metric 222 for each bin 216 is included. In some examples, the initial bin-by-bin mismatch metric 222 for the selected bin is the difference between the amount of values in the selected bin ( [Number] and the result of dividing the first predetermined number 212 (N) of simulations 204 for the variable by the second predetermined number 210 (M) of layers:

[0046] [Number] and includes.

[0047] Referring again to FIG. 3, the initial Monte Carlo simulation data includes 10 values, and the unit interval of the CDF 208 is divided into 10 bins 216A - 216J. In a uniform distribution, one simulation is placed in each bin. [Number] Therefore, the initial bin-by-bin mismatch metric 222 measures the uniformity of the initial Monte Carlo simulation data across bins 216A - 216J.

[0048] In some examples, the initial mismatch score 220 includes the maximum bin-by-bin mismatch metric from a plurality of bins 216:

[0049] [Number]

[0050] By this method, the initial mismatch score 220 represents the maximum mismatch between the initial simulation result 206 and the CDF 208.

[0051] In other examples, the initial mismatch score 220 includes the sum of the initial bin-by-bin mismatch metrics 222 for the plurality of bins 216:

[0052]

Number

[0053] In this way, the initial mismatch score 220 represents the aggregated mismatch for the initial simulation result 206.

[0054] In some applications, it is determined that the accuracy in one or both tails of the CDF is greater than the accuracy of the initial Monte Carlo simulation data at or near the mean value (e.g., within one standard deviation). Thus, in some examples, the computing system 202 is configured to weight the initial bin-by-bin mismatch metrics for each of the plurality of bins based on the proximity to the tail of the CDF of the layer corresponding to the bin. For example, in FIG. 3, the initial bin-by-bin mismatch metric for the tenth bin 216J may be assigned a greater weight (e.g., 0.99 within the range of [0-1]) than the initial bin-by-bin mismatch metrics for bins 216A-216I. In this way, the initial mismatch score places greater emphasis on the accuracy in the tail (e.g., the tenth bin 216J) than on other parts of the distribution.

[0055] The computing system 202 is further configured to determine an initial sum 226 of the initial mismatch score 220. In some examples, the initial mismatch score 220 is summed over all variables (e.g., dimension j of the initial simulation result 204 and any other dimension). As described in more detail below, the initial sum 226 serves as an optimization metric for selecting a representative simulation result value 218 for output.

[0056] The initial sum 226 of the initial inconsistency score 220 is compared with the optimization threshold 228. As described in more detail below, if a set of simulation results meets the optimization threshold 228, the set of simulation results is output as the representative simulation result 218. In some examples, "meeting the optimization threshold" means that the initial sum 226 of the initial inconsistency score 220 is less than or equal to the optimization threshold 228. In this way, the arithmetic system 202 ensures that the representative simulation result 218 closely reflects the CDF 208.

[0057] In some examples, the optimization threshold 228 is derived from simulation and statistical parameters. For example, the optimization threshold 228 can be a parameter specified by the user based on the above first predetermined number 212 (N) and the number of layers 210 of the second predetermined number (M). In some examples, the optimization threshold is in the range of 0.001N / M to 0.5N / M, representing a deviation of 0.1% to 50% from a uniform distribution. In other examples, the optimization threshold is in the range of 0.01N / M to 0.25N / M. In still other examples, the optimization threshold is in the range of 0.01N / M to 0.5N / M. In this way, the optimization can end when the initial sum 226 of the initial inconsistency score 220 is within the percentage of the optimally specified value by the user (e.g., representing a uniform distribution across the entire layer M).

[0058] It will also be understood that the optimization threshold 228 may be defined in any other suitable way. Other suitable examples of the optimization threshold 228 include the inconsistency value specified by the user. In this way, the optimization can end when the initial sum 226 of the initial inconsistency score 220 is less than or equal to the inconsistency value specified by the user.

[0059] In yet other examples, the optimization process ends additionally or alternatively based on reaching or exceeding a user-specified execution time or a user-specified number of iterations. In such examples, the computing system 202 may proceed as if the optimization threshold 228 is met and output the current set of simulation results as a plurality of representative simulation result values 218. The computing system 202 may additionally or alternatively output a notification to the user that a user-specified execution time or a user-specified number of iterations has been reached or exceeded.

[0060] If the set of simulation results does not meet the optimization threshold, the set of simulation results is modified as shown at 230. To generate the modified set of simulation results 230, the computing system 202 is configured to remove at least one of the plurality of initial simulations 204 based on a determination that the initial sum 226 of the initial mismatch scores 220 is not within the optimization threshold 228. For example, FIG. 4 shows a plot of the initial Monte Carlo simulation data of FIG. 3, where one of the initial Monte Carlo simulation data values 206A has been removed. In some examples, the computing system 202 of FIG. 2 randomly selects the value 206A to be removed. This can help reduce simulation error through the statistical effect achieved via random sampling. In other examples, the computing system 202 selects the value 206A to be removed based on a determination that the value 206A contributes to a mismatch score that is above a threshold (e.g., the optimization threshold 228). In this way, the process of modifying the set of initial simulation results 204 is explicitly driven to reduce the initial mismatch score. In yet other examples, the value 206A to be removed is selected by the user in response to, for example, receiving a prompt from the computing system 202 indicating that the optimization threshold 228 is not met. By selectively removing data in this way, representative simulation results 218 that meet the user-specified goal (e.g., conform to a distribution pattern not defined by the computing system 202) can be obtained.

[0061] At least one other simulation result is added to the remaining one or more initial simulation results. It will be recognized that "added" means that at least one other simulation result is included in the set containing these results, rather than being numerically added to the remaining one or more initial simulation results. For example, another simulation value 206B is added to the remaining value from the initial simulation result 206 after the value 206A is removed. In the examples of FIGS. 3-4, one other value 206B is added for each removed value 206A. By this method, the quantity 212 of the initial simulation result 204 is equal to the target number of representative simulation result values 218 for output. It will also be understood that in other examples, the number of representative simulation result values 218 can be greater than or less than the first predetermined number 212. By this method, the arithmetic system 202 can increase or decrease the sample size respectively to generate a representative sample.

[0062] In some examples, at least one other simulation result value 206B is derived from a pre-computed Monte Carlo simulation 232. The pre-computed simulation 232 can be selected from a pool of pre-computed Monte Carlo simulation results that also includes the initial simulation result 206. By pre-computation, the arithmetic system 202 can pre-estimate the CDF 208 for the aggregated value of all pre-computed simulation results (including the simulation selected as the representative simulation 218 and the simulations not selected). Thereby, the arithmetic system 202 can compare the same CDF 208 with the initial simulation result 204 and the modified simulation result value 230 without re-computing a new CDF 208 for the modified simulation result value 230.

[0063] Similar to the selection of the removed value 206A, in some examples, the computing system 202 generates or selects at least one other simulation result value 206B via a randomized process. In this way, the simulation result 206B can reduce the simulation error in the corrected data 230 via random sampling. In other examples, the simulation result value 206B is explicitly selected by either the computing system 202 or the user to drive the corrected data 230 towards the optimization threshold 228.

[0064] As described above, the computing system 202 is configured to generate an updated disagreement score 234 for each variable. The updated disagreement score is generated using the amount of values within one or more bins where one or more simulation results 206A are removed, and the amount of values within one or more bins where one or more other simulation result values 206B are added. As described in more detail below, by formulating this updated disagreement score 234, the computing system 202 does not need to recompute the initial bin-by-bin disagreement metric 222 for each bin 216A - 216J. Instead, the computing system 202 can reuse the initial bin-by-bin disagreement metric 222 for non-updated bins (e.g., bins 216A - 216H), and the updated bin-by-bin disagreement metric 236 is computed for bins where one or more simulation results are added or removed (e.g., bins 216J and 216I, respectively). As a result, each time a simulation result is added and / or removed, the disagreement score 234 is updated more quickly while reducing the processor time and memory allocation requirements compared to recomputing the initial bin-by-bin disagreement metric 222 for all bins 216A - 216J.

[0065] For example, referring again to FIG. 4, the updated bin-by-bin disagreement metric is determined for bin 216I corresponding to the removed value 206A and for bin 216J corresponding to the added value 206B. In some examples, determining the updated bin-by-bin disagreement metric involves decrementing the amount of values (e.g., 206A) in one or more bins (e.g., bin 216I) corresponding to the removed simulation and incrementing the amount of values (e.g., 206B) in one or more bins (e.g., bin 216J) corresponding to the added simulation. For example, for the removed sample

Number

Number

Number

[0066]

Number

[0067]

Number

[0068] The computing system further determines the bin count of

Number

Number

[0069] The simulation value to be added

Number

Number

Number

[0070]

Number

[0071]

Number

[0072] The arithmetic system further has the dimension

Number

Number

Number

[0073] In some examples, the updated mismatch score 234 is determined using the same operations as the initial mismatch score 220. For example, the updated mismatch score 234 may be generated using Equation (6) or (7). In this way, the updated mismatch score 234 can be made comparable to the initial mismatch score 220.

[0074] Referring back to FIG. 2, the arithmetic system 202 is further configured to determine the updated sum 238 of the updated mismatch scores 234. In this way, the updated mismatch scores 234 can be made comparable to the optimization threshold 228. By this comparison, the arithmetic system 202 can determine whether the corrected simulation result 230 meets the optimization threshold 228 for output as the representative simulation result value 218, or start an additional optimization loop (e.g., by further correcting the corrected simulation result 230).

[0075] In some examples, the computing system 202 is configured to accept or reject at least one other simulation result value 206B for potential inclusion in the event-based model based on the updated sum 238 of the disagreement scores 234. For example, the computing system 202 may reject at least one other simulation result value 206B if the updated sum 238 of the disagreement scores 234 is greater than the initial sum 226 of the initial disagreement scores 220. The computing system 202 may additionally or alternatively reject the removal of the value 206A if the other simulation result value 206B is rejected. In this way, the computing system 202 is configured to drive the selection of the representative simulation result value 218 towards the optimization threshold 228.

[0076] The computing system 202 is configured to output a plurality of representative simulation result values 218 that represent the CDF 208 across the layer based on the updated sum 238 of the disagreement scores 234. For example, if the modified simulation result 230 meets the optimization threshold 228, the modified simulation result 230 is output as the representative simulation result 218. If the modified simulation result 230 does not meet the optimization threshold 228, the computing system 202 is configured to iteratively modify the simulation result and update the sum of the disagreement scores until the optimization threshold is met.

[0077] FIG. 5 shows a plot of representative simulation result values 218 obtained from at least six iterations of swapping simulation values. In each iteration, one simulation value is removed and one simulation value is added. The resulting approximate distribution 240 formed by the representative simulation results 218 is closer to the CDF 208 than the initial simulation results 206 of FIG. 3. As a result, the representative simulation results 218 may be considered a more accurate model than the initial simulation results.

[0078] Referring now to FIGS. 6A - 6C, a flowchart is shown illustrating an exemplary method 600 for selecting a representative sequence of simulations based on stratification. The following description of method 600 is provided with reference to the above software and hardware components and is shown in FIGS. 1 - 5 and 16, and the method steps in method 600 are described with reference to the corresponding portions of FIGS. 1 - 5 and 16 below. It will be recognized that method 600 may also be implemented in other contexts using other suitable hardware and software components.

[0079] It will be recognized that the following description of method 600 is provided by way of example and is not intended to be limiting. Various steps of method 600 may be omitted or performed in an order different from that described, and methods 6A - 6C may include additional and / or alternative steps with respect to those shown in FIGS. 6A - 6C without departing from the scope of the present disclosure.

[0080] Referring first to FIG. 6A, at 602, method 600 includes receiving, for a plurality of correlation variables, a first predetermined number of simulations from Monte Carlo simulation samples. Each simulation includes a plurality of initial simulation results for a plurality of variables. Method 600 further includes receiving, based on a correlation probability variable, one or more target statistics 213. Method 600 further includes receiving a second predetermined number of layers for each variable and target statistic, and a cumulative distribution function for each variable and target statistic. For example, the computing system 202 of FIG. 2 receives the initial simulation results 206, a first predetermined number 212 of simulations 204, a second predetermined number 210 of layers for each variable and target statistic, and the CDF 208. With the first predetermined number 212 of simulations, the second predetermined number 210 of layers, and the CDF 208, the computing device 202 can evaluate the initial simulation results 206 for selection error and output a plurality of representative simulations 218 that are within the optimization threshold of the CDF 208 across the layers.

[0081] At 604, method 600 includes, for each variable and for each of the one or more target statistics, dividing the unit interval of the cumulative distribution function into the second predetermined number of layers, and dividing the support of the cumulative distribution function into a plurality of bins such that each bin of the plurality of bins corresponds to one of the layers. For example, FIG. 3 shows the CDF 208 divided into ten bins 216A - 216J corresponding to the deciles k of the cumulative probability. By dividing the CDF 208, the computing system 202 can evaluate the distribution of the initial simulation results 206. The CDF 208 is also stratified based on the target statistic (e.g., an aggregated value, a minimum value, a maximum value, or any other target statistic for which an operation based on a dependent variable is desired).

[0082] Method 600 further includes, at 606, determining an initial mismatch score based on the amount of values in each bin, a first predetermined number of simulations, and a second predetermined number of layers for a variable. For example, the computing system 202 is configured to determine an initial mismatch score 220 based on the amount of values in each bin 216, a first predetermined number 212 of simulations 204, and a second predetermined number 210 of layers. As described above, the initial mismatch score measures the deviation between the initial Monte Carlo simulation data and the CDF.

[0083] In some examples, at 608, determining the initial mismatch score includes determining an initial bin - specific mismatch metric for each of a plurality of bins. For example, the computing system 202 is configured to determine an initial bin - specific mismatch metric 222 for each bin 216. The initial bin - specific mismatch metric indicates how evenly the initial simulation results 206 are distributed among bins 216A - 216J.

[0084] At 610, in some examples, determining the initial bin - specific mismatch metric for a selected bin includes determining the difference between the amount of values in the selected bin and the value obtained by dividing the first predetermined number of simulations by the second predetermined number of layers for the variable or a target statistic. For example, the initial bin - specific mismatch metric 222 can be expressed in the form of the above formula (5). By this method, the initial bin - specific mismatch metric 222 measures the mismatch between the number of values in each bin and the number of values if the initial simulation were evenly distributed across the entire bin.

[0085] In some examples, at 612, determining the initial bin-by-bin disagreement metric for the selected bins at 612 includes determining the maximum bin-by-bin disagreement metric, or determining the sum of the initial bin-by-bin disagreement metrics for multiple bins. For example, the initial disagreement score 220 may include the maximum bin-by-bin disagreement metric described in Equation (6), or the sum of the initial bin-by-bin disagreement metrics described in Equation (7). In this way, the initial disagreement score represents, respectively, the maximum disagreement between the initial Monte Carlo simulation data and the CDF, or the aggregated disagreement with respect to the initial Monte Carlo simulation data.

[0086] At 614, in some examples, determining the initial disagreement score includes weighting the initial bin-by-bin disagreement metric for each of the multiple bins based on the proximity to the tail of the cumulative distribution function of the layer corresponding to the bin. For example, in FIG. 3, a larger weight (e.g., 0.99 within the range of [0-1]) may be assigned to the initial bin-by-bin disagreement metric for insurance losses of $33 million or more than a value within the range of $15 million to $25 million (e.g., 0.15 within the range of [0-1]). In this way, the initial disagreement score can be weighted to place greater emphasis on the tail of the CDF or any other suitable portion of the CDF (e.g., the mean of the CDF).

[0087] Now, referring to 616 in FIG. 6B, method 600 includes determining an initial sum of the initial disagreement scores. For example, the computing system 202 determines an initial sum 226 of the initial disagreement scores 220. As described above, the initial sum 226 is used as the optimization metric, which is compared with the optimization threshold 228 to select the representative simulation 218.

[0088] For example, at 618, method 600 includes removing at least one of a plurality of initial simulations based on a determination that an initial sum of initial disagreement scores is not within an optimization threshold. For example, in FIG. 4, one simulation result value 206A is removed from a plurality of initial simulation result values. Removal of at least one of the plurality of initial simulation result values can help reduce simulation error via random sampling or via an explicit selection of removing outliers from the CDF.

[0089] At 620, method 600 includes adding at least one other simulation to the remaining one or more initial simulations. For example, simulation result value 206B is added to the values remaining after value 206A is removed. Similar to the removal of simulation result value 206A, the addition of simulation result value 206B can reduce simulation error in the modified data 230 via random sampling or via an explicit selection of a simulation value closer to the CDF than the removed value 206A. Additionally, the addition of at least one other simulation result helps achieve or maintain a target number of representative simulation results for output.

[0090] In some examples, at 622, at least one other simulation includes a precomputed simulation. For example, computing system 202 is configured to precompute one or more Monte Carlo simulations 232. This enables computing system 202 to pre-estimate the CDF 208 and reuse the CDF 208 for the modified simulation result values 230 instead of recomputing a new CDF 208 at each step of the optimization process.

[0091] Referring now to 624 in FIG. 6C, method 600 includes generating an updated disagreement score for each variable and target statistic using the amount of values within one or more bins corresponding to at least one of a plurality of initial simulation results and the amount of values within one or more bins corresponding to at least one other simulation result. For example, computing system 202 is configured to generate an updated disagreement score 234 for each variable and target statistic. This updated disagreement score 234 is compared to an optimization threshold 228 to determine whether the optimization threshold 228 is met or whether further removal and addition of simulation values is to be performed.

[0092] In some examples, at 626, generating an updated disagreement score includes determining an updated per-bin disagreement metric for one or more bins corresponding to at least one of a plurality of initial simulations and for one or more bins corresponding to at least one other simulation. The updated per-bin disagreement metric is used for one or more bins corresponding to at least one of a plurality of initial simulations and for one or more bins corresponding to at least one other simulation. The initial per-bin disagreement metric is used for each of the remaining one or more initial simulations to generate an updated disagreement score. For example, computing system 202 samples

Number

Number

Number

Number

[0093] In some examples, at 628, generating an updated mismatch score can include decrementing the amount of values in one or more bins corresponding to at least one of a plurality of initial simulations; and incrementing the amount of values in one or more bins corresponding to at least one other simulation. For example, the calculation system can be configured to decrement the bin count of in dimension k by the number of removed simulation values (e.g., 1). The calculation system can further be configured to increment the bin count of in dimension k by the number of added simulation values (e.g., 1). The amount of updated values in each of bins 216I and 216J is used to generate the updated bin-by-bin mismatch metric.

Number

Number

Number

Number

[0094] Method 600 further includes, at 630, determining an updated sum of the updated disagreement scores. For example, the computing system 202 is configured to calculate an updated sum 238. The updated sum is a measure of the aggregated disagreement that can be compared to the same optimization threshold 228 as the initial sum 226. Based on this comparison, the computing system 202 can start an additional optimization loop (e.g., by further correcting the corrected simulation result 230), or determine whether the corrected simulation result 230 meets the optimization threshold 228 for output as the representative simulation result value 218.

[0095] In some examples, at 632, method 600 includes adopting or rejecting at least one other simulation based on the updated sum of the updated disagreement scores. For example, the computing system 202 may reject the additional simulation result value 206B if the updated sum 238 of the disagreement scores 234 is greater than the initial sum 226 of the initial disagreement scores 220. Thereby, the corrected simulation result 230 can be driven towards the optimization threshold 228.

[0096] At 634, method 600 includes outputting a plurality of representative simulation results representing a cumulative distribution function across the layer based on the updated sum of the updated disagreement scores. For example, the computing system 202 is configured to output the corrected simulation result 230 as the representative simulation result 218 if the corrected simulation result 230 meets the optimization threshold 228.

[0097] The above-described system and method can be used to select a representative sequence of simulations from a pool of Monte Carlo simulation data. At least one simulation is removed from a plurality of initial simulations, based at least in part on an initial disagreement score, and at least one other simulation is added to the remaining values. Thereby, the disagreement of the selected sequence of simulations can be reduced compared to the initial disagreement score, either through the statistical effect achieved via random sampling or by explicitly replacing outliers from the CDF of the Monte Carlo simulation data values. Additionally, adding at least one other simulation helps to achieve or maintain the target number of simulations for output. An updated disagreement score is generated and used to select a plurality of representative simulations for output. The updated disagreement score is generated by using the updated bin-by-bin disagreement metric for any bins that have been modified and reusing the initial bin-by-bin disagreement metric for any bins that have not been modified. This formulation of the updated disagreement score can be updated more quickly and at a reduced computational cost compared to recalculating the initial bin-by-bin disagreement metric for each bin. As a result, the above-described system and method enhance the predictive power of models based on a relatively small pool of representative samples by reducing the sampling error for the initial simulation results while maintaining the same accuracy as models based on more samples.

[0098] As discussed above with reference to FIG. 2, when the arithmetic system 202 receives the initial simulation result 204 and the respective CDFs 208 of the variables and target statistics of the initial simulation result 204, the arithmetic system 202 may further receive the quantity 210 of layers into which the CDF 208 can be divided. Thus, the CDF 208 can be divided into a number of bins 216 equal to the quantity 210. By classifying the initial simulation result 204 into the bins 216, the arithmetic system 202 may be able to implement the above-described stratified sampling technique to generate a representative simulation result 218 for the CDF 208 across the entire layer. The representative simulation result 218 may have a reduced total discrepancy compared to the initial simulation result, thereby reducing the redundancy when the representative simulation result 218 is used as an input to the Monte Carlo simulation.

[0099] When using the above-described stratified sampling technique, determining the position of the boundaries between the bins 216 is one of the possible issues. Depending on the shape of the CDF 208, the position within the unit interval of the boundaries between the bins 216 can vary between different sets of the initial simulation result 204. To obtain a set of representative simulation results 218 with reduced discrepancies with respect to the initial simulation result 204, the arithmetic system 202 is configured to utilize a surrogate cumulative distribution model, as described in more detail below.

[0100] Section 3. Modeling of the Aggregate Distribution Using the Surrogate Cumulative Distribution Model and Compression of the Monte Carlo Simulation

[0101] In addition to modeling the shape of the CDF208, the surrogate cumulative distribution model described below can also be used when modeling the dependent variable of the initial simulation results 204. One such target statistic that is particularly notable in applications such as insurance, inventory management, and energy production is the aggregation of correlated probability variables. The aggregation can be, for example, the aggregated losses by an insurance company, the aggregated quantity of products sold, or the aggregated quantity of energy generated. As described below, the surrogate cumulative distribution model can be used when generating low-discrepancy samples of the aggregated distribution.

[0102] FIG. 7 shows an exemplary computing system 702 configured to execute an event simulation model 704, a surrogate cumulative distribution model 714, and a resampling module 724, according to an example. The computing system 702 can be, for example, an implementation of the computing system 102 of FIG. 1. Alternatively, the computing system 702 can be a separate computing system. The aggregated distribution modeling 134 and the Monte Carlo simulation compression 136 shown in FIG. 1 can be executed in the computing system 702.

[0103] The computing system 702 may be configured to receive a plurality of simulation results 712 for a plurality of correlation probability variables 706. For example, as shown in FIG. 7, the plurality of simulation results 712 may be generated in the event simulation model 704 and may be included in the plurality of simulations 711. In some examples, the event simulation module 704 is included in the correlator 124 of FIG. 1. When the computing system 702 executes the event simulation module 704, the computing system may be configured to generate a simulation sample 710 that includes a plurality of simulations 711, where each simulation 711 includes a plurality of simulation results 712. Each of the simulation results 712 may be a value of a correlation probability variable 706, and each simulation 711 may include respective simulation results 712 for each of the correlation probability variables 706. Each of the simulations 711 may include the same number of simulation results 712. In some examples, the plurality of simulation results 712 may be included in the initial simulation results 204 shown in FIG. 2.

[0104] In some examples, the plurality of simulation results 712 may include a plurality of aggregate values across the plurality of correlation probability variables 706. In other examples, the plurality of simulation results 712 may include a plurality of minimum or maximum values across the plurality of correlation probability variables 706. Additionally or alternatively, other target statistics of the correlation probability variables 706 may be included in the plurality of simulation results 712.

[0105] In some examples, the computing system 702 can be configured to generate a plurality of simulation results 712 for a plurality of correlated probability variables 706 by, at least in part, executing an Iman-Conover algorithm 750 in an event simulation model 704. When the Iman-Conover algorithm 750 is executed in the event simulation module 704, the computing system 702 can be configured to sample simulation results 712 for the plurality of correlated probability variables 706 in a way that maintains dependencies among these simulation results 712. Thus, the computing system 702 can be configured to generate a simulation 711. The dependencies between correlated events 706 can be represented by a copula 118 and / or a correlation matrix 122 as discussed above with reference to FIG. 1.

[0106] The event simulation module 704 can be configured to receive and / or compute each CDF 708 associated with a correlated probability variable 706. The event simulation module 704 can further be configured to use the CDF 708 as an input when computing simulation results 712 included in the simulation 711. The CDF 708 can be, for example, the CDF 208 and can be estimated as shown in the examples of FIGS. 2-5.

[0107] The computing system 702 can further be configured to generate a surrogate cumulative distribution model 714 configured to model the CDF 708. The surrogate cumulative distribution model 714 can be, for example, a user-defined custom function, an aggregation of a plurality of variables or some other statistic of the variables, and can be configured to model the CDF 708 of the minimum or maximum value. As will be described below, the surrogate cumulative distribution model 714 may have a plurality of surrogate model parameters 718, and generating the surrogate cumulative distribution model 714 can include estimating the plurality of surrogate model parameters 718 based at least in part on the plurality of simulation results 712.

[0108] In some examples, the surrogate cumulative distribution model 714 can be a mixture Erlang model that includes a plurality of Erlang distributions 716. Thus, in such examples, the surrogate cumulative distribution model 714 can approximate the CDF 708 as a weighted sum of the plurality of Erlang distributions 716 each having a respective mixing weight included in the plurality of surrogate model parameters 718. Each of the plurality of Erlang distributions can be parameterized by parameters

Number

Number

Number

Number

Number

Number

[0109] In some examples, the surrogate cumulative distribution model 714 may further include one or more alternative tail region distributions 720. The one or more alternative tail region distributions 720 may be configured to replace one or more respective tail regions in one or more of the plurality of Erlang distributions 716, and may be different from the one or more Erlang distributions 716 within one or more respective tail regions. One or more tail regions configured to be replaced by the one or more alternative tail region distributions 720 may include the lower tail and / or the upper tail of the mixed Erlang model. Thus, when the surrogate cumulative distribution model 714 includes one or more alternative tail region distributions 720, the surrogate cumulative distribution model 714 may be a sum of the plurality of Erlang distributions 716 and a piecewise function of the one or more alternative tail region distributions 720, and may have one or more respective thresholds that specify one or more cut-off points between the sum of the plurality of Erlang distributions 716 and the one or more alternative tail region distributions 720. In an example where the surrogate cumulative distribution model 714 includes one or more alternative tail region distributions 720, the computing system 702 may be configured to normalize the surrogate cumulative distribution model 714 such that the integral value of the surrogate cumulative distribution model 714 over an interval

Number

[0110] In some examples, the surrogate cumulative distribution model 714 may be an empirical model. In such examples, the computing system 702 may be configured to estimate the surrogate model parameters 718 based at least in part on empirical data included in a plurality of simulations 711. The plurality of simulations 711 may include both empirically collected data and synthetic data generated programmatically in such examples.

[0111] By modeling the CDF 708 with a surrogate cumulative distribution model 714 represented as the sum of a plurality of Erlang distributions 716, the computing system 702 can flexibly model the shape of the CDF 708 while using only a few parameters. Additionally, by including one or more alternative tail region distributions 720 in the surrogate cumulative distribution model 714, the computing system 702 may be able to more accurately represent a light-tailed or heavy-tailed distribution. Since heavy-tailed distributions are particularly important in risk modeling applications (e.g., insurance for extreme weather events), a mixed Erlang model with an alternative tail region may be able to more accurately model regions of the distribution that are particularly relevant to a user's decision-making.

[0112] FIG. 8 shows an exemplary process in which the computing system 702 may be configured to estimate a plurality of surrogate model parameters 718 when generating the surrogate cumulative distribution model 714. In the example of FIG. 8, the computing system 702 is configured to estimate a plurality of surrogate model parameters 718 by at least partially performing iterative expectation maximization. When the computing system 702 performs iterative expectation maximization, the computing system 702 may be configured to compute respective expected values 734 of the surrogate cumulative distribution model 714 in each of a plurality of parameter update iterations 736 when simulation results 712 included in simulation samples 710 are input to the surrogate cumulative distribution model 714.

[0113] In some examples, as shown in FIG. 8, iterative expectation maximization may be performed using a generalized expectation maximization (GEM) algorithm in which a plurality of parameter update iterations 736 include an E-step 738, an M-step 740, and a cross-validation step 742. The E-step 738, M-step 740, and cross-validation step 742 may each include a plurality of iterative steps. During the E-step 738, the computing system 702 may be configured to compute an expectation value 734 as a conditional log-likelihood expectation value. The expectation value 734 may be computed as a function of the simulation results 712 and the current values of the surrogate model parameters 718.

[0114] In the M-step 740, the computing system 702 may further be configured to compute an estimated argmax of the expectation value 734 as a function of the surrogate model parameters 718. In some examples, the argmax may be estimated at least in part by performing a probabilistic search algorithm such as simulated annealing, simulated quantum annealing, population annealing, or parallel tempering. Additionally, or alternatively, the computing system 702 may be configured to compute at least in part the argmax of the expectation value 734 by repeatedly executing a 3-opt algorithm to update the shape parameter k of the Erlang distribution 716 to the estimated argmax value.

[0115] In cross-validation step 742, the computing system 702 may further be configured to select such that the surrogate cumulative distribution model 714 includes a small number of Erlang distributions 716 in order to avoid overfitting. Performing the cross-validation step 742 may include selecting a training set 710A and a validation set 710B each including a simulation 711 included in the simulation sample 710. The computing system 702 may use the simulations 711 included in the training set 710A as inputs when computing a density function fitted with the surrogate model parameters 718. The computing system 702 may further be configured to use the simulations 711 included in the validation set 710B as inputs together with the fitted density function when computing a cross-validation score. The computing system 702 may further be configured to compute respective cross-validation scores for different numbers of Erlang distributions 716 and select the number of Erlang distributions 716 that maximizes the cross-validation score. Thus, the computing system 702 may be configured to iteratively compute the parameters of a mixed Erlang model that accurately represents the CDF 708 while also including a small number of Erlang distributions 716.

[0116] Returning to the example of FIG. 7, based at least in part on the surrogate cumulative distribution model 714 including the surrogate model parameters 718, the computing system 702 may further be configured to select one or more subsets 722 of the plurality of simulations 711. In some examples, the computing system 702 may be configured to select the plurality of simulations 711 included in the subset 722 using a random or pseudo-random process.

[0117] The computing system 702 can be configured to stratify an image of the surrogate cumulative distribution model 714 into multiple layers. These layers can be quantiles. For example, the quantiles can be quintiles, deciles, percentiles, or some other subsets obtained by dividing the range of the surrogate cumulative distribution model into equally sized layers. The computing system 702 can be configured to estimate the positions of the boundaries between the layers by computing the positions of the boundaries between the quantiles of the surrogate cumulative distribution model 714. In an example where the positions of the boundaries are computed, the computing system can further be configured to select one or more subsets 722 of the plurality of simulations 711 such that the simulation results 712 included in the simulations 711 included in the one or more subsets 722 are evenly distributed across the entire layer.

[0118] The computing system 702 can further be configured to perform one or more resampling iterations 748 in the resampling module 724. During each of the resampling iterations 748, the computing system 702 can be configured to compute a respective mismatch score 726 for each of the one or more subsets 722. Each mismatch score 726 can be computed as described above with reference to FIG. 2. Additionally, the computing system 702 can further be configured to compute the sum of the one or more mismatch scores 726. The one or more resampling iterations 748 can be repeatedly executed until it is determined that the sum of the mismatch scores 726 for the one or more subsets 722 is less than a predetermined mismatch threshold 728. Alternatively, some other optimization threshold can be used as the endpoint of the one or more resampling iterations 748. The optimization threshold can be selected as described above with reference to FIG. 2, for example.

[0119] During each of the one or more resampling iterations 748, based at least in part on the disagreement score 726, the computing system 702 may further be configured to sample one or more resampled simulations 730 for the plurality of correlation probability variables 706. The resampled simulations 730 may be sampled from a plurality of simulations 711 that are included in the simulation samples 710 but not yet included in the one or more subsets 722. Thus, the computing system 702 may precompute the plurality of simulations 711 and resample the simulations 711 included in the one or more subsets 722 without the need to generate additional simulations 711 in the event simulation module 704. The resampling iteration 748 may thus be executed more quickly in cases where the event simulation module 704 takes a significant amount of time to compute the simulations 711.

[0120] After generating one or more resampled simulations 730 in each resampling iteration 748, the computing system 702 may further be configured to replace one or more simulations 711 included in one or more subsets 722 with one or more resampled simulations 730. If one or more of the plurality of simulations 711 included in one or more subsets 722 are resampled, the computing system 702 may be configured to select one or more simulations 711 included in simulation samples 710 that are not yet included in the subset 722. Therefore, one or more simulations 711 used for estimating the surrogate distribution model 714 may be reused, thereby reducing the amount of computation performed in generating one or more resampled simulations 730. Thus, if subsequent resampling iterations 748 follow the resampling iteration 748 in which one or more resampled simulations 730 are generated, the computing system 702 may be configured to use an updated subset 722 that includes one or more resampled simulations 730 in a subsequent resampling iteration 748 when computing one or more mismatch scores 726.

[0121] In some examples, the computing system 702 may be configured to at least partially replace one or more simulations 711 with one or more resampled simulations 730 by executing a quantum-inspired algorithm. The quantum-inspired algorithm may be a Markov chain Monte Carlo algorithm such as, for example, simulated annealing, simulated quantum annealing, population annealing, or parallel tempering. The sum of the mismatch scores 726 may be used as a loss that the resampling module 724 may be configured to estimate a minimum value in such examples.

[0122] In some examples, when one or more of the plurality of simulations 711 are resampled, the computing system 702 may be configured to sample at least partially one or more resampled simulations 730 by executing the Iman-Conover algorithm 750. The computing system 702 may be configured to execute the Iman-Conover algorithm 750 in the resampling module 724, with the simulations 711 included in the simulation samples 710 as input.

[0123] FIG. 9 shows an exemplary resampling iteration 748. During the resampling iteration 748, as shown in FIG. 9, the computing system 702 may be configured to compute a plurality of layers 744 of the surrogate cumulative distribution model 714 of the statistic of interest. The computing system 702 may further be configured to select one or more subsets 722 such that the simulations 711 included in the one or more subsets 722 have simulation results 712 related to the statistic of interest that are evenly distributed among the plurality of layers 744. In the example of FIG. 9, the values of the statistic of interest for the plurality of simulation results 712 have a range 746 divided into five quantiles 744.

[0124] When resampling iteration 748 is executed, the computing system 702 may replace simulation 711 with resampled simulation 730. In this case, simulation result 712 included in simulation 711 is replaced with resampled simulation result 731 included in resampled simulation 730. This resampled simulation result 731 may be in layer 744 where simulation result 712 is located. Alternatively, resampled simulation 730 may be in a layer different from layer 744 where simulation result 712 is located. Simulation results 712 included in other simulations 711 of simulation sample 710 may remain unchanged. FIG. 9 shows one simulation 711 resampling during resampling iteration 748, but multiple simulations 711 may be resampled in resampling iteration 748.

[0125] Returning to FIG. 7, after executing one or more resampling iterations 748, the computing system 702 may further be configured to output one or more subsets 722 of simulations 711. One or more subsets 722 may be output to additional computational process 732 where simulations 711 included in one or more subsets 722 can be used as input.

[0126] In some examples, the additional arithmetic process 732 can be a Monte Carlo algorithm 732A. In such examples, one or more subsets 722 can be used as one or more compressed inputs to the Monte Carlo algorithm 732A for which the arithmetic system 702 can compute an estimated solution to the optimization problem. The Monte Carlo algorithm 732A can be configured to compute an estimated solution to the optimization problem with a small input set for the complete simulation sample 710. However, since the subset 722 is resampled to have a sum of one or more mismatch scores 726 less than a predetermined mismatch threshold 728, the accuracy of the Monte Carlo simulation can be maintained while reducing the number of inputs. Therefore, by compressing the simulation sample 710 using the surrogate cumulative distribution model 714, the Monte Carlo simulation can be made more efficiently executable.

[0127] In some examples, as shown in FIG. 10, a graphical user interface (GUI) 800 can be implemented in the arithmetic system 702. The GUI 800 can be displayed, for example, on a display device included in the arithmetic system 702. In the GUI 800, the arithmetic system 702 can be configured to receive an indication of input data for which the surrogate cumulative distribution model 714 and one or more subsets 722 are to be generated. The user may, for example, specify the number of simulations 711 initially generated in the event simulation module 704.

[0128] In the GUI 800, the arithmetic system 702 can further be configured to generate the surrogate cumulative distribution model 714 in response to receiving a surrogate model type selection 802 in the GUI 800. The surrogate model type selection 802 can include a selected type of function configured to be used as the surrogate cumulative distribution model 714. In the example of FIG. 10, an Erlang mixture model is selected. Additionally, the surrogate model type selection 802 can include one or more specifications of one or more alternative tail region distributions 720.

[0129] In the GUI 800, the computing system 702 may further be configured to generate one or more subsets 722 of the plurality of simulations 711 in response to receiving a simulation generation command 804 in the GUI 800. As shown in the example of FIG. 10, the simulation generation command 804 may include the number of simulations 711 to include in the total across one or more subsets 722. Alternatively, the user may specify, in the simulation generation command 804, the number of simulations 711 to include in each subset 722. The simulation generation command 804 may further include a plurality of layers 744 that divide a range 746 of simulation results 712 included in the simulations 711 of the simulation samples 710.

[0130] The computing system 702 may further be configured to output one or more subsets 722 of the simulations 711 to the GUI 800. The illustrated GUI 800 of FIG. 10 includes an option to “display compressed samples”.

[0131] FIG. 11A shows a flowchart of a method 900 used in the computing system. The method 900 may be executed, for example, by the computing system 702 of FIG. 7. In step 902, the method 900 may include receiving a simulation sample 710 that includes a plurality of simulations 711 for a plurality of correlation probability variables 706. Each simulation may include a plurality of simulation results 712 that may be sampled values of the correlation probability variables 706. In some examples, the plurality of simulation results 712 may include a plurality of aggregated values for the plurality of correlation probability variables 706. Alternatively, the plurality of simulation results 712 may be the minimum value, the maximum value, or the values of some other function of the correlation probability variables 706.

[0132] A plurality of simulations 711 may be received from the event simulation module 704. The event simulation module 704 may be configured to receive the CDF 708 as an input. In some examples, a plurality of simulation results 712 may be computed using the Iman-Conover algorithm 750.

[0133] In step 904, method 900 may further include generating a surrogate cumulative distribution model 714 for a plurality of correlation probability variables 706. Generating the surrogate cumulative distribution model 714 may include estimating a plurality of surrogate model parameters 718 based at least in part on the plurality of simulation results 712. The plurality of surrogate model parameters 718 may be estimated at least in part, for example, by performing iterative expectation maximization. For example, the GEM algorithm may be used. In some examples, the surrogate cumulative distribution model 714 may be an empirical model where the surrogate model parameters 718 are estimated based at least in part on empirical data included in the plurality of simulations 711.

[0134] In some examples, the surrogate cumulative distribution model 714 may be a mixed Erlang model that includes a plurality of Erlang distributions 716. The plurality of surrogate model parameters 718 may include, in such examples, the parameters of the Erlang distributions 716 and the mixing weights for the Erlang distributions 716. The surrogate cumulative distribution model 714 may further include one or more alternative tail region distributions 720 configured to replace one or more respective tail regions in one or more of the plurality of Erlang distributions 716. The one or more alternative tail region distributions 720 may replace the upper and / or lower tails of the mixed Erlang model and may also be different from the one or more Erlang distributions 716 within one or more of the respective tail regions.

[0135] In step 906, method 900 may further include selecting one or more subsets 722 of the plurality of simulations 711 based at least in part on a surrogate cumulative distribution model 714 that includes surrogate model parameters 718. In some examples, the one or more subsets may be selected to be of equal size. Alternatively, the plurality of subsets may have a plurality of different sizes.

[0136] Steps 908, 910, and 912 of method 900 may be performed in each of one or more resampling iterations 748. These steps may be performed until it is determined that the sum of the one or more disagreement scores 726 for one or more of the subsets 722 meets an optimization threshold. For example, the sum of the disagreement scores 726 may meet the optimization threshold if this sum is less than a predetermined disagreement threshold 728. In step 908, method 900 may include computing one or more disagreement scores 726 for one or more of the subsets 722. In step 910, method 900 may further include sampling one or more resampled simulations 730 for the plurality of correlation events 706 based at least in part on the sum of the one or more disagreement scores 726. The one or more resampled simulations 730 may be included in the simulation samples 710 and may be sampled from a plurality of simulations 711 that are not yet included in the one or more subsets 722. In step 912, method 900 may further include replacing one or more simulations 711 included in the one or more subsets 722 with the one or more resampled simulations 730. Thus, the sum of the one or more disagreement scores 726 for the one or more subsets 722 may be reduced during the course of the plurality of resampling iterations 748.

[0137] In step 914, method 900 may further include outputting simulations 711 included in one or more subsets 722 after performing one or more resampling iterations 748. In some examples, one or more subsets 722 may be output to a Monte Carlo algorithm 732A. The one or more subsets 722 may be one or more compressed subsets of the input to the Monte Carlo algorithm 732A for which the sum of one or more disagreement scores 726 is less than that of the initial simulation sample 710. Thus, by using subset 722 as input, Monte Carlo algorithm 732A may more efficiently compute a solution to the optimization problem.

[0138] Figures 11B - 11D illustrate additional steps of method 900 that may be performed in some examples. In step 916 of FIG. 11B, method 900 may further include computing the number of layers 744 of the surrogate cumulative distribution model 714. In step 918, method 900 may further include selecting one or more subsets 722 such that the simulation results 712 include simulations 711 included in the one or more subsets 722 that are evenly distributed across the plurality of layers 744. Thus, the compressed samples may more accurately model the variables associated with the surrogate cumulative distribution model 714 in the sparse regions of the CDF 708 and avoid high redundancy in the dense regions.

[0139] Figure 11C shows additional steps of a method 900 that may be performed when sampling a plurality of resampled simulations 730. In step 920, method 900 may further include executing an Iman-Conover algorithm 750. The Iman-Conover algorithm may be executed when selecting one or more resampled simulations 730 from simulation samples 710. In some examples, an initial simulation 711 may also be generated using the Iman-Conover algorithm 750. In step 922, method 900 may further include executing a quantum-inspired algorithm. The quantum-inspired algorithm may be a Markov chain Monte Carlo algorithm that may be, for example, simulated annealing, simulated quantum annealing, population annealing, or parallel tempering. For example, the Markov chain Monte Carlo algorithm may use the sum of one or more disagreement scores 726 as a loss function for which an estimated minimum value is computed.

[0140] FIG. 11D shows additional steps of a method 900 that can be performed in an example where the GUI 800 is presented to a user. At step 924, the method 900 may include generating a surrogate cumulative distribution model in response to receiving a surrogate model type selection 802 in the GUI 800. The surrogate model type selection 802 may specify a type of function for which a plurality of surrogate model parameters 718 are computed to generate the surrogate cumulative distribution model 714. At step 926, the method 900 may further include generating one or more subsets 722 of the plurality of simulations 711 in response to receiving a simulation generation command 804 in the GUI 800. The simulation generation command 804 may indicate, for example, the number of simulations 711 included in the compressed samples and the number of layers into which the range 746 of the surrogate cumulative distribution model is divided. At step 928, the method 900 may further include outputting to the GUI 800 one or more subsets 722 containing the simulations 711 after performing resampling.

[0141] By using the system and method described above, a compressed subset of Monte Carlo simulation inputs can be generated, thereby enabling the Monte Carlo simulation to be performed more efficiently without significantly sacrificing accuracy. For example, the variables shown in the simulation results included in the compressed subset can be aggregate values of dependent variables related to multiple different types of events. The aggregate values can be, for example, the aggregate loss by an insurance company, the aggregate quantity of product inventory, or the value of the aggregate quantity of generated energy. The system and method described above can thus facilitate the use of the Monte Carlo method for computing estimates of such quantities.

[0142] Section 4. Stratification of the Event-Driven Model

[0143] Referring now to FIG. 12, an example regarding the stratification of an event-driven model is disclosed. The event-driven model may include a plurality of different statistics. For example, in the scenario of "individual excess loss" reinsurance, an insurance company purchases insurance (referred to as "reinsurance") that compensates for payments from a threshold (also known as the "attachment point") up to a limit for individual claims or events (which may be referred to as a "layer" of reinsurance). In this scenario, both the insurance company and the party providing the reinsurance (the "reinsurance company") will be interested in the statistics regarding claims within this layer; these may be used by the reinsurance company to determine the price of compensation and by the insurance company to evaluate the value and the downside protection provided (and, for example, may be contrasted with the cost of holding additional capital to cover potential losses). In this setting, a model of individual claim amounts, rather than the total loss which is the sum over all claims, is required for analysis. In other examples related to reinsurance, in addition to deductions or limits on individual recoveries, various features may limit, for example, the total amount for which the reinsurance company is obligated to pay the insurance company (these amounts are referred to as "recoveries"), and in the model, multiple other contractual features may introduce highly non-linear functions and make it difficult to analytically calculate the statistics of interest, thus making Monte Carlo simulation frequently necessary. Further, these layers may be triggered only by rare events (especially when the attachment point is high), which means that Monte Carlo simulation is rather inefficient. Even if a large number of simulations (and thus high computational costs) are performed on the underlying claims, in many simulations the recovery amount in reinsurance will be zero, and in the few simulations showing a recovery amount, a sample of sufficient size may not be obtained for stable computation of the statistics of interest.If the statistical quantity of interest cannot be calculated with stable results, it may be inaccurate or lead to a noisy decision-making; for example, in the market, choices may be made for inaccurate or unplanned prices (resulting in loss-making or unprofitable enterprises, and ultimately, the risk of insolvency or regulatory intervention).

[0144] In contrast to the case where the number of variables is fixed and known and simulation results requiring stratification are given, in an event-driven model of a type called a frequency-severity model, the number of variables in each simulation changes according to a specified frequency distribution. Therefore, stratification needs to be done conditional on the number of variables, and thus it is possible to determine additional statistics to be stratified using information on the frequency distribution. This means that it is necessary to consider multiple conditional cumulative distribution functions corresponding to the distribution conditional on a given number of events. In the constraint of the sum of the discrepancy scores, a term is added for each of these conditional distributions.

[0145] Using the characteristics of the frequency distribution, it is determined how many terms to add (and this typically includes a cut-off value because many frequency distributions do not have a finite support, which may be based on a very low probability threshold and may be configured to select one simulation from each layer, and thus, based on the number of simulations, the related number of events to be conditioned is maximized).

[0146] Terms may be added to the sum of the discrepancy scores to represent the dependent variable of the simulation data such as the aggregated quantity in the frequency-severity model. The CDF of the dependent variable can be calculated using techniques such as the fast Fourier transform (which can be done using the slope).

[0147] To address these issues, examples related to stratification of an event-driven model are disclosed. Briefly, a conditional cumulative distribution model is approximated for multiple simulation results from multiple simulations. As described herein, the conditional cumulative distribution model refers to the distribution of quantities derived for each simulation within an input data set. The conditional cumulative distribution model is used to determine the sum of the disagreement scores for multiple simulations that satisfy the conditions associated with the conditional cumulative distribution model. Multiple simulation results within a selected simulation are replaced with another multiple simulation results from a resampled simulation based on a policy. An updated sum of the disagreement scores is generated for one or more remaining simulations and the resampled simulation. Thus, the results of one or more selected simulations are more representative of the conditional cumulative distribution model. The results of one or more selected simulations can be used as input to more general stratification and / or correlation methods (e.g., in the correlator 124 of FIG. 1A). Since the results of one or more selected simulations are more representative samples than the initial simulation, stratification of the event-driven model can increase the accuracy for downstream processing without requiring a larger sample size (e.g., improving the accuracy of the output data 116).

[0148] FIG. 12 shows an example of an arithmetic system 1002 for selecting a representative sequence of simulations based on stratification. In some examples, the arithmetic system 1002 embodies the arithmetic system 102 of FIG. 1A. In other examples, the arithmetic system 1002 is a separate arithmetic system.

[0149] The computing system 1002 is configured to receive a simulation sample 1004 that includes a plurality of simulations 1005. Each simulation may include simulation results 1006, and the amount of simulation results within one simulation may vary between simulations. The simulation results may be correlated. The simulation results may represent a dependent variable with respect to other simulation results. For example, Table 1 shows an example of a plurality of simulations of storm patterns, where one simulation result (the total amount of electricity generated by a wind power plant) is the aggregated value of other simulation results (the amount of electricity generated during individual storms).

[0150] In some examples, the computing system 1002 receives 50,000 to 250,000 samples obtained from Monte Carlo simulations. In other examples, and as will be described in more detail below, the simulation sample 1004 includes a smaller number of simulations (e.g., fewer than 50,000 simulations). As described above, a larger number of simulation results (e.g., 50,000 to 250,000) provides higher accuracy than a smaller number of simulation results. However, performing such a large number of simulations may result in high computational costs.

[0151] Continuing to refer to FIG. 12, each simulation 1005 includes a plurality of simulation results 1006. Table 1 shows a simplified example of 10 simulations of wind power generation (MWh) generated by a storm.

[0152] [Table 1]

[0153] Each row in Table 1 represents a simulation. Each simulation is tagged with the amount of events occurring in that simulation, which is represented in the column "Number of Events" in Table 1. The amount of events follows a discrete distribution function (e.g., Poisson distribution). Wind power generation is distributed according to a lognormal distribution, Pareto distribution, or beta distribution. The aggregated values across all events are shown in the column "Total Power". However, for each scenario where one event is modeled in total, the power generated by that event is relatively low compared to the maximum power generated by a single storm shown in Table 1 (e.g., 4355.40 MWh) and does not represent the overall distribution of wind power generation.

[0154] In other examples, at least a portion of Simulation 1005 includes empirical data (e.g., measured wind power generation obtained from actual storms). This can lead to a more realistic starting point through the optimization of simulation results because a randomly generated set of initial simulation results may contain unrealistically high (e.g., 10 GWh for one event) or low (e.g., 1 kwh for five events) data. As a result, the use of at least some empirical data can lead to the achievement of a more representative sample of simulation results in fewer iterations of the optimization process than the use of randomly selected Monte Carlo simulations.

[0155] The computing system 1002 is further configured to receive one or more cumulative distribution models 1008. These cumulative distribution models can be for dependent variables such as the power generated by individual storms shown in Table 1. In some examples, the cumulative distribution model 1008 is based on statistics such as "Total Power" shown in Table 1. In this example, the cumulative distribution model 1008 can be calculated based on the values of the variables from the simulation results in Table 1.

[0156] Other simulation results 1006 may be introduced to yield an additional cumulative distribution model 1008. For example, in a supply chain application, the inventory levels of sports equipment in a region are adjusted considering the sharp increase in demand before and after a sports event. For example, the sales volume of sports equipment in a region is associated with victory as described above. When the number of victories exceeds a threshold or the team advances to the postseason, the previously predicted inventory levels may no longer be accurate. Here, the number of victories may also represent simulation results in addition to or instead of the aggregated sales volume.

[0157] The computing system 1002 is configured to approximate a conditional cumulative distribution model 1012 related to the cumulative distribution model 1008 for multiple simulation results 1006 in multiple simulations. In some examples, the conditional cumulative distribution model 1012 for each of the multiple cumulative distribution models 1008 is conditioned on a predetermined amount of events occurring. In this way, the conditional cumulative distribution model 1012 reflects the simulation results 1006 based on a predetermined amount of correlated events (e.g., the average power generation when at least three storms occur in Table 1).

[0158] In some examples, the conditional cumulative distribution model 1012 is approximated using a Fourier transform. The approximation of the conditional cumulative distribution model 1012 may include numerical techniques such as slopes additionally or alternatively. Thereby, the computing system 1002 can quickly compute an accurate conditional cumulative distribution model for the target statistic.

[0159] In other examples, the conditional cumulative distribution model 1012 is approximated using a Monte Carlo method. In yet other examples, the conditional cumulative distribution model 1012 is approximated using a Fourier transform based on Monte Carlo simulation data. For example, while operating the conditional cumulative distribution model 1012 itself using a Fourier transform, a boundary can be brought about using a Monte Carlo method to estimate how the cumulative distribution model 1008 supports the conditional cumulative distribution model 1012. Although the approximation of the conditional cumulative distribution model 1012 becomes non-deterministic by using the Monte Carlo method, by using Monte Carlo-based support data, the computing system 1002 can compute the conditional cumulative distribution model 1012 more quickly and at a similar or higher accuracy level compared to using only deterministic methods.

[0160] Accordingly, in some examples, the cumulative distribution model 1008 and the conditional cumulative distribution model 1012 are stratified into multiple layers. For example, the simulation results 1006 can include values across all events that occur, which are reflected by the cumulative distribution model 1008. The conditional cumulative distribution model 1012 can reflect values that are conditional on the occurrence of a predetermined number of events, values that are conditional on the occurrence of at least a predetermined number of events, or the probability of values that are conditional on the occurrence of a predetermined number of events. As described in more detail below, by stratifying each of the cumulative distribution model 1008 and the conditional cumulative distribution model 1012 by the quantity of these events, the computing system 1002 can optimize the simulation results to obtain a representative sample that accurately reflects the number of events and the values associated with these events.

[0161] Using one or more cumulative distribution models and one or more conditional cumulative distribution models 1012 associated with each, a sum 1014 of the disagreement scores for a plurality of simulation results 1006 for each simulation 1004 is computed. For example, the sum of the disagreement scores 1014 can be determined as described above with reference to FIGS. 2-6. The disagreement score measures the deviation between the simulation results 1006 and the conditional cumulative distribution model 1012.

[0162] As described in more detail below, the computing system 1002 is operably configured to perform one or more resampling iterations until it is determined that the sum of one or more respective mismatch scores meets the optimization threshold 1016. As described in more detail below, if a set of simulations meets the optimization threshold 1016, the computing system 1002 is configured to output an event-driven model that includes the selected simulation, as shown at 1018. In some examples, "meeting the optimization threshold" means that the sum of the mismatch scores 1014 is less than or equal to the optimization threshold 1016. In this way, the computing system 1002 ensures that a set of simulations closely approximates the conditional cumulative distribution model 1012.

[0163] If the sum of the mismatch scores 1014 does not meet the optimization threshold, the set of simulations is modified as shown at 1020. To generate the modified set of simulations 1020, the computing system 1002 is configured to replace one of the simulations 1004 with one or more resampled simulations. For example, FIG. 13 shows the simulation results 1004 associated with one limiting cumulative distribution model and the modified set of simulation results 1020 in the form of one-dimensional Monte Carlo simulation values 1022. In other examples, it will be understood that the Monte Carlo simulation values 1022 may have any other number of dimensions (e.g., 10 or more dimensions) to represent any suitable number of limiting cumulative distribution models. For simplicity, the set of simulation results 1004 and the modified set of simulation results 1020 shown in the example of FIG. 13 each include 20 simulation results 1022.

[0164] In the example shown in FIG. 13, one of the simulation results 1022A is removed. In some examples, the computing system 1002 of FIG. 12 randomly selects the simulation to be removed. This can help reduce simulation error through the statistical effect achieved via random sampling. In other examples, the computing system 1002 selects the simulation to be removed based on a determination that the simulation results 1022A included therein contribute to the sum of the disagreement scores 1014 being greater than a threshold (e.g., the optimization threshold 1016). In this way, the process of modifying the set of simulations 1004 is explicitly driven to reduce the sum of the disagreement scores 1014. In still other examples, the simulation to be removed is selected by the user in response to receiving, for example, a prompt from the computing system 1002 indicating that the optimization threshold 1016 is not met, such that a goal specified by the user is achieved (e.g., conforms to a distribution pattern not defined by the computing system 1002).

[0165] At least one resampled simulation result 1022B is added to the remaining values from the initial simulation results 1006 after the simulation results 1022A are removed. In the example of FIG. 13, one resampled simulation result 1022B is added for each removed simulation 1022A. It will also be understood that in other examples, any other suitable number of simulation results may be added or removed. In this way, the computing system 1002 can increase or decrease the sample size respectively to generate a representative sample.

[0166] In some examples, the resampled simulation results 1022B are included in the precomputed simulation 1024. For example, the precomputed simulation 1024 can be selected from a pool of precomputed Monte Carlo simulations that also include simulation 1004. As described above, through precomputation, the computing system 1002 can pre-estimate the cumulative distribution model 1008 for the aggregated values of all simulation results.

[0167] Similar to the selection of the removed simulation results 1022A, in some examples, the computing system 1002 generates or selects resampled simulations via a randomized process. In this way, other simulation results 1022B can reduce the simulation error in the modified data 1020 through random sampling. In other examples, other simulation results 1022B are explicitly selected by either the computing system 1002 or the user to drive the modified data 1020 towards the optimization threshold 1016.

[0168] The computing system 1002 is further configured to generate an updated sum 1026 of the disagreement scores for one or more remaining simulations and the resampled simulations 1022B. In some examples, the updated sum of the disagreement scores 1026 is generated as described above with reference to FIGS. 2-6. For a target statistic conditioned on a predetermined amount of events, each simulation result corresponding to the number of these events is replaced, and a new sum of the disagreement scores is computed based on the added and / or removed simulation results. Using the updated sum of the disagreement scores, it is determined how much the modified simulation 1020 deviates from the conditional marginal distribution 1012.

[0169] As described in more detail below, the computing system 1002 is configured to apply policy 1028 and adopt or reject the resampled simulation 1022B based on the updated sum of the disagreement scores 1026. For example, the computing system 1002 may reject the resampled simulation 1022B if the updated sum of the disagreement scores 1026 is greater than the initial sum of the disagreement scores 1014. The computing system 1002 may additionally or alternatively reject the removal of the simulation if the resampled simulation is rejected. In this way, the computing system 1022 is configured to drive the disagreement score towards the optimization threshold 1016.

[0170] In some examples, the policy 1028 is implemented in a Markov Chain Monte Carlo (MCMC) agent 1030. FIG. 14 shows a schematic diagram of an exemplary MCMC agent 1030 configured to evaluate the simulation results 1020 modified based on the policy 1028. The policy 1028 is used to evaluate the energy parameters of a set of simulations (e.g., the modified simulation 1020) at a temperature 1034 associated with an iteration of the resampling loop 1032. Updates to the simulation results (e.g., in the form of the removed simulation results 1022A and / or the added simulation results 1022B) are adopted or rejected based on the temperature 1034 (e.g., using the Metropolis-Hastings method). As described in more detail below, at higher temperatures 1034, the policy 1028 of the MCMC agent 1030 can explore the solution surface more, while at lower temperatures 1034, the policy 1028 is constrained to adopt modified simulations 1020 that reduce the updated sum 1026 of the disagreement scores.

[0171] During one or more iterations of the resampling loop 1032, the MCMC agent 1030 is configured to conditionally adopt a set of modified simulations 1020 having a higher evaluation cost than the previous pass of the resampling loop 1032, such that ease increases at higher temperatures 1034 and ease decreases at lower temperatures 1034.

[0172] In some examples, the computing system 1002 is configured to reduce the temperature parameter 1034 over a series of steps 1036. As the temperature 1034 decreases over successive passes through the resampling loop 1032, the policy 1028 becomes further constrained to seek lower-cost solutions and ultimately converges towards a local minimum on the solution surface. As a result, the updated sum 1026 of the disagreement scores is minimized.

[0173] The MCMC agent 1030 is used to repeatedly evaluate the modified simulation 1020 (including one or more remaining simulations) when the simulation is replaced and to adjust the temperature parameter 1034. FIG. 14 shows an example of the formulation of a policy 1028 where the modified simulation is adopted when δE < 0, as shown at 1038. The modified simulation can be conditionally adopted by the MCMC agent 1030 when δE > 0, as shown at 1040. In other examples, the modified simulation 1020 is rejected by the MCMC agent 1028 when δE > 0, as shown at 1042. For example, a set of modified simulations that were conditionally adopted at a higher temperature can be rejected at a lower temperature. As a result, the MCMC agent 1030 outputs the corresponding status update 1044 to the computing system 1002. The update 1044 includes a simulation data structure update 1046 indicating whether the modified simulation 1020 was adopted or rejected, and an annealing temperature update 1048. In this way, the computing system 1002 can proceed to complete the optimization process (e.g., if the optimization threshold 1016 is met) or move to another iteration of the resampling loop 1032.

[0174] For each step 1036 passing through the resampling loop 1032, the value of the temperature parameter 1034 is determined by the MCMC agent 1030 according to a temperature function (e.g., the temperature function of the quantum inspired algorithm as described in more detail below) that tends to decrease over time. In this example, the value of the temperature (K) for the first step I through the optimization loop is set to 5, the value for the second step II (K) is set to 3, the value for the third step III (K) is set to 2, and the value for the fourth step IV (K) is set to 1. The selected set of the resampled simulations 1020 modified from the computing system 1002 is conditionally adopted by the MCMC agent 1030 with the value of the temperature parameter for each optimization loop. For example, in the first step I through the optimization loop, as the updated sum of the mismatch scores 1026 decreases, the modified simulation 1020 is adopted unconditionally until a local minimum is reached. On the other hand, the modified simulation 1020 is conditionally adopted according to the value of the temperature (K = 5) after the local minimum as the updated sum of the mismatch scores 1026 increases. In the second step II, the updated sum of the mismatch scores 1026 decreases to a local minimum and then increases with the value of the temperature (K = 3). However, the increase in the updated sum of the mismatch scores 1026 in the second step II is smaller than the increase in the updated sum of the mismatch scores 1026 in the first step I because the value of the temperature (K) has decreased from 5 to 3. Similarly, the increase in the third step III is smaller than the increase in the second step II, and the increase in the fourth step IV is smaller than the increase in the third step III. At the end of the fourth step IV, an estimate of the solution, which is the lowest point of the updated sum of the mismatch scores 1026, is determined. By this method, an optimized solution related to the set of modified simulations 1020 having the smallest updated sum of the mismatch scores 1026 can be calculated with reasonable accuracy in an efficient number of optimization steps.

[0175] In some examples, policy 1028 is adjusted according to the temperature parameter of quantum-inspired algorithm 1050 to transition from a first optimization threshold to a second updated optimization threshold. As used herein, the term "quantum-inspired algorithm" refers to an algorithm executed on conventional computing hardware that emulates one or more features of quantum mechanics for computational advantages. In particular, a quantum-inspired optimization algorithm emulates quantum tunneling, which is an effect that provides an advantage over adiabatic quantum optimization algorithms executed on quantum computers. Annealing algorithms are commonly included in quantum-inspired algorithms, and the additional randomness whose intensity is controlled by a temperature that decreases as the algorithm progresses provides additional computational advantages and is widely utilized by practitioners in the art. Some examples of quantum-inspired algorithms include, but are not limited to, Quantum Monte Carlo, Substochastic Monte Carlo, population annealing, and parallel tempering. In some examples, quantum-inspired algorithm 1050 is at least partially implemented on a classical computing device that simulates quantum behavior. In other examples, quantum-inspired algorithm 1050 is at least partially implemented on a quantum computer. Quantum-inspired algorithms provide the ability to escape local minima of the solution surface through a tunneling-like effect. Thus, quantum-inspired algorithms can enable computational system 1002 to explore the solution surface more efficiently than classical annealing and can prevent modified simulation 1020 from being trapped in local minima.

[0176] The calculation system 1002 is further configured to output an event-driven model 1018 including the selected simulation results. Table 2 shows an output example of the calculation system 1002 for the scenario presented with reference to Table 1. Table 2 shows 10 selected simulations related to wind power generation (MWh) generated by a storm. Compared with Table 1, the output values in Table 2 represent more the overall distribution of wind power generation that can be brought about by a given number of storms.

[0177]

Table 2

[0178] As described above, in some examples, the calculation system 1002 is configured to output an event-driven model 1018 (e.g., Table 2) to the correlator 124 in FIG. 1A. For example, the output result can be the first input data 104, the second input data 106, and / or the third input data 108 in FIG. 1A. Since the output result 1018 is a more representative sample of correlated events than the initial simulation data 1004, by using the output result 1018 as an input to the correlator 124, the accuracy of the output data 116 can be improved without requiring a larger sample size.

[0179] Referring now to FIGS. 15A - 15B, a flowchart showing an exemplary method 1300 for stratifying an event-driven model is shown. The following description of method 1300 is provided with reference to the above software and hardware components and is shown in FIGS. 1 - 14 and FIG. 16, and the method steps in method 1300 are described with reference to the corresponding parts of FIGS. 1 - 14 and FIG. 16 below. It will be recognized that method 1300 can also be implemented in other contexts using other suitable hardware and software components.

[0180] The following description of method 1300 is provided as an example and is not intended to be limiting. It will be recognized that various steps of method 1300 may be omitted or performed in an order different from that described, and that methods 15A - 15B may include additional and / or alternative steps compared to those shown in FIGS. 15A - 15B without departing from the scope of the present disclosure.

[0181] Referring first to FIG. 15A, at 1302, method 1300 includes receiving a plurality of simulations. Each simulation includes a plurality of simulation results. Method 1300 also includes receiving a discrete distribution function and one or more cumulative distribution models. For example, computing system 1002 is configured to receive a simulation sample 1004 that includes a plurality of simulations 1005 and one or more cumulative distribution models 1008.

[0182] In some examples, at 1304, method 1300 includes computing a discrete distribution function based on the number of simulation results included in each simulation. For example, computing system 1002 may compute a discrete distribution function for a range of the number of storms as shown in Tables 1 and 2. For example, at 1306, the discrete distribution function may form a Poisson distribution. By stratifying the CDF by the amount of these events, the computing system can optimize the simulation results to be representative samples for target statistics.

[0183] At 1308, method 1300 includes generating one or more conditional cumulative distribution models based at least in part on the discrete distribution function and one or more cumulative distribution models. For example, computing system 1002 is configured to generate one or more conditional cumulative distribution models 1012. By approximating the conditional marginal distribution, the computing device can generate representative samples of simulation result values that conform to that distribution.

[0184] In some examples, at 1310, the conditional cumulative distribution model is conditioned on a predetermined quantity of simulation results. For example, the conditional cumulative distribution model 1012 may reflect simulation results 1006 based on a predetermined amount of correlated events (e.g., in Table 1, the average power generation when at least three storms occur).

[0185] At 1312, in some examples, the approximation of the conditional cumulative distribution model includes one or more uses of Fourier transform, Monte Carlo method, or Fourier transform based on Monte Carlo simulation data. For example, while computing the conditional cumulative distribution model 1012 itself using Fourier transform, the Monte Carlo method can be used to yield a boundary for estimating how the cumulative distribution model 1008 supports the conditional cumulative distribution model 1012. Although the Monte Carlo method is not deterministic, by using Monte Carlo-based support data, the computing system 1002 can compute the conditional cumulative distribution model 1012 more quickly and at a similar or higher accuracy level compared to using only deterministic methods.

[0186] The method 1300 further includes, at 1314, stratifying the ranges of one or more cumulative distribution models and one or more conditional cumulative distribution models into a plurality of layers. For example, the computing system 1002 is configured to stratify the cumulative distribution model 1008 and the conditional cumulative distribution model 1012. This enables the computing system to optimize the simulation to generate representative samples for target statistics, which can improve the accuracy of the target statistics for downstream processing.

[0187] At 1316, method 1300 includes computing the sum of the disagreement scores for a plurality of simulations, based at least in part on a cumulative distribution model and a conditional cumulative distribution model. For example, computing system 1002 is configured to determine the sum of the disagreement scores 1014 for simulation results 1006 using cumulative distribution model 1008 and conditional cumulative distribution model 1012. As described above, the disagreement score measures the deviation between the simulation results 1006 and the conditional cumulative distribution model 1012.

[0188] Method 1300 further includes, at 1318, one or more resampling iterations. Each of the one or more resampling iterations may be performed until the sum of the respective disagreement scores (e.g., the sum of the disagreement scores 1014 or the updated sum of the disagreement scores 1026) is determined to satisfy an optimization threshold (e.g., optimization threshold 1016).

[0189] In each of the one or more resampling iterations, method 1300 includes, at 1320, generating one or more resampled simulations based at least in part on one or more cumulative distribution models. For example, at 1322, the resampled simulations may be derived from pre-computed simulations.

[0190] Each of the one or more resampling iterations further includes, at 1324, replacing one or more of the plurality of simulations with one or more resampled simulations, based on a policy. For example, FIG. 13 shows an example where one simulation result 1022A is replaced with a resampled simulation result 1022B to generate a set of modified simulation results 1020. This may help reduce simulation error via random sampling or via an explicit selection to remove outliers.

[0191] In some examples, at 1326, the policy is implemented in a Markov Chain Monte Carlo (MCMC) agent. For example, the policy 1028 of FIG. 12 is implemented by the MCMC agent 1030. The MCMC agent 1030 is configured to execute the policy and adopt or reject the resampled simulation results based on the energy parameters for the modified set of simulation results. Thereby, the computing system can optimize the updated sum of the disagreement scores.

[0192] At 1328, each of the one or more resampling iterations further includes generating an updated sum of the disagreement scores for a plurality of simulations in which one or more simulations are replaced with one or more resampled simulations. For example, the computing system 1002 is configured to generate an updated sum of the disagreement score 1026 for the modified simulation result 1020. This updated sum of the disagreement scores is compared with an optimization threshold to determine whether to perform another resampling iteration.

[0193] In some examples, at 1330, the resampling iteration is defined by a quantum inspired algorithm. For example, at 1332, the quantum inspired algorithm may include Quantum Monte Carlo, Substochastic Monte Carlo, population annealing, or parallel tempering algorithms. These quantum inspired algorithms can prevent the modified simulation results from being trapped in local minima of the solution surface and can enable the computing system to explore solutions more efficiently than other algorithms such as classical annealing.

[0194] In 1334, method 1300 includes outputting a plurality of simulations after performing one or more resampling iterations. For example, computing system 1002 is configured to output a plurality of simulations 1018 based on modified simulation results 1020 that satisfy an optimization threshold 1016. Because the output results are based on an optimized set of modified simulation results, the output event-driven model can have at least the same accuracy as results obtained using a substantially larger (e.g., at least 10 to 100 times larger) set of Monte Carlo simulations.

[0195] An event-driven model can be stratified using the above-described system and method. For example, the ranges of one or more cumulative distribution models and one or more conditional cumulative distribution models are stratified. By stratifying the one or more cumulative distribution models and the one or more conditional cumulative distribution models, the computing system can generate representative samples of the simulation. During an iterative resampling process, one or more simulations are replaced with one or more resampled simulations based on a policy. This policy allows the computing system to optimize the modified simulation results (e.g., by minimizing a disagreement score) such that the event-driven model based on one or more selected simulation result values represents the cumulative distribution model and the conditional cumulative distribution model better than the event-driven model based on the initial simulation results. In some examples, the resampling iteration is defined by a quantum-inspired algorithm. This allows the computing system to explore the solution surface more efficiently than other algorithms such as classical annealing while also preventing the modified simulation results from being trapped at a local minimum of the solution surface. The computing system is configured to output a plurality of simulations after performing one or more resampling iterations. As a result of the above-described system and method, the output simulations can be more representative samples than the initial simulation result values. This can also improve the accuracy of downstream processing (e.g., that in the correlator 124) without requiring a larger sample size for Monte Carlo simulation.

[0196] In some embodiments, the methods and processes described herein can be coupled to a computing system of one or more computing devices. In particular, such methods and processes can be implemented as a computer-application program or service, an application-program interface (API), a library, and / or other computer-program products.

[0197] FIG. 16 schematically shows a non-limiting embodiment of an arithmetic system 1400 that can implement one or more of the above methods and processes. The arithmetic system 1400 is shown in a simplified form. The arithmetic system 1400 may embody the above arithmetic system 102 illustrated in FIG. 1. The components of the arithmetic system 1400 may be embodied in one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, video game devices, mobile computing devices, mobile communication devices (e.g., smartphones) and / or other computing devices, and wearable computing devices such as smart wristwatches and head-mounted augmented reality devices.

[0198] The arithmetic system 1400 includes a logic processor 1402, a volatile memory 1404, and a non-volatile storage device 1406. The arithmetic system 1400 may optionally include a display subsystem 1408, an input subsystem 1410, a communication subsystem 1412, and / or other components not shown in FIG. 16.

[0199] The logic processor 1402 includes one or more physical devices configured to execute instructions. For example, the logic processor may be configured to execute instructions that are part of one or more of an application, program, routine, library, object, component, data structure, or other logical construct. Such instructions may be implemented to perform a task, implement a data type, transform the state of one or more components, achieve a technical effect, or reach a desired result in other ways.

[0200] A logical processor may include one or more physical processors (hardware) configured to execute software instructions. Additionally, or alternatively, a logical processor may include one or more hardware logic circuits or firmware devices configured to execute logic or firmware instructions implemented in hardware. The processor of the logical processor 1402 may be single-core or multi-core, and the instructions executed thereby may be configured for sequential processing, parallel processing, and / or distributed processing. The individual components of the logical processor may optionally be remotely located and / or distributed among two or more separate devices configured for cooperative processing. Aspects of the logical processor may be virtualized and may be executed by remotely accessible network-connected computing devices configured in a cloud-computing configuration. In such cases, it will be understood that these virtualized aspects are executed on different physical logical processors of various different machines.

[0201] The volatile memory 1404 may include a physical device including random access memory. The volatile memory 1404 is typically utilized by the logical processor 1402 to temporarily store information during the processing of software instructions. It will be recognized that the volatile memory 1404 typically does not continue to store instructions when the power to the volatile memory 1404 is turned off.

[0202] The non-volatile storage device 1406 includes one or more physical devices configured to hold instructions executable by the logical processor to implement the methods and processes described herein. When such methods and processes are implemented, the state of the non-volatile storage device 1406 may change, for example, to hold different data.

[0203] The non-volatile memory device 1406 may include a physical device that is removable and / or embedded. The non-volatile memory device 1406 may include optical memory (e.g., CD, DVD, HD-DVD, Blu-Ray Disc, etc.), semiconductor memory (e.g., ROM, EPROM, EEPROM, FLASH memory, etc.), and / or magnetic memory (e.g., hard disk drive, floppy disk drive, tape drive, MRAM, etc.), or other mass storage device technologies. The non-volatile memory device 1406 may include devices that are non-volatile, dynamic, static, read-writeable, read-only, sequentially accessible, location-addressable, file-addressable, and / or content-addressable. It will be appreciated that the non-volatile memory device 1406 may be configured to retain instructions even when the power to the non-volatile memory device 1406 is turned off.

[0204] Aspects of the logical processor 1402, volatile memory 1404, and non-volatile memory device 1406 may be integrated into one or more hardware-logic components. Such hardware-logic components may include, for example, field-programmable gate arrays (FPGA), program- and application-specific integrated circuits (PASIC / ASIC), program- and application-specific standard products (PSSP / ASSP), system-on-chips (SOC), and complex programmable logic devices (CPLD).

[0205] The terms "module", "program", and "engine" can be used to describe aspects of the computing system 1400 typically implemented in software by a processor to execute certain functions using a portion of volatile memory, which functions include transformation processes that specially configure the processor to execute this function. Thus, a module, program, or engine can be instantiated using a portion of volatile memory 1404 via a logic processor 1402 that executes instructions held by a non-volatile storage device 1406. It will be appreciated that different modules, programs, and / or engines can be instantiated from the same application, service, code block, object, library, routine, API, function, etc. Similarly, the same module, program, and / or engine can be instantiated by different applications, services, code blocks, objects, routines, APIs, functions, etc. The terms "module", "program", and "engine" can include individual or multiple ones of executable files, data files, libraries, drivers, scripts, database records, etc.

[0206] When included, the display subsystem 1408 can be used to represent a visual representation of data held by the non-volatile storage device 1406. The visual representation can be in the form of a graphical user interface (GUI). The methods and processes described herein change data held by the non-volatile storage device and thus transform the state of the non-volatile storage device, and accordingly, the state of the display subsystem 1408 can also be transformed to visually represent the underlying data change. The display subsystem 1408 can include one or more display devices that utilize virtually any type of technology. Such display devices may be combined with the logic processor 1402, volatile memory 1404, and / or non-volatile storage device 1406 in a shared housing, or such display devices may be peripheral display devices.

[0207] When included, the input subsystem 1410 may include or interface with one or more user-input devices such as a keyboard, mouse, touch screen, or game controller. In some embodiments, the input subsystem may include or interface with selected natural user input (NUI) components. Such components may be built-in or peripheral, and the conversion and / or processing of input actions may be performed on-board or off-board. Exemplary NUI components may include a microphone for speech and / or voice recognition; an infrared camera, color camera, stereo camera, and / or depth camera for machine vision and / or gesture recognition; a head tracker, eye tracker, accelerometer, and / or gyroscope for motion detection and / or intent recognition; and / or an electric field sensing component for evaluating brain activity; and / or any other suitable sensor.

[0208] When included, the communication subsystem 1412 may be configured to communicatively couple the various computing devices described herein to each other and to other devices. The communication subsystem 1412 may include wired and / or wireless communication devices that are compatible with one or more different communication protocols. By way of non-limiting example, the communication subsystem may be configured for communication via a wireless telephone network or a wired or wireless local or wide area network such as an HDMI (registered trademark) connection via Wi-Fi. In some embodiments, the communication subsystem may enable the computing system 1400 to send and / or receive messages to and from other devices via a network such as the Internet.

[0209] In the following paragraphs, multiple aspects of the present disclosure are described. One aspect is a processor that, for a plurality of correlation variables, performs a first predetermined number of simulations from Monte Carlo simulation samples, where each simulation includes a plurality of initial simulation results for a plurality of variables, one or more target statistics, a second predetermined number of layers for each variable and each target statistic, and cumulative distribution functions for each variable and each target statistic; receives the cumulative distribution functions; for each variable and for each of the one or more target statistics, divides the unit interval of the cumulative distribution function into the second predetermined number of layers, and divides the support of the cumulative distribution function into a plurality of bins such that each bin of the plurality of bins corresponds to one of the layers, determines an initial disagreement score based on the amount of values in each bin, the first predetermined number of simulations, and the second predetermined number of layers for the variable; determines an initial sum of the initial disagreement scores; based on the determination that the initial sum of the initial disagreement scores is not within an optimization threshold, removes at least one of the plurality of initial simulations; adds at least one other simulation to the remaining one or more initial simulations; generates an updated disagreement score using the amount of values in one or more bins corresponding to at least one of the plurality of initial simulations and the amount of values in one or more bins corresponding to the at least one other simulation for each variable and for each of the one or more target statistics; generates an updated sum of the updated disagreement scores; and outputs a plurality of representative simulations representing the cumulative distribution function across the layers based on the updated sum of the updated disagreement scores. A potential technical advantage of such a configuration is that a representative sequence of simulations is selected from a pool of Monte Carlo simulation data.

[0210] In addition to this aspect, in some examples, the processor is additionally or alternatively configured to adopt or reject at least one other simulation based on the updated sum of the updated disagreement scores. A potential technical advantage of such a configuration is that the selection of representative simulations is driven towards an optimization threshold.

[0211] In addition to this aspect, in some examples, the processor is additionally or alternatively configured to determine a layer of values for each initial simulation result for each variable and for each of one or more target statistics, and to place the values into one of a plurality of bins based on the layer. A potential technical advantage of such a configuration is that the arithmetic system can identify the bin into which the values are placed.

[0212] In addition to this aspect, in some examples, at least one other simulation additionally or alternatively includes pre-computed simulations. A potential technical advantage of such a configuration is that the CDF for the aggregate value of all pre-computed simulations can be pre-estimated.

[0213] In addition to this aspect, in some examples, the processor is additionally or alternatively configured to determine an initial bin-by-bin disagreement metric for each bin of the plurality of bins. A potential technical advantage of such a configuration is that the arithmetic system can measure the uniformity of the initial simulation results across the plurality of bins.

[0214] In addition to this aspect, in some examples, the initial bin-by-bin disagreement metric for the selected bin additionally or alternatively includes the difference between the amount of values in the selected bin and the value obtained by dividing the first predetermined number of simulations by the second predetermined number of layers for the variable or the target statistic. A potential technical advantage of such a configuration is that the initial bin-by-bin disagreement metric can be computed using arithmetic operations.

[0215] In addition to this aspect, in some examples, the initial disagreement score additionally or alternatively includes the maximum bin-by-bin disagreement metric or the sum of the initial bin-by-bin disagreement metrics for a plurality of bins. A potential technical advantage of such a configuration is that the initial disagreement score represents the maximum disagreement between the initial Monte Carlo simulation data and the CDF or aggregated disagreement for the initial Monte Carlo simulation data.

[0216] In addition to this aspect, in some examples, the processor is configured to additionally or alternatively weight the initial bin-by-bin disagreement metric for each bin of the plurality of bins based on the proximity to the tail of the cumulative distribution function of the layer corresponding to the bin. A potential technical advantage of such a configuration is that the initial disagreement score places more emphasis on the accuracy in the tail than in other regions of the distribution.

[0217] In addition to this aspect, in some examples, the processor is configured to additionally or alternatively: determine updated bin-by-bin disagreement metrics for one or more bins corresponding to at least one of the plurality of initial simulations and one or more bins corresponding to at least one other simulation; and determine an updated disagreement score using the updated bin-by-bin disagreement metrics for one or more bins corresponding to at least one of the plurality of initial simulations and one or more bins corresponding to at least one other simulation and the initial bin-by-bin disagreement metrics for each of the remaining one or more initial simulations. A potential technical advantage of such a configuration is that this formulation of the updated disagreement score does not require an arithmetic system to recalculate the initial bin-by-bin disagreement metric for each bin.

[0218] In addition to this aspect, in some examples, the processor is further configured, additionally or alternatively, for each variable and target statistic to: decrement the amount of values in one or more bins corresponding to at least one of a plurality of initial simulations; and increment the amount of values in one or more bins corresponding to at least one other simulation. A potential technical advantage of such a configuration is that the updated disagreement score can be determined using the same operations as the initial disagreement score.

[0219] In another aspect, in a computing device: for a plurality of correlation variables, a first predetermined number of simulations from Monte Carlo simulation samples, where each simulation includes a plurality of initial simulation results for a plurality of variables, one or more target statistics, a second predetermined number of layers for each variable and each target statistic, and receiving cumulative distribution functions for each variable and each target statistic; for each variable and for each of the one or more target statistics, dividing the unit interval of the cumulative distribution function by the second predetermined number of layers, and dividing the support of the cumulative distribution function into a plurality of bins such that each bin of the plurality of bins corresponds to one of the layers, determining an initial disagreement score based on the amount of values in each bin, the first predetermined number of simulations, and the second predetermined number of layers for the variable, and determining an initial sum of the initial disagreement scores; based on a determination that the initial sum of the initial disagreement scores is not within an optimization threshold, removing at least one of the plurality of initial simulations; adding at least one other simulation to the remaining one or more initial simulations; for each variable and for each of the one or more target statistics, generating an updated disagreement score using the amount of values in one or more bins corresponding to at least one of the plurality of initial simulation results and the amount of values in one or more bins corresponding to at least one other simulation result; determining an updated sum of the updated disagreement scores; and based on the updated sum of the updated disagreement scores, outputting a plurality of representative simulations representing the cumulative distribution function across the layers. A potential technical advantage of such a configuration is that a representative sequence of simulation results is selected from a pool of Monte Carlo simulation data.

[0220] In addition to this aspect, in some examples, this method further includes, additionally or alternatively, adopting or rejecting at least one other simulation based on the updated sum of the updated disagreement scores. A potential technical advantage of such a configuration is that the selection of representative simulation results is driven towards an optimization threshold.

[0221] In addition to this aspect, in some examples, at least one other simulation further includes, additionally or alternatively, pre-computed simulations. A potential technical advantage of such a configuration is that the cumulative distribution function (CDF) for the aggregated values of all pre-computed simulation results can be estimated in advance.

[0222] In addition to this aspect, in some examples, determining the initial disagreement score further includes, additionally or alternatively, determining the initial bin-by-bin disagreement metric for each of a plurality of bins. A potential technical advantage of such a configuration is that the initial bin-by-bin disagreement metric measures the uniformity of the initial simulation results across the plurality of bins.

[0223] In addition to this aspect, in some examples, determining the initial bin-by-bin disagreement metric for a selected bin further includes, additionally or alternatively, determining the difference between the amount of values in the selected bin and the value obtained by dividing the first predetermined number of simulations by the second predetermined number of layers for the variable or the target statistic. A potential technical advantage of such a configuration is that the initial bin-by-bin disagreement metric can be computed using arithmetic operations.

[0224] In addition to this aspect, in some examples, determining the initial bin - by - bin mismatch metric for a selected bin may additionally or alternatively include determining the maximum bin - by - bin mismatch metric, or determining the sum of the initial bin - by - bin mismatch metrics for a plurality of bins. A potential technical advantage of such a configuration is that the initial mismatch score represents the maximum mismatch between the initial Monte Carlo simulation data and the CDF or aggregated mismatch for the initial Monte Carlo simulation data.

[0225] In addition to this aspect, in some examples, determining the initial mismatch score may additionally or alternatively include weighting the initial bin - by - bin mismatch metric for each bin of a plurality of bins based on proximity to the tail of the cumulative distribution function of the layer corresponding to the bin. A potential technical advantage of such a configuration is that the initial mismatch score places more emphasis on accuracy in the tail than in other regions of the distribution.

[0226] In addition to this aspect, in some examples, generating an updated mismatch score may additionally or alternatively include determining updated bin - by - bin mismatch metrics for one or more bins corresponding to at least one of a plurality of initial simulation results and one or more bins corresponding to at least one other simulation; and determining the updated mismatch score using the updated bin - by - bin mismatch metrics for one or more bins corresponding to at least one of a plurality of initial simulation results and one or more bins corresponding to at least one other simulation result, and the initial bin - by - bin mismatch metrics for each of the remaining one or more initial simulations. A potential technical advantage of such a configuration is that this formulation of the updated mismatch score does not require the initial bin - by - bin mismatch metrics to be recalculated for each bin.

[0227] In addition to this aspect, in some examples, generating an updated disagreement score may additionally or alternatively include decrementing the amount of values in one or more bins corresponding to at least one of a plurality of initial simulations; and incrementing the amount of values in one or more bins corresponding to at least one other simulation. A potential technical advantage of such a configuration is that the updated disagreement score can be determined using the same operations as the initial disagreement score.

[0228] Another aspect is: for a plurality of correlation variables, a first predetermined number of simulations from Monte Carlo simulation samples where each simulation includes a plurality of initial simulation results for a plurality of variables, one or more target statistics, a second predetermined number of layers for each variable and each target statistic, and a cumulative distribution function for each variable and each target statistic are received; for each variable and for each of the one or more target statistics, the unit interval of the cumulative distribution function is divided into the second predetermined number of layers, and the support of the cumulative distribution function is divided into a plurality of bins such that each bin of the plurality of bins corresponds to one of the layers, the amount of values in each bin of the plurality of bins is counted, and an initial disagreement score is determined based on the difference between the amount of values in each bin and the value obtained by dividing the amount of the initial simulations by the second predetermined number of layers for the variable or the target statistic; an initial sum of the initial disagreement scores is determined; based on the determination that the initial sum of the initial disagreement scores is not within an optimization threshold, at least one of the plurality of initial simulation results is removed; at least one other simulation is added to the remaining one or more initial simulations; for each variable and target statistic, the amount of values in one or more bins corresponding to at least one of the plurality of initial simulation results is decremented, the amount of values in one or more bins corresponding to at least one other simulation result is incremented, and an updated disagreement score is generated using the amount of values in one or more bins corresponding to at least one of the plurality of initial simulation results and the amount of values in one or more bins corresponding to at least one other simulation result; an updated sum of the updated disagreement scores is generated; and an arithmetic system is provided comprising a processor configured to output a plurality of representative simulations representing the cumulative distribution function across the layers based on the updated sum of the updated disagreement scores. A potential technical advantage of such a configuration is that a representative sequence of simulation results is selected from a pool of Monte Carlo simulation data.

[0229] According to one aspect of the present disclosure, there is provided an arithmetic system including a processor configured to receive a simulation sample including a plurality of simulations for a plurality of correlation probability variables. Each simulation may include a plurality of simulation results. The processor may further be configured to at least partially generate a surrogate cumulative distribution model by estimating a plurality of surrogate model parameters based at least in part on the plurality of simulation results. Based at least in part on the surrogate cumulative distribution model having the surrogate model parameters, the processor may further be configured to select one or more subsets of the plurality of simulations. In each of one or more resampling iterations, the processor may further be configured to calculate one or more disagreement scores for the one or more subsets until it is determined that the sum of the one or more disagreement scores for the one or more subsets satisfies an optimization threshold. In each of one or more resampling iterations, based at least in part on the sum of the one or more disagreement scores, the processor may further be configured to sample one or more resampled simulations for the plurality of correlation probability variables from a plurality of simulations included in the simulation sample and not yet included in the one or more subsets. In each of one or more resampling iterations, the processor may further be configured to replace one or more simulations included in the one or more subsets with the one or more resampled simulations. The processor may further be configured to output the simulations included in the one or more subsets after performing the one or more resampling iterations. A potential technical advantage of such a configuration is that one or more subsets can be compressed relative to the simulation sample, thereby enabling the Monte Carlo algorithm to be executed more efficiently when the one or more subsets are used as input.

[0230] According to this aspect, for multiple quantiles of multiple simulation results, the processor may further be configured to calculate multiple layers of the surrogate cumulative distribution model. The processor may further be configured to select one or more subsets of the simulations such that the simulation results included in the one or more subsets are uniformly distributed among the multiple layers. A potential technical advantage of such a configuration is that the distribution of the correlation probability variables in the sparse regions of the range of the simulation results can be accurately represented with a small number of simulations.

[0231] According to this aspect, the processor may be configured to at least partially replace one or more simulations with one or more resampled simulations by executing a quantum-inspired algorithm. A potential technical advantage of such a configuration is that the sum of one or more mismatch scores can be reduced such that it can quickly converge to a value less than the optimization threshold.

[0232] According to this aspect, the processor may be configured to at least partially generate multiple simulation results for multiple correlation probability variables by executing the Iman-Conover algorithm. A potential technical advantage of such a configuration is that the event simulation module can efficiently generate the simulation results.

[0233] According to this aspect, the processor may be configured to at least partially sample one or more resampled simulations by executing the Iman-Conover algorithm. A potential technical advantage of such a configuration is that the processor can efficiently resample the resampled simulations.

[0234] According to this aspect, the surrogate cumulative distribution model can be a mixture Erlang model including a plurality of Erlang distributions. A potential technical advantage associated with such a configuration is that the surrogate cumulative distribution model can accurately model the cumulative distribution function with a small number of parameters.

[0235] According to this aspect, the surrogate cumulative distribution model may further include one or more alternative tail region distributions configured to replace one or more respective tail regions of the plurality of Erlang distributions. The one or more alternative tail region distributions may be different from the one or more Erlang distributions in the one or more respective tail regions. A potential technical advantage associated with such a configuration is that the surrogate cumulative distribution model can more accurately model a heavy-tailed distribution or a light-tailed distribution.

[0236] According to this aspect, the surrogate cumulative distribution model can be an empirical model in which a processor is configured to estimate surrogate model parameters based at least in part on empirical data included in a plurality of simulations. A potential technical advantage associated with such a configuration is that the surrogate cumulative distribution model can accurately model empirical data.

[0237] According to this aspect, the processor can be configured to at least partially estimate a plurality of surrogate model parameters by performing iterative expectation maximization. A potential technical advantage associated with such a configuration is that the processor can set the values of the surrogate model parameters such that the surrogate cumulative distribution model accurately models the cumulative distribution function.

[0238] According to this aspect, the plurality of simulation results can include a plurality of aggregated values, minimum values, or maximum values across a plurality of correlated probability variables. A potential technical advantage associated with such a configuration is that it can model quantities that may be of high interest in fields such as energy production, inventory management, and insurance.

[0239] According to this aspect, the processor may further be configured to generate a surrogate cumulative distribution model in response to receiving a surrogate model type selection in a graphical user interface (GUI). The processor may further be configured to generate one or more subsets of a plurality of simulations in response to receiving a simulation generation command in the GUI. The processor may further be configured to output one or more subsets of the simulations to the GUI. A potential technical advantage of such a configuration is that the GUI may enable a user to specify properties of the surrogate cumulative distribution model and one or more subsets to display one or more subsets.

[0240] According to another aspect of the present disclosure, a method for use in a computing system is provided. The method may be implemented on a computer. The method may include receiving a simulation sample including a plurality of simulations for a plurality of correlation probability variables. Each simulation may include a plurality of simulation results. The method may further include at least partially generating a surrogate cumulative distribution model by estimating a plurality of surrogate model parameters based at least in part on the plurality of simulation results. Based at least in part on the surrogate cumulative distribution model having the surrogate model parameters, the method may further include selecting one or more subsets of the plurality of simulations. The method may further include, in each of one or more resampling iterations, computing one or more disagreement scores for the one or more subsets until it is determined that the sum of the one or more disagreement scores for the one or more subsets meets an optimization threshold. In each of the one or more resampling iterations, the method may further include sampling one or more resampled simulations for the plurality of correlation probability variables from a plurality of simulations included in the simulation sample and not yet included in the one or more subsets, based at least in part on the sum of the one or more disagreement scores. The method may further include, in each of the one or more resampling iterations, replacing one or more simulations included in the one or more subsets with the one or more resampled simulations. The method may further include outputting the simulations included in the one or more subsets after performing the one or more resampling iterations. A potential technical advantage of such a configuration is that one or more subsets may be compressed with respect to the simulation sample, thereby enabling a Monte Carlo algorithm to be executed more efficiently when the one or more subsets are used as input.

[0241] According to this aspect, the method may further include calculating a plurality of layers of the surrogate cumulative distribution model for a plurality of quantiles of a plurality of simulation results. The method may further include selecting one or more subsets of the simulations such that the simulation results included in the one or more subsets are uniformly distributed among the plurality of layers. A potential technical advantage of such a configuration is that the distribution of the correlation probability variables in the sparse regions of the range of the simulation results can be accurately represented with a small number of simulations.

[0242] According to this aspect, replacing one or more simulations with one or more resampled simulations may include executing a quantum-inspired algorithm. A potential technical advantage of such a configuration is that the sum of one or more mismatch scores can be reduced such that it can quickly converge to a value less than the optimization threshold.

[0243] According to this aspect, sampling one or more resampled simulations may further include executing the Iman-Conover algorithm. A potential technical advantage of such a configuration is that the resampled simulations can be efficiently resampled.

[0244] According to this aspect, the surrogate cumulative distribution model may be a mixture Erlang model including a plurality of Erlang distributions. A potential technical advantage of such a configuration is that the surrogate cumulative distribution model can accurately model the cumulative distribution function with a small number of parameters.

[0245] According to this aspect, the surrogate cumulative distribution model may be an empirical model in which surrogate model parameters are estimated based at least in part on empirical data included in a plurality of simulations. A potential technical advantage of such a configuration is that the surrogate cumulative distribution model can accurately model the empirical data.

[0246] According to this aspect, estimating a plurality of surrogate model parameters may include performing iterative expectation maximization. A potential technical advantage of such a configuration is that the processor can set the values of the surrogate model parameters such that the surrogate cumulative distribution model accurately models the cumulative distribution function.

[0247] According to this aspect, the plurality of simulation results may include a plurality of aggregated values, minimum values, or maximum values over a plurality of correlated probability variables. A potential technical advantage of such a configuration is that it can model quantities that may be of high interest in fields such as energy production, inventory management, and insurance.

[0248] According to another aspect of the present disclosure, there is provided an arithmetic system including a processor configured to receive a simulation sample including a plurality of simulations for a plurality of correlation probability variables. Each simulation may include a plurality of simulation results. Based at least in part on the plurality of simulation results, the processor may further be configured to generate a surrogate cumulative distribution model. Based at least in part on the surrogate cumulative distribution model, the processor may further be configured to select a compressed subset of the plurality of simulations. In each of one or more resampling iterations, the processor may further be configured to calculate a mismatch score of the compressed subset until it is determined that the mismatch score of the compressed subset is less than a predetermined mismatch threshold. In each of one or more resampling iterations, based at least in part on the mismatch score, the processor may further be configured to sample one or more resampled simulations for the plurality of correlation probability variables from a plurality of simulations included in the simulation sample and not yet included in the subset. In each of one or more resampling iterations, the processor may further be configured to replace one or more simulations included in the compressed subset with one or more resampled simulations. The processor may further be configured to output the simulations included in the compressed subset after performing one or more resampling iterations. A potential technical advantage of such a configuration is that one or more subsets can be compressed with respect to the simulation sample, whereby the Monte Carlo algorithm can be executed more efficiently when one or more subsets are used as input.

[0249] Another aspect is as follows: Each simulation receives a plurality of simulations including a plurality of simulation results, a discrete distribution function, and one or more cumulative distribution models; generates one or more conditional cumulative distribution models based at least in part on the discrete distribution function and the one or more cumulative distribution models; stratifies the ranges of the one or more cumulative distribution models and the one or more conditional cumulative distribution models into a plurality of layers; calculates the sum of the disagreement scores for the plurality of simulations based at least in part on the cumulative distribution models and the conditional cumulative distribution models; in each of the one or more resampling iterations, until it is determined that the sum of the respective disagreement scores of the one or more satisfies an optimization threshold: generates one or more resampled simulations based at least in part on the one or more cumulative distribution models; replaces one or more of the plurality of simulations with one or more of the resampled simulations based on a policy, and generates an updated sum of the disagreement scores for the plurality of simulations in which one or more of the simulations are replaced with one or more of the resampled simulations; and after performing the one or more resampling iterations, provides an arithmetic system comprising a processor configured to output the plurality of simulations. A potential technical advantage of such a configuration is that a set of modified simulation results is generated that represents one or more cumulative distribution models and one or more conditional cumulative distribution models, rather than the initial simulation results.

[0250] In addition to this aspect, in some examples, the discrete distribution function is additionally or alternatively calculated based on the number of simulation results included in each simulation. A potential technical advantage of such a configuration is that representative samples of Monte Carlo simulations are generated that represent conditional marginal distributions rather than the initial simulation results.

[0251] In addition to this aspect, in some examples, the discrete distribution function additionally or alternatively forms a Poisson distribution. A potential technical advantage associated with such a configuration is that the Poisson distribution represents the distribution of the discrete amount of events included in each simulation.

[0252] In addition to this aspect, in some examples, the conditional marginal distribution for each of a plurality of target statistics is additionally or alternatively conditioned on a predetermined quantity of simulation results. A potential technical advantage associated with such a configuration is that the conditional marginal distribution reflects the target statistics based on a predetermined amount of events.

[0253] In addition to this aspect, in some examples, the processor is configured to approximate the conditional marginal distribution using a Fourier transform, a Monte Carlo method, or a Fourier transform based on Monte Carlo simulation data, additionally or alternatively. A potential technical advantage associated with such a configuration is that an accurate conditional marginal distribution can be calculated quickly for the target statistics.

[0254] In addition to this aspect, in some examples, other simulation results are additionally or alternatively derived from pre-computed simulations. A potential technical advantage associated with such a configuration is that the statistical characteristics of the aggregated values of all pre-computed simulation results can be estimated in advance.

[0255] In addition to this aspect, in some examples, the acceptance / rejection policy is additionally or alternatively implemented in a Markov chain Monte Carlo agent. A potential technical advantage associated with such a configuration is that the updated sum of the disagreement scores can be optimized.

[0256] In addition to this aspect, in some examples, the resampling iteration is additionally or alternatively defined by a quantum-inspired algorithm. Potential technical advantages associated with such a configuration are that the solution surface can be explored more efficiently compared to classical annealing and that modified simulation results can be prevented from being trapped in local minima.

[0257] In addition to this aspect, in some examples, the quantum-inspired algorithm additionally or alternatively includes Quantum Monte Carlo, Substochastic Monte Carlo, population annealing, or parallel tempering algorithms. Potential technical advantages associated with such a configuration are that the solution surface can be explored more efficiently compared to classical annealing and that modified simulation results can be prevented from being trapped in local minima.

[0258] In addition to this aspect, in some examples, the processor is configured to apply a policy to minimize one or more respective mismatch scores. A potential technical advantage associated with such a configuration is that a set of simulations that accurately approximate the conditional cumulative distribution model is generated.

[0259] In another aspect, in an arithmetic system: receiving, in each simulation, a plurality of simulations including a plurality of simulation results, a discrete distribution function, and one or more cumulative distribution models; generating, at least partially based on the discrete distribution function and the one or more cumulative distribution models, one or more conditional cumulative distribution models; stratifying the ranges of the one or more cumulative distribution models and the one or more conditional cumulative distribution models into a plurality of layers; calculating, at least partially based on the cumulative distribution models and the conditional cumulative distribution models, the sum of the mismatch scores for the plurality of simulations; in each of one or more resampling iterations, generating, at least partially based on the one or more cumulative distribution models, one or more resampled simulations until it is determined that the sum of the respective mismatch scores of the one or more satisfies an optimization threshold; replacing, based on a policy, one or more of the plurality of simulations with the one or more resampled simulations; and generating an updated sum of the mismatch scores for the plurality of simulations in which one or more of the simulations have been replaced with the one or more resampled simulations; and after performing the one or more resampling iterations, outputting the plurality of simulations. A potential technical advantage of such a configuration is that a set of modified simulation results is generated that represents the one or more cumulative distribution models and the one or more conditional cumulative distribution models better than the initial simulation results.

[0260] In addition to this aspect, in some examples, the method additionally or alternatively includes calculating a discrete distribution function based on the number of simulation results included in each simulation. A potential technical advantage of such a configuration is that representative samples of Monte Carlo simulations are generated that represent the conditional marginal distribution better than the initial simulation results.

[0261] In addition to this aspect, in some examples, the discrete distribution function additionally or alternatively forms a Poisson distribution. A potential technical advantage of such a configuration is that the Poisson distribution represents the distribution of the discrete amount of events included in each simulation.

[0262] In addition to this aspect, in some examples, the conditional cumulative distribution model is additionally or alternatively conditioned on a predetermined quantity of simulation results. A potential technical advantage of such a configuration is that the conditional marginal distribution reflects the target statistic based on a predetermined amount of events.

[0263] In addition to this aspect, in some examples, this method additionally or alternatively includes approximating the conditional cumulative distribution model using a Fourier transform, a Monte Carlo method, or a Fourier transform based on Monte Carlo simulation data. A potential technical advantage of such a configuration is that the accurate conditional marginal distribution can be calculated quickly for the target statistic.

[0264] In addition to this aspect, in some examples, this method additionally or alternatively includes deriving resampled simulation results from pre-computed simulations. A potential technical advantage of such a configuration is that the statistical characteristics of the aggregate value of all pre-computed simulation results can be estimated in advance.

[0265] In addition to this aspect, in some examples, this method additionally or alternatively includes implementing a policy on a Markov chain Monte Carlo agent. A potential technical advantage of such a configuration is that the updated sum of the disagreement scores can be optimized.

[0266] In addition to this aspect, in some examples, the resampling iteration is additionally or alternatively defined by a quantum-inspired algorithm. Potential technical advantages associated with such a configuration are that the solution surface can be explored more efficiently compared to classical annealing and that modified simulation results can be prevented from being trapped in local minima.

[0267] In addition to this aspect, in some examples, the quantum-inspired algorithm additionally or alternatively includes Quantum Monte Carlo, Substochastic Monte Carlo, population annealing, or parallel tempering algorithms. Potential technical advantages associated with such a configuration are that the solution surface can be explored more efficiently compared to classical annealing and that modified simulation results can be prevented from being trapped in local minima.

[0268] Another aspect is: each simulation receives a plurality of simulations including a plurality of simulation results, a discrete distribution function, and one or more cumulative distribution models; generates one or more conditional cumulative distribution models based at least in part on the discrete distribution function and the one or more cumulative distribution models; stratifies the ranges of the one or more cumulative distribution models and the one or more conditional cumulative distribution models into a plurality of layers; calculates the sum of the disagreement scores for the plurality of simulations based at least in part on the cumulative distribution models and the conditional cumulative distribution models; in each of one or more resampling iterations defined by a quantum-inspired algorithm, until it is determined that the sum of the respective disagreement scores of the one or more satisfies an optimization threshold; generates one or more resampled simulations based at least in part on the one or more cumulative distribution models; replaces one or more of the plurality of simulations with one or more of the resampled simulations based on a policy configured to minimize each of the one or more respective disagreement scores, and generates an updated sum of the disagreement scores for the plurality of simulations in which one or more simulations are replaced with one or more resampled simulations; and provides an arithmetic system comprising a processor configured to output a plurality of simulations after performing one or more resampling iterations. A potential technical advantage of such a configuration is that a set of modified simulation results is generated that represents one or more cumulative distribution models and one or more conditional cumulative distribution models rather than the initial simulation results.

[0269] Features described in the context of separate aspects and embodiments of the present invention may be used together and / or interchangeably. Similarly, where features are described in the context of a single embodiment for the sake of brevity, these may also be provided individually or in any suitable sub-combination. Features described in relation to a system may have corresponding features definable in relation to a method, and vice versa, and these embodiments are specifically contemplated.

[0270] As used herein, "and / or" is defined as an inclusive logical disjunction or (∨) as defined by the following truth table.

[0271]

Table 3

[0272] It will be understood that the configurations and / or approaches described herein are of an illustrative nature, and these specific embodiments or examples are not to be construed in a limiting sense because numerous variations are possible. The specific routines or methods described herein may represent one or more of any number of processing strategies. Accordingly, the various acts illustrated and / or described may be performed in the order illustrated and / or described, in other orders, in parallel, or may be omitted. Similarly, the order of the above processes may be changed.

[0273] The subject matter of the present disclosure includes all novel and non-obvious combinations and sub-combinations of any and all of the various processes, systems, and configurations disclosed herein, as well as other features, functions, acts, and / or properties, and their equivalents.

[0274] Furthermore, it will be recognized that the terms "includes", "including", "has", "contains", their variations, and other similar terms used in the embodiments or claims for carrying out the invention are intended to be open transitional terms that do not exclude any additional elements or other elements, just like the term "comprising".

Claims

1. An arithmetic system comprising: receiving a simulation sample including a plurality of simulations, each simulation including a plurality of simulation results for a plurality of correlation probability variables; at least partially generating a surrogate cumulative distribution model by estimating a plurality of surrogate model parameters based at least in part on the plurality of simulation results; selecting one or more subsets of the plurality of simulations based at least in part on the surrogate cumulative distribution model having the surrogate model parameters; in each of one or more resampling iterations, until it is determined that the sum of one or more respective disagreement scores of the one or more subsets satisfies an optimization threshold calculating the one or more disagreement scores of the one or more subsets; sampling one or more resampled simulations for the plurality of correlation probability variables from the plurality of simulations included in the simulation sample and not yet included in the one or more subsets, based at least in part on the sum of the one or more disagreement scores; and replacing one or more simulations included in the one or more subsets with the one or more resampled simulations; and outputting the simulations included in the one or more subsets after performing the one or more resampling iterations An arithmetic system comprising a processor configured as such.

2. The arithmetic system according to claim 1, wherein for a plurality of layers of the plurality of simulation results, the processor: calculates a plurality of layers of the surrogate cumulative distribution model; and selects the one or more subsets of simulations such that the simulation results included in the simulations included in the one or more subsets are uniformly distributed among the plurality of layers An arithmetic system further configured as such.

3. The arithmetic system according to claim 1, wherein the processor is configured to at least partially replace the one or more simulations with the one or more resampled simulations by executing a quantum-inspired algorithm.

4. The arithmetic system according to claim 1, wherein the processor is configured to at least partially generate the plurality of simulation results for the plurality of correlation probability variables by executing the Iman-Conover algorithm.

5. The arithmetic system according to claim 1, wherein the processor is configured to at least partially sample the one or more resampled simulations by executing the Iman-Conover algorithm.

6. The arithmetic system according to claim 1, wherein the surrogate cumulative distribution model is a mixture Erlang model including a plurality of Erlang distributions.

7. The surrogate cumulative distribution model further includes one or more alternative tail region distributions configured to replace one or more respective tail regions of the plurality of Erlang distributions; and The arithmetic system according to claim 6, wherein the one or more alternative tail region distributions are different from the one or more Erlang distributions in the one or more respective tail regions.

8. The arithmetic system according to claim 1, wherein the surrogate cumulative distribution model is an empirical model configured such that the processor estimates the surrogate model parameters based at least in part on empirical data included in the plurality of simulations.

9. The arithmetic system according to claim 1, wherein the processor is configured to at least partially estimate the plurality of surrogate model parameters by executing iterative expectation maximization.

10. The arithmetic system according to claim 1, wherein the plurality of simulation results comprise a plurality of aggregate values, minimum values, or maximum values across the plurality of correlation probability variables.

11. The processor further: generates the surrogate cumulative distribution model in response to receiving a surrogate model type selection in a graphical user interface (GUI); generates the one or more subsets of the plurality of simulations in response to receiving a simulation generation command in the GUI; and outputs the one or more subsets of the simulations to the GUI The arithmetic system according to claim 1, configured as such.

12. A method for use in an arithmetic system, comprising: Receiving a simulation sample that includes a plurality of simulations, where each simulation includes a plurality of simulation results for a plurality of correlation probability variables; Generating a surrogate cumulative distribution model at least in part by estimating a plurality of surrogate model parameters based at least in part on the plurality of simulation results; Selecting one or more subsets of the plurality of simulations based at least in part on the surrogate cumulative distribution model having the surrogate model parameters; In each of one or more resampling iterations, until it is determined that the sum of one or more respective disagreement scores of the one or more subsets satisfies an optimization threshold: Calculating one or more disagreement scores of the one or more subsets; Sampling one or more resampled simulations for the plurality of correlation probability variables from the plurality of simulations that are included in the simulation sample and not yet included in the one or more subsets, based at least in part on the sum of the one or more disagreement scores; and Replacing one or more simulations included in the one or more subsets with the one or more resampled simulations; and After performing the one or more resampling iterations, outputting the simulations included in the one or more subsets A method comprising.

13. For a plurality of layers of the plurality of simulation results: Calculating a plurality of layers of the surrogate cumulative distribution model; and Selecting the one or more subsets of the simulations such that the simulation results included in the simulations included in the one or more subsets are uniformly distributed among the plurality of layers The method according to claim 12, further comprising.

14. The method according to claim 12, wherein replacing the one or more simulations with the one or more resampled simulations includes performing a quantum-inspired algorithm.

15. The method according to claim 12, wherein sampling the one or more resampled simulations further includes performing an Iman-Conover algorithm.

16. The method according to claim 12, wherein the surrogate cumulative distribution model is a mixture Erlang model including a plurality of Erlang distributions.

17. The method according to claim 12, wherein the surrogate cumulative distribution model is an empirical model in which the surrogate model parameters are estimated based at least in part on empirical data included in the plurality of simulations.

18. The method according to claim 12, wherein estimating the plurality of surrogate model parameters includes performing iterative expectation maximization.

19. The method according to claim 12, wherein the plurality of simulation results include a plurality of aggregated values, minimum values, or maximum values over the plurality of correlation probability variables.

20. An arithmetic system comprising: Receiving a simulation sample including a plurality of simulations in which each simulation includes a plurality of simulation results for a plurality of correlation probability variables; Generating a surrogate cumulative distribution model based at least in part on the plurality of simulation results; Selecting a compressed subset of the plurality of simulations based at least in part on the surrogate cumulative distribution model; In each of one or more resampling iterations, until it is determined that a disagreement score of the compressed subset is less than a predetermined disagreement threshold: Calculating the disagreement score of the compressed subset; Sampling one or more resampled simulations for the plurality of correlation probability variables from the plurality of simulations included in the simulation sample and not yet included in the compressed subset, based at least in part on the disagreement score; and Replacing one or more simulations included in the compressed subset with the one or more resampled simulations; and Outputting the simulations included in the compressed subset after performing the one or more resampling iterations A processor configured as An arithmetic system comprising.