An ensemble approach to predictive damage analysis
A hybrid predictive model combining probabilistic and AI/machine learning methods addresses the uncertainty in damage prediction by using AI to identify item attributes and probabilistic models to simulate events, resulting in a comprehensive and accurate damage forecast.
Patent Information
- Application Number
- JP2025521402
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-11
- Filing Date
- 2023-10-11
- Publication Date
- 2025-10-17
AI Technical Summary
Current predictive models for potential damage to items, such as buildings or people, suffer from significant uncertainty and variability in predicting outcomes due to their reliance on probabilistic probability models or artificial intelligence/machine learning approaches, which either fail to quantify damage accurately or provide inconsistent results.
A combined approach that balances probabilistic probability models and AI/machine learning models, utilizing a bottom-up AI method to identify patterns in individual item attributes and a top-down probabilistic model to simulate events, generating a comprehensive and accurate prediction of damage likelihood.
This hybrid method provides a more reliable and precise assessment of potential damage by leveraging the strengths of both models, offering a balanced and multifaceted forecast of future damage outcomes.
Smart Images

Figure 2025534732000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 415,280, entitled "Ensemble Approach to Predictive Damage Analysis," filed October 11, 2022 (Attorney Reference No. 016267-001PV1), the entire contents of which are incorporated herein by reference as if fully set forth herein.
[0002] The present disclosure relates to systems and methods that use predictive analytics modeling, such as to identify risk profiles that apply to potential damage to property, people, or other items. [Background technology]
[0003] Risk profiles may be used for any number of purposes, including, but not limited to, risk mitigation, risk identification, insurance pricing, and disaster recovery planning. Any collection of similarly categorized items may be subject to damage from multiple forces, events, or occurrences. Being able to forecast such potential damage is important in many fields in order to predict or assess physical, financial, or life-related impacts. For example, FIG. 1 is a schematic diagram illustrating examples of multiple items subject to damage or other impacts due to external forces. Examples of damaged items include, but are not limited to, building structures, furniture, artwork, people, computer systems, and networks (101, 102). Such items may be subjected to external adverse forces (103). Adverse forces may include natural catastrophic events (including excessive rainfall causing flooding, strong winds from tropical cyclones, tornadoes, tsunamis, and earthquakes), man-made events (such as cyber-attacks in the case of computer systems), the failure of dams or levees, or disease events (such as pandemics that cause illness in affected humans). These influences can then cause damage to the affected item (104). Summary of the Invention
[0004] The present disclosure relates to systems and methods that use predictive analytics modeling, such as to identify risk profiles that apply to potential damage to property, people, or other items.
[0005] In one implementation, for example, an artificial intelligence / machine learning (AI) model or engine is adapted to be used as a mechanism for achieving predictive analysis of unknown adverse future events for one or more items. The AI model is adapted to evaluate multiple known factors for one or a set of evaluation items. The AI model is adapted to build an AI training set based on aggregated data, refine multiple clusters from the aggregated data, use AI on individual items in each cluster, and compare outputs between the clusters to find one or more ranked factors; and identify individual items through characteristics to provide the training set. The training set is used to evaluate and identify all items in the overall dataset, resulting in a score or other indicator for each individual item that predicts the likelihood of an adverse or favorable future impact.
[0006] In another implementation, a plurality of probabilistic probability models are provided, each developed from a set of past events, where each such event has a footprint that has previously occurred or is predicted to occur, and the probabilistic probability models are adapted to simulate a plurality of actual or simulated events for a plurality of items.
[0007] An artificial intelligence / machine learning (AI) model or engine is adapted to be used as a mechanism for achieving predictive analysis of unknown adverse future events for one or more items. The AI model is adapted to evaluate multiple known factors for one or a set of evaluation items. The AI model is adapted to build an AI training set based on aggregated data, narrow down multiple clusters from the aggregated data, use AI on individual items in each cluster, and compare outputs between clusters to find one or more ranked factors; and identify individual items through characteristics to provide the training set.
[0008] The training set is used to evaluate and identify all items in the overall data set, resulting in a score or other indicator for each individual item that predicts the likelihood of adverse or favorable future impact.
[0009] In yet another implementation, a method or system for discretizing aggregate data is provided, which includes using AI to build an AI training set to narrow down clusters based on broad features; using AI on individual items in each cluster (positive and negative clusters); comparing outputs between clusters to find ranked factors that drive positive or negative outcomes; and identifying individual items through characteristics that comprise the training set.
[0010] These and other aspects, features, details, utilities, and advantages of the present invention will become apparent from a reading of the following description and claims, and from a review of the accompanying drawings. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a schematic diagram illustrating examples of items subject to damage or other impact due to external forces.
[0012] [Figure 2] FIG. 1 is a schematic diagram illustrating an example of an event footprint.
[0013] [Figure 3] FIG. 1 is a schematic diagram showing a group of past events used to create an event catalog.
[0014] [Figure 4] FIG. 1 is a schematic diagram illustrating an example of a simulation of the adverse effects of one or more events on each of multiple items input into the model.
[0015] [Figure 5] 5 is a schematic diagram illustrating an exemplary chart 501 showing aggregated data 508 available for each of multiple geographic regions or areas 502-507, including multiple features 509 located within the multiple geographic regions or areas.
[0016] [Figure 6] FIG. 6 is a schematic diagram illustrating an exemplary chart 601 illustrating how clusters identified within distinct geographic regions or areas can be used to generate effective AI training sets.
[0017] [Figure 7] FIG. 7 is a schematic diagram illustrating an example of operations for identifying predicted features such as those shown in FIG. 6 and constructing a training set from the predicted features.
[0018] [Figure 8] 8 is a schematic diagram illustrating an example of applying features 802 to a global set of items 803 to obtain a score 804 for each individual item.
[0019] [Figure 9] We present examples of systems and methods for performing analysis of predicted events based on aggregated data using an AI / machine learning engine or model.
[0020] [Figure 10]1 illustrates an exemplary computing system or electronic device for implementing examples of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0021] In one embodiment, predictive analytics modeling is used to identify one or more risk profiles for one or more items for harm from one or more influences that have the potential to harm the item.
[0022] The resulting damage from an impact on an item can vary greatly. For example, two buildings may experience the same or similar level of impact from a single event, such as a tropical cyclone, and the resulting physical damage may be zero percent, 100 percent, or any value in between. The problem of predictive analysis is complex, involving significant levels of uncertainty in expected outcomes, because the nature of an adverse impact and the ability of a given item to withstand such an impact cannot be predicted in a sufficiently accurate manner.
[0023] There are many computerized modeling approaches designed to predict the impact of damage or loss on items affected by adverse forces. These include probabilistic probability models, artificial intelligence or machine learning models, and many others. Each of these classes of models attempts to predict potential future damage to items based on past events, characteristics, environmental attributes, structural features, and other types of data inputs. As an example, a disease that reaches pandemic levels may adversely affect many people; however, some may become ill while others may remain unill. Furthermore, those adversely affected by the disease may be affected to varying degrees, from asymptomatic to very severe symptoms and even death. Such models use computerized modeling based on multiple data inputs to attempt to predict the impact on a given set of items of a particular category, such as people, buildings, objects, or other assets, utilizing past and known inputs.
[0024] Two main categories of predictive models for systems and methods currently proposed are probabilistic probability models and artificial intelligence / machine learning models.
[0025] Probabilistic probability models are typically developed from a set of past events, where each such event has a footprint (202) that occurred before. Figure 2 is a schematic diagram illustrating an example of an event footprint. In this example, the scope of the footprint (202) is limited in some way to a portion of a more general area or universe of items within the entire possible set (201). The footprint (202) may be geographic (a specific region), demographic (a specific category of people), or taxonomic (such as the type of computer network affected by a previous attack), among other things. The past event footprint (202) has a detailed cataloged set of impacts (203) that can be converted into severities indicating the likelihood of adverse damage from the impacts (203) to items affected by the event footprint (202). In some cases, the amount of adverse damage (205) to specific items (204) in past events is known, along with several different attributes associated with those items.
[0026] FIG. 3 is a schematic diagram illustrating a collection of past events used to create an event catalog. In this example, a study of past event footprints (202) and their damaging effects is used to create an event catalog (301), or a probabilistic set of events (302) based on various data inputs influenced by the judgment, experience, and expertise of the modeler, modeling human, or humans. Typically, the resulting events are based on past actual events, with applicable variations. For example, in a catastrophic event affecting a building, the variations in the resulting events may be geographic, involving a shift in the event footprint location. The resulting events may also vary in severity or any number of other factors according to the model design. Each resulting event in the catalog may include a set of variable parameters, including the geographic location of the event, the magnitude and vector of a given impact, the potential frequency of such an event, and the direction or severity of the impact. The probabilistic event catalog groups actual past and / or resulting events into pseudo-random patterns for use in future simulations of damage to one or more items submitted to the model. These items may be adversely affected by such events in the future, and the purpose of the model is to predict or estimate the potential outcomes. This is the fundamental predictive mechanism and purpose of probabilistic models.
[0027] The event catalog may be organized by a set of simulation periods, each containing one or more events from the catalog. Organizing by simulation period allows each such simulation period to be run against one or more items submitted to the model for assessment. The model developer may select a standardized number of simulation periods for the model, for example, 10,000 or 100,000 simulation periods. A simulation period typically represents a randomized, exemplary year or other possible time period that could affect one or more items submitted to the model. Each simulation period is used in the sampling process (described below). The number of simulation periods is directly related to analytical calculations performed on sample damage values from the model. The number of simulation periods can then be fed into analytical statistics to predict the probability or likelihood of damage for one or more items. An example of such output from a probabilistic model is an estimated probability of damage. This type of output is often referred to as a return period, or probability estimate of how likely the predicted damage is in any given year, e.g., 1 in 100, or 1 percent, chance of damage estimate, or 1 in 1000, or 0.1 percent, chance of damage, in any given actual year.
[0028] The next phase in probabilistic probability models is to simulate damage to a new, as yet unknown, set of items, which are input into the model. Each item input into the model has multiple attributes or features that may make the item more or less prone to damage from an event. In the field of disease, these attributes may relate to individual human prior conditions, such as age, comorbidities, physical conditions, or other factors. In the field of predicting damage to buildings from catastrophic events, attributes typically include the location of the structure, its construction method, the building's height, the type and value of the contents within the building, and many others. The output of the event simulation is typically a set of curve parameters that indicate the probability of item damage between 0 and 1 (zero percent to 100 percent).
[0029] Such models typically include a damage function or algorithm for estimating potential damage to items from any given event based on the attributes or features of the items. Events can potentially be included in a simulation period along with other events. For example, a building constructed from concrete and steel can withstand higher impact forces from a tropical cyclone compared to one constructed from a wood frame. Continuing the example, higher floors may experience higher wind speeds, potentially damaging windows and contents in a disproportionate manner compared to the rest of the structure. Humans with pre-conditions such as high blood pressure or narrowed arteries may be more susceptible to the effects of a given disease. As another example, when considering flooding as a cause of damage, artwork or mechanical equipment may be less susceptible to damage when located on higher floors or levels within a building, as long as the building in question can withstand the impact forces from the flood without collapsing. Any combination of multiple attributes related to structures, humans, or the environment may be considered when estimating damage from a given event. Furthermore, some attributes, when considered, may result in greater or lesser impacts. All of these physical and environmental factors contribute to complex simulations and analyses that can then be applied to predict future real-world events.
[0030] The model then simulates the adverse impact of each event in the event catalog (401) for each of the items (403, 404, 405, 406, 407) input into the model. Figure 4 is a schematic diagram illustrating an example of simulating the adverse impact of one or more events for each of multiple items input into the model. One or more items may be in one or more of the event footprints (402), and if so, there is a possibility of a simulated damage outcome. It should be noted that some items (407) may not be present within any of the event footprints (402) in the catalog (401); in such cases, damage simulation is not possible, and the predicted damage for items from such event footprints will always be zero or zero percent. Damage can only be simulated for items (403, 405, 406) that belong within one or more event footprints (402).
[0031] Further uncertainty handling can be performed by the model, as any given event only provides a probabilistic likelihood of damaging items (403, 405, 406) within its footprint (402). For each event footprint (402) with a given severity, a set of attributes or features of each item or set of items is used to predict the susceptibility of any given item to that damage.
[0032] The model then generates output, often in the form of a probability curve of potential damage for an item or set of items, ranging from zero (no damage) to one (complete damage). In some probabilistic probability models, this curve is then processed using a randomized sampling phase to select discrete potential damage points and values from the sampled points along the curve. This output can then be used to analyze multiple discrete points to develop a useful view of the potential damage resulting from the model simulation and sampling process.
[0033] The number of samples per event curve can effectively double the number of discrete damage output values from the model. For example, a model utilizing a set of 100,000 simulation periods can generate 10 or 100 samples for each potential damage-generating event. Thus, the effective number of simulation periods in the sampled output can be 1,000,000 or 10,000,000 simulation periods, respectively.
[0034] The output of simulation and damage curve sampling from such models when using this sampling approach is a set or plurality of discrete damage points, typically expressed as a percentage of damage for each target item or items being assessed. As explained above, the discrete damage values are generated by the model from randomized points along the damage curve from events that adversely affect one or more items. There may be thousands or even millions of such damage points generated in the sampled damage value results.
[0035] Various statistical calculations can then be performed on this set of damage points to indicate potential damage (percentage of probability) at various return periods, or an average annual number of damages, which is the average of all simulated damage points divided by the number of simulation periods. These analytical indices can be used for various purposes to provide an insight into the potential risk of damage to the item or items being evaluated. In particular, the average annual damage can be used to calculate the monetary cost of repairing the damage, or a numerical score or other statistical derivation to be used as the model output. Thus, a given catastrophe model run with a set of items as its input can simulate damage and useful analytical indices based on these input items. The model output can represent potential damage to one or more items that can be submitted for processing.
[0036] Different probabilistic probability models can produce widely varying results, even when the same set of inputs is used. This can be due to variations in past events or the interpretation of those events, different methods for creating new simulated events, the amount of input data available when developing the model, or variations in the interpretation of severity or event severity impact. These are just some of the potential variations between given models, even when the intended purpose and category of those items, item attributes, and event classifications are similar. Thus, simulated damage for input items and resulting outputs can vary widely. This variation reduces the usefulness of the model and makes it more difficult for model users to interpret meaningful results.
[0037] In summary, probabilistic probability models utilize a top-down approach in their analysis, simulating multiple events for multiple items. Such models are useful in quantifying risk into meaningful metrics, such as the potential severity of illness or the percentage of damage to buildings or other physical assets of potential value.
[0038] Artificial intelligence / machine learning (AI) can be used as an alternative mechanism to achieve predictive analysis from as-yet-unknown adverse future events for one or more items. In contrast to probabilistic probability models, AI utilizes a bottom-up approach that considers everything known about an item or set of items and compares it through various pattern recognition to learn algorithms to predict potential future damage or outcomes.
[0039] It is important to note that when using the above modeling approaches to assess the percentage of quantitative potential damage for a given item, probabilistic models are very powerful. This is achieved through decades of development and refinement in these models. AI approaches are qualitative in nature; they can predict the likelihood of loss in alternative ways, but are not useful in quantifying expected damage. Probabilistic models are more useful in quantifying potential damage, while AI models can indicate the likelihood of damage, or lack thereof, for a given item.
[0040] In one embodiment, the system or method provided utilizes a balance of both probabilistic and AI model approaches to provide predictive analysis from events to a set of items, resulting in improved output and interpretation compared to either method alone.
[0041] An AI approach first identifies a set of items using known outcomes from various events, usually past events. The items have true (positive) or false (negative) harm from the known events. This set of items is known as the AI training set. Furthermore, the training set may include identified severities of harm from various events. For example, considering a disease in a pandemic, the result of training would be to identify sets of people who either became ill or did not become ill from the disease. After this, further grading of harm, such as how severe the harm was for those adversely affected, may be applied.
[0042] Rather than the model developer constructing a damage function or algorithm, the AI model "learns" through many simulations and pattern recognition, which may number in the millions, billions, or even trillions. Simulations are first run against a training set using all known attributes or features for each item. The computer "learns" from the identified simulations and patterns without direction from the model developer. Learning from the training set produces a derived set of attributes that are more or less likely to cause harm, without the model developer feeding such parameters into the model.
[0043] As an example, consider a set of buildings as a training set, where damage or lack thereof is identified, and create a training dataset of buildings that sustained damage, or positive items, and those that did not sustain damage, or negative items. In this example, for an AI model to "learn" from the training set, the actual damage or lack thereof must be known in advance in the training dataset. Many attributes of each item in the training set are considered. If buildings are the target of the AI model, the attributes or features could be geographic location, architectural type, building height, and potentially dozens or hundreds of others. No assumptions about which attributes may contribute to damage or lack thereof are provided to the model; rather, the AI model runs many simulations for the attributes of each item in the training set, then looks for patterns among the set of items to determine which attributes contribute more or less to damage. These learnings from the model processing can then be used to interpret the likelihood of damage based on known positive or negative inputs provided in the training set.
[0044] As previously mentioned, a training set of items is evaluated against many characteristics or attributes that may be known about each item. Such attributes may also relate to a given item's environment. Continuing with the pandemic example, preconditions, heart health, nutritional history, genetic history, socioeconomic status, and many other factors may be considered as features for each item or person in the training set. The known illnesses or lack thereof of each item or person in the training set must be provided to a model that can then "learn" which attributes or features drive the likelihood of illness without direction from the modeler. Thus, the relevance, completeness, and accuracy of the training set of items are critical to the success and usefulness of an AI model approach.
[0045] When applying this approach to buildings that are adversely affected by extreme weather events, the training set would include a set of known buildings that did or did not suffer adverse damage from the type of event being considered, e.g., a rain-induced flood. The degree of damage can be considered for each building, either positive or damaged as a result of the adverse event. Conversely, the training set would include a set of buildings that did not sustain damage, or negative items. Many attributes about the buildings would be collected, such as geographic location, topography, weather history, construction type, building height, number of floors, or type of use.
[0046] By running many algorithmic matches across all attributes, potentially millions, billions, or trillions, through AI, patterns of attributes may emerge, or "learn" which attributes contribute to the likelihood of harm, along with the relative importance of each attribute in determining such harm. Conversely, features or attributes, or the lack thereof, that indicate resilience to harm may be identified. In this manner, the AI model generates a new set of features or attributes that are useful in predicting harm or lack thereof. The model may also provide "learned" weights for specific attributes compared to other attributes; in other words, it may develop which attributes in the training set are most predictive of harm or lack thereof for a particular item. Such attributes or patterns are compared by the AI engine among the many attributes in the training set. This approach is referred to as artificial intelligence or machine learning because the computerized algorithm itself derives correlations between attributes and their associated patterns, including statistical weights or importance of the attributes or features in a given set. The resulting output of multiple attributes or features generated by the model provides new insight into what drives harm, or lack thereof.
[0047] This allows for the recognition and discovery of patterns and associations not possible through human analysis alone. AI uses correlations of known attributes among a set of known positive / negative items to determine the likelihood of harm, producing as its output a set of weighted attributes and features with specific correlations. This is the opposite approach to probabilistic probability models, which focus on studying an event in detail and then determining through random simulation of such an event a set of potentially harmed items to generate a set of probabilistic indicators.
[0048] As previously mentioned, in order for an AI model to function and "learn" which features are those that drive harm, it needs a valid training data set that includes multiple positive and negative individual items, along with a set of known attributes or features for each item. The AI model (or models) then performs multiple pattern matches against each item's known attributes or features, thus "learning" or determining those attributes or features that contribute to harm, or lack thereof.
[0049] Once an AI model has been trained on a valid training dataset of sufficient size and validity, along with a set of discriminatory features or attributes for each item in the training set, it can be used to predict the likelihood of harm for new items previously unknown to the model that are fed into the model. Here, this is done by applying the set of discriminatory features or attributes "learned" from the training set to these new items. All that is required for a new item is the identification of the item's features or attributes, from which the AI model, based on its prior training, predicts the likelihood of harm, or lack thereof.
[0050] A common obstacle to AI approaches is the unavailability of a valid training set with a sufficient number of individual items and their plausibility for positive and negative events. However, in many cases, valid aggregate or summary data is available (502, 508) that represents a set of unidentified individual items. Figure 5 is a schematic diagram illustrating an exemplary chart 501 showing aggregate data 508 available for each of multiple geographic regions or areas 502-507, including multiple features 509 located within the multiple geographic regions or areas.
[0051] Continuing with the disease example of this application, it is often possible to obtain aggregate data regarding cumulative statistics and damage resulting from disease outbreaks. In such cases, only aggregate data is provided to protect the anonymity of the individual people adversely affected by the disease. Similarly, in the real estate example, there may be data sources that provide aggregate damage percentages by geographic region or area (508), number of damaged buildings, monetary costs, and other summary data without identifying individual buildings or properties.
[0052] As previously mentioned, AI approaches utilize training sets composed of a sufficient number of individual items, each of which is identified in the set as either positive (adversely affected) or negative (not adversely affected). Aggregate data sets do not meet this requirement.
[0053] While it is possible to identify attributes or features of individual items from various sources (509), such data does not correlate with harm or lack thereof from past events. Furthermore, in one instance, individual items in aggregate data are identified along with their attributes or features (509). Therefore, individual item attributes or features that are not correlated with harm cannot be used as a valid training set. The identity of individual items that are positive (harmed) or negative (not harmed) must be determined.
[0054] Continuing with the real estate example, in the available dataset, geographic regions (502, 503...507) are identified to show the total damage, or lack thereof, across all items in the specified area over a period of time. While it is possible to identify all properties or items within each region, it is not known which items have suffered damage, or lack thereof. All that is known is the cumulative damage, or lack thereof, across the overall geographic area and the entire set of properties within that region.
[0055] FIG. 6 is a schematic diagram illustrating an exemplary chart 601 showing how clusters identified within individual geographic regions or areas can be used to generate an effective AI training set. In one example, an approach is utilized to generate a set of items representing likely clusters (606) of positive or negative items from aggregated, summarized data. For example, this can be done by considering various factors, such as the density of real estate within a geographic region as a subset of the overall area described by the aggregated data. Similarly, in the disease example, population density is used to identify such clusters of likely positive or negative people or items. Other factors considered could be the value of such real estate, or proximity to coast or water, in the case of assessing tropical cyclones for flood risk; other salient factors related to the cause of potential damage could also be assessed. Using these initial assumptions, a set of clusters containing a preponderance of positive or damaged items and, conversely, a lack of negative or damaged items is generated. After generating clusters likely to contain a preponderance of damaged or undamaged items, the aggregated data is then applied as a percentage to each cluster. In this exemplary approach, such exploratory analysis would be used to identify multiple clusters (606, 607) of properties within each set of items described by the aggregate summary statistics or data. Each cluster (606, 607) would consist of multiple or sets of items or properties included in the aggregate statistics. At this point, the requirements for a viable training set of individual items that are positive or negative are not yet known, only multiple clusters of items that have a high likelihood of having endured past damage, or lack thereof. In other words, which individual items within a given cluster have endured damage, or lack thereof, is unknown.
[0056] To further illustrate this real estate example, a particular geographic area (602), such as 100 square kilometers, is described in the aggregated data set. The geographic area includes a plurality of buildings, e.g., 10,000 individual buildings. The aggregated data indicates that a total of 500 buildings (5%) were damaged over a specified time period, with a total cost of damage of $10,000,000. Using factors such as the geographic building density in the particular geographic area (602) or other salient data elements that may be known about the buildings, clusters (606, 607) of buildings, perhaps 100 to 1,000 buildings, that are likely to experience a preponderance of damage or lack of damage can be identified from the aggregated set. Conversely, the approach may be applied to geographic areas that have collectively endured little or no damage in the past, identifying clusters of buildings that are likely not to have been damaged by a past event. In this case, the generation of clusters of items that include a preponderance of items likely to be affected or not affected by the known aggregated data is critical to the described approach. In our building example, using aggregate data and cluster identification, as many as 500,000 such clusters were generated from the initial attributes used to narrow the cluster area.
[0057] At this point, clusters of interest are identified and aggregate percentages of damage are applied to the particular clusters. Additionally, each item within each is identified, and the attributes and features of each item are known as collected from various data sources associated with the items. The individual items within a cluster that did or did not sustain damage are still unknown, and therefore new data must be generated to meet the requirements of a useful training dataset.
[0058] Now that clusters of buildings (606, 607) have been identified, AI can be applied to build an effective training dataset for an AI model. In one case, AI can be run on each item in the cluster (606) using multiple attributes or features. For example, 100 or more features can be evaluated. AI processing on an individual cluster (606) still cannot identify which particular buildings likely survived damage (positive) or did not survive damage (negative). The example can continue with running multiple AI methods on multiple clusters, each derived from a known aggregate dataset. In the described real estate example, this could require as many as 500,000 clusters to be so evaluated.
[0059] For example, Figure 7 is a schematic diagram illustrating operations 701 for identifying predictive features as shown in Figure 6 and constructing a training set from the predictive features. As shown in Figure 7, aggregated data 704 is provided in an aggregated data source corresponding to one or more geographic regions or areas 702, 703, and based on the provided data and other filters, one or more clusters 706, 707 are identified as sets of likely primarily positive or negative features within one or more of the geographic regions or areas 702, 703. Predictive features 708 are identified from the clusters across the entire aggregated data set 706, 707 within the geographic regions or areas 702, 703. A training set of individual items 706 is constructed from the predictive features 708, and these features are used to assign likely positive or negative cases. This process requires extensive comparison of clusters and their contained items across a set of clusters.
[0060] The AI model can then be used to compare patterns of the derived attributes for each individual item within a cluster, and then compare the generated patterns to all other clusters. This requires potentially trillions of AI model iterations to perform pattern matching and identification, utilizing many AI model approaches and 100 or more attributes or features, where such features are known for each individual item within each cluster. Each feature or attribute can further be weighted and prioritized for its importance as a contribution to harm. The data generated from this step are multiple features or attributes correlated across all clusters so identified. In this case, sets of features or attributes are prioritized and weighted to indicate harm or lack thereof.
[0061] For example, Figure 8 is a schematic diagram showing operations 801 in which features 802 are applied to a global set of items 803 to obtain a score 804 for each individual item. Thus, an outcome regarding the likelihood of a future event affecting a given item is predicted for each item.
[0062] Continuing with the real estate example for tropical cyclone risk, height above ground level may be found to be the highest driving feature for predicting damage, perhaps with an importance weighting of 12%. This may be followed by building height, ranked second with an importance weighting of 10% in predicting likelihood of damage. After this, all other attributes may be ranked, with attributes such as building surface being identified by the AI model as unimportant or low-ranking attributes.
[0063] The data generated from this operation is a reduced set of AI approaches that predict damage, or lack thereof, and a reduced set of features and attributes (608), which are prioritized and weighted as described above to rank their impact on contributing to the likelihood of loss.
[0064] The AI approach and the reduced set of identified features and attributes (608) are then used to evaluate each building or item in each identified cluster against the overall dataset described by the aggregated data. Each set of likely positive or negative items in each cluster is compared or correlated by the AI model with sets of likely positive or negative items from each other cluster. This correlation or comparison is used to generate sets of individual items that have or have not sustained damage. The result is a valid training set generated by the AI model. This training set, the ranked items and features that drive damage or lack thereof, can then be applied to items outside the identified cluster in future evaluations, providing a useful prediction of the likelihood of damage.
[0065] By applying the ranked and weighted attributes and features to each item in each cluster, a composite score for each item can be generated that can be used as an indicator of the likelihood of damage, or lack thereof. This score is generated for each individual item or building across the buildings identified in all clusters in the dataset (608). These generated individual buildings and their scores then comprise a valid training set for the AI, which together generate a training output: prioritized, weighted features, attributes (608), and an AI model approach that predicts damage for any individual item, regardless of its specificity.
[0066] In this example, the approach can be validated by checking the model results against real-world events that cause harm. Another method used to validate a model can be to reserve and isolate a large percentage of item clusters from the initial training dataset, run an AI model on the items in this cluster, and use the AI model to predict aggregate statistics provided by the original data source. The reserved segment of the training set is completely independent of the training set segment used to train each model; the purpose is to test the validity of a given modeling approach. The predicted aggregate statistics are then compared to the original data source, and the percentage of such validation is referred to in AI terms as the area under the curve, or the success of the AI model in predicting future susceptibility to harm, or lack thereof, as the as-yet-unknown item to be evaluated.
[0067] A number of AI modeling approaches are first used to identify correlations or relationships between attributes among a set of individual items. The results of any given AI modeling approach can be compared for its ability to predict future harm by testing against a set of reserved segments of items from the training set. Through simulation and pattern recognition across many items, item attributes, and modeling approaches, and through testing the output of each modeling approach against a set of reserved segments of the training set, a subset of modeling approaches is generated that more accurately predict known outcomes in the training set.
[0068] Thus, AI models look at individual items and their properties, or the properties of the environment for each item, based on a known training set (bottom-up); probabilistic probability models can be said to do a similar job, albeit from a different perspective and approach, using a catalog of events (top-down).
[0069] Once the AI training set and all available known attributes have been pattern matched and processed via multiple modeling approaches, the resulting output is a set of characteristics of items (608) along with a specific AI simulation approach known to predict harm among the training set of items. This AI output generated from model training can then be applied to new, previously unassessed items, using the learnings from the training set to determine whether or not they are susceptible to harm from a similar type of event, based solely on their attributes.
[0070] The training set is of paramount importance to the success of an AI modeling approach. The example approach discussed herein provides a useful AI training set derived using only aggregated summary data of generalized datasets and the uncorrelated known attributes and features of the individual items within those generalized datasets as input. The AI process itself, through the described approach, determines predictive weighted features and attributes and resulting item scores.
[0071] The AI approach starts with each item and builds until learning is complete. Probabilistic probability approaches work from a base of events of specified intensity, then apply simulations to new sets of items to predict future potential damage. As mentioned above, probabilistic probability models are good at predicting quantitative potential damage, but the models vary widely depending on the quality and approach used to create the underlying event set. The AI approach provides a strong qualitative measure of the likelihood of damage for individual items.
[0072] While both approaches are useful for predictive analysis of damage to one or more items in a given classification or category, they can produce very different results due to their fundamental differences in approach. Each type of predictive model has specific advantages. Probabilistic probability models are useful for predicting adverse damage for a broad set of input items. Furthermore, such models are also useful for predicting the extent or quantitative projection of damage based on various probabilistic and randomized methods used to determine potential damage. Such models may also assess potential damage to various characteristics or segments of a given item, for example, the contents contained within a building. While this approach can be used to predict potential damage across an entire set of items, it is less accurate or predictive of damage to any single item in the input set.
[0073] The strength of AI approaches is based on their ability to evaluate many attributes and features of any single item against a training set of known sets. With sufficient simulation and pattern matching, the outputs and correlations from AI approaches can be very useful in predicting the likelihood or susceptibility of harm for any individual item. However, AI approaches cannot predict the extent of harm or the recurrence period (probability) of harm as easily as probabilistic probability approaches. They also cannot estimate potential harm for an entire set of items when compared to probabilistic probability approaches.
[0074] The proposed method presented here balances the utility of both approaches, yielding better and more accurate predictive results for the consumer of the model when applied to determining potential future damage from adverse events when a given set of items are evaluated or assessed. The approach considers the strengths and shortcomings of both probabilistic probability models and AI, resulting in an output that provides a more comprehensive and multifaceted forecast of future damage outcomes.
[0075] In the examples and cases presented, we show how multiple probabilistic models can be intelligently weighted with scores or other outputs from an AI approach in a combined view of the multiple probabilistic model outputs that more accurately approximates potential harm by comparing the output results of each model in the multiple models and weighting the results based on scores or other metrics developed by the AI approach.
[0076] Probabilistic probability models developed by different modelers can vary widely in their results. For example, in assessing real estate risk, loss estimates can vary by a factor of 10 or more between any two models for the same geographic location or building being assessed. These differences arise due to the various events included in each model's event catalog (Figure 4), the method or methods used to develop those events, the availability of historical data, the modeler's expertise, the availability of resources to the modeler, the model's focus and objectives, the assigned frequency of a given event in the catalog, and many other factors. Furthermore, given the vast number of potential items in any universe for a given classification, it is not feasible for any single model to adequately cover all potential items. When considering potential damage to real estate or buildings, for example, one model may provide more accurate data and assessments in one geographic region or event type compared to another that may focus on a different region or event type. Given the range and amount of data available or processed in constructing the event catalog, each model has gaps in its coverage. However, despite these differences, the probabilistic approach results in a damage assessment with the strengths of the approach previously described. Each given probabilistic model can be thought of as outputting a view or opinion of risk, given the approach taken to develop its event catalog.
[0077] To balance the output of differences in risk perception and consider the best assessment capabilities, in one example, the proposed method utilizes multiple, at least two, probabilistic models when predicting potential harm. The multiple probabilistic models are referred to as a model set. The output and potential harm generated by each individual model are then considered for assessment, and the models are compared with each other for validity and accuracy.
[0078] In some cases, event-based models are largely in agreement in their view of the risk they present; this is usually the case for potential damage sources for a particular area or item, where the field has matured and the modeling community has adopted a common underlying data source for event development. One example of a commonly used data source for such potential damage is for tropical cyclones along our eastern Gulf Coast of the United States. The underlying data source for such models is often the Sea, Lake and Overland Surges from Hurricanes (SLOSH) model provided by the National Weather Service. Results can be very similar when considering one or more models whose events are based on the SLOSH model.
[0079] In other cases, any two probabilistic models can differ widely in their assessment of potential damage for a given item, especially in newer types of assessments that lack extensive or common data sets for the development of their event sets. Examples of newer, less documented, or standardized sources of damage are terrorism or cyber risk. These fields have low data availability and cannot benefit from common data sets such as the SLOSH model; therefore, independent model developers can differ widely in their development of event catalogs, contributing to large differences in damage predictions.
[0080] As previously mentioned, to accommodate these variations, the proposed method in one example considers multiple probabilistic models when predicting potential damage to one or more items and recognizes the differences in assessments from each model. This is done on a sample output of each probabilistic model, which is a large amount of data that requires advanced computational processing to evaluate.
[0081] It is important to consider that using 100,000 simulation periods for a given model in a model set can result in much more sampling, effectively increasing the number of simulation periods. Each simulation period includes one or more events from the event catalog, each with its own curve parameters for predicting potential damage between 0 and 1, as described above. Each event curve, in turn, can be sampled to generate a separate, discrete loss value. Based on the event hit rate for the items in the set being assessed, there may be more or fewer events that generate potential damage. Each event that generates potential loss is then sampled. For example, a given model may generate 100 separate samples for each simulation period, effectively resulting in 10,000,000 simulation period outputs. This sampling increases the granularity of the data for a more detailed picture of predicted loss and smooths the final analysis results.
[0082] In one implementation, the method combines the sampled discrete damage value results of each model in the model set. If one model in the model set utilizes a 100,000 simulation-year event catalog with 100 samples, it may output probabilities for 10,000,000 discrete samples. Another model may use a smaller number of simulation years, for example, a 20,000-year event catalog with 10 events for an output of 200,000 years. The goal is to combine the results of these varying models into a single assessment.
[0083] In order to assess and evaluate the model set, the model outputs need to be normalized to a standardized number of simulation years. Continuing with this implementation, 10,000,000 simulation years may be chosen as the standardized value. Outputs from models with smaller numbers of simulation periods need to be normalized so that the outputs from the models can be combined with other models in the model set.
[0084] One approach to achieving this period normalization involves replicating sample damage values from a model with a smaller set of simulation periods to match the standardized number of simulation periods. Another method is to utilize the curve parameters from a given model and sample the curve with the number of samples needed to match that of the standardized simulation period output. Although the amount of data processed is larger in this implementation, it can be realized that it produces greater accuracy in damage prediction.
[0085] In this case, the method then assigns a weight to each discrete probabilistic sample damage output from each model in the model set. By performing weighting at the detailed sample data level, greater control over the assessment of differences is achieved. These weighted samples are then combined by standardizing the simulation period size across the model set.
[0086] The described implementation may then combine the sampled outputs from each model in the model set using various methods. One such method is to sort the sampled outputs of each model in the set from highest to lowest, and then add the sorted sets together, resulting in a combined set of sampled losses that match the normalized simulation period outputs. Another method may be to concatenate the sample loss values from each model into a larger set.
[0087] This is the output of the overall process, which applies the AI-generated scores or other data for each individual item to the sampled probabilistic model outputs for each model, then combines the sampled outputs into a single new result of integrated probabilistic output. This is a new set of probabilistic data outputs that contains the integrated view as influenced by the AI approach. This output is used for analytical processing, with the same computational methods used for any individual probabilistic model.
[0088] A standardized number of simulation periods allows analytical calculations to be performed on combined standardized samples from multiple models in the model set, resulting in a more accurate representation of the projected damage.
[0089] As previously explained, the output of the ensemble approach can be used for many purposes. One example is risk management for an association with many real estate or other tangible assets deployed in various geographic locations. This creates a great need to predict risks so that they can be appropriately managed. Many types of disasters are involved, including supply chain shortages, natural disasters, and even economic catastrophes. Risks are quantified so that they can be estimated, with complete financial and operational plans developed based on the analysis described in the presented embodiments. Using accurate risk prediction and quantification, alternative business operational plans and financial provisions can be allocated to ensure business or operational continuity. Then, when a disaster actually occurs, the risk managers of such organizations are prepared to respond by implementing such plans with prior knowledge of the likelihood of damage to given assets and their associated costs. This preparation can include alternative facilities, backup systems in the event of a cyberattack, alternative suppliers to overcome supply chain shortages, insurance coverage, and financial estimates of the costs of any given disaster.
[0090] Furthermore, the output of the presented modeling approach can be used to harden assets and systems to make them more resilient to exposure to damage. As an example, this can include engineering improvements to buildings to reduce the impact from various types of events. Similarly, risk prediction tools can be used to enable computer systems and processes to reduce the likelihood of cyber attacks or harden systems against such attacks.
[0091] Another example is a government organization at the national, state, or community level that wishes to plan improvements, through legislation, tax assessments, or investments, to improve the resilience of an item or items within their jurisdiction. A local community, for example, may undertake a project to add flood walls, seawalls, or levees to protect the community based on prior knowledge of the likelihood of damage to a particular property from a flood-related disaster. The engineering and resulting costs of such improvements can be better estimated with accurate predictions of the likelihood of such damage.
[0092] In the fields of healthcare and disease management, government or medical officials may use the model's predictions to identify characteristics of individuals most susceptible to a given disease. This information may then be used to inform, warn, and potentially protect such individuals, reducing the impact of a pandemic or other outbreak.
[0093] The processing required for AI involves large amounts of data, computing power, and multiple processors, all working in parallel. In one implementation, it took six months of processing across many parallel computer processes, each running on a single "core" server, performing trillions of calculations. This implementation required trillions of calculations to identify clusters, compare the clusters with each other, and run multiple AI approaches on the data; over 70 were tested. This trained an AI model based on the prevalence of items with high positive or negative likelihood. This resulted in identifying clusters, identifying which AI approaches would work with the aggregated dataset, and generating a list of significant item features with importance (out of dozens of item features). These significant features were then further processed across clusters—over 100,000 such clusters in this implementation—resulting in a training set of multiple individual items, each individually identified as having high positive or negative likelihood. This resulted in a viable AI training dataset that may not have been created from the original aggregated data. In addition to the training items and clusters, further processing was then required, involving a large number of processes running on multiple cores, to apply significant features to the overall dataset. The final result from this example was a numerical score indicating which items were more or less likely to be affected by an adverse event. This resulted in a dataset, and all of the intermediate datasets were stored for future access to utilize the numerical scores in predicting risk for this implementation. Finally, the same cluster of computer processors and processes was used in the computational scheme to backtest and validate the results and compare them to real-world events. The conclusion in all cases was a 92%+ accuracy rate in predicting the numerical scores. This large dataset, involving many steps of computational analysis and AI approaches, is not possible by conventional or manual means, and requires the cluster of processes to be continuously run on the input data to produce the final results.
[0094] FIG. 9 illustrates a system and method 901 for performing an analysis of predicted events based on aggregated data using an AI / machine learning engine or model. In this embodiment, the AI / machine learning engine or model is adapted to perform a series of operations corresponding to those shown in FIGS. 5-8. In one operation corresponding to FIG. 5, aggregated data is available for each of multiple geographic regions or areas, including multiple features for individual items located within the multiple geographic regions or areas. At this stage, it is not possible to identify which item features contribute to a prediction of likely impact or lack of impact from a future event.
[0095] Clusters are identified within distinct geographic regions or areas in separate operations, as shown in Figure 6. Predictive features are identified within geographic regions or areas from extensive AI processing and comparison across multiple clusters, and then used to generate valid AI training sets for individual items, as shown in Figure 7.
[0096] The identified features are then applied to the overall set of items to generate a predictive score, as shown in Figure 8, where the score predicts the likelihood, or lack thereof, of a potential future event occurring. Thus, the overall process uses extensive and heavy AI processing to move through only the aggregate data and predict the likelihood of impact, or lack thereof, for each individual item in the overall dataset.
[0097] The process operations shown in FIGS. 5-9 are processed on one or more servers 903, 904, each of which includes multiple cores 905 that run multiple processes in parallel to perform the operations of FIGS. 5-9.
[0098] 10 illustrates an exemplary computing system or electronic device for implementing examples of the present disclosure. System 1000 may include known components such as, but not limited to, a central processing unit (CPU) 1001, storage 1002, memory 1003, a network adapter 1004, a power supply 1005, an input / output (I / O) controller 1006, an electrical bus 1007, one or more displays 1008, one or more user input devices 1009, and other external devices 1010. It will be understood by those skilled in the art that system 1000 may include other well-known components, which may be added, for example, via expansion slot 712 or by any other method known to those skilled in the art. Such components may include, but are not limited to, hardware redundancy components (e.g., dual power supplies or data backup units), cooling components (e.g., fan- or water-based cooling systems), additional memory and processing hardware, and the like.
[0099] System 1000 may, for example, be in the form of a client-server computer that can be connected to and / or facilitate the operation of multiple workstations or similar computer systems through a network. In another embodiment, system 1000 may be connected to one or more workstations through an intranet or Internet network, thus facilitating communication with a larger number of workstations or similar computer systems. Still further, system 1000 may, for example, include a main workstation or main general-purpose computer that allows a user to interact directly with a central server. Alternatively, a user may interact with system 1000 through one or more remote or local workstations 1013. As will be appreciated by those skilled in the art, there may be any practical number of remote workstations for communicating with system 1000.
[0100] The CPU 1001 may include one or more processors, such as an Intel® Core® G7 processor, an AMD FX® series processor, or other processors, as will be understood by those skilled in the art (including, for example, graphical processing unit (GPU)-type dedicated computing hardware used, among other things, for machine learning applications, such as training and / or executing the machine learning algorithms of the present disclosure; such GPUs may include, for example, an NVIDIA Tesla® K80 processor). The CPU 1001 may further communicate with an operating system, such as Microsoft Corporation's Windows NT® operating system, the Linux® operating system, or a Unix®-like operating system. However, those skilled in the art will appreciate that similar operating systems may also be utilized. The storage 1002 (e.g., a non-transitory computer-readable medium) may include one or more types of storage known to those skilled in the art, such as a hard disk drive (HDD), a solid-state drive (SSD), a hybrid drive, and the like. In one example, the storage 1002 is utilized to persistently retain data for long-term storage. Memory 1003 (e.g., non-transitory computer-readable medium) may include one or more types of memory known to those skilled in the art, such as random access memory (RAM), read-only memory (ROM), hard disk or tape, optical memory, or a removable hard disk drive. Memory 1003 may be used for short-term memory access, such as, for example, loading software applications or handling transient system processes.
[0101] As will be appreciated by those skilled in the art, storage 1002 and / or memory 1003 may store one or more computer software programs. Such computer software programs may include logic, code, and / or other instructions that enable processor 1001 to perform the tasks, operations, and other functions described herein (e.g., probabilistic modeling, AI engine, data disaggregation, compilation of one or more AI training sets, as described herein), as well as additional tasks and functions that will be appreciated by those skilled in the art. Operating system 1002 may further work in conjunction with firmware, as is well known in the art, to enable processor 1001 to coordinate and execute the various functions and computer software programs described herein. Such firmware may reside in storage 1002 and / or memory 1003.
[0102] The I / O controller 1006 may also include one or more devices for receiving, transmitting, processing, and / or interpreting information from external sources, as known by those skilled in the art. In one embodiment, the I / O controller 1006 may include functionality for facilitating connection to one or more user devices 1009, such as one or more keyboards, mice, microphones, trackpads, touchpads, or the like. For example, the I / O controller 1006 may include a serial bus controller, a universal serial bus (USB) controller, a FireWire® controller, and the like, for connecting to any suitable user device. The I / O controller 1006 may also enable communication with one or more wireless devices, for example, via technologies such as near field communication (NFC) or Bluetooth®. In one embodiment, the I / O controller 1006 may include circuitry or other functionality for connecting to other external devices 1010, such as a modem card, a network interface card, a sound card, a printing device, an external display device, or the like. Additionally, I / O controller 1006 may include controllers for various display devices 1008 known to those skilled in the art. Such display devices may visually convey information to a user or users in the form of pixels, which may be logically arranged on the display device to allow the user to perceive information rendered on the display device. Such display devices may be in the form of touch screen devices, traditional non-touch screen display devices, or any other form of display device, as will be appreciated by those skilled in the art.
[0103] Additionally, CPU 1001 may further communicate with I / O controller 1006 to render a graphical user interface (GUI), for example, on one or more display devices 1008. In one example, CPU 1001 may access storage 1002 and / or memory 1003 to execute one or more software programs and / or components to enable a user to interact with the system as described herein. In one embodiment, the GUI described herein includes one or more icons or other graphical elements with which a user may interact to perform various functions. For example, GUI 1007 may be displayed on a touchscreen display device 1008, whereby a user interacts with the GUI via the touchscreen, for example, by physically contacting the screen with the user's finger. As another example, the GUI may be displayed on a conventional non-touch display, whereby a user interacts with the GUI via a keyboard, mouse, and other conventional I / O components 1009. The GUI may reside at least in part as a set of software instructions in storage 1002 and / or memory 1003, as will be understood by those skilled in the art. Also, the GUI is not limited to the methods of interaction described above. Those skilled in the art will appreciate any of a variety of means by which one may interact with the GUI, such as voice-based or other voice-based methods of interaction with a computing system.
[0104] The network adapter 1004 may also enable the device 1000 to communicate with a network 1011. The network adapter 1004 may be a network interface controller, such as a network adapter, a network interface card, a LAN adapter, or the like. As will be appreciated by those skilled in the art, the network adapter 1004 may enable communication with one or more networks 1011, such as, for example, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a cloud network (IAN), or the Internet.
[0105] The one or more workstations 1013 may include known components such as, for example, a CPU, storage, memory, network adapters, power supplies, I / O controllers, electrical buses, one or more displays, one or more user input devices, and other external devices. Such components may be the same as, similar to, or equivalent to those described above with respect to system 1000. It will be understood by those skilled in the art that the one or more workstations 1013 may include other well-known components, including, but not limited to, redundant hardware components, cooling components, additional memory / processing hardware, and the like.
[0106] The foregoing description of the present invention is provided as an enabling teaching of the invention in its best, currently known embodiment. To this end, those skilled in the art will recognize and understand that many changes can be made to various aspects of the invention described herein while still obtaining the beneficial results of the present invention. It will also become apparent that some of the desired advantages of the present invention can be obtained by selecting some of the features of the present invention without utilizing other features. Thus, those skilled in the art will recognize that many modifications and adaptations to the present invention are possible and may even be desirable in certain circumstances and are a part of the present invention. Accordingly, the following description is offered as an illustration of the principles of the present invention, and not as a limitation thereof.
[0107] As used throughout, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, reference to "a" component may include two or more such components unless the context indicates otherwise. Additionally, the words "proximal" and "distal" are used to describe items or portions of items that are located closer to or further from a user or operator, such as a surgeon, respectively. Thus, for example, the tip or free end of a device may be referred to as the distal end, while the opposite end or handle, generally, may be referred to as the proximal end.
[0108] All directional references (e.g., top, bottom, upward, downward, left, right, leftward, rightward, top, bottom, up, down, vertical, horizontal, clockwise, and counterclockwise) are used for identification purposes only to aid the reader in understanding the present invention and are not intended to create any particular limitation on the location, orientation, or use of the present invention. Joined references (e.g., attached, coupled, and connected, etc.) should be interpreted broadly and may include intermediate members between the connection of elements and relative movement between elements. As such, joined references do not necessarily infer that two elements are directly connected and in a fixed relationship to each other.
[0109] Ranges may be expressed herein as from "about" one particular value, and / or to "about" another particular value. When such a range is expressed, another aspect includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the preceding term "approximately," it is to be understood that the particular value forms another aspect. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint.
[0110] As used herein, the term "optionally" or "optionally" means that the event or circumstance subsequently described may or may not occur, and that the description includes instances where the event or circumstance occurs and instances where it does not occur.
[0111] As used herein, the term "effectively" may be applied to modify any quantitative expression that may be acceptably varied without resulting in a change in the basic function associated therewith. (Other possible items) (Item 1) An artificial intelligence / machine learning (AI) model or engine adapted to be used as a mechanism for achieving predictive analysis from as yet unknown adverse future events for one or more items, wherein the AI model is adapted to evaluate multiple known factors for one or a set of evaluation items; wherein the artificial intelligence / machine learning (AI) model uses AI to: constructing an AI training set based on the aggregated data and narrowing down a plurality of clusters from the aggregated data; Use AI on each individual item in each cluster, Compare the output between the clusters to find one or more ranked elements; and Identifying individual items through their characteristics to provide the training set It is adapted to wherein the training set is used to evaluate and identify all items across the entire dataset, resulting in a score or other indicator that predicts the likelihood of adverse or favorable future impact for each individual item; A system comprising: (Item 2) Item 1, wherein the AI model is adapted to identify a set of items with known outcomes from a variety of events, optionally past events and item characteristics, to provide an artificial intelligence training set. (Item 3) Item 3. The system of item 2, wherein the training includes identified likelihoods of potential events from various item characteristics. (Item 4) Item 3. The system of item 3, wherein further evaluation of items in the clusters is performed to identify item features or characteristics that drive the likelihood of potential events affecting individual items. (Item 5) Item 3. The system of item 2, wherein for each item, a simulation is run on the training set using multiple known attributes or features. (Item 6) Item 6. The system of item 5, wherein a set of derived attributes is generated, the attributes being more or less likely to cause harm, and the developer does not supply such parameters to the model. (Item 7) Item 10. The system of item 1, wherein the AI model is trained on a training data set along with a set of discriminatory features or attributes for each item in the training set and is used to predict the likelihood of an event occurring for new items fed into the model. (Item 8) Item 8. The system of item 7, wherein the new item was previously unknown to the model. (Item 9) 9. The system of claim 7 or 8, wherein the identified set of features or attributes learned from the training set is applied to the new items. (Item 10) Item 10. The system of item 9, wherein the features or attributes of the new item are used by the AI model to predict the likelihood of an event occurrence or lack thereof based on pre-training of the AI model. (Item 11) a plurality of probabilistic probability models each developed from a set of past events, where each such event has a footprint that has previously occurred or is predicted to occur, where the probabilistic probability models are adapted to simulate a plurality of actual or simulated events for a plurality of items; and An artificial intelligence / machine learning (AI) model or engine adapted to be used as a proxy mechanism for achieving predictive analysis from as yet unknown adverse future events for one or more items, wherein the AI model is adapted to evaluate multiple known factors for one or a set of evaluation items; wherein the artificial intelligence / machine learning (AI) model is adapted to: use AI to build an AI training set based on aggregated data, refine a plurality of clusters from the aggregated data, use AI for individual items in each cluster, and compare outputs between clusters to find one or more ranked elements; and identify individual items through characteristics to provide the training set. wherein the training set is used to evaluate and identify all items across the entire dataset, resulting in a score or other indicator for each individual item that predicts the likelihood of adverse or favorable future impact; wherein the score or index is used to weight each probabilistic probability model and adjust its predictions accordingly based on the score. A system comprising: (Item 12) Item 12. The system of item 11, wherein the combination of the probabilistic probability model and the AI model is adapted to provide predictive analysis from events to one or a set of the endpoints, resulting in improved output and interpretation compared to either method alone. (Item 13) Item 12. The system of item 11, wherein the footprint comprises a geographic footprint, a demographic footprint, or a classification footprint. (Item 14) Item 12. The system of item 11, wherein the footprint of past events comprises a detailed cataloged set of influences and possible impacts on items in the set submitted to the model. (Item 15) Item 15. The system of item 14, wherein the set of impact factors is converted into a severity indicating the likelihood of adverse damage from the impact factors to items affected by the event footprint. (Item 16) Item 16. The system of item 15, wherein the amount of adverse damage to specific items in the past events is known along with a number of different attributes associated with those items. (Item 17) 10. The system of claim 1, wherein the footprints of past events and their damaging impacts are used to create an event catalog, or probabilistic set of events, based on multiple data inputs. (Item 18) A system of any of the preceding items, wherein an event catalog is organized by set of simulation periods, each simulation period containing one or more events from said catalog. (Item 19) Item 19. The system of item 18, wherein the arrangement of the simulation periods into sets is adapted to enable each such simulation period to be run for one or more items submitted to the model for assessment. (Item 20) 20. The system of claim 19, wherein each simulation period is used in a sampling process. (Item 21) 21. The system of any of items 17 to 20, wherein the number of simulation periods is fed into analytical statistics to predict the probability or likelihood of damage to the one or more items. (Item 22) Item 22. The system of item 21, wherein the output from the probabilistic probability model includes an estimated probability of damage. (Item 23) Item 12. The system of item 11, wherein the AI model is adapted to identify a set of items with known outcomes from a variety of events, optionally past events and item characteristics, to provide an artificial intelligence training set. (Item 24) Item 24. The system of item 23, wherein the training includes identified likelihoods of potential events from various item characteristics. (Item 25) Item 25. The system of item 24, wherein further evaluation of items in the clusters is performed to identify item features or characteristics that drive the likelihood of potential events affecting individual items. (Item 26) Item 24. The system of item 23, wherein for each item, a simulation is run on the training set using a plurality of known attributes or features. (Item 27) 27. The system of claim 26, wherein a set of derived attributes is generated, the attributes being more or less likely to cause harm, and the model developer does not supply such parameters to the model. (Item 28) Item 12. The system of item 11, wherein the AI model is trained on a training data set along with a set of discriminatory features or attributes for each item in the training set and is used to predict the likelihood of an event occurring for new items fed into the model. (Item 29) Item 29. The system of item 28, wherein the new item was previously unknown to the model. (Item 30) 30. The system of claim 28 or 29, wherein the identified set of features or attributes learned from the training set is applied to the new items. (Item 31) Item 21. The system of item 20, wherein the features or attributes of the new item are used by the AI model to predict the likelihood of an event occurrence or lack thereof based on pre-training of the AI model. (Item 32) 1. A method or system for discretizing aggregate data, comprising: Use AI to build an AI training set based on rough features and refine the clusters; Using AI for individual items in each cluster (positive and negative clusters); Comparing outputs across clusters to find ranked factors that drive positive or negative outcomes; and Identifying individual items through characteristics comprising said training set. 1. A method or system comprising: (Item 33) Item 33. The method of item 32, wherein the score is derived from a weighting of features and / or attributes. (Item 34) Item 34. The method of item 33, wherein each item is scored based on the training set and the identified features. (Item 35) 2. The system of claim 1, wherein the AI model or engine is processed across multiple parallel processes running on multiple processors or cores. (Item 36) 36. The system of claim 35, wherein individual clusters of the plurality of clusters are processed using different processors or cores of the plurality of processors or cores. (Item 37) Using an artificial intelligence / machine learning (AI) model or engine as a mechanism for achieving predictive analysis of unknown adverse future events for one or more items, wherein the AI model is adapted to evaluate multiple known factors for one or a set of evaluation items; The artificial intelligence / machine learning (AI) model uses AI to: constructing an AI training set based on the aggregated data; and refining a plurality of clusters from the aggregated data; We use AI for each individual item in each cluster. Compare the output between the clusters to find one or more ranked elements, and identifying individual items via characteristics to provide said training set; wherein the training set is used to evaluate and identify all items across the entire dataset, resulting in a score or other indicator that predicts the likelihood of adverse or favorable future impact for each individual item; A method for providing (Item 38) Item 38. The method of item 37, wherein the AI model is adapted to identify a set of items with known outcomes from a variety of events, optionally past events and item characteristics, to provide an artificial intelligence training set. (Item 39) 39. The method of claim 38, wherein the training includes discriminated likelihoods of potential events from various item characteristics. (Item 40) 40. The method of claim 39, wherein further evaluation of items in the clusters is performed to identify item features or characteristics that drive the likelihood of potential events affecting individual items. (Item 41) Item 39. The method of item 38, wherein for each item, a simulation is run on the training set using a plurality of known attributes or features. (Item 42) Item 42. The method of item 41, wherein a set of derived attributes is generated, said attributes being more or less likely to cause harm, and the developer does not supply such parameters to the model. (Item 43) 38. The method of claim 37, wherein the AI model is trained on a training data set together with a set of discriminatory features or attributes for each item in the training set and is used to predict the likelihood of an event occurring for new items fed to the model. (Item 44) Item 44. The method of item 43, wherein the new item was previously unknown to the model. (Item 45) 45. The method of claim 43 or 44, wherein the identified set of features or attributes learned from the training set is applied to the new items. (Item 46) Item 46. The method of item 45, wherein the features or attributes of the new item are used by the AI model to predict the likelihood of an event occurrence or lack thereof based on pre-training of the AI model. (Item 47) The operations of constructing an AI training set based on aggregated data and narrowing down a number of clusters from the aggregated data include: Identifying clusters within multiple regions; and identifying predicted features within the plurality of regions; Item 38. The method according to Item 37, comprising: (Item 48) 48. The method of claim 47, wherein the identified predictive features are applied to the overall set to generate a predictive score. (Item 49) Item 49. The method of item 48, wherein the operations of identifying clusters, identifying predictive features, and applying the identified predictive features to the overall set to generate a predictive score are processed on multiple cores of multiple servers. (Item 50) simulating a plurality of actual or simulated events for a plurality of items using a plurality of probabilistic probability models, each developed from a set of past events, where each such event has a footprint that has previously occurred or is predicted to occur; and using an artificial intelligence / machine learning (AI) model or engine adapted to be used as a proxy mechanism for achieving predictive analysis from as yet unknown adverse future events for one or more items, wherein the AI model is adapted to evaluate multiple known factors for one or a set of evaluation items; Using said artificial intelligence / machine learning (AI) model, AI can: constructing an AI training set based on the aggregated data; and narrowing down a plurality of clusters from the aggregated data; We use AI for each individual item in each cluster. Comparing the outputs between the clusters to find one or more ranked elements; and identifying individual items via characteristics to provide said training set; wherein the training set is used to evaluate and identify all items across the entire dataset, resulting in a score or other indicator for each individual item that predicts the likelihood of adverse or favorable future impact; wherein the score or index is used to weight each probabilistic probability model and adjust its predictions accordingly based on the score. A method for providing (Item 51) 51. The method of claim 50, wherein the combination of the probabilistic probability model and the AI model is adapted to provide predictive analysis from events to one or a set of endpoints, resulting in improved output and interpretation compared to either method alone. (Item 52) 51. The method of claim 50, wherein the footprint comprises a geographic footprint, a demographic footprint, or a classification footprint. (Item 53) 51. The method of claim 50, wherein the footprint of past events comprises a detailed cataloged set of influences and possible effects on items in a set submitted to the model. (Item 54) Item 54. The method of item 53, wherein the set of impact factors is converted into a severity that indicates the likelihood of adverse damage from the impact factors to items affected by the event footprint. (Item 55) Item 55. The method of item 54, wherein the amount of adverse damage for specific items in the past events is known along with a number of different attributes associated with those items. (Item 56) 10. The method of any of the previous items, wherein the footprints of past events and their damaging effects are used to create an event catalog, or probabilistic set of events, based on multiple data inputs. (Item 57) 10. The method of any of the preceding items, wherein an event catalog is organized by set of simulation periods, each simulation period containing one or more events from said catalog. (Item 58) Item 58. The method of item 57, wherein the arrangement of the simulation periods into sets is adapted to enable each such simulation period to be run against one or more items submitted to the model for assessment. (Item 59) Item 59. The method of item 58, wherein each simulation period is used in the sampling process. (Item 60) 60. The method of any of items 56 to 59, wherein the number of simulation periods is fed into analytical statistics to predict the probability or likelihood of damage to the one or more items. (Item 61) Item 61. The method of item 60, wherein the output from the probabilistic probability model includes an estimated probability of damage.
Claims
1. An artificial intelligence / machine learning (AI) model or engine adapted to be used as a mechanism for achieving predictive analysis from as yet unknown adverse future events for one or more items, wherein the AI model is adapted to evaluate multiple known factors for one or a set of evaluation items; wherein the artificial intelligence / machine learning (AI) model uses AI to: constructing an AI training set based on the aggregated data and narrowing down a plurality of clusters from the aggregated data; Use AI for each individual item in each cluster, Compare the outputs between the clusters to find one or more ranked elements; and Identifying individual items through their characteristics to provide the training set It is adapted to wherein the training set is used to evaluate and identify all items across the entire dataset, resulting in a score or other indicator that predicts the likelihood of adverse or favorable future impact for each individual item; A system comprising:
2. 10. The system of claim 1, wherein the AI model is adapted to identify sets of items with known outcomes from a variety of events, optionally past events and item characteristics, to provide an artificial intelligence training set.
3. The system of claim 2 , wherein the training set includes identified likelihoods of potential events from various item characteristics.
4. The system of claim 3 , wherein further evaluation of items in the clusters is performed to identify item features or characteristics that drive the likelihood of potential events affecting individual items.
5. The system of claim 2 , wherein for each item, a simulation is run on the training set using multiple known attributes or features.
6. The system of claim 5 , wherein a set of derived attributes is generated, the attributes being more or less likely to cause harm, and the developer does not supply such parameters to the model.
7. 10. The system of claim 1, wherein the AI model is trained on a training data set along with a set of discriminatory features or attributes for each item in the training set and is used to predict the likelihood of an event occurring for new items fed to the model.
8. The system of claim 7 , wherein the new item was previously unknown to the model.
9. The system of claim 7 or 8, wherein the identified set of features or attributes learned from the training set is applied to the new item.
10. 10. The system of claim 9, wherein features or attributes of the new item are used by the AI model to predict the likelihood of an event occurrence or lack thereof based on pre-training of the AI model.
11. a plurality of probabilistic probability models each developed from a set of past events, where each such event has a footprint that has previously occurred or is predicted to occur, where the probabilistic probability models are adapted to simulate a plurality of actual or simulated events for a plurality of items; and An artificial intelligence / machine learning (AI) model or engine adapted to be used as an alternative mechanism for achieving predictive analysis from as yet unknown adverse future events for one or more items, wherein the AI model is adapted to evaluate multiple known factors for one or a set of evaluation items; wherein the artificial intelligence / machine learning (AI) model is adapted to: use AI to build an AI training set based on aggregated data, refine a plurality of clusters from the aggregated data, use AI on individual items in each cluster, and compare outputs between clusters to find one or more ranked elements; and identify individual items via characteristics to provide the training set. wherein the training set is used to evaluate and identify all items across the entire dataset, resulting in a score or other indicator for each individual item that predicts the likelihood of adverse or favorable future impact; wherein the score or index is used to weight each probabilistic probability model and adjust its predictions accordingly based on the score. A system comprising:
12. 12. The system of claim 11, wherein the combination of the probabilistic probability model and the AI model is adapted to provide predictive analysis from events to one or a set of endpoints, producing improved output and interpretation compared to either method alone.
13. The system of claim 11 , wherein the footprint comprises a geographic footprint, a demographic footprint, or a categorical footprint.
14. The system of claim 11 , wherein the footprint of past events comprises a detailed cataloged set of influences and possible impacts on items in a set submitted to the model.
15. The system of claim 14 , wherein the set of impact factors is converted into a severity that indicates the likelihood of adverse damage from the impact factors to items affected by the event footprint.
16. 16. The system of claim 15, wherein the amount of adverse damage for particular items in the past events is known along with a number of different attributes associated with those items.
17. The system of claim 11 , wherein the footprints of past events and their damaging effects are used to create an event catalog, or probabilistic set of events, based on multiple data inputs.
18. The system of claim 11 , wherein the event catalog is organized by a set of simulation periods, each simulation period including one or more events from the event catalog.
19. 20. The system of claim 18, wherein the arrangement of the simulation periods into sets is adapted to enable each such simulation period to be run against one or more items submitted to the model for assessment.
20. 20. The system of claim 19, wherein each simulation period is used in a sampling process.
21. 21. The system of any one of claims 18 to 20, wherein the number of simulation periods is fed into analytical statistics to predict the probability or likelihood of damage to the one or more items.
22. 22. The system of claim 21, wherein output from the probabilistic probability model comprises an estimated probability of damage.
23. 12. The system of claim 11, wherein the AI model is adapted to identify sets of items with known outcomes from a variety of events, optionally past events and item characteristics, to provide an artificial intelligence training set.
24. 24. The system of claim 23, wherein the training set includes identified likelihoods of potential events from various item characteristics.
25. 25. The system of claim 24, wherein further evaluation of items in the clusters is performed to identify item features or characteristics that drive the likelihood of potential events affecting individual items.
26. 24. The system of claim 23, wherein for each item, a simulation is run on the training set using multiple known attributes or features.
27. 27. The system of claim 26, wherein a set of derived attributes is generated, the attributes being more or less likely to cause harm, and the model developer does not supply such parameters to the model.
28. 12. The system of claim 11, wherein the AI model is trained on a training data set along with a set of discriminatory features or attributes for each item in the training set and is used to predict the likelihood of an event occurring for new items fed to the model.
29. 30. The system of claim 28, wherein the new item was previously unknown to the model.
30. 30. The system of claim 28 or 29, wherein the identified set of features or attributes learned from the training set is applied to the new item.
31. 30. The system of claim 28, wherein features or attributes of the new item are used by the AI model to predict the likelihood of an event occurrence or lack thereof based on pre-training of the AI model.
32. 1. A method for discretizing aggregate data, comprising: Use AI to build an AI training set based on rough features and refine the clusters; Using AI for individual items in each cluster (positive and negative clusters); Comparing outputs between clusters to find ranked factors that drive positive or negative outcomes; and Identifying individual items through characteristics comprising said training set. A method for providing
33. 33. The method of claim 32, wherein the score is derived from a weighting of features and / or attributes.
34. 34. The method of claim 33, wherein individual items are scored based on the training set and the identified features.
35. The system of claim 1 , wherein the AI model or engine is processed across multiple parallel processes running on multiple processors or cores.
36. 36. The system of claim 35, wherein individual clusters of the plurality of clusters are processed using different processors or cores of the plurality of processors or cores.
37. Using an artificial intelligence / machine learning (AI) model or engine as a mechanism for achieving predictive analysis of as-yet-unknown adverse future events for one or more items, wherein the AI model is adapted to evaluate multiple known factors for one or a set of evaluation items; Using the artificial intelligence / machine learning (AI) model AI: constructing an AI training set based on the aggregated data and narrowing down a plurality of clusters from the aggregated data; Use AI for each individual item in each cluster, Comparing the outputs between the clusters to find one or more ranked elements; and identifying individual items via characteristics to provide said training set; wherein the training set is used to evaluate and identify all items across the entire dataset, resulting in a score or other indicator that predicts the likelihood of adverse or favorable future impact for each individual item; A method for providing
38. 38. The method of claim 37, wherein the AI model is adapted to identify sets of items with known outcomes from a variety of events, optionally past events and item characteristics, to provide an artificial intelligence training set.
39. 39. The method of claim 38, wherein the training set includes identified likelihoods of potential events from various item features.
40. 40. The method of claim 39, wherein further evaluation of items in the clusters is performed to identify item features or characteristics that drive the likelihood of potential events affecting individual items.
41. 39. The method of claim 38, wherein for each item, a simulation is run on the training set using multiple known attributes or features.
42. 42. The method of claim 41, wherein a set of derived attributes is generated, the attributes being more or less likely to cause harm, and the developer does not supply such parameters to the model.
43. 38. The method of claim 37, wherein the AI model is trained on a training data set together with a set of discriminatory features or attributes for each item in the training set and is used to predict the likelihood of an event occurring for new items fed to the model.
44. 44. The method of claim 43, wherein the new item was previously unknown to the model.
45. 45. The method of claim 43 or 44, wherein the identified set of features or attributes learned from the training set is applied to the new item.
46. 46. The method of claim 45, wherein features or attributes of the new item are used by the AI model to predict the likelihood of an event occurrence or lack thereof based on pre-training of the AI model.
47. The operations of constructing an AI training set based on aggregated data and narrowing down a plurality of clusters from the aggregated data include: Identifying clusters within the plurality of regions; and identifying predicted features within the plurality of regions; 38. The method of claim 37, comprising:
48. 48. The method of claim 47, wherein the identified predictive features are applied to the entire set to generate a predictive score.
49. 49. The method of claim 48, wherein the operations of identifying clusters, identifying predictive features, and applying the identified predictive features to the overall set to generate a predictive score are processed on multiple cores of multiple servers.
50. simulating a plurality of actual or simulated events for a plurality of items using a plurality of probabilistic probability models, each developed from a set of past events, where each such event has a footprint that has previously occurred or is predicted to occur; and using an artificial intelligence / machine learning (AI) model or engine adapted to be used as a proxy mechanism for achieving predictive analysis from as yet unknown adverse future events for one or more items, wherein the AI model is adapted to evaluate a plurality of known factors for one or a set of evaluation items; Using the artificial intelligence / machine learning (AI) model, using AI: Constructing an AI training set based on the aggregated data, and narrowing down a plurality of clusters from the aggregated data; Use AI for each individual item in each cluster, Comparing the outputs between the clusters to find one or more ranked elements; and identifying individual items via characteristics to provide said training set; wherein the training set is used to evaluate and identify all items across the entire dataset, resulting in a score or other indicator for each individual item that predicts the likelihood of adverse or favorable future impact; wherein the score or index is used to weight each probabilistic probability model and adjust its predictions accordingly based on the score. A method for providing
51. 51. The method of claim 50, wherein the combination of the probabilistic probability model and the AI model is adapted to provide predictive analysis from events to one or set of endpoints, producing improved output and interpretation compared to either method alone.
52. 51. The method of claim 50, wherein the footprint comprises a geographic footprint, a demographic footprint, or a classification footprint.
53. 51. The method of claim 50, wherein the footprint of past events comprises a detailed cataloged set of influences and possible effects on items in a set submitted to the model.
54. 54. The method of claim 53, wherein the set of impact forces is converted into a severity indicating the likelihood of adverse damage from the impact forces to items affected by the event footprint.
55. 55. The method of claim 54, wherein the amount of adverse damage for particular items in the past events is known along with a number of different attributes associated with those items.
56. 51. The method of claim 50, wherein the footprints of past events and their damaging effects are used to create an event catalog, or probabilistic set of events, based on multiple data inputs.
57. 51. The method of claim 50, wherein the event catalog is organized by a set of simulation periods, each simulation period including one or more events from the event catalog.
58. 58. The method of claim 57, wherein the arrangement of the simulation periods into sets is adapted to enable each such simulation period to be run against one or more items submitted to the model for assessment.
59. 59. The method of claim 58, wherein each simulation period is used in the sampling process.
60. 60. The method of any one of claims 57 to 59, wherein the number of simulation periods is fed into analytical statistics to predict the probability or likelihood of damage to the one or more items.
61. 61. The method of claim 60, wherein output from the probabilistic probability model comprises an estimated probability of damage.