System and method for determining model suitability and stability of model deployment in automated model generation.

JP2026139635APending Publication Date: 2026-09-01ORACLE INT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026074892
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-01-27
Filing Date
2026-04-28
Publication Date
2026-09-01

Smart Images

  • Figure 2026139635000001_ABST
    Figure 2026139635000001_ABST
Patent Text Reader

Abstract

This invention provides a system, method, and non-temporary computer-readable storage medium for use with a computing environment to provide determination of model suitability and stability for model deployment and automated model generation. [Solution] A system comprising a model fit and stability component that provides one or more functions to support model selection, use of model deployability scores and deployability flags, and mitigation of model drift risk in order to determine model fit and stability for a specific application, and which is used with analytical applications, data analytics or other types of computing environments to provide directly actionable risk prediction in financial applications or other types of applications.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Claim of Priority This application claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 142,826, filed on January 28, 2021, entitled "SYSTEM AND METHOD FOR DETERMINATION OF MODEL FITNESS AND STABILITY FOR MODEL DEPLOYMENT IN AUTOMATED MODEL GENERATION", and U.S. Patent Application No. 17 / 586,639, filed on January 27, 2022, entitled "SYSTEM AND METHOD FOR DETERMINATION OF MODEL FITNESS AND STABILITY FOR MODEL DEPLOYMENT IN AUTOMATED MODEL GENERATION", the contents of each of the foregoing applications are incorporated herein by reference.

[0002] Technical Field Embodiments described herein generally relate to systems and methods for providing determination of model fitness and stability for data models, data analytics environments, model deployment and automated model generation. [Background Art]

[0003] Background With respect to systems for supporting data analytics and processes addressing specific customer requirements, such as processes for predicting accounts receivable in a customer's financial application, it may be conjectured that different customers may require generation of different models approximating the characteristics of their underlying data-generating business processes.

[0004] Such models may differ for similar processes in different departments within a client company. Furthermore, over time, data generation business processes may change, and the characteristic distribution of inputs to those processes may also change. [Overview of the project] [Means for solving the problem]

[0005] summary According to one embodiment, this specification describes a system and method used with a computing environment for providing model fit and stability determination for model deployment and automated model generation. The model fit and stability component may provide one or more functions to support model selection, the use of model deployability scores and deployability flags, and mitigation of model drift risk in order to determine model fit and stability for a particular application. For example, the embodiment may be used with an analytical application, data analytics, or other type of computing environment to provide directly actionable risk prediction, for example, in a financial application or other type of application. [Brief explanation of the drawing]

[0006] [Figure 1] This figure shows an exemplary data analytics environment according to one embodiment. [Figure 2] This figure further illustrates an exemplary data analytics environment according to one embodiment. [Figure 3] Further illustrating is an exemplary data analytics environment according to one embodiment. [Figure 4] This figure further illustrates an exemplary data analytics environment according to one embodiment. [Figure 5] This figure further illustrates an exemplary data analytics environment according to one embodiment. [Figure 6]This figure shows a method for determining model fit and stability, used in relation to a data analytics environment according to one embodiment. [Figure 7] This figure shows comparative examples of probability scores for various models according to one embodiment. [Figure 8] This figure shows a process or method for determining model fit and stability according to one embodiment. [Figure 9] This figure further illustrates a process or method for determining model fit and stability according to one embodiment. [Figure 10] This is a diagram illustrating a sorted list of invoices according to one embodiment. [Figure 11] This figure illustrates the output of a data analysis model according to one embodiment. [Figure 12] This flowchart shows a method for determining model suitability and stability in model deployment during automated model generation according to one embodiment. [Modes for carrying out the invention]

[0007] Detailed explanation As mentioned above, with regard to systems that support data analytics and processes that address specific customer requirements (for example, predicting accounts receivable in a customer's financial application), it can be assumed that different customers may require the generation of different models that approximate the characteristics of their underlying data generation business processes.

[0008] Such models may differ for similar processes in different departments within a client company. Furthermore, over time, data generation business processes may change, and the characteristic distribution of inputs to those processes may also change.

[0009] According to one embodiment, the foregoing describes a system and method used with a computing environment for providing model suitability and stability determination for model deployment and automatic model generation. The model suitability and stability component can provide one or more functions to support model selection, use of model deployability scores and deployability flags, and mitigation of model drift risk in order to determine the model suitability and stability of a particular application.

[0010] Depending on the various embodiments, the approach described can be used to address a variety of considerations, for example, such as the following:

[0011] Model fitting benefits from automation because manual methods are prohibitively expensive in terms of time and money. When systems and methods use data samples to create classes of models for a company, the system doesn't have the opportunity to manually tune models with expert data scientists using case-by-case customer data, because with thousands of customers and the unpredictable changes in each dataset, manually examining and tuning models based on the data is prohibitively expensive. The approach described here uses a broad set of specific model classes to represent the maximum identifiable content that can be extracted from customer datasets. The suitability can be systematically identified.

[0012] Furthermore, the use of scores necessitates the automated generation of new models that account for changes over time and across sectors, automatically filtering thousands of potential model candidates using appropriate metrics without requiring human intervention, and finding the most important actionable insights based on predictions. The approach described addresses this particular problem in the space of binary classification models and can be extended to multi-class classification.

[0013] Model drift risk should be mitigated. While model accuracy metrics can change significantly due to drift between the training distribution and the test distribution, systems and methods cannot simply use model accuracy metrics as a criterion for model selection. When the input distribution or the distribution of samples of a population collected on a specific day or week changes, it is expected that the new model will exhibit large drift in the decision boundary, to the extent that classification is reversed for multiple instances—for example, an invoice classified as highly likely to have been paid yesterday is classified as highly likely to be unpaid today. The approach described can be used to investigate how much the scoring distribution has shifted from the training distribution, and how much the shift between training distributions has changed over time.

[0014] Models must be stable. If it is detected that a model is unstable to the extent that the decision boundary drifts significantly every day, this indicates that there are multiple problems with the model fitting. In such cases, classification decisions will continue to change day to day to the extent that they reverse previous day's predictions, even if the data for individual instances does not change. The approach described can be used to detect such instability.

[0015] Data analytics environment Generally described data analytics enable computer-based investigation or analysis of large volumes of data to derive conclusions or other information from the data. In contrast, business intelligence (BI) tools provide business users of an organization with information describing enterprise data in a format that enables business users to make strategic business decisions.

[0016] As examples of data analytics environments and business intelligence tools / servers include Oracle Business Intelligence Server (OBIS), Oracle Analytics Cloud (OAC), and Oracle Fusion Analytics Warehouse (FAW), which support functions such as data mining or analytics, and analysis applications.

[0017] Figure 1 illustrates an exemplary data analytics environment according to one embodiment. The exemplary embodiment shown in Figure 1 is provided for the purpose of illustrating an example of a data analytics environment in which various embodiments described herein can be used. According to other embodiments and examples, the approach described herein can be used with other types of data analytics, database, or data warehouse environments. The components and processes illustrated in Figure 1 and further described herein with respect to various other embodiments may be provided, for example, as software or program code executable by a cloud computing system or other appropriately programmed computer system.

[0018] As shown in Figure 1, according to one embodiment, data analytics environment 100 includes computer hardware (e.g., processor, memory) 101, and a control plane 10 2, and operates as a data plane 104, may be provided by or otherwise operate on a computer system including one or more software components that provide access to a data warehouse, a data warehouse instance 160 (a database 161, or other types of data sources).

[0019] According to one embodiment, the control plane operates to provide control over cloud or other software products provided within the context of a SaaS or cloud environment, such as the Oracle® Analytics Cloud environment or other types of cloud environments. For example, according to one embodiment, the control plane may include a console interface 110 that enables access by customers (tenants) and / or the cloud environment having provisioning components 111.

[0020] According to one embodiment, the console interface can enable access by customers (tenants) operating a graphical user interface (GUI) and / or a command-line interface (CLI) or other interface, and / or may include interfaces used by the provider of the SaaS or cloud environment and its customers (tenants). For example, according to one embodiment, the console interface can provide an interface that enables customers to provision services for use within the SaaS environment and configure those provisioned services.

[0021] According to one embodiment, a customer (tenant) can request the provisioning of a customer schema within a data warehouse. The customer can also provide a number of attributes related to the data warehouse instance, including required attributes (e.g., login credentials) and optional attributes (e.g., size or speed), via a console interface. The provisioning component can then provision the requested data warehouse instance containing the customer schema and populate the data warehouse instance with the appropriate information provided by the customer.

[0022] According to one embodiment, the provisioning component can also be used to update or edit data warehouse instances and / or ETL processes operating in the data plane, for example, by changing or updating the frequency at which requested ETL processes are executed for a particular customer (tenant).

[0023] According to one embodiment, the data plane may include a data pipeline or process layer 120 and a data transformation layer 134, which together process operational or transactional data from an organization's enterprise software applications or data environment, such as business productivity software applications provisioned in a customer's (tenant's) SaaS environment. The data pipeline or process may include various functions for extracting transactional data from business applications and databases provisioned in the SaaS environment and loading the transformed data into a data warehouse.

[0024] According to one embodiment, the data transformation layer is used to transform transactional data received by the system from business applications provisioned in a SaaS environment and corresponding transactional databases into a model format that can be understood by the data analytics environment, such as a knowledge model (K This may include data models such as M) or other types of data models. The model format can be provided in any data format suitable for storage in a data warehouse. According to one embodiment, the data plane may also include a data configuration user interface, as well as a mapping configuration database.

[0025] According to one embodiment, the data plane is responsible for performing extract, transform, and load (ETL) operations, which include extracting transactional data from an organization's enterprise software applications or data environment, such as business productivity software applications and corresponding transactional databases provided in a SaaS environment, transforming the extracted data into a model format, and loading the transformed data into a customer schema in a data warehouse.

[0026] For example, according to one embodiment, each customer (tenant) in the environment can associate with its own customer tenancy within the data warehouse (associated with its own customer schema), and can be provided with read-only access to the data analytics schema, which can be updated periodically or otherwise by a data pipeline or process, such as an ETL process.

[0027] According to one embodiment, a data pipeline or process can be scheduled to run at intervals (e.g., hourly / daily / weekly) to extract transactional data from enterprise software applications or data environments, such as business productivity software applications and corresponding transactional databases 106 provisioned in a SaaS environment.

[0028] According to one embodiment, the extraction process 108 can extract transactional data, and during extraction, the data pipeline or process can insert the extracted data into a data staging area, which can function as a temporary staging area for the extracted data. Data quality components and data protection components can be used to ensure the integrity of the extracted data. For example, according to one embodiment, the data quality component can perform validation on the extracted data while the data is temporarily held in the data staging area.

[0029] According to one embodiment, once the extraction process has completed the extraction, the data transformation layer can be used to initiate a transformation process that converts the extracted data into a model format to be loaded into the customer schema of the data warehouse.

[0030] According to one embodiment, a data pipeline or process can work with a data transformation layer to transform data into a model format. A mapping configuration database can store metadata and data mappings that define the data models used by the data transformation. A data configuration user interface (UI) can facilitate access to and modification of the mapping configuration database.

[0031] According to one embodiment, the data transformation layer can transform the extracted data into a format suitable for loading into a customer schema of a data warehouse, for example, according to a data model. During the transformation, the data transformation may perform dimension generation, fact generation, and aggregation generation as appropriate. Dimension generation may include generating dimensions or fields for loading into a data warehouse instance.

[0032] According to one embodiment, after the extracted data has been transformed, the data pipeline or process can execute a warehouse load procedure 150 to load the transformed data into the customer schema of the data warehouse instance. After loading the transformed data into the customer schema, the transformed data can be analyzed and used in various additional business intelligence processes.

[0033] Customers with different data analytics environments may have different requirements regarding how their data is classified, aggregated, or transformed, whether for the purpose of providing data analytics or business intelligence data, or for the purpose of developing software analytics applications. According to one embodiment, to support such different requirements, the semantic layer 180 may contain data that defines a semantic model of the customer's data (which is useful in supporting users in understanding and accessing the data using commonly understood business terminology), and may provide custom content to the presentation layer 190.

[0034] According to one embodiment, the semantic model can be defined, for example, as a BI repository (RPD) file in an Oracle environment, and has metadata that defines logical schemas, physical schemas, physical-to-logical mappings, aggregation table navigation, and / or other constructs that realize various physical, business model and mapping, and presentation layer aspects of the semantic model.

[0035] According to one embodiment, a customer can modify the data source model to support specific requirements by, for example, adding custom facts or dimensions related to data stored in a data warehouse instance, and the system can extend the semantic model accordingly.

[0036] According to one embodiment, the presentation layer can enable access to data content using, for example, software analytics applications, user interfaces, dashboards, key performance indicators (KPIs), or other types of reports or interfaces that may be provided by products such as Oracle Analytics Cloud or Oracle Analytics for Applications.

[0037] According to one embodiment, the query engine 18 (for example, OBIS) operates like a collaborative query engine, and via SQL, for example, Oracle Analytics It provides analytical queries within a cloud environment, pushes the operation down to supported databases, and translates business user queries into the appropriate database-specific query language (e.g., Oracle SQL, SQL Server SQL, DB2 SQL, or Essbase MDX). The query engine (such as OBIS) also supports internal execution of SQL operators that cannot be pushed down to the database.

[0038] According to one embodiment, a user / developer can interact with a client computer device 10 which includes computer hardware 11 (e.g., processor, storage, memory), a user interface 12, and an application 14. A query engine or business intelligence server such as OBIS generally handles inbound requests to a database model, such as SQL requests, constructs and executes one or more physical database queries, processes data appropriately, and responds to requests. It operates in a way that returns data.

[0039] To achieve this, according to one embodiment, a query engine or business intelligence server may include various components or functions, such as a logical model or business model or metadata that describes the data available as the scope of the query, a request generator that receives incoming queries and translates them into physical queries used by connected data sources, and a navigator that receives incoming queries, navigates the logical model, and generates physical queries that best return the data required for a particular query.

[0040] For example, according to one embodiment, a query engine or business intelligence server can employ a logical model mapped to data in a data warehouse, which allows users to query data as if it originated from a single source, by creating a simplified star schema business model on various data sources. The information can then be returned to the presentation layer as a subject area according to business model layer mapping rules.

[0041] According to one embodiment, a query engine (e.g., OBIS) can process queries against a database according to a query execution plan 56 which may include various child (leaf) nodes commonly called RqLists in various embodiments, and which generates one or more diagnostic log entries. Within the query execution plan, each execution plan component (RqList) represents a block of queries in the query execution plan, which is generally translated into SELECT statements. An RqList may have nested child RqLists, just as a SELECT statement can select from nested SELECT statements.

[0042] According to one embodiment, during operation, the query engine or business intelligence server can create a query execution plan, which can then be further optimized, for example, to perform data aggregation necessary to respond to a request. For example, the data can be joined and calculations applied before the results are returned to the calling application via the ODBC interface.

[0043] According to one embodiment, a complex multipath request requiring multiple data sources may require a query engine or business intelligence server to break down the query, determine which sources, multipath calculations, and aggregations can be used, and generate a logical query execution plan spanning multiple databases and physical SQL statements. The results are then returned by the query engine or business intelligence server and further joined or aggregated.

[0044] Figure 2 further illustrates an exemplary data analytics environment according to one embodiment. As shown in Figure 2, according to one embodiment, the provisioning component may also include a provisioning application programming interface (API) 112, a number of workers 115, a metering manager 116, and a data plane API 118, as further described below. The console interface can communicate with the provisioning API by making API calls when commands, instructions, or other inputs are received at the console interface, for example, to provision a service within a SaaS environment or to make configuration changes to a provisioned service.

[0045] According to one embodiment, a data plane API can communicate with the data plane. For example, according to one embodiment, a service provided by the data plane can be directed to the data plane. Lobbying and configuration changes can be communicated to the data plane via the data plane API.

[0046] According to one embodiment, the metering manager provides a control plane via the provision It may include various functions for measuring the service and usage of the metered services. For example, according to one embodiment, the metering manager may record, for billing purposes, the time-series usage of processors provisioned via the control plane for a particular customer (tenant). Similarly, the metering manager may record, for billing purposes, the amount of storage space in a data warehouse partitioned for use by customers in a SaaS environment.

[0047] According to one embodiment, the data pipeline or process provided by the data plane may include a monitoring component 122, a data staging component 124, a data quality component 126, and a data projection component 128, as further described below.

[0048] According to one embodiment, the data transformation layer may include a dimension generation component 136, a fact generation component 138, and an aggregation generation component 140, as further described below. The data plane may also include a data configuration user interface 130 and a mapping configuration database 132.

[0049] According to one embodiment, the data warehouse may include a default data analytics schema (referred to herein as the analytics warehouse schema, according to some embodiments) 162 and a customer schema 164 for each customer (tenant) of the system.

[0050] According to one embodiment, in order to support multiple tenants, the system can enable the use of multiple data warehouses or data warehouse instances. For example, according to one embodiment, a first warehouse customer tenancy for a first tenant may include a first database instance, a first staging area, and a first data warehouse instance of multiple data warehouses or data warehouse instances, and a second customer tenant for a second tenant may include a second database instance, a second staging area, and a second data warehouse instance of multiple data warehouses or data warehouse instances.

[0051] According to one embodiment, based on a data model defined in a mapping configuration database, a monitoring component can determine the dependencies between multiple different datasets to be transformed. Based on the determined dependencies, the monitoring component can determine which of the multiple different datasets should be transformed into the model format first.

[0052] For example, according to one embodiment, if the first model dataset does not include dependencies on other model datasets, and the second model dataset does include dependencies on the first model dataset, the monitoring component may decide to transform the first dataset before the second dataset to accommodate the dependency of the second dataset on the first dataset.

[0053] For example, according to one embodiment, the dimension may include categories of data such as "name," "address," or "age." Fact generation can take data This involves generating values, or "scales." These facts can then be associated with appropriate dimensions within the data warehouse instance. Aggregation generation involves creating data mappings that compute aggregations of the transformed data against existing data in the customer schema of the data warehouse instance.

[0054] According to one embodiment, once any transformation is performed (as defined by the data model), a data pipeline or process can read the source data, apply the transformation, and push the data to a data warehouse instance.

[0055] According to one embodiment, data transformations can be expressed by rules, and once the transformation is performed, the values ​​can be held intermediately in a staging area. Data quality components and data projection components can verify and check the integrity of the transformed data before the data is uploaded to the customer schema in the data warehouse instance. Monitoring can be provided when the extraction, transformation, and loading processes are performed, for example, on a large number of compute instances or virtual machines. Dependencies can also be maintained during the extraction, transformation, and loading processes, and the data pipeline or process can handle such ordering decisions.

[0056] According to one embodiment, after the extracted data has been transformed, the data pipeline or process can execute a warehouse load procedure to load the transformed data into the customer schema of the data warehouse instance. After loading the transformed data into the customer schema, the transformed data can be analyzed and used in various additional business intelligence processes.

[0057] Figure 3 further illustrates an exemplary data analytics environment according to one embodiment. As shown in Figure 3, according to one embodiment, data can be supplied, for example, from a customer's (tenant's) enterprise software application or data environment (106) or as custom data 109 supplied from one or more customer-specific applications 107 using a data pipeline process, and in some examples, it can be loaded into a data warehouse instance, which includes the use of object storage 105 for storing the data.

[0058] For example, in an analytics environment such as Oracle Analytics Cloud (OAC), a user can create a dataset using tables from different connections and schemas. The system then uses the relationships defined between these tables to create relationships or joins within the dataset.

[0059] According to one embodiment, for each customer (tenant), the system prepopulates the customer's data warehouse instance based on an analysis of data within the customer's enterprise application environment and customer tenancy 117, using a data analytics schema maintained and updated by the system within the system / cloud tenancy 114. In this way, the data analytics schema maintained by the system enables a data pipeline or process to retrieve data from the customer's environment and load it into the customer's data warehouse instance.

[0060] According to one embodiment, the system also provides a customer schema for each customer of the environment, which is easily modifiable by the customer and allows the customer to supplement and utilize data within its own data warehouse instance. For each customer, the resulting data warehouse instance operates as a database whose contents are partially controlled by the customer and partially controlled by the environment (system).

[0061] For example, according to one embodiment, a data warehouse (e.g., Oracle Autonomous Data Warehouse: ADW) may include a data analytics schema and, for each customer / tenant, a customer schema supplied from its enterprise software application or data environment. Data provisioned in a data warehouse tenancy (e.g., an ADW cloud tenancy) is accessible only to that tenant, while simultaneously enabling access to various shared environment functions, such as ETL-related or other functions.

[0062] According to one embodiment, in order to support multiple customers / tenants, the system allows the use of multiple data warehouse instances, for example, a first customer tenancy may include a first database instance, a first staging area, and a first data warehouse instance, and a second customer tenancy may include a second database instance, a second staging area, and a second data warehouse instance.

[0063] According to one embodiment, for a specific customer / tenant, when extracting its data, the data pipeline or process can insert the extracted data into the tenant's data staging area, which can function as a temporary staging area for the extracted data. The integrity of the extracted data can be ensured, for example, by performing validation of the extracted data while it is temporarily held in the data staging area, using data quality and data protection components. Once the extraction process completes the extraction, a data transformation layer can be used to initiate a transformation process, converting the extracted data into a model format that can be loaded into the customer schema of the data warehouse.

[0064] Figure 4 further illustrates an exemplary data analytics environment according to one embodiment. As shown in Figure 4, according to one embodiment, for example, the process of extracting data as customer data supplied from a customer (tenant) enterprise software application or data environment, or from one or more customer-specific applications, using the data pipeline process described above, loading the data into a data warehouse instance, or refreshing the data within the data warehouse, generally includes three main stages, which are performed by an ETP service 160 or process including one or more extraction services 163, a transformation service 165, and a load / publish service 167, and are performed by one or more compute instances 170.

[0065] For example, according to one embodiment, a list of view objects for extraction can be submitted to an Oracle BI Cloud Connector (BICC) component, for example, via a REST call. The extracted files can be uploaded to an object storage component, such as an Oracle Storage Service (OSS) component, for storing the data. The transformation process ingests the data files from the object storage component (such as OSS) and loads them into a target data warehouse (such as an ADW database) that is internal to the data pipeline or process and not exposed to the customer (tenant), while applying business logic. The load / publish service or process ingests the data from the ADW database or warehouse, for example, and exposes it to a data warehouse instance accessible to the customer (tenant).

[0066] Figure 5 further illustrates an exemplary data analytics environment according to one embodiment. Figure 5 shows the operation of a system with multiple tenants (customers) according to one embodiment. For example, data can be supplied from each of the enterprise software applications or data environments of multiple customers (tenants) and loaded into a data warehouse instance using the data pipeline process described above.

[0067] According to one embodiment, a data pipeline or process maintains a data analytics schema for each of several customers (tenants), for example, customer A180, customer B182, which is periodically updated by the system in accordance with best practices for specific analytics use cases.

[0068] According to one embodiment, for each of several customers (e.g., customer A, customer B), the system prepopulates the customer's data warehouse instances based on an analysis of data within the customer's enterprise application environments 106A, 106B and within each customer's tenancy (e.g., customer A tenancy 181, customer B tenancy 183), using data analytics schemas 162A, 162B maintained and updated by the system, thereby retrieving data from the customer's environment by a data pipeline or process and loading it into the customer's data warehouse instances 160A, 160B.

[0069] According to one embodiment, the data analytics environment also provides, for each of the multiple customers of the environment, customer schemas (e.g., Customer A schema 164A, Customer B schema 164B) that are easily modifiable by the customer and allow the customer to supplement and utilize data within its own data warehouse instance.

[0070] As described above, according to one embodiment, for each of the multiple customers of the data analytics environment, the resulting data warehouse instance operates as a database whose contents are partially controlled by the customer and partially controlled by the data analytics environment (system), and the database will have appropriate data taken from the enterprise application environment pre-populated to accommodate various analytics use cases. Once the extraction processes 108A,108B for a particular customer have completed the extraction, a data transformation layer can be used to initiate a transformation process to convert the extracted data into a model format that will be loaded into the customer schema of the data warehouse.

[0071] According to one embodiment, the startup plan 186 can be used to control the operation of a data pipeline or process service for a customer in a specific functional area in order to meet the specific needs of the customer (tenant).

[0072] For example, according to one embodiment, a launch plan can define a number of extraction, transformation, and load (publish) services or steps to be executed in a specific order, during a specific time period, and within a specific time window.

[0073] According to one embodiment, each customer can be associated with its own launch plan(s). For example, the launch plan for a first customer A could determine which tables to retrieve from that customer's enterprise software application environment (e.g., an Oracle Fusion Applications environment), or how services and their processes should be executed in sequence. Similarly, the launch plan for a second customer B could determine which tables to retrieve from that customer's enterprise software application environment, or how services and their processes should be executed in sequence.

[0074] Determination of model suitability and stability According to one embodiment, the system may include means for determining model suitability and stability for model deployment and automated model generation.

[0075] Figure 6 shows a model fit and stability determination used in relation to a data analytics environment according to one embodiment.

[0076] For example, as shown in Figure 6, according to one embodiment, the system may include one or more data models 230. A packaged (initial, no additional configuration) model 232 can be used to load data from a customer's enterprise software application or data environment into a data warehouse instance by providing packaged content 234 based on the use of ETL or other data pipelines or processes as described above, and the packaged model can then be used to provide the packaged content to the presentation layer 240. A custom model 236 can be used to extend the packaged model or to provide custom content 238 to the presentation layer.

[0077] According to one embodiment, the presentation layer can enable access to data content using software analytics applications, user interfaces, dashboards, key performance indicators (KPIs), or other types of reports or interfaces, such as those provided by products like Oracle Analytics Cloud or Oracle Analytics for Applications.

[0078] As further shown in Figure 6, according to one embodiment, the system includes a model suitability and stability component 250 and can provide one or more functions to support model selection 252, the use of model deployability score and deployability flag 254, and mitigation of model drift risk 256 in order to determine model suitability and stability for a particular application, as described below.

[0079] In one embodiment, to meet a customer's business need for the automated generation of new models that take into account cross-departmental and temporal changes, the system enables the automatic filtering of thousands of potential model candidates using appropriate metrics, without requiring human intervention, and to identify the most important actionable insights based on predictions.

[0080] Model scoring and selection As described above, according to one embodiment, the system includes a model compatibility and stability component that can provide one or more features to support model selection in order to determine model compatibility and stability for a particular application.

[0081] For example, in binary classification problems such as whether a customer will pay accounts receivable on time, model selection is crucial. In such environments, various classes of metrics can be used to determine the fit of the model.

[0082] According to one embodiment, the first class of metrics tends to weight the upper (e.g., p=[0.8,0.9]) and lower (e.g., p=[0.1,0.2]) parts of the distribution, which are generated by different algorithms without calibration, or they are unevenly distributed such that the highest probability bin (e.g., p=[0.9,1]) has a lower proportion of cases than successively lower probability bins, or This addresses the problem of distorted probability bins, which can result in uneven, sawtooth patterns.

[0083] Model success criteria When a well-calibrated model is deployed, the expected outcome is that instance membership in the stochastic bins decreases steadily and rapidly (e.g., exponentially) from higher bins to lower bins. This indicates that the model classifies most cases with high confidence and only a small number of cases with low confidence.

[0084] According to one embodiment, in order to remove and deploy only such models, the system employs a metric for finding models that satisfy the above criteria and for removing models that exhibit sawtooth frequency features of instances within a probability bin.

[0085] Score based on probability bin Figure 7 shows a comparison of probability scores for various models according to one embodiment.

[0086] As shown in Figure 7, according to one embodiment, the score can be based on probability bins, and the number of correct classifications decreases sharply from higher probability bins to lower probability bins.

[0087] According to one embodiment, as shown in Figure 7, scores are generated for two different models, namely models 710 and 720. Each of models 710 and 720 is an example of a model that can be used to determine whether an invoice is payable or not. As shown in the figure, the model is divided into 10 probability bins. The scoring model shows not only the number of multiple correct classifications but also the number of misclassifications, and the weights of each associated scoring mechanism are also provided. As shown in the figure, model 710 has a nearly linear decrease between each probability bin, while model 720 has an exponential decrease from high probability (0.9-1) to low probability.

[0088] According to one embodiment, result scores 711 and 721 can be determined for each model, indicating that models with an exponential decrease in probability are scored higher, indicating good models that predict the correct outcome with a high probability.

[0089] According to one embodiment, the exemplary scoring function shown below represents a class of functions having a modified stepwise shape, having a descending penalty for non-decreasing number of correctly classified instances from high-probability bins to low-probability bins, and a penalty for all bins for misclassification, normalized by the total number of classified instances:

[0090]

number

[0091] According to one embodiment, a system programmed according to (Equation 1) above takes the following into consideration:

[0092]

number

[0093]

number

[0094] According to one embodiment, another example of a scoring function can be shown as follows:

[0095]

number

[0096] Figure 8 shows a process or method for determining model fit and stability according to one embodiment.

[0097] As shown in Figure 8, according to one embodiment, a score can be determined for a given model of a dataset. When the model generates probabilities (for example, the probability of whether or not an invoice will be paid), the model's output can be grouped into probability "bins," that is, groups of probability ranges. For example, the model's output can be grouped into a probability bin of 10. If this is the case, the bin ranges will be 0-0.1, 0.1-0.2, 0.2-0.3, 0.3-0.4, 0.4-0.5, 0.5-0.6, 0.6-0.7, 0.7-0.8, 0.8-0.9, and 0.9-1.0. The model can be examined by comparing the model's output with the actual result (e.g., whether the invoice was actually paid or not) to determine the number of correct and incorrect classifications for each probability bin.

[0098] In one embodiment, the example discussed and illustrated herein utilizes 10 probability bins to demonstrate the scoring process described herein, but more or fewer probability bins may be used (for example, 100 probability bins, each covering a probability range of 0.01).

[0099] According to one embodiment, in step 810, the scoring process can determine the sequential difference of the positive classifications and, for the low probability bins, apply weights to each of the sequentially lower probability bins. The weights applied to each probability bin can be automatically generated, for example, if the model predicts the outcome with a high probability, the high probability bins can be made heavier so that greater importance is placed on the model being correct.

[0100] According to one embodiment, in step 820, the scoring process may then apply a penalty each time a classification fails.

[0101] According to one embodiment, in step 830, the scoring process can apply weights to the penalties evaluated in step 820 for each probability bin. The weights can be higher for high-probability bins and even exponentially higher, as in step 810. Such penalty weights can also be generated automatically. A higher penalty is applied to failed classifications in high-probability bins, since misclassifications in high-probability bins must similarly result in a lower score compared to failed classifications in low-probability bins.

[0102] According to one embodiment, in step 840, the scoring process can normalize the generated scores by the number of classified samples. That is, for example, normalization may be performed by dividing the generated scores by the number of samples.

[0103] According to one embodiment, in step 850, the scoring process may selectively consider other possibilities, for example, by Monte Carlo simulation, and eliminate low-quality scoring techniques.

[0104] Deployability score and deployability flag As described herein, according to one embodiment, the system includes a model suitability and stability component capable of providing one or more functions that support the use of a model deployability score and deployability flag to determine model suitability and stability for a particular application.

[0105] According to one embodiment, the following approach can be used to determine the deployability score and deployability flag:

[0106]

number

[0107] Here, the system programmed according to (Equation 2) above takes the following into consideration:

[0108]

number

[0109] In one embodiment, the deployability score (ψ) is on a scale of -10 to +10: in the case of perfect classification, ψ is above 10, and in the case of perfect misclassification, ψ is below -10.

[0110] According to one embodiment, the model deployability flag can be defined based on Heaviside's step function as follows:

[0111]

number

[0112] Here, the system programmed according to (Equation 3) above takes the following into consideration:

[0113]

number

[0114] According to one embodiment, the deployability score can be achieved as follows. The system can use the Matthews correlation coefficient (MCC), as shown by (Equation 5) below, in place of all other measures of positive classification, such that precision, recall, accuracy, and F1 score are all explained by the MCC.

[0115] The stochastic bin score is a prerequisite for model hygiene, and if it exceeds the base threshold, it is added to the overall deployability.

[0116] After exceeding the basic hygiene coefficient, the deployability score shows a high correlation with the MCC and improves along with the probability bin score.

[0117] If the deployability score exceeds the threshold in (Equation 3) above, the system will deploy the model. It can be considered possible to use Roy.

[0118] For the initial model deployment, the system can determine that the deployability score > T above.

[0119]

number

[0120] According to one embodiment, to check the deployability of a new model following the original model, the system is deployable as long as the new model continues to have a deployability score > T and is not worse than 1 times the original deployability score when judged by a human. This roughly corresponds to a shift of 0.1 or less in the MCC, F1 score, and Area Under the Curve of the Receiver Operator Characteristic (AUC of ROC).

[0121] According to one embodiment, a second class of known metrics can be used to determine how well the classes are distinguished, such as by determining the relative counts of true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN), such as an F1 score in which type I and type II errors are equally weighted.

[0122] According to one embodiment, the described approach allows the customer to choose to weight recall over precision, and if the customer wants recall to be higher than precision, β can be set to a value greater than 1, and if the customer wants precision to be higher than recall, β can be set to a value less than 1:

[0123]

number

[0124] However, these F-scales and related scales are skewed due to class imbalances, particularly when the actual classes of interest, such as non-payment cases, are rarer than the cases where payment was completed. To address this class imbalance, the system can filter the model using the Matthews correlation coefficient (MCC).

[0125]

number

[0126] The above judgment is close to 1 for perfect correct classification, close to -1 for misclassification, and close to 0 for random classification. According to one embodiment, if the above score is met, a model with an MCC greater than 0.5 can be accepted. The Matthews correlation coefficient is also well applicable to cases of multi-class classification.

[0127] Mitigating the risk of model drift As described above, according to one embodiment, the system comprises a model compatibility and stability component that can provide one or more functions to support the mitigation of model drift risk in order to determine model compatibility and stability for a particular application.

[0128] According to one embodiment, the risk of model drift can be mitigated using model stability detection. While model accuracy metrics can vary significantly due to drift in the training and test distributions, the system and method do not simply use model accuracy metrics as a criterion for model selection. When the input distribution, or the distribution of a population sample taken on a particular day or week, changes, it is expected that there will be significant drift in the decision boundary of the new model, to the extent that classifications are reversed in multiple instances, such as classifying an invoice that was classified as likely to have been paid yesterday as likely not to have been paid today.

[0129] According to one embodiment, the system and method should expect that the forecast will change when new data regarding the same invoice comes in, but should not expect a significant change in the forecast regarding the same invoice if the independent variable remains substantially the same compared to the previous period.

[0130] However, according to one embodiment, a shift in the decision boundary may occur if there is a significant shift in the training distribution over time. These shifts can be explicitly detected and communicated to the end user, for example, when two distributions are substantially divergent from each other to such an extent that their measures of central tendency and variance are statistically significant.

[0131] According to one embodiment, if the system detects that the model is unstable to the extent that the decision boundary drifts significantly (for example, daily), this indicates several problems in model fitting. In such cases, the classification decision will continue to change day by day, to the point of reversing the previous day's prediction, even if the data for individual instances remains unchanged.

[0132] According to one embodiment, the approach described herein can be used to assess the stability of a model using a sensitivity metric such that if a small, random perturbation (less than 5% of the standard deviation of the independent variable) occurs in a subset of class instances of interest and a significant shift in the classification is detected, it can be concluded that a model instability scenario has been reached, or the system and method may be dealing with instances close to the decision boundary. The system can use a normalized distance metric to distinguish between instances close to the decision boundary and instances within a cluster of instances in a given classification.

[0133] According to one embodiment, the change in classification probability jumps significantly, so instability is expected to be observed even in instances close to the center of the class cluster.

[0134] According to one embodiment, the system and method can determine and examine how much the scoring distribution is shifted from the training distribution, and how much the shift between training distributions is over time. For this purpose, the approach described can use the following two combinations of scores.

[0135] Model and Distribution Drift: According to one embodiment, a decrease in the F1 score (a measure of accuracy) and the Matthews correlation coefficient (MCC) is a direct indicator of drift, and whenever the F1 score falls below a threshold (e.g., 0.6) or the MCC falls below a boundary (e.g., 0.35), the system can automatically flag an alarm to request retraining of the model. Kullback-Leibler Divergence Alternatively, evaluate a Bhattacharya distance type scale to train the By determining the shift in the distribution of the input independent variables from the dataset to the scoring dataset, it is possible to determine how much the input distribution deviates from the past training data.

[0136] Model Stability: According to one embodiment, the approach described can be used to provide a scoring mechanism for changes in classification despite negligible changes in the input independent variables.

[0137] Figure 9 further illustrates a process or method for determining model fit and stability according to one embodiment.

[0138] As shown in Figure 9, according to one embodiment, the process can be used to determine whether the model is drifting and whether mitigation is necessary. This process can also be used to determine the risk to the stability of the model.

[0139] For example, the process in Figure 9 can be used to determine if the model is shifting / reversing its predictions (for example, if the number of predictions reverses from "paid" to "unpaid" from one day to the next—this could be an indication of model instability or degradation).

[0140] According to one embodiment, in step 910, the process can detect one or more signals of model degradation under distribution drift. For example, the process can track MCC and AUC scores to determine whether the scores are decreasing. Losses exceeding a threshold can be considered as the model drifting or drifting significantly (e.g., a loss threshold of 0.1 or more). Furthermore, the process can evaluate measures of the type Kullback-Leibler Divergence (also known as relative entropy) or Bhattacharya Distance to determine the shift in the distribution of the input independent variables from the training dataset to the scoring dataset.

[0141] According to one embodiment, in step 920, the process can start the model stability detection and scoring process.

[0142] According to one embodiment, in step 930, the process can determine the distance of each instance (e.g., an invoice) from the nearest neighbor cluster (e.g., 30 neighbors) having the same preclassification. The distance can be calculated, for example, by determining the Mahalanobis distance of each invoice or instance from its closest 30 neighbor clusters having the same preclassification.

[0143] According to one embodiment, in step 940, if the process determines that at least one of these nearest neighbors reverses the classification in the new version of the model, it may add this to the count of reversed classifications.

[0144] According to one embodiment, in step 950, the process can determine the proportion or ratio of such inverted classifications out of the total number of instances to be classified.

[0145] According to one embodiment, in step 960, if such inverted classifications exceed a threshold (for example, 2 percent of the total number of instances where there is no corresponding increase in MCC), the process may flag the model as slightly unstable.

[0146] According to one embodiment, in step 970, if such inverted classifications exceed a second threshold (for example, 10% of the total number of instances where there is no corresponding increase in MCC), the process flags the model as unstable.

[0147] According to one embodiment, the thresholds described above can be set, modified, and / or changed based on input received by the system, for example, by a user or administrator.

[0148] According to one embodiment, the described approach uses a Mahalanobis distance-based measure of the standard deviation-normalized distance between invoices or instances by converting all numerical independent variables (e.g., amount, number of days overdue, number of follow-ups) into z-scores, converting all categorical independent variables (e.g., customer industry, location, invoice type, invoice item type) into entropy-coded renormalized z-scores, and finding the Euclidean distance (if the covariance matrices are identical) between the current invoice and clusters of different invoice types or customer types.

[0149] For example, according to one embodiment, if the invoice distance from a paid invoice is greater than that from an unpaid invoice, the process can assign it to a high-risk category. As shown in Figure 10, the system can present a sorted list of invoices to the user by risk.

[0150] Figure 10 illustrates a sorted list of invoices according to one embodiment. As shown in Figure 10, an exemplary screenshot 1000 may be provided, for example, through the system's user interface. Based on the model selected due to the scoring system described above, various metrics can be provided through the user interface. These include, but are not limited to, the top 10 riskiest invoices and their amounts, the top 10 paid invoices and their amounts, the top 20% of invoices and total risk, the top 20% of invoices and total paid amounts, and so on.

[0151] Figure 11 illustrates the output of a data analysis model according to one embodiment. As shown in Figure 11, an exemplary screenshot 1100 may be provided, for example, via the system's user interface. Based on the model selected due to the scoring system described above, various metrics related to the probability bins can be provided via the user interface. The system can generate such a chart by creating bins with equal probability intervals and creating correlations (e.g., Pearson correlations) between the columns of bins for all numerical variables. The top number (e.g., 5) of the interrelated variables can then be determined.

[0152] Following such a decision, according to one embodiment, the system can determine whether the bin mean of these variables differs from the population mean by at least a certain percentage (e.g., 50%). If the bin mean differs from the population mean by at least, for example, 50%, this variable can be displayed along with a list of descriptions.

[0153] Figure 12 is a flowchart of a method for determining the model suitability and stability of a model deployment in automated model generation, according to one embodiment.

[0154] According to one embodiment, in step 1210, the method can provide a computer including one or more microprocessors, and a data analytics cloud or other computing environment running on the computer.

[0155] According to one embodiment, in step 1220, the method is to perform data analytics class In Udo, multiple models can be provided.

[0156] According to one embodiment, in step 1230, the method can score a set of multiple models based on a set of data in a data analytics cloud.

[0157] According to one embodiment, in step 1240, the method can select a model from a set of multiple models based on scoring.

[0158] According to one embodiment, in step 1220, the method can monitor the model for signs of instability or drift.

[0159] According to various embodiments, the teachings herein can be conveniently implemented using one or more conventional general-purpose or specialized computers, computing devices, machines, or microprocessors, including one or more processors, memory, and / or computer-readable storage media programmed in accordance with the teachings herein. Appropriate software coding can be readily produced by a skilled programmer based on the teachings herein, as will be apparent to those skilled in the art of software technology.

[0160] In some embodiments, the teachings herein may include a computer program product which is one or more non-temporary computer-readable storage media storing instructions, which can be used to program a computer to perform any of the processes of these teachings. Examples of such storage media include, but are not limited to, hard disk drives, hard disks, hard drives, fixed disks, or other electromechanical data storage devices, floppy disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic or optical cards, nanosystems, or other types of storage media or devices suitable for non-temporary storage of instructions and / or data.

[0161] The foregoing descriptions are provided for illustrative and explanatory purposes only and are not intended to be exhaustive or to limit the scope of protection to the exact form disclosed. Many modifications and variations will be apparent to those skilled in the art. For example, some of the examples provided herein illustrate use in a cloud environment such as Oracle Analytics Cloud, but according to various embodiments, the systems and methods described herein can be used in other types of enterprise software applications, cloud environments, cloud services, cloud computing, or other computing environments.

[0162] The embodiments have been selected and described to best illustrate the principles of this teaching and their practical application. This will enable those skilled in the art to understand the various embodiments and the various modifications suitable for specific conceivable uses. The scope is intended to be defined by the appended claims and their equivalents.

Claims

1. A system for determining model suitability and stability in model deployment during automated model generation, A computer comprising one or more microprocessors, and a data analytics cloud or other computing environment running on said computer, The one or more microprocessors described above are: In the aforementioned data analytics cloud, multiple models are provided, Based on the set of data in the aforementioned data analytics cloud, the set of multiple models is scored, Based on the scoring, select a model from the set of multiple models, A system that operates to monitor the model for signs of instability or drift.

2. Scoring the set of the aforementioned multiple models means that for each of the set of the aforementioned multiple models, The predictions of the aforementioned model are automatically assigned to a probability bin within a set of probability bins, Determining the sequential difference of positive classifications between sequential probability bins, This includes applying a weight to each of the sequential differences of the positive classifications between sequential probability bins, The system according to claim 1, wherein the weights applied to each of the sequential differences of the positive classifications are dependent on the probability bin to which the weights are applied.

3. The system according to claim 2, wherein the weight increases with the probability of a bottle being higher.

4. Scoring the set of the aforementioned models further involves, for each of the set of the aforementioned models, A penalty will be applied for each failed classification in the probability bins, The system according to claim 3, further comprising applying a penalty weight to each of the penalties applied for each failed classification.

5. The system according to claim 4, wherein the weight of the penalty increases with the probability of a bin being higher.

6. Scoring the set of the aforementioned models further involves, for each of the set of the aforementioned models, The system according to claim 5, comprising normalizing the generated scores by the number of classified samples.

7. Monitoring the model for signs of instability or drift is, Detecting one or more signals of model degradation, Determining the distance of each instance generated by the model to a cluster of instances having the same preclassification, In the new version of the aforementioned model, it is determined that at least one of the nearest neighbors has inverted classification, The proportion of the inverted classification among the total number of instances generated by the aforementioned model is determined. If the determined ratio exceeds a first threshold, the model is tagged as slightly unstable. If the determined ratio exceeds a second threshold, the model is tagged as unstable. The system according to claim 1, which includes the act of doing the following.

8. A method for determining model suitability and stability in model deployment during automated model generation, To provide a computer including one or more microprocessors, and a data analytics cloud or other computing environment that operates on said computer, In the aforementioned data analytics cloud, multiple models are provided, Based on the set of data in the aforementioned data analytics cloud, the computer scores the set of the multiple models, The computer selects a model from the set of multiple models based on the scoring, A method comprising the computer monitoring the model for signs of instability or drift.

9. Scoring the set of the aforementioned multiple models means that for each of the set of the aforementioned multiple models, The predictions of the aforementioned model are automatically assigned to a probability bin within a set of probability bins, Determining the sequential difference of positive classifications between sequential probability bins, This includes applying a weight to each of the sequential differences of the positive classifications between sequential probability bins, The method according to claim 8, wherein the weight applied to each of the sequential differences of the positive classifications is dependent on the probability bin to which the weight is applied.

10. The method according to claim 9, wherein the weight increases with the probability of a bottle being higher.

11. Scoring the set of the aforementioned models further involves, for each of the set of the aforementioned models, A penalty will be applied for each failed classification in the probability bins, The method according to claim 10, further comprising applying a penalty weight to each of the penalties applied for each failed classification.

12. The method according to claim 11, wherein the weight of the penalty increases with the probability of a bin being higher.

13. Scoring the set of the aforementioned models further involves, for each of the set of the aforementioned models, The method according to claim 12, comprising normalizing the generated score by the number of classified samples.

14. Monitoring the model for signs of instability or drift is, Detecting one or more signals of model degradation, Determining the distance of each instance generated by the model to a cluster of instances having the same preclassification, In the new version of the aforementioned model, it is determined that at least one of the nearest neighbors has inverted classification, The proportion of the inverted classification among the total number of instances generated by the aforementioned model is determined. If the determined ratio exceeds a first threshold, the model is tagged as slightly unstable. The method of claim 8, comprising tagging the model as unstable if the determined ratio exceeds a second threshold.

15. A non-temporary computer-readable storage medium that, when read and executed by one or more computers, stores instructions causing the one or more computers to execute a method, wherein the method is To provide a computer including one or more microprocessors, and a data analytics cloud or other computing environment that operates on said computer, In the aforementioned data analytics cloud, multiple models are provided, Scoring the set of multiple models based on the set of data in the aforementioned data analytics cloud, Based on the aforementioned scoring, select a model from the set of multiple models, A non-transient computer-readable storage medium, including monitoring the model for signs of instability or drift.

16. Scoring the set of the aforementioned multiple models means that for each of the set of the aforementioned multiple models, The predictions of the aforementioned model are automatically assigned to a probability bin within a set of probability bins, Determining the sequential difference of positive classifications between sequential probability bins, This includes applying a weight to each of the sequential differences of the positive classifications between sequential probability bins, The non-temporary computer-readable storage medium according to claim 15, wherein the weights applied to each of the sequential differences of the positive classifications are dependent on the probability bin to which the weights are applied.

17. The non-temporary computer-readable storage medium according to claim 16, wherein the weight increases with the probability of a bin being higher.

18. Scoring the set of the aforementioned models further involves, for each of the set of the aforementioned models, A penalty will be applied for each failed classification in the probability bins, This includes applying a penalty weight to each of the penalties applied for each failed classification, The non-temporary computer-readable storage medium according to claim 17, wherein the weight of the penalty increases with higher probability bins.

19. Scoring the set of the aforementioned models further involves, for each of the set of the aforementioned models, A non-temporary computer-readable storage medium according to claim 18, comprising normalizing the generated score by the number of classified samples.

20. Monitoring the model for signs of instability or drift is, Detecting one or more signals of model degradation, Determining the distance of each instance generated by the model to a cluster of instances having the same preclassification, In the new version of the aforementioned model, it is determined that at least one of the nearest neighbors has inverted classification, The proportion of the inverted classification among the total number of instances generated by the aforementioned model is determined. If the determined ratio exceeds a first threshold, the model is tagged as slightly unstable. A non-temporary computer-readable storage medium according to claim 15, comprising tagging the model as unstable if the determined ratio exceeds a second threshold.