System and method for securely auditing results of machine learning model
The method uses state estimation and MSET to create a tamper-proof dataset for auditing machine learning model results, addressing data integrity and compliance issues in regulatory environments.
Patent Information
- Application Number
- JP2025029274
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-03-23
- Filing Date
- 2025-02-26
- Publication Date
- 2025-07-01
AI Technical Summary
The challenge lies in ensuring the integrity and reliability of sensor data stored in databases used for machine learning models, particularly in industries like power and utilities, where regulatory compliance requires assurance that the data has not been tampered with, as fines can be imposed for non-compliance.
A method involving state estimation calculations, reconstruction of time series data, and comparison to ensure data integrity, using techniques like multivariate state estimation (MSET) to generate a tamper-proof dataset that can be audited for compliance.
Ensures data integrity and compliance by providing a reliable audit mechanism that detects any modifications or tampering, reducing adversarial relationships and ensuring regulatory compliance without requiring hardware upgrades.
Smart Images

Figure 2025098009000001_ABST
Abstract
Description
Background Art
[0001] Background With the reduction of the cost of sensor and processor technologies, sensors are added to or associated with components of electrical systems, mechanical systems, distribution systems, and other systems in order to enable sensor-based automation. These sensors generate a large amount of information that describes the behavior of the components. This large amount of information is stored in a database for at least a certain period of time. By applying machine learning algorithms to the detected information, prediction, anomaly detection / discovery, and predictive maintenance of system components monitored by the sensors can be enabled.
Summary of the Invention
Problems to be Solved by the Invention
[0002] Examples of sensor-based automation can be found in the power industry, oil and gas, environmental and water quality monitoring, data processing, manufacturing, passenger and cargo transportation, and even the financial services sector. In these fields, the behavior of equipment or systems and / or the judgments made by machine learning processes may be subject to inspection by regulatory agencies. Such inspections may be based on stored sensor information, and large fines may be imposed on entities that violate regulations based on stored sensor information that exhibits non-compliant behavior. Therefore, it would be beneficial if the regulatory entities and regulatory agencies could ensure that the stored sensor information is not damaged or tampered with.
[0003] For the above reasons and other reasons, there is a need for a technique to effectively and efficiently ensure that the results of machine learning models to be guaranteed can be audited.
Means for Solving the Problems
[0004] Summary In one embodiment, a computer-implemented method for auditing the results of a machine learning model includes retrieving a set of state estimate values for original time series data values from a database under audit, each of the state estimate values being generated by a state estimation calculation for one of the time series data values. The method further includes, for each of the state estimate values, reversing the state estimation calculation to generate a reconstructed time series data value for each of the state estimate values; retrieving the original time series data values from the database under audit; comparing the original time series data values with the reconstructed time series data values pair by pair to determine whether the original time series and the reconstructed time series match; and generating a signal that the database under audit has not been modified when the original time series and the reconstructed time series match and has been modified when the original time series and the reconstructed time series do not match.
[0005] In one embodiment, a computer-implemented method for auditing the results of a machine learning model further includes training a multivariate state estimation model with a set of training values selected from the original time series values; decomposing a matrix associated with the multivariate state estimation module into a set of eigenvectors; selecting a subset of principal eigenvectors from the set of eigenvectors; and creating a restricted multivariate state estimation model from the subset of principal eigenvectors. Reversing the state estimation calculation further includes generating a reverse of the calculation in which the restricted multivariate state estimation model forms the state estimate value. The method further includes generating a reverse of the calculation in which the restricted multivariate state estimation model forms the state estimate value.
[0006] In one embodiment, a computer-implemented method for auditing the results of a machine learning model further includes excluding training data values from the original time series data values from the database under audit, preprocessing the remaining original time series data values to reduce the impact of sensor faults, and generating a model that records changes to the remaining original time series data, and using the restricted multivariate state estimation model to perform state estimation on the preprocessed remaining original time series data values to create a compressed time series database.
[0007] In one embodiment, a computer-implemented method for auditing the results of a machine learning model further includes preprocessing the original time series values to reduce the impact of sensor faults on the quality of the original time series values, and generating a data preprocessing model that records one or more changes to the original time series values during the preprocessing, and the inversion of the state estimation calculation further includes retrieving the data preprocessing model and reversing the preprocessing described by the data preprocessing model for each of the state estimation values.
[0008] In one embodiment, a computer-implemented method for auditing the results of a machine learning model further includes generating an electronic data report data structure, the electronic data report data structure including a preprocessing model that records one or more changes to the original time series values during preprocessing to reduce the impact of sensor faults on the quality of the original time series values, and a compressed time series database, the compressed time series database being generated by performing state estimation with a restricted multivariate state estimation model and excluding values that are not parameters of the restricted multivariate state estimation model from the compressed time series database, the electronic data report data structure further including one or more of (i) a set of training values selected from the original time series values and used to train the restricted multivariate state estimation model, and (ii) the restricted multivariate state estimation model.
[0009] In one embodiment, a method implemented by a computer for auditing the results of a machine learning model, wherein the inversion of the state estimation calculation further includes identifying an inverse state estimation calculation that cancels the steps performed by the state estimation calculation to form the state estimation value from the original time series data values, and generating a set of inverse state estimation values for the original time series data from the set of state estimation values, each of the inverse state estimation values being generated by performing the inverse state estimation for one of the set of state estimation values, and the reconstructed time series data values for each of the state estimation values being based on the inverse state estimation values.
[0010] In one embodiment, a method implemented by a computer for auditing the results of a machine learning model further includes: (i) generating an electronic verification report message indicating that, in response to a signal that the database under audit has not been modified, the database under audit is not damaged and has been certified as not having been tampered with; (ii) generating an electronic verification report message indicating that the database under audit is damaged or has been tampered with in response to a signal that the database under audit has been modified; and transmitting the generated electronic verification report to a computing device to cause the verification report message to be stored or displayed by the computing device.
[0011] In one embodiment, computer-executable instructions for auditing the results of a machine learning model A non-transitory computer-readable medium for storing computer-executable instructions that, when executed by at least a processor of a computer, cause the computer to retrieve a set of state estimate values for original time-series data values from a database under audit, where the state estimate values are generated by state estimation calculations for each of the time-series data values, and further, for each of the state estimate values, reverse the state estimation calculation to generate a reconstructed time-series data value for each of the state estimate values, retrieve the original time-series data values from the database under audit, compare the original time-series data values with the reconstructed time-series data values pair by pair to determine whether the original time-series and the reconstructed time-series match, and cause the database under audit to generate a signal that (i) has not been modified when the original time-series and the reconstructed time-series match, and (ii) has been modified when the original time-series and the reconstructed time-series do not match.
[0012] In one embodiment, a non-transitory computer-readable medium, the instructions further cause the computer to train a multivariate state estimation model with a set of training values selected from the original time series values, decompose a matrix associated with the multivariate state estimation module into a set of eigenvectors, select a subset of eigenvectors having the largest eigenvalue from the set of eigenvectors, create a restricted multivariate state estimation model from the subset of eigenvectors, preprocess the original time series values to reduce the impact of sensor failures on the quality of the original time series values, create a data preprocessing model that records one or more changes to the original time series values during the preprocessing, perform state estimation on the original time series data values using the restricted multivariate state estimation model to create a compressed time series database, generate a report including the preprocessing model, the compressed time series database, and one or more of (i) the set of training values and (ii) the restricted multivariate state estimation model, the set of state estimation values being the set of state estimation values included in the compressed time series database within the report, and the reconstructed time series data being generated based on the report.
[0013] In one embodiment, a computing system for auditing the results of a machine learning model, the system including a processor, a memory operably connected to the processor, a sensor interface operably connected to the processor and the memory, and a non-transitory computer-readable medium operably connected to the processor and the memory and storing computer-executable instructions, the computer-executable instructions, when executed by at least the processor of the computer, causing the computer to retrieve a set of state estimate values for original time-series data values received via the sensor interface from a database under audit, the state estimate values being generated by state estimation calculations for each of the time-series data values, and further, for each of the state estimate values, reversing the state estimation calculation to generate, for each of the state estimate values, a reconstructed time-series data value, retrieving the original time-series data values from the database under audit, comparing the original time-series data values with the reconstructed time-series data values pair by pair to determine whether the original time series and the reconstructed time series match, and generating a signal that the database under audit has not been modified when (i) the original time series and the reconstructed time series match and has been modified when (ii) the original time series and the reconstructed time series do not match.
[0014] In one embodiment, a computing system for auditing the results of a machine learning model, the non-transitory computer-readable medium further including instructions that, when executed by at least the processor, cause the computing system to train a multivariate state estimation model with a set of training values selected from the original time-series values, and a matrix associated with the multivariate state estimation module to a set of eigenvectors Decompose it, select a subset of the principal eigenvectors from the set of the eigenvectors, create a restricted multivariate state estimation model from the subset of the principal eigenvectors, and the inversion of the state estimation calculation further includes generating an inversion of the calculation in which the restricted multivariate state estimation model forms a state estimation value.
[0015] In one embodiment, a computing system for auditing the results of a machine learning model, the non-transitory computer-readable medium further includes instructions that, when executed by at least the processor, cause the computing system to exclude training data values from the original time series data values from the database under audit, preprocess the remaining original time series data values to reduce the impact of sensor failures, generate a model that records changes to the remaining original time series data, perform state estimation on the preprocessed remaining original time series data values using the restricted multivariate state estimation model to create a compressed time series database.
[0016] In one embodiment, a computing system for auditing the results of a machine learning model, wherein the instructions for the inversion of the state estimation calculation further cause the computing system to generate an electronic data reporting data structure, the electronic data reporting data structure including a preprocessing model that records one or more changes to the original time series values during preprocessing to mitigate the impact of sensor faults on the quality of the original time series values, and a compressed time series database, the compressed time series database being generated by performing state estimation with a restricted multivariate state estimation model and excluding values that are not parameters of the restricted multivariate state estimation model from the compressed time series database, the electronic data reporting data structure further including one or more of (i) a set of training values selected from the original time series values and used to train the restricted multivariate state estimation model, and (ii) the restricted multivariate state estimation model, and the instructions for the inversion of the state estimation calculation further cause the computer to cancel the steps performed by the state estimation calculation using an inverse state estimation calculation and to invert the preprocessing described by the preprocessing model for each of the state estimation values.
[0017] In one embodiment, a computing system for auditing the results of a machine learning model, wherein the non-transitory computer-readable medium further includes instructions that, when executed at least by the processor, cause the computing system to identify an inverse state estimation calculation that cancels the steps performed by the state estimation calculation to form the state estimation values from the original time series data values, generate a set of inverse state estimation values for the original time series data from the set of state estimation values, each of the inverse state estimation values being generated by performing an inverse state estimation for one of the set of state estimation values, and the reconstructed time series data values for each of the state estimation values being based on the inverse state estimation values.
[0018] In one embodiment, a computing system for auditing the results of a machine learning model, the non-transitory computer-readable medium further includes instructions that, when executed at least by the processor, cause the computing system to: (i) in response to a signal that the database under audit has not been modified, generate an electronic verification report message indicating that the database under audit is proven to be intact and not tampered with; (ii) in response to a signal that the database under audit has been modified, generate an electronic verification report message indicating that the database under audit is damaged or tampered with, and transmit the generated electronic verification report to a computing device to cause the verification report message to be stored or displayed by the computing device.
[0019] Brief Description of the Drawings The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate various systems, methods, and other embodiments of the disclosure. It will be recognized that the boundaries of the elements illustrated in the figures (e.g., boxes, groups of boxes, or other shapes) represent one embodiment of these boundaries. In some embodiments, one element may be implemented as multiple elements, or multiple elements may be implemented as one element. In some embodiments, an element shown as an internal component of another element may be implemented as an external component, and vice versa. Further, the elements may not be shown to scale.
Brief Description of the Drawings
[0020]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8A
Figure 8B
Figure 9
Figure 10
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0021] Detailed Description A system and method for enabling reliable auditing of the results of a machine learning model are described herein.
[0022] Sensors such as Internet-of-Things (IoT) sensors can be added to physical devices to monitor the operation of these devices. These sensors can be numerous, especially in high-density sensor industries such as utilities, oil and gas, and manufacturing. For example, an oil refinery can be equipped with over one million sensors. The power grid can be equipped with far more than one million sensors, especially when considering sensors for supervisory control and data acquisition (SCADA) for utility assets such as power plants and transformer substations, and sensors for advanced metering infrastructure (AMI). The data from these sensors can be stored as a time-series, a series of data points indexed in chronological order, or pairs of values and associated times. Time-series data may be stored in a time-series database, a database system optimized to store and provide time-series. Thus, a utility can generate a very large time-series database containing sensor readings on the order of petabytes or more.
[0023] Note that it should be noted that a system that captures a large amount of time-series signals generated from a system equipped with a physical transducer may collect incorrect time-series values from a failed physical transducer / sensor. Some percentage of sensors may be degraded and / or may malfunction due to a "stuck-at" fault that continuously signals one value despite receiving input. Also, intermittent problems may occur in the sensors or upstream data acquisition and / or data aggregation electronics, and individual time series may contain missing values. If these types of anomalies are not detected when ingesting time-series data into a time-series database, these anomalies will affect the subsequent use of the data stored in the time-series database. For example, if time-series data is used for product development or other scientific purposes, the accuracy of data analysis and, in some cases, the soundness of the conclusions drawn will be adversely affected by anomalies of the type described above. The systems and methods described herein also make it possible to solve this problem by providing signal verification and sensor operability verification for time-series databases generated from sensors that monitor critical assets.
[0024] Machine learning (ML) algorithms can be applied to sensor data stored in a time-series database to enable prediction, anomaly detection / discovery, and predictive maintenance for devices monitored by sensors. This can be referred to as automated predictive monitoring. One aspect that is becoming increasingly important, particularly in the high-density sensor industry for utilities, is the ability to explain to regulatory agencies, insurance investigators, and end customers the reasons why a predictive ML algorithm flagged one or more signals as anomalous and, equally importantly, the reasons why ML did not issue a warning. Accordingly, the systems and methods described herein provide "auditability assurance" for pattern recognition ML decisions.
[0025] Industries such as public utilities that utilize automated predictive monitoring systems must meet legal requirements and are subject to strict government regulations. For example, public utilities in the United States are subject to regulations at the regional, county, state, and federal government levels. These regulations include additional federal regulations from the United States Nuclear Regulatory Commission (USNRC) if the public utility operates a nuclear power plant, as well as additional regulations from the North America Electric Reliability Corporation (NERC) and the Federal Energy Regulatory Commission (FERC). In principle, such requirements and regulations may require the submission of data for legal or regulatory purposes. However, it is impossible to submit the entire mass of petabyte-scale data that forms the basis of the decisions made by these large-scale automation systems. This is not only because the size of the dataset is huge, but also because most of the data is not continuously stored in a manner that would enable such submission. States Nuclear Regulatory Commission:USNRC) and additional regulations from the North American Electric Reliability Corporation (NERC) and the Federal Energy Regulatory Commission (FERC). In principle, such requirements and regulations may require the submission of data for legal or regulatory purposes. However, it is impossible to submit the entire mass of petabyte-scale data that forms the basis of the decisions made by these large-scale automation systems. This is not only because the size of the dataset is huge, but also because most of the data is not continuously stored in a manner that would enable such submission.
[0026] Technologies for extracting data for legal and regulatory purposes often suffer from some or all of the following disadvantages. (i) The technology may select data on an ad hoc basis, as needed, or using arbitrary criteria, without a rigorous theoretical basis. (ii) The selected dataset for the technology is insufficient to reconstruct the original source data. (iii) The technology may be unable to reconstruct the original source data, potentially allowing violations of regulatory requirements to go undetected. In fact, it has been reported that this has actually occurred in practice with regard to the third disadvantage. Regarding this, it has been reported that this has actually occurred in practice.
[0027] In particular, the lack of the ability to restore the original source data creates an unnecessary adversarial relationship between the industry and the government. For example, electric utilities operating in California are required to maintain a minimum 10% generation overhead cushion each month. That is, even if there is demand during any given day and time, there is at least 10% overhead or "spare" capacity, so that even if a generation asset or transmission network asset suddenly fails, the likelihood of an overall power outage due to the overhead capacity is reduced. The utility is required to submit to the California Public Utilities Commission (PUC) time-series data collected monthly to demonstrate that the utility has maintained a minimum 10% generation overhead cushion each month. Each time the utility's generation capacity falls below the 10% generation overhead cushion, the PUC imposes a large fine (this can happen multiple times in a month ). The PUC scrutinizes the time-series signals very carefully. This is because the PUC habitually assumes that the utility has an incentive to adjust some of the time-series values in order to make it appear that the generation capacity does not fall below the 10% threshold. Similarly, the utility habitually assumes that the PUC has an incentive to adjust some of the time-series values in order to extract a large fine from the utility by making it appear that the generation capacity is just a few percent below the 10% threshold.
[0028] These disadvantages (including the mistrust on both sides) are eliminated by the systems and methods described herein that enable the results of machine learning models to be reliably audited. In one embodiment for ensuring auditability, the following information should be retained, for example, in a journal of changes to a database.
[0029] · Detected time-series data (including, for example, timestamped sensor readings and, if any, the user associated with the sensor input).
[0030] ·Description of the sensor (e.g., the detected asset associated with the sensor, the operating state of the sensor during measurement).
[0031] ·Description of the detected asset (e.g., the version of the asset at the time of measurement, a record of the relationship between the asset and other assets, e.g., an instance graph related to time or a parts list related to time).
[0032] ·"Knowledge" state at the time of measurement (e.g., one or more specific machine learning models, specific conditions, and specific data used during analysis).
[0033] ·Identification of the version of the machine learning model creation software and the dataset used to develop a specific machine learning model and analyze sensor readings.
[0034] Note that it should be noted that the detected time-series data may be extremely large. In one embodiment, instead of extracting a large amount of data, a relatively small set of data that is notable (or most prominent or important) can be saved, and from this set of data, data that is very close to the original data can be recovered. This is sufficient for legal requirements and regulatory requirements. A set of saved data with such a recoverability characteristic (i.e., the recoverability of approximate data about the original data from the saved data) is referred to herein as a "tamper-proof" dataset. In one embodiment, the creation of the tamper-proof dataset was by inventors D. Gawlick, K. C. Gross, Z. H. Liu, and A. Ghoneimy filed on March 19, 2018, titled "Intelligent Preprocessing of Multi-Dimensional Time-Series Data (Inte lligent Preprocessing of Multi-Dimensional Time-Series Data)", the entire It may include one or more of the technologies described in U.S. Patent Application Serial No. 15 / 925,427, which is incorporated herein by reference. Although it is not necessarily prevented that a malicious actor modifies the original data or even modifies the anti-tampering data set, the systems and methods disclosed herein positively guarantee that such tampering will be discovered in an audit conducted in accordance with the procedures described herein.
[0035] In one embodiment, for this reason, the systems and methods described herein enable the creation of data that is very close to the original source data with a relatively small amount of processed data (anti-tampering data set) that can be used to meet legal and regulatory requirements. Further, while the anti-tampering data set can be used to meet these requirements, it can still accommodate machine learning signal verification and sensor operability verification techniques.
[0036] In one embodiment, the processing techniques for extracting the anti-tampering data set exhibit three advantageous characteristics.
[0037] · Determinism: The processed data of the anti-tampering data set is uniquely determined by the input data without irregularity.
[0038] · Compression: The processed data of the anti-tampering data set is relatively small compared to the input data. In fact, the processed data of the anti-tampering data set may be several orders of magnitude smaller.
[0039] · Reversibility: It is possible to recover approximate data that is relatively close to the original source data from the processed data of the anti-tampering data set.
[0040] In fact, the degree of approximation is determined by legal and regulatory requirements and can be adjusted to meet these requirements.
[0041] Processing techniques such as Multivariate State Estimation Technique (MSET) exhibit the characteristics of determinism, compression, and reversibility, and are thus used in one embodiment of the present invention. However, any processing technique for extracting a dataset that satisfies the above three characteristics can be used to create an anti-tampering dataset according to the systems and methods described herein.
[0042] In one embodiment, in addition to the above three characteristics, there is an automation technique for determining whether legal requirements and regulatory requirements are met based on an anti-tampering dataset and history location information (which tracks changes to the database but does not guarantee data integrity). If these requirements are not met, the system may determine appropriate legal responses or government fines. By making the legal and regulatory processes completely clear to all parties, the motivation to attempt to conceal violations and adversarial relationships can be significantly reduced or even eliminated.
[0043] In one embodiment, the systems and methods described herein improve the existing ML monitoring systems to which they are applied, further enhancing the accuracy of warnings and reducing the false alarm rate for ML predictive anomaly detection. It should be noted that these improvements can be achieved by the implementation examples of the systems and methods described herein and do not require hardware upgrades anywhere within the systems in which they are implemented. Therefore, the systems and methods described herein have direct backward compatibility with any existing IoT system. This is particularly advantageous in the power utilities, oil and gas, manufacturing, and aviation industries, where conventional sensor data collection systems are already in place and would require significant effort to upgrade.
[0044] The systems and methods described herein are described in the context of the electric utility sector, but will clearly have any application where IoT sensor time series data is collected and used, for example, in the oil and gas, manufacturing, and aviation sectors. In one embodiment, the systems and methods described herein can be applied in a process for streaming digitized data about utility assets (such as, for example, coal power plants, oil power plants, nuclear power plants, wind turbines, geothermal generators, gas turbine power plants, etc., as well as critical assets within the power distribution network, such as transformers, substations, and SCADA systems) within a power generation facility.
[0045] -Auditability in Time Series Processes for Utility Predictions- In one embodiment, the methods and processes for ensuring the audibility of the results of a machine learning model include features such as (i) anti-tampering, (ii) snapshot isolation, (iii) journaling, (iv) activity journaling, and (v) data provenance recording.
[0046] In one embodiment, anti-tampering is a process or system configuration to ensure that malicious data modification (tampering) or accidental data modification (corruption) is detected. In one embodiment, an audit report can include a compact record that will indicate any changes to the original data.
[0047] Outliers in time series data can be identified by ML processing of the raw time series data, generating estimated values indicating what these values "should be" in the context of the surrounding data. A neural network (NN) and a support vector machine (SVM) can be employed for anomaly detection. MSET (and variants such as the Oracle (registered trademark)'s own advanced MSET pattern recognition "MSET2") can also be employed for anomaly detection. All three approaches (NN, SVM, and MSET) are, at the black box level, Nonlinear nonparametric (NLNP) regression algorithms. NLNP regression is mainly for time series data employed for prediction, anomaly discovery / detection, and predictive maintenance in the flow process. Because NLNP machine learning techniques do not make hypotheses about linear or non-linear relationships between time series "signals", but instead learn those relationships empirically.
[0048] Of these three NLNP machine learning approaches, both NN and SVM adopt a probabilistic (or clearly random) process for weight optimization. In the case of NN, probabilistic optimization of weights is performed between perceptron layers. In the case of SVM, probabilistic optimization is performed by convex quadratic programming optimization of regularization parameters to maintain the balance between bias and variance in SVM estimates. For example, in both cases (NN and SVM), when pattern recognition is trained with data from Monday and when it is trained with data from Tuesday, the relationship between the output estimate and the raw input signal will be very similar. However, when "looking into" the black box by comparing the intermediate weights for Monday's calculation with the intermediate weights for Tuesday's calculation, these intermediate weights will be significantly different. When applying pattern recognition only experimentally and practically, as long as the output of the black box is an accurate estimate of the underlying time series, it is not a problem that the weights within the black box can be substantially different each time the black box is executed. However, when applying auditability, the irregularities introduced by NN and SV The irregularities introduced by M make them unsuitable for extracting anti-tampering data sets. This is because they are not deterministic and thus not reversible.
[0049] When MSET is applied for anomaly detection (as well as related prediction and predictive maintenance), an estimated value can be obtained that is stored along with the original raw time - series telemetry values. In contrast to NN and SVM, MSET is a deterministic (albeit complex) mathematical algorithm, and the MSET estimated values are reversible as described above, which is the key to preventing forgery in auditability assurance. Thus, in one embodiment, MSET is applied for anomaly detection in a time - series data process to enable forgery prevention as part of auditability assurance. If any of the original raw data streams generated by this anomaly detection process are modified, changed, replaced, or transformed either by a malicious user or accidentally through any data corruption error within the storage medium, the changes to the original raw data values can be detected based on the accompanying MSET estimated values. Or, if the original raw data stream has not been forged or damaged, the original raw data values can be verified or confirmed as unchanged based on the accompanying estimated values. Proof of the absence of such forgery can be performed at any time after the creation of the MSET estimated values. This proof of the absence of such forgery is based on incorporating the deterministic and reversible MSET algorithm into the anomaly detection process described herein.
[0050] Thus, a compact data set for estimating the original values of time - series data is captured at any point in time, stored along with other information regarding changes made during the lifetime of the time - series database, and can be employed in an auditability process.
[0051] In one embodiment, snapshot isolation forms part of an auditability process. A database with respect to time provides the ability to store and retrieve records of any version. These versions are identified by the exact ascending order of transaction time (the time when the version was available or visible for query and subsequent processing). A particular version may sometimes be referred to as a snapshot. In one embodiment, an anti-tampering dataset created at a specific time may be used to verify the data values of the snapshots available at that creation time.
[0052] In one embodiment, a specific creation time for the anti-tampering dataset may be the creation of a journal entry. The journaling process is configured to track all changes to the database. Each change within the database results in the creation of a journal entry in the journal associated with the database. The journal typically has high resilience against data loss. For example, in a journaled database, there are no changes to the data on permanent media with respect to the database, and no external notification regarding the database is permitted until the journal data describing the changes is stored. This enables auditing of the database even if there are failures. The journal enables the system to reconstruct any snapshot of the database, sacrificing a large performance burden in snapshot queries even without time support. All databases have a copy of the latest snapshot, and the database with respect to time provides many snapshots going back in time. For immutable data such as sensor readings, the snapshot and the journal can be stored as a single copy.
[0053] However, in one embodiment, the audit includes additional knowledge about the user's activities (e.g., "who saw what information when (in response to a query)?", or "who inserted, updated, or deleted what information when?", and "the same transaction The answer to the question "What operations were performed in Yon (how are the operations related)?" may be used. Such information can be recorded in an activity journal entry. The activity journal entry is synchronized with the standard journal entry. In this way, the activity journal provides additional context information that describes the situation leading to the database change and the accompanying creation of the journal entry.
[0054] In one embodiment, information describing the data location may also be included in the standard journal entry or synchronized with the standard journal entry to provide yet another context information that describes the situation leading to the database change and the accompanying creation of the journal entry. The data location information associates the derived data with the corresponding input, processing step, and physical processing environment. For example, the location information identifies the data that forms the basis of the query result. In one embodiment, the location process may be configured to rewrite the query to determine these data. In an auditing situation, since the location metadata holds a record of all the data that formed the basis of the query result, it substantially reduces the data that needs to be considered. In one embodiment, the location information may include yet other information such as a description of the sensor that provides the data, the operational state of the sensor, the assets associated with the sensor, the version of the assets, the relationships between the assets, etc.
[0055] In one embodiment, the anti-tampering dataset is created in response to changes to the database simultaneously with the creation of journal entries. In one embodiment, the anti-tampering dataset is incorporated into the journal in the same way as other metadata associated with the changes (change metadata including activity journaling and data location metadata). In one embodiment, the anti-tampering dataset is incorporated for each change to the database, thereby ensuring that any subsequent data changes will be detected in the audit process conducted in accordance with the systems and methods described herein (and also ensuring that the machine learning software configuration in operation at the time of the change is incorporated with the remaining journal information and made audit-able). Thus, the systems and methods disclosed herein enable audits to be performed on the original data in any snapshot of the data, even though the data in the time-series database evolves over time later. For this reason, the auditability of the results of machine learning in a time-series database is ensured by enabling the incorporation of all the information necessary for each change to the database (i.e., journal entries, activity journal entries, data location, and anti-tampering machine learning records), thereby tracking the entire history location of the signal's database.
[0056] -Exemplary Environment- FIG. 1 shows one embodiment of a system 100 related to enabling reliable auditing of the results of a machine learning model.
[0057] In one embodiment, system 100 includes a time series data service 105 and a corporate network 110 connected by a network 115 such as the Internet. The time series data service 105 is directly connected to a sensor (such as sensor 120) or a remote terminal unit (RTU) via a network 125, or indirectly connected to a sensor (such as sensor 130) or an RTU via one or more upstream devices 135. In one embodiment, networks 115 and 125 are the same network, and in another embodiment, networks 115 and 125 are separate networks.
[0058] In one embodiment, the time series data service 105 may include a machine learning audit assurance system 140, a sensor interface server 145, a prediction, anomaly detection, and predictive maintenance system 150, a web interface server 155, and a data store 160. Each of these systems 140 - 160 is interconnected by a server - side network 165. Each of these systems 140 - 160 is composed of logic by various software modules for performing functions described as what they execute. In one embodiment, systems 140 - 160 are realized by dedicated computing devices. In one embodiment, one or more of systems 140 - 160, although represented as separate units in FIG. 1, may be realized by a common (or shared) computing device. Among other systems. Each of these systems 140 - 160 is interconnected by a server - side network 165. Each of these systems 140 - 160 is composed of logic by various software modules for performing functions described as what they execute. In one embodiment, systems 140 - 160 are realized by dedicated computing devices. In one embodiment, one or more of systems 140 - 160, although represented as separate units in FIG. 1, may be realized by a common (or shared) computing device.
[0059] In one embodiment, the time series data service 105 may be hosted by a third party and / or may be operated by a third party for the benefit of multiple account owners / tenants, each of whom operates a business and each of whom has an associated enterprise network 110. In one embodiment, the time series data service 105 is associated with a utility entity such as an electric utility or is associated with major utility assets such as power generation facilities, substations, or other major power grid components. In one embodiment, the time series data service 105 is composed of logic such as software modules for operating the time series data service 105 in order to (i) create and export a time series database and / or (ii) audit a time series database according to the systems and methods described herein.
[0060] In one embodiment, sensors 120, 130 can be attached or configured to detect the performance of one or more components of a device or system. The device or system generally includes any type of machinery or equipment having components that perform measurable activities. Sensors 120, 130 can include, but are not limited to, voltage sensors, current sensors, temperature sensors, pressure sensors, rotational speed sensors, flow meter sensors, vibration sensors, microphones, electromagnetic radiation sensors, proximity sensors, gyroscopes, inclinometers, accelerometers, global positioning system (GPS) sensors, torque sensors, strain sensors, nuclear radiation detectors, or any of a variety of other sensors or transducers for generating electrical signals that describe a detected or sensed physical behavior.
[0061] In one embodiment, sensors 120 and 130 are connected to sensor interface server 145 via network 125. In one embodiment, sensor interface server 145 is composed of logic such as a software module for collecting readings from sensors 120 and 130 and storing these readings as time-series observations in, for example, data store 160. Sensor interface server 145 exposes one or more application programming interfaces (APIs) configured to receive readings from sensors using, for example, sensor data formats and communication protocols applicable to various sensors 120 and 130 Thereby, it is configured to interact with sensors. The sensor data format is generally defined by the sensor device. The communication protocol can be a custom protocol (such as a legacy protocol prior to IoT implementation examples), or any of various IoT or machine-to-machine (M2M) protocols, for example, Constrained Application Protocol (CoAP), Data Distribution Service (DDS), Devices Profile for Web Services (DPWS), Hypertext Transport Protocol / Representational State Transfer (HTTP / REST), MQ Telemetry Transport (MQTT), Universal Plug and Play (UPnP) , Extensible Messaging and Presence Protocol (XMPP), ZeroMQ, and transmitted by the Transmission Control Protocol Other communication protocols that can be used, i.e., it can be an Internet protocol or a User Datagram Protocol (TCP / IP or UDP) transfer protocol. SCADA protocols, such as OLE for Process Control Unified Architecture (OPC UA), Modbus RTU, RP-570, Profibus, Conitel, IEC60870-5-101 or 104, IEC61850, and DNP3, etc., can also be adopted when extended to operate via TCP / IP or UDP. In one embodiment, the sensor interface server 145 polls the sensors 120, 130 to retrieve sensor readings. In one embodiment, the sensor interface server passively receives sensor readings actively transmitted by the sensors 120, 130.
[0062] In one embodiment, the enterprise network 110 may be associated with a utility entity such as an electric utility. In one embodiment, the enterprise network 110 may be associated with a regulatory entity such as a government. For the sake of brevity and clarity of explanation, the enterprise network 110 is represented by an on-site local area network 170. One or more personal computers 175 or servers 180 are operably connected to the local area network 170, along with one or more remote user computers 185, and the one or more remote user computers 185 are connected to the enterprise network 110 through a network 115 or other suitable communication network or combination of networks. The personal computers 175 and the remote user computers 185 can be, for example, desktop computers, laptop computers, tablet computers, smartphones, or other devices that have the ability to connect to the local area network 170 or the network 115 or have other synchronization capabilities. The computers of the enterprise network 110 interface with the time series data service 105 via a network 115 or another suitable communication network or combination of networks.
[0063] In one embodiment, a remote computing system (such as the remote computing system of enterprise network 110) can access information or applications provided by time-series data service 105 through web interface server 155. For example, computers 175, 180, 185 of enterprise network 110 may request a time-series database from time-series data series data service 105. Or, for example, computers 175, 180, 185 of enterprise network 110 may perform an audit of the time-series database according to the systems and methods described herein. In one embodiment, the remote computing system can send a request to web interface server 155 and receive a response from web interface server 155. In one example, access to the information or application may be achieved by using a web browser on personal computer 175 or remote user computer 185. In one example, these communications may be exchanged between web interface server 155 and server 180, for example, a remote representational state transfer (REST) request using JavaScript (registered trademark) object notation (JSON) as a data exchange format, or in the form of a simple object access protocol (SOAP) request to and from an XML server.
[0064] In one embodiment, data store 160 includes one or more time-series databases configured to store and provide time-series data from sensors 120, 130 received by sensor interface server 145. In one embodiment, the time-series database is an Oracle (registered trademark) database configured to store and provide time-series data. In some exemplary configurations, data store 160 is a network connection strain It can be implemented using a dual (network-attached storage: NAS) device and / or other dedicated server devices.
[0065] In one embodiment, the upstream device 135 can be a third-party service for managing IoT-connected devices. Or, in one embodiment, the upstream device 135 may be a gateway device configured to enable the sensor 130 to communicate with the sensor interface server 145 (thus, for example, if the sensor 130 is not IoT-compatible, it cannot communicate directly with the sensor interface server 145).
[0066] -Exemplary method for ML model audit assurance- In one embodiment, each step of the computer-implemented method described herein can be performed by a processor of one or more computing devices (such as processor 1110 as illustrated and described with reference to FIG. 11). The processor of the one or more computing devices is configured with logic to (i) access a memory (such as memory 1115 and / or other computing device components as illustrated and described with reference to FIG. 11), and (ii) cause the system to perform steps of the method (such as machine learning audit assurance logic 1130 as illustrated and described with reference to FIG. 11). For example, the processor accesses, reads from, or writes to the memory to perform steps of the computer-implemented method described herein. These steps can include (i) steps of retrieving any necessary information, (ii) steps of calculating, determining, generating, classifying, or creating any data, and (iii) steps of storing any data that has been calculated, determined, generated, classified, or created. When referring to storage or memory, this refers to storage as a data structure in the memory or storage / disk of the computing device (such as memory 1115 as illustrated and described with reference to FIG. 11, or storage / disk 1135 of computing device 1105, or remote computer 1165, etc.).
[0067] In one embodiment, each subsequent step of the method is initiated in response to analyzing a received signal or retrieved stored data indicating that the previous step has been performed to at least the extent necessary to initiate the subsequent step. Generally, the received signal or retrieved stored data indicates completion of the previous step.
[0068] FIG. 2 shows one embodiment of a method 200 related to enabling reliable auditing of the results of a machine learning model. In one embodiment, the steps of method 200 are performed by either the machine learning auditing assurance system 140 or one of the computers 175, 180, 185 within the enterprise network 110 (as illustrated and described with reference to FIG. 1). In one embodiment, either the machine learning auditing assurance system 140 or one of the computers 175, 180, 185 is a dedicated computing device (such as computing device 1105) configured with machine learning auditing assurance logic 1130.
[0069] Method 200 may be initiated based on various triggers, which may include, for example, receipt of a signal on the network, or analysis of stored data, such as (i) the time series data service 105 or a user (or administrator) of the computers 175, 180, 185 initiating method 200, (ii) being scheduled to be initiated at a defined time or time interval by method 200, (iii) a user (associated with a utility or regulatory entity) of the time series data service 105 or the computers 175, 180, 185 requesting an audit of the time series database, or (iv) any other trigger indicating that method 200 should be initiated. Method 200 is initiated at start block 205 in response to analyzing a received signal or retrieved stored data and determining that the signal or stored data indicates that method 200 should be initiated. Processing proceeds to process block 210.
[0070] At process block 210, the processor retrieves a set of state estimate values for the original time series data values from the database under audit. The state estimate values are generated by state estimation calculations for each of the time series data values. When the processing at process block 210 is complete, the processing proceeds to process block 215.
[0071] In process block 215, the processor reverses the state estimation calculation for each of the state estimation values to generate a reconstructed time series data value for each of the state estimation values. When the processing in process block 215 is completed, the processing proceeds to process block 220.
[0072] In process block 220, the processor retrieves the original time series data values from the database under audit. When the processing in process block 220 is completed, the processing proceeds to process block 225.
[0073] In process block 225, the processor compares the original time series data values with the reconstructed time series data values pair by pair to determine whether the original time series and the reconstructed time series match. When the processing in process block 225 is completed, the processing proceeds to decision block 230.
[0074] In decision block 230, the processor evaluates whether the original time series matches the reconstructed time series. If the original time series matches the reconstructed time series (YES), the processing in decision block 230 is completed and the processing proceeds to process block 235. If the original time series does not match the reconstructed time series (NO), the processing in decision block 230 is completed and the processing proceeds to process block 240.
[0075] In process block 235, since the original time series and the reconstructed time series match, the processor generates a signal indicating that the database under audit has not been modified. When the processing in process block 235 is completed, the processing proceeds to end block 245 and process 200 ends.
[0076] In process block 240, since the original time series and the reconstructed time series do not match, the processor generates a signal indicating that the database under audit has been modified. When the processing in process block 240 is completed, the processing proceeds to end block 245 and process 200 ends.
[0077] For each of the foregoing process blocks of method 200, it will be described in more detail elsewhere in this specification.
[0078] - Preparation for Reconstructing the Original Time-Series Data from the MSET Estimate - FIG. 3 shows a flowchart of an embodiment of an auditability process 300 related to enabling reliable auditing of the results of a machine learning model. MSET is incorporated into the data flow process, and the MSET operation is inverted to reconstruct the original time-series data for the purpose of auditability. This figure shows the auditability for this application example of MSET.
[0079] In one embodiment, at a high level, the auditability process 300 is initiated when the system retrieves the MSET estimate from the time-series database. Next, the system reverses the MSET calculation to generate the reconstructed time-series data from the MSET estimate. As described above, MSET is a reversible process, which means that the original time-series data will be reconstructed by performing the inverse MSET calculation on the MSET estimate. Then, the system is provided with the original time-series data from the database being audited. Finally, the system compares the reconstructed time-series data with the original time-series data, and if the comparison result shows a match, it proves that there was no intentional tampering (or accidental data corruption) in the original time-series data. As described in more detail above, the operations described with reference to the auditability process 300 can be performed by a processor of one or more computing devices that access the memories, storage, and / or other computing device components illustrated and described with reference to FIGS. 1 and 11.
[0080]
[0081] In one embodiment, the archived time series database 305 is presented for auditing to determine whether the original time series data in the database has been corrupted or whether the original time series data in the database has been damaged or tampered with. The archived time series database 305 may include one or more time series signals that are sequences of time series values for all observations of the time series. The archived time series database 305 is presented for the paired difference analysis 310 between the original time series data values from the archived time series database 305 and the reconstructed time series values derived from the MSET modeling of the archived database, and for the MSET estimation of these values.
[0082] In one embodiment, the archived time series database 305 is a snapshot of the database when a change is made to the database. Additionally, at the time a change is made to the database, a journal entry describing the change is created, and a set of state estimate values is created for later paired difference analysis 310 under the auditing of the archived time series database 305. In one embodiment, the set of state estimate values is stored in association with the journal entry for the database under audit, i.e., the archived time series database 305. For example, the set of state estimate values may be generated in response to one or more commands indicating changes to the database under audit and stored in a data structure associated with the journal entry describing the change. Other metadata, including location metadata identifying the data on which the query results (resulting from the journaled changes) are based, may also be stored in association with the journal entry.
[0083] Sensor impairments, including non-calibrated bias, intermittent degeneracy faults, gain change drift, and transient spikiness, degrade signal quality and are a major cause of false alarms (type I errors) and missed alarms (type II errors) in ML prediction for IoT applications. The original time-series data values in the archived time-series database 305 may include some of these sensor impairments. In one embodiment, the archived time-series database 305 undergoes a series of intelligent data preprocessing 315 steps to cleanse the data and prepare the cleansed data for MSET modeling and estimation. The preprocessing serves to mitigate the undesired effects of sensor impairments. In one embodiment, the intelligent data preprocessing 315 is repeated for one or more time series included in the archived time-series database 305.
[0084] In one embodiment of the intelligent data preprocessing 315, all observations from all signals within the archived time-series database 305 are first preprocessed and then optimally resampled and “harmonized,” for example, by using an Oracle® analytical resampling process (ARP). Thereby, an updated database for the signals that have been cleansed and optimally resampled / synchronized is generated. In one embodiment, the ARP may include one or more of the techniques described in U.S. Patent Application Serial No. 16 / 168,193, filed October 23, 2018, by inventors K.C. Gross and G.C. Wang, entitled “Automated Analytic Resampling Process for Optimally Synchronizing Time-Series Signals,” which is hereby incorporated by reference in its entirety. into this specification.
[0085] In one embodiment, the archived time series database 305 goes through a missing value imputation process 320 during intelligent data preprocessing 315. The missing value imputation process 320 analyzes the original time series signal using a missing value check for each time series signal within the archived time series database 305. In response to analyzing the missing values in the time series, the missing value imputation process 320 fills the missing values with estimated values. In one embodiment, the estimated value is simple interpolation. In one embodiment, the estimated value is a very accurate estimated value based on the serial correlation relationship derived by MSET and the cross-correlation relationship with the existing values, rather than simple interpolation. The updated and filled time series is stored for future processing. The missing value imputation process 320 may also store a record of the positions of the missing values in the time series filled with estimated values for later reversal. In one embodiment, the ARP was filed on June 11, 2018 by inventors G.C. Wang, K.C. Gross, and D.G. Gawlick and titled "Missing Value Imputation to Facilitate Prognostic Analysis of Time-Series Sensor Data", which is hereby incorporated by reference in its entirety. It may include one or more of the technologies described in U.S. Patent Application Serial No. 16 / 005,495.
[0086] In one embodiment, the intelligent data preprocessing 315 includes a despiking process 325. In one example, the despiking process 325 may be performed after missing values in the original time series are replaced by a missing value imputation process 320. In the despiking process 325, the updated time series signal is analyzed by an outlier check to detect and remove data "spikes", sudden short-lived variations that do not represent accurate sensor readings. The outlier check detects spikes in the signal by repeatedly characterizing various statistical distributions for the signal (generating descriptive parameters for describing the characteristics and behavior of the statistical distributions). Time series data values that are outliers based on these characteristics are flagged as data value spikes. The captured spikes are temporarily replaced with the signal average. The updated and despiked time series is stored for later processing. The despiking process 325 may also store a record of the values and positions within the time series of the detected spikes. In one embodiment, the despiking process 325 is the "Synthesizing High-Fidelity Signals with Spikes for Prognostic Surveillance Applications" filed on December 10, 2018 by inventors G.C. Wang and K.C. Gross, and the entire thing is incorporated herein by reference and may include one or more of the techniques described in U.S. Patent Application Serial No. 16 / 215,345. Applications" and is incorporated herein by reference in its entirety and may include one or more of the techniques described in U.S. Patent Application Serial No. 16 / 215,345, which is incorporated herein by reference in its entirety. It may include one or more of the techniques described in U.S. Patent Application Serial No. 16 / 215,345, which is incorporated herein by reference in its entirety.
[0087] In one embodiment, the intelligent data preprocessing 315 includes a dequantization process 330. For example, the dequantization process 330 may be performed after data spikes are detected and removed from the time series by the despiking process 325. In the dequantization process 330, the updated signal is analyzed through a "quantization" check to determine whether the signal values are rapidly and repeatedly switching between adjacent representative values by data quantization (a non-reversible data compression technique in which data intervals are grouped or summarized into a single representative value). In other words, the quantization check identifies time series intervals where the observation points are repeatedly going back and forth between a certain number of observation upper limits. Data quantization can be caused by low-resolution "quantized" transducers or sensors. Further, the dequantization process 330 converts the quantized values detected in the time series into high-precision continuous signals. These continuous signals are very similar to what the signal values would have been if detected using a higher-resolution transducer. The updated non-quantized time series is saved for later processing. The dequantization process 330 may also store a record of the values and their positions within the time series of the detected quantized values. In one embodiment, the dequantization process 330 is the one described in U.S. Patent No. 10,496,084, titled "Dequantizing Low-Resolution IOT Signals to Produce High-Accuracy Prognostic Indicators," issued to inventors M. Li and K. C. Gross on December 3, 2019, which is hereby incorporated by reference in its entirety and may include one or more of the techniques described therein.
[0088] In one embodiment, the intelligent data preprocessing 315 includes a non-staircasing process 335. For example, the non-staircasing process 335 may be performed after a quantized value is detected and replaced in time series by the non-quantization process 330. The staircasing is due to a sampling rate mismatch between the recording system and the detection system. In this case, the slower sampling rate signals simply repeat their last measured values at the higher sampling rate, so that all measured signals will result in a uniform sampling rate. The time series values in the case of a slower sampling rate sensor have a sequence of flat segments similar to step steps. Staircasing is a common problem for commercial data history archives that use a simple algorithm to collect low sampling rate data into a higher sampling rate time series. In one embodiment, the non-staircasing process 335 analyzes the time series to identify any that exist with staircased values. The non-staircasing process 335 "fills in" the staircased portions of the signal with higher sampling rate signals. The higher sampling rate signal values may be derived using MSET estimates in one embodiment. The updated and filled time series is stored for future processing. The non-staircasing process 335 may also store a record of the values and positions within the time series for the detected staircased values for later inversion. In one embodiment, the non-staircasing process 335 was filed on September 11, 2018 by inventors K.C. Gross and G.C. Wang and titled "Prognostic Replacing Stair-Stepped Values in Time-Series Sensor Signals With Inferential Values to Facilitate Prognostic Surveillance Operations", the entirety of which may include one or more of the techniques described in U.S. Patent Application Serial No. 16 / 128,071, which is incorporated herein by reference in its entirety.
[0089] In one embodiment, the intelligent data preprocessing 315 includes a uniform sampling process 340. In one embodiment, the uniform sampling process 340 may be performed after the stepped values are detected and replaced in the time series by the non-stepped process 335. In the uniform sampling process 340, the signals are analyzed by a sampling rate check to identify whether the sampling rates of the signals are different. If the signals indicate different sampling rates (e.g., having different numbers of observed values over the same period), the observed values of the slower signal will be resampled to match the highest sampling rate of the signals. The updated resampled time series is stored for future processing. The uniform sampling process 340 may also store a record of the original non-resampled values and their placement in the time series for later inversion.
[0090] In one embodiment, the intelligent data preprocessing 315 includes a phase synchronization process 345. For example, the phase synchronization process 345 may be performed after the uniform sampling process 340. Or, in one embodiment, the phase synchronization process 345 may be performed in parallel with the uniform sampling process 340. In the phase synchronization process 345, the updated signals are analyzed by a correlation check to detect phase-shifted observed values. The phase-shifted observed values (or the time signal values associated with incorrect time indices in the time series) may be due to, for example, clock synchronization mismatches in measurement devices such as sensors. The phase synchronization process 345 shifts the phase-shifted values in the time domain to align them with the correct time indices. The updated phase-synchronized time series is stored for future processing. The phase synchronization process 345 may also store a record of the original placement of the phase-shifted values in the time series for later inversion.
[0091] In one embodiment, one or more of the uniform sampling process 340 and the phase synchronization process are performed using an Oracle (registered trademark) Analytic Resampling Process (ARP), and include one or more of the techniques described in "Automated Analytic Resampling process for Optimally Synchronizing Time-Series Signals", which is incorporated herein by reference. Note that signal sampling can vary among the time series signals in the archived time series database 305. For example, one time series may have a high but regular sampling rate or interval between observations. Another time series may have a low but regular sampling rate or interval between observations. Another time series may have an irregular or unequal interval between observations. In one embodiment, the phase of the time series signal can be adjusted so that the set of time series signals in the archived time series database 305 are aligned with respect to the observation time. In one embodiment, the data values of the time series signal can be resampled at a new sampling interval by interpolating the data values estimated for the observations at the new sampling interval within the time series signal. This may be performed on multiple synchronized time series signals to provide a common sampling interval for the time series signals.
[0092] The six intelligent data preprocessing 315 procedures described above result in a high-quality "cleaned" or "extended" version of the archived time series database 305. This time series database 305 can be stored and retrieved, and can further be used for subsequent machine learning processes. Note that other data preprocessing techniques may be applied in the intelligent data preprocessing 315, or a smaller number of these six techniques may be performed. All observations from all signals are preprocessed here, optimally resampled, and then "harmonized" using ARP.
[0093] The procedures of the intelligent data preprocessing 315 correct common problems in a typical machine learning dataset. However, each of the steps modifies some of the original data. Records must be maintained for each modification so that the original data can be reconstructed. As described in the explanation of each of the six procedures of the intelligent data preprocessing 315 above, each procedure will remember a record indicating the changes made to the dataset. In one embodiment, the change records for all procedures of the intelligent data preprocessing 315 applied to the archived time series database 305 will be stored in a single electronic data structure called an intelligent data preprocessing (IDP) model.
[0094] Referring now to FIG. 4, FIG. 4 shows a schematic diagram 400 of one embodiment that stores the change records from the intelligent data preprocessing 315 in an exemplary IDP model 405. The IDP model 405 is a complete ordered record of all operations to be performed on the archived time series database in preparation for machine learning operations.
[0095] The missing value imputation process 320 stores a record (also referred to as a mark or marker) of the location of missing values in the time series filled with estimated values within the missing value mark data structure 410 in the IDP model 405. In one embodiment, the missing value mark data structure includes a set of arrays associated with the time series within the archived time series database 305. In the archived time series database 305, there is a missing value mark array associated with each time series that contains missing values. The missing value mark array may be an array of observation index (e.g., time) values, and each index value or time value indicates the observation value in the associated time series where the missing value was filled with an estimated value.
[0096] The non-spiking process 325 stores in the spike data structure 415 within the IDP model 405 a record of the values and positions within the time series of the detected spikes. In one embodiment, the spike data structure includes a set of arrays associated with the time series within the archived time series database 305. There is a spike array associated with each time series in which a spike is detected. These arrays can be an array of tuples that include an observation index value (e.g., time) and the associated amplitude value of the spike at that observation value (i.e., the incorrect sensor reading value).
[0097] The non-quantization process 330 stores in the smoothing model data structure 420 within the IDP model 405 a record of the values and positions within the time series of the detected quantization values. In one embodiment, the smoothing model 420 includes a set of quantized observed value arrays associated with the time series within the archived time series database 305. There is a quantized observed value array associated with each time series in which a quantized value is detected. These arrays can be an array of tuples that include an observation index value (e.g., time) and the associated quantization value at that observation value.
[0098] The non-staging process 335 stores in the smoothing model data structure 420 a record of the values and positions within the time series of the detected staging values. In one embodiment, the smoothing model 420 includes a set of staged observed value arrays associated with the time series within the archived time series database 305. There is a staged observed value array associated with each time series in which a staged value is detected. These arrays can be an array of tuples that include an observation index value (e.g., time) and the associated staged (original or not yet staged) value at that observation value. In one embodiment, the quantized observed value array and the staged observed value array are a single array that includes all of the original quantized and staged values indexed by their positions within the time series. In this situation, the non-staging process may make the staged value tuples into the quantized observed value array created by the non-quantization process 330.
[0099] The uniform sampling process 340 stores in the timestamp sequence data structure 425 in the IDP model 405 the records of the original non-resampled values and their arrangements within the time series. Also, the phase synchronization process 345 stores in the timestamp sequence data structure 425 the records of the original arrangements of the out-of-phase values within the time series. In one embodiment, the timestamp sequence data structure 420 includes a set of timestamp arrays associated with the time series in the archived time series database 305. There is a timestamp array observation value array associated with each time series in which out-of-phase observed values are detected. These arrays include observation index values (e.g., time) for each observed value of the time series. In one embodiment, there are timestamp arrays created for both the uniform sampling process 340 and the phase synchronization process 345. This can be applicable when the processes are executed continuously. When the uniform sampling process 340 and the phase synchronization process 345 are executed together, the results are stored together in a single timestamp array obtained.
[0100] Accordingly, in one embodiment, the processor preprocesses the original time series values, as shown in the intelligent data preprocessing 315, to reduce the impact of sensor failures on the quality of the original time series values. Over the course of that intelligent data preprocessing, the processor generates a data preprocessing model that records, for example, as shown by the IDP model 405, one or more changes to the original time series values during preprocessing. In one embodiment, the inversion of the state estimation calculation described with reference to process block 215 of FIG. 2 also includes the step of retrieving a data preprocessing model, such as the IDP model 405, from storage or memory, and for each of the state estimation values, the step of inverting the preprocessing described by the data preprocessing model.
[0101] Each of the six intelligent data preprocessing 315 procedures described above may be reversed or canceled by retrieving the original time series values held in the IDP model 405 and replacing the corresponding extended values in the extended time series database with the original values. In one embodiment, the reversal process for the various processes of the intelligent data preprocessing 135 is performed in an order reverse to the order in which the intelligent data preprocessing 135 processes were executed.
[0102] -Model Training- Referring again to FIG. 3, after the intelligent data preprocessing 315, an extended time series signal is used to train the MSET model 350. FIG. 5 shows a schematic diagram of one embodiment of an MSET model training process 500 related to enabling reliable auditing of the results of a machine learning model. The archived time series database 305 is provided to a selective training data process 505. In the selective training data process 505, data is selected from the time series data to form a training data set. In one embodiment, a subset of the time series observation vectors is selected to form the training set. The selected training data set is then stored in a training data structure 510.
[0103] Once the selection of the training data 505 is complete and stored in the training data structure 510, as described above with reference to FIGS. 3 and 4, the intelligent data preprocessing 315 is initiated and the IDP model 405 is created and stored. In one embodiment, the training data 505 is selected before the intelligent data preprocessing 315, and the training data is selected from the unextended data. In another embodiment, the training data 505 is selected after the intelligent data preprocessing 315, and the training data is selected from the extended data.
[0104] When the selection of the training data 505 and the intelligent data preprocessing 315 are completed, the loop of the training vector selection process 515 and the MSET model training process 520 is started. During vector selection, a set of training vectors is selected from the training data and provided to the MSET model for training. For example, 50 or 100 vectors may be selected. In one embodiment, the preprocessed training data is split into two parts to improve the model. For example, even observations may form the first part of the training set and odd observations may form the second part of the training set. During the iteration of the selection process, the set of training vectors is selected from the odd observations during the even / odd "hopscotch" vector selection and then from the even observations. and then selected from the even observations.
[0105] As described above, training an MSET model such as the MSET model 350 is , is a deterministic mathematical procedure. MSET training 520 learns the correlation relationships between time series signals using time series signals representing the time series recorded in the archived time series database. The MSET model training process 520 uses the extended signals to identify within the archived time series database 305 all signals having any degree of association with any other signal in the archived time series database 305. In one embodiment, the identification of the association between signals is performed both over the entire universe of signals within the archived time series database 305 as a whole and separately for clusters of signals. This empirical clustering approach recognizes that the archived time series database of signals can be derived from separate systems within a utility entity's facility or individual assets within a fleet of utility devices and systems of a utility entity. What the MSET training process 520 ultimately outputs is the trained MSET model 350. In one embodiment, the training process is the MSET2 training process and the MSET model is the MSET2 model.
[0106] In one embodiment, vector selection 515 and the MSET training process 520 can be repeated within a loop (of one or more iterations) until the MSET model 350 stabilizes, as shown in decision block 525. If the MSET model 350 does not change very much as much as the training dataset is changed, the model 350 is stable (YES), training is complete, and the trained MSET model 350 is stored as a data structure for further use. If the MSET model 350 still changes significantly when the training dataset is changed, the model 350 is not yet stable (NO), at 515 an additional set of training vectors is selected, and at 520 the MSET model 350 is further trained with the additional set of training vectors.
[0107] Referring back to FIG. 3, the trained MSET model 350 is used to calculate the MSET state estimate 355 of each signal in the archived time series database 305 based on the empirical correlation pattern learned during model training 520 between each signal in the archived time series database 305 and other signals. Although the MSET estimates are very accurate, their accuracy may vary based on the degree of training of the MSET model.
[0108] In process block 360, the MSET estimates, MSET parameters, and training vectors are stored in a reporting data structure. This reporting data structure may be associated with a journal entry that describes a change to the archived time series database 305 in a database or data structure (e.g., a change that resulted in the archived state of the archived time series database). In one embodiment, processes blocks 320 through 355 are initiated and executed in response to a change to the archived time series database so as to be included with a journal entry that describes the change in order to reliably detect data tampering in the archived time series database 305. In one embodiment, activity journal metadata that describes one or more queries preceding the change and location metadata that describes the data underlying the query results are stored along with the journal entry and the reporting data structure.
[0109] Note that since a report can be generated in response to any change to the database under audit, a series of anti-tampering data sets related to each change to the database can be created. For this reason, historical location information (including the change, the activity journal that describes the activity preceding the change, and the location information that describes the basis of the information presented in response to the query associated with the change) and a record for reconstructing the original data are both recorded in each journal entry. Thus, for each database under audit The complete history of the database is recorded and can be reviewed for audits, and the record is much more compact than maintaining additional snapshots of the database, resulting in significantly improved performance and improved portability of the data used for audits.
[0110] In one embodiment, the MSET estimates 355 are stored in an archived time series database 305 along with the original raw signals (MSET parameters). Also stored in the archived time series database 305 are "sensor operability flags" with a "1" for fully verified or a "0" for signals where an anomaly was found in the sensor that measured the original raw signal. In one embodiment, these sensor operability flags are evaluated using sequential probability ratio tests to analyze the residual between the MSET estimates and the original raw signal values and label the values as anomalous or not anomalous. The fault detection estimation is based on the results of the fault detection ratio test (SPRT). If a threshold number of outliers occur in the time series data for a sensor, the sensor may be flagged as partially or completely inoperable ("0") in the time series. An inoperable flag may indicate a sensor / transducer failure or degradation, a stuck-at fault, or a signal from an intermittent problem with the sensor / transducer or the upstream data collection electronics or network. If the number of outliers occurring in the time series data for a particular sensor does not exceed the threshold number, the sensor may be flagged as verified as operational ("1") in the time series. In one embodiment, the threshold may be as low as meeting or exceeding one anomalous reading in the time series, or as high as desired. In one embodiment, the threshold may be set by a machine learning analysis of the time series, such as an MSET analysis. In one embodiment, these flags may be set in a data structure that includes the time series.
[0111] -Model Restrictions- In some examples, the trained MSET model 350 may not be very compact. Therefore, for portability, it may be desirable to limit the size of the MSET model 350. Referring now to FIG. 6, FIG. 6 shows a schematic diagram of one embodiment of an MSET model limiting process 600 related to enabling reliable auditing of the results of a machine learning model.
[0112] In one embodiment, the system first trains a multivariate state estimation model, such as the MSET model 350, with a set of training values selected from the original time series values, as illustrated and described with reference to FIG. 5 for example. The complexity of the trained multivariate state estimation model can be reduced by performing a principal component analysis of the matrix for the trained multivariate state estimation model and limiting the trained multivariate state estimation model to the principal components of the matrix. For example, the MSET model 350 can be reduced by performing a principal component analysis of the MSET matrix of the model 350 and limiting the model 350 to the principal components of the matrix. Note that the MSET model does not consist of just the MSET matrix alone, and it should be noted that the MSET matrix forms the core of the MSET algorithm. As described above, the MSET model 350 can be an MSET2 model.
[0113] In one embodiment, the system decomposes the matrix associated with the multivariate state estimation module into a set of eigenvectors. For example, the complexity of the MSET model 350 is reduced by performing a singular value decomposition (SVD) 605 of the MSET matrix included in the MSET model 350. In one embodiment, the SVD 605 consists of a sequence of eigenvectors and their associated eigenvalues. In one embodiment, the eigenvectors resulting from the SVD 605 are sorted in descending order by their respective associated eigenvalues. The system then stores the sorted sequence of eigenvectors and their associated eigenvalues 610 as an SVD data structure, for further processing for example.
[0114] In one embodiment, the system selects a subset of the principal eigenvectors from the set of eigenvectors stored in the SVD data structure. For example, the system selects the principal eigenvectors 615 from the SVD 605 based on the sorted eigenvectors and eigenvalues 610. The eigenvector with the largest eigenvalue is the principal eigenvector. These principal eigenvectors are the main factors of the largest variation in the sensor data over time and are thus most beneficial for state estimation modeling. In one embodiment, the processor selects the "top" N eigenvectors having the largest associated eigenvalues as the principal eigenvectors 620 having the associated principal eigenvalues. The processor then removes all eigenvectors and eigenvalues within the SVD data structure except for the top N eigenvectors, thereby removing all eigenvectors that are not principal eigenvectors (along with their associated eigenvalues).
[0115] In one embodiment, the system creates a "restricted" multivariate state estimation model from a subset of the principal eigenvectors. For example, the retained top N principal eigenvectors and their associated principal eigenvalues 620 are provided to the MSET model limiter 625. The MSET model 350 is also provided to the MSET model limiter 625. The MSET model limiter 625 operates to limit the MSET model 350 to the top N principal eigenvectors and their eigenvalues 620. This substantially reduces the amount of data required to encode the results of the MSET algorithm. In one embodiment, the MSET model limiter 625 constructs a restricted MSET matrix from the principal eigenvectors and associated eigenvalues 620. The model limiter 625 replaces the original MSET matrix in the MSET model 350 with the restricted MSET matrix to create a restricted MSET model 630. The restricted MSET model 630 is stored as a data structure in memory or storage.
[0116] Referring again to FIG. 2, in one embodiment, the inversion of the state estimation calculation also includes generating an inversion of the calculations by which a restricted multivariate state estimation model (such as the restricted MSET model 630) forms the state estimate.
[0117] -Data Compression- In some examples, the entire archived time series database 305 can be very large. Thus, in some cases it may be desirable to limit the size of the database to the parameters of the restricted MSET model 630 for portability. FIG. 7 shows a schematic diagram of one embodiment of a data compression process 700 related to enabling reliable auditing of the results of a machine learning model. In the data compression process 700, the archived time series database 305 is reduced to a minimum size suitable for auditing the uncompressed or original archived time series database 305.
[0118] In one embodiment, in process block 705, the system excludes training data values from the original time series values within the database under audit, i.e., the archived time series database 305. For example, in one embodiment, the system creates a copy of the archived time series database 305 that does not include the training data used to train the MSET model. In one embodiment, the training data for the MSET model, such as the training data 510 for the MSET model 350, is removed from a copy of the archived time series database 305. For example, all the observations used to train the MSET model 350 can be deleted from a copy of the archived time series database 305. In another embodiment since a copy of the archived time series database 305 is initially created without training data, there is no need to delete it. The system stores the reduced copy of the archived time series database for subsequent processing.
[0119] In one embodiment, in process block 707, the system preprocesses the remaining original time series data values to reduce the impact of sensor failures and generates a model that records changes to the remaining original time series data. For example, intelligent data preprocessing 315 (as described above) may be performed on a reduced copy of the archived time series database 305. The system forms an IDP operation model 710 similar to the IDP model 405, where the IDP operation model 710 is only for the remaining observation records remaining in the reduced copy of the archived time series database 305 after excluding the training observations. The system stores the preprocessed reduced copy of the archived time series database for subsequent processing.
[0120] In one embodiment, in the MSET operation block 715, the system performs state estimation (such as MSET state estimation) for the remaining original time series data values that have been preprocessed using a limited multivariable state estimation model to create a compressed time series database. The limited MSET model 630 and a copy of the archived time series database 305 from which the training data has been excluded are used to perform the MSET operation 715. In one embodiment, as described above, the MSET operation 715 and the limited MSET model 630 are the MSET2 operation and the MSET2 model. The MSET operation 715 (i) forms a state estimation value for each observation based on the remaining parameters, and (ii) removes time series variables that are not parameters of the limited MSET model 630 from a copy of the archived time series database. In this way, data values that do not inform the MSET state estimation (data values that do not significantly affect the value of the MSET state estimation) are deleted from a copy of the archived time series database, further reducing the size of the data. The MSET state estimation values formed for each observation and the remaining parameter values for each observation form the compressed data 720. The compressed data 720 is stored in memory or storage as a data structure.
[0121] Thus, in one embodiment, when the archived time series database is preprocessed, the MSET algorithm is executed using the restricted MSET model 630. The restricted MSET model 630 constitutes a compressed version of the original time series database (compressed data 720), together with the MSET parameters for the restricted MSET model 630. MSET estimates are stored along with the MSET model to represent the original time series data in a reduced size. If it is necessary to identify whether any of the original raw data streams have been modified, changed, or replaced, the system will be able to invert the stored MSET estimates and MSET model to reconstruct the original time series data, and further, use this to validate or invalidate the intended original time series data as discussed below.
[0122] - Data Format - Figures 8A and 8B show two exemplary data reporting formats. Each of these two formats contains sufficient information to perform an audit of a database under audit, such as the original archived time series database 305. In one embodiment, the system generates an electronic data reporting data structure according to one of the two exemplary data reporting formats.
[0123] In one embodiment, the system generates an electronic data reporting data structure. The electronic data reporting data structure is used to mitigate the impact of sensor failures on the quality of the original time series values , including a preprocessing model that records one or more changes to the original time series values during preprocessing, and a compressed time series database, the compressed time series database performing state estimation using a restricted multivariate state estimation model trained with a set of training values selected from the original time series values, and excluding from the compressed time series database values that are not parameters of the restricted multivariate state estimation model, and the electronic data reporting data structure further includes (i) a set of training values and (ii) one or more of the restricted multivariate state estimation models.
[0124] When comparing the two formats, the first data reporting format 800 specifies (i.e., includes in the data report) all of the training data 510 (rather than the restricted MSET2 model 630), while the second data reporting format 850 specifies the restricted MSET2 model 630 (rather than the training data 510). Both the first data reporting format 800 and the second data reporting format 850 specify the IDP operation model 710 and the compressed data 720.
[0125] Accordingly, in one embodiment, the electronic data reporting data structure includes a preprocessing model (IDP operation model 710) that records one or more changes to the original time series values during preprocessing to mitigate the impact of sensor faults on the quality of the original time series values (as illustrated and described with reference to FIGS. 3, 4, and 7). The electronic data reporting data structure also includes a compressed time series database (compressed data 720). The compressed time series database is generated by performing state estimation with a limited multivariate state estimation model and excluding values that are not parameters of the limited multivariate state estimation model (as illustrated and described with reference to FIG. 7) from the compressed time series database. The electronic data reporting data structure also includes one or more of (i) a set of training values selected from the original time series values and used to train the limited multivariate state estimation model, such as in the first data reporting format 800, and (ii) the limited multivariate state estimation model in a second data reporting format 850, etc.
[0126] One exemplary advantage of the first data reporting format 800 is that the details of the MSET (or MSET2) auditing algorithm are not revealed. One exemplary advantage of the second data reporting format (850) is that the amount of data that must be provided in the report is reduced. Note that the amount or volume of the compressed data 720 can generally affect the size of the report in both the first data reporting format 800 and the second data reporting format 850, and thus it should be noted that in practice, there may be no significant difference in size between the two formats in some cases.
[0127] Referring back to FIG. 3, in one embodiment, the system reverses the MSET calculation to reconstruct the raw data, as shown in process block 365. For example, a data report may be created and used to reconstruct data for auditing the archived time series database 305 according to one of data report formats 800, 850. Note that these reports may be stored with journal entries, for example, to await evaluation under the circumstances of auditing the archived time series database 305. In some embodiments, these reports may be stored for a significant period of time before being used to reconstruct the raw data. In some embodiments, the archived time series database 305 may not be audited, and the data reports may not be retrieved and are used to reconstruct the raw data in process block 365. Also, the reports may be distributed to a third party for external auditing of a third party copy of the archived time series database.
[0128] -Data Reconstruction Using an Exemplary First Data Report Format- FIG. 9 shows a schematic diagram of one embodiment of a data reconstruction process 900 using the first data report format 800, and the data reconstruction process 900 is further associated with enabling reliable auditing of the results of a machine learning model.
[0129] In one embodiment, in the data reconstruction process 900, the restricted MSET (or MSET2) model 630 is first trained using the IDP model 405 and the training data 510 as shown in process block 905 (e.g., as illustrated and described with reference to FIG. 5), and then restricted (e.g., as illustrated and described with reference to FIG. 6). The IDP model 405 is retrieved from storage and provided to the training and restriction process 905. The training data 510 is read from the data report 800. Since there is no restricted MSET model 630 from the first data report format 800, it is necessary to create the restricted MSET model 630.
[0130] In one embodiment, the system identifies an inverse state estimation calculation that cancels the steps performed by the state estimation calculation to form a state estimate value from the original time series data values. For example, as shown in process block 910, the MSET (or MSET2) algorithm is inverted and applied to the compressed data 720. To invert the MSET algorithm, the steps of forming the MSET estimate values performed by the trained model 630 are analyzed, and a sequence of discrete operations is recorded. For each of the discrete operations in the sequence, an inverse operation that cancels the discrete operation is identified and recorded in the sequence of inverse operations. In one embodiment, the sequence of inverse operations should be in the reverse order of the sequence of discrete operations. For example, if a discrete operation should be performed first in the sequence of discrete operations, the inverse operation of the discrete operation should be performed last in the sequence of inverse operations, the inverse operation of the second discrete operation should be performed second to last in the sequence of inverse operations, and so on, such that the inverse operations of the discrete operations are performed in the reverse order of the discrete operations. The system stores the sequence of inverse operations for subsequent processing.
[0131] In one embodiment, the system then generates a set of inverted state estimates for the original time series data from the set of state estimates. Each of the inverted state estimates is generated by performing an inverted state estimate for one of the set of state estimates, as indicated, for example, by a sequence of inversion operations. For example, the system performs an inversion of the MSET calculation used to create the state estimates stored in the compressed data 720 on those state estimates. The inverted MSET estimates of the original time series data values for each observation are created from the state estimates of those values and the observations of other parameters stored in the compressed data 720. The sequence of inversion operations is performed on the estimated data values for each observation, creating an inverted MSET estimate for each of the estimated data values.
[0132] In one embodiment, as shown in process block 915, the intelligent data preprocessing 315 is also inverted and applied to the inverted MSET estimates. In one embodiment, the IDP operation model 710 is read from the data report 800, and the original time series values held in the IDP operation model 710 are replaced with any corresponding values within the inverted MSET estimates, thereby inverting the intelligent data preprocessing 315 for the compressed data to form reconstructed data, approximated original data 920. The reconstructed data 920 is stored for later use in an auditing process, including, for example, the verification process illustrated and described with reference to blocks 310 and 370 - 380 of FIG. 3. Thus, the reconstructed time series data values for each of the state estimates are based on the inverted state estimates.
[0133] Note that the resulting reconstructed time-series data approximates the original data in the archived time-series database 305 and may be slightly different from the original data. Nevertheless, the approximation can be used to verify that the archived time-series database 305 has not been tampered with or damaged. Thus, using this approximation, it can be verified that the archived time-series database 305 indicating compliance or violation of regulations accurately indicates such compliance or violation.
[0134] -Data Reconstruction Using an Exemplary Second Data Reporting Format- FIG. 10 shows a schematic diagram of an embodiment of a data reconstruction process 1000 using a second data reporting format 850, and the data reconstruction process 1000 is further associated with enabling the results of a machine learning model to be reliably audited. The reconstruction process 1000 generally follows the same process steps as process 900, except that the limited MSET model 630 does not need to be calculated first because it is already stored in the second data reporting format 850. Note that in one embodiment, the order of performing the inversion 1010 of the intelligent data preprocessing 315 and the MSET inversion 1015 may be switched in process 1000 from the order in which the MSET inversion 910 is performed prior to the inversion 915 of the intelligent data preprocessing 315 in process 900.
[0135] In one embodiment, as shown in process block 1010, intelligent data preprocessing 315 is first inverted and applied to the MSET estimated values within the compressed data 720. In one embodiment, the IDP operation model 710 is read from the data report 800, and the original time series values held within the IDP operation model 710 are replaced with any corresponding values within the MSET estimated values in the compressed data, thereby inverting the intelligent data preprocessing 315 for the compressed data 720. Thus, the reconstructed time series data values for each of the state estimated values are based on the inverted state estimated values as described with reference to process block 1015.
[0136] In one embodiment, the MSET (or MSET2) algorithm is inverted and applied to the compressed data 720 in a manner similar to that described with reference to process block 910 above, as shown in process block 1015. For example, the system performs an inversion of the MSET calculations used to create the state estimated values stored in the compressed data 720 on those state estimated values. The inverted MSET estimated values of the original time series data values for each observation are created from the state estimated values of those values and the observed values of other parameters stored in the compressed data 720. The inverted MSET estimated values and the replaced intelligent data preprocessing values from process block 1010 form the reconstructed data, i.e., the approximated original data 1020. The reconstructed data 1020 is stored for later use in an auditing process that may include a verification process illustrated and described with reference to blocks 310 and 370 - 380 of FIG. 3. Similar to the process 900 above, the resulting time series data approximates the original data within the archived time series database 305 and may differ slightly from the original data, but can be used to verify that the archived time series database 305 has not been tampered with or damaged.
[0137] -Tampering Report- If there is no intentional tampering (or accidental data corruption) in the "original" time series data (as shown by a match in a statistical sense through comparison), the system proves that the original data is intact; otherwise, the original data is corrupted.
[0138] Referring again to FIG. 3, the system then proceeds to audit the (intact or corrupted / tampered) state of the original time series within the archive 305. In one embodiment, under the circumstances of this audit, the accuracy of the reported time series is verified using the reconstructed data. The original data is compared to the reconstructed (or rebuilt) data, e.g., the reconstructed data 920 obtained from process 900 or the reconstructed data 1020 obtained from process 1000. The two data streams, the original time series and the reconstructed time series, are compared pairwise by a pair-difference analyzer as shown in process block 310. The pair-difference analyzer 310 compares the original time series data value with the reconstructed time series data value for each observation of the two time series to confirm whether the reconstructed value matches (or closely approximates within a threshold range) the original value.
[0139] In one embodiment, the pair-difference analyzer compares the original time series data value with the reconstructed time series data value for each observation to determine whether it changes beyond a pre-set threshold amount, e.g., a percentage amount. In another embodiment, the pair-difference analyzer compares each original time series data value with the reconstructed time series data value for each observation to determine whether to trigger a fault detection using a trained fault detection model included in the trained limited MSET2 model 630. The fault detection model can adopt a sequential probability ratio test (SPRT) to analyze the residuals between the original time series data value and the reconstructed time series data value for each observation to determine whether the intended original time series data is abnormal.
[0140] The result of each comparison is evaluated in decision block 370. If the two data streams match (YES), the verification is passed and the process proceeds to process block 375. If the two data streams do not match (NO), the verification fails and the process proceeds to process block 380.
[0141] In process block 375, a "passed" verification report indicating that the original time series data has not been corrupted is generated. In one embodiment, the passed verification report is, at the time of analysis, a signal indicating that the original time series data has not been corrupted. In one embodiment, the passed verification report is a human-readable document indicating that the original time series data has not been corrupted. In one embodiment, the system generates, in response to the passed verification report, an instruction to display an indication that the original time series data has not been corrupted on a graphical user interface (GUI), and executes it or sends it for execution. The GUI may be associated with a utility entity or a regulatory agency entity. Then, the processing in process 300 ends.
[0142] In process block 380, a "failed" verification report indicating that the original time series data is damaged or tampered with is generated. In one embodiment, the failed verification report is, at the time of analysis, a signal indicating that the original time series data is damaged or tampered with. In one embodiment, the failed verification report is a human-readable document indicating that the original time series data is damaged or tampered with. In one embodiment, the system generates, in response to the failed verification report, an instruction to display an indication that the original time series data is damaged or tampered with on the GUI, and executes it or sends it for execution. The GUI may be associated with a utility entity or a regulatory agency entity. Then, the processing in process 300 ends.
[0143] In this way, in response to a signal indicating that the database under audit has not been modified, the system generates an electronic verification report message indicating that the database being verified is not damaged and has not been tampered with. In response to a signal indicating that the database under audit has been modified, the system generates an electronic verification report message indicating that the database under audit is damaged or has been tampered with. Next, the system sends the generated electronic verification report to a computing device so that the verification report message can be stored or displayed by the computing device. In one embodiment, the computing device may be associated with a utility entity or a regulatory agency entity.
[0144] -Embodiments of Software Modules- Generally, software instructions are designed to be executed by a properly programmed processor. These software instructions may include, for example, computer-executable code and source code that can be compiled into the computer-executable code. These software instructions may also include instructions written in an interpreted programming language such as a scripting language.
[0145] In complex systems, such instructions are typically arranged into program modules, each of which performs a specific task, process, function, or operation. The entire set of modules can be controlled or coordinated during their operation by an operating system (OS) or other form of organized platform.
[0146] In one embodiment, one or more of the components, functions, methods, or processes described herein are configured as modules stored in a non-transitory computer-readable medium. These modules, when executed by at least a processor that accesses a memory or storage, are composed of stored software instructions that cause a computing device to perform the corresponding functions described herein.
[0147] -Cloud system, multi-tenant, and enterprise embodiments- In one embodiment, the system is a computing / data processing system that includes an application or a collection of distributed applications for an enterprise organization. The application and the computing system may be configured to operate with or be realized as a cloud-based networking system, a software as a service (SaaS) architecture , or other types of networked computing solutions. In one embodiment, the system is a centralized server-side application that provides at least the functions disclosed herein and is accessed by many users via a computing device / terminal that communicates with a computing system (functioning as a server) via a computer network.
[0148] -Embodiments of computing devices- FIG. 11 shows an exemplary computing device 1100 configured and / or programmed with one or more of the exemplary systems and methods described herein and / or equivalent examples. The exemplary computing device can be a computer 1105 that includes a processor 1110 operably connected by a bus 1125, a memory 1115, and an input / output port 1120. In one example, the computer 1105 can include machine learning audit assurance logic 1130 configured to facilitate ensuring the auditing of the results of a machine learning model, similar to the logic and systems shown in FIGS. 1-10. In various examples, the logic 1130 can be implemented in hardware, a non-transitory computer-readable medium storing instructions, firmware, and / or combinations thereof. Although the logic 1130 is shown as a hardware component attached to the bus 1125, it is recognized that in other embodiments, the logic 1130 can be implemented in the processor 1110, stored in the memory 1115, or stored on the disk 1135. In one embodiment, the logic 1130 or the computer is a means (e.g., structure: hardware, non-transitory computer-readable medium, firmware) for performing the described operations. In some embodiments, the computing device can be a server operating in a cloud computing system, a server configured in a software as a service (SaaS) architecture, a smartphone, a laptop, a tablet computing device, etc.
[0149] The means can be implemented, for example, as an ASIC programmed to facilitate ensuring the auditing of the results of a machine learning model. The means can also be presented to the computer 1105 as stored computer-executable instructions temporarily stored in the memory 1115 and further executed by the processor 1110 as data 1140.
[0150] Logic 1130 may provide means (e.g., hardware, a non-transitory computer-readable medium storing executable instructions, firmware) to enable reliable auditing of the results of a machine learning model.
[0151] To briefly describe an exemplary configuration of computer 1105, processor 1110 may be various processors including dual microprocessor and other multiprocessor architectures. Memory 1115 may include volatile memory and / or non-volatile memory. The non-volatile memory may include, for example, ROM, PROM, etc. The volatile memory may include, for example, RAM, SRAM, DRAM, etc.
[0152] Storage disk 1135 may be operably connected to computer 1100 via, for example, an input / output (I / O) interface (e.g., card, device) 1145 and an I / O port 1120. Disk 1135 may be, for example, a magnetic disk drive, a solid state disk drive, a floppy (registered trademark) disk drive, a tape drive, a Zip drive, a flash memory card, a memory stick, etc. Further, disk 1135 may be a CD-ROM drive, a CD-R drive, a CD-RW drive, a DVD ROM, etc. Memory 1115 may be able to store, for example, process 1150 and / or data 1140. Disk 1135 and / or memory 1115 may store an operating system that controls and allocates the resources of computer 1105.
[0153] Computer 1105 can interact with input / output (I / O) devices via I / O interface 1145 and input / output port 1120. The I / O devices can be, for example, keyboard 1180, microphone 1184, pointing and selection device 1182, camera 1186, video card, display 1170, scanner 1188, printer 1172, speaker 1174, disk 1135, network device 1155, etc. Input / output port 1120 can include, for example, serial ports, parallel ports, and USB ports.
[0154] Computer 1105 is operable in a network environment and can thus be connected to network device 1155 via I / O interface 1145 and / or I / O port 1120. Computer 1105 can interact with network 1160 through network device 1155. Computer 1105 can be logically connected to remote computer 1165 via network 1160. The networks with which computer 1105 can interact can include, but are not limited to, LANs, WANs, and other networks.
[0155] -Definitions and Other Embodiments- In another embodiment, the methods described and / or their equivalents can be implemented using computer-executable instructions. Thus, in one embodiment, a non-transitory computer-readable / storage medium is configured to store computer-executable instructions of an algorithm / executable application that, when executed by a machine, cause the machine (and / or related components) to execute the method. Exemplary machines include, but are not limited to, processors, computers, servers operating in cloud computing systems, servers configured in a software as a service (SaaS) architecture, smartphones, etc. In one embodiment, a computing device is realized using one or more executable algorithms configured to execute any of the disclosed methods.
[0156] In one or more embodiments, the disclosed methods or their equivalents are implemented by computer hardware configured to execute the methods, or by computer instructions embodied in modules stored on a non-transitory computer-readable medium. Here, the instructions are configured as executable algorithms that, when executed by at least a processor of a computing device, execute the methods.
[0157] For the sake of brevity, the exemplary methods shown in the figures are illustrated and described as a series of blocks of an algorithm, but it should be understood that the methods are not limited by the order of the blocks. Some blocks may be performed in a different order than shown and described, and / or concurrently with other blocks. Additionally, fewer blocks than all of the illustrated blocks may be used to implement the exemplary methods. The blocks may be combined or separated into multiple operations / components. Further, additional and / or alternative methods may use additional operations not shown in the blocks.
[0158] The following includes definitions of selected terms used herein. These definitions include various examples and / or forms of components that are within the scope of the terms and can be used for implementation. These examples are not intended to be limiting. Both the singular and plural forms of the terms can be within the scope of the definitions.
[0159] When referring to "one embodiment", "an embodiment", "an example", "an instance", etc., it indicates that the embodiment or example so described may include a particular feature, structure, characteristic, property, element, or limitation, but not all embodiments or examples necessarily include that particular feature, structure, characteristic, property, element, or limitation. Further, repeated use of the expression "in one embodiment" does not necessarily refer to the same embodiment, but may be the same embodiment and may include the following.
[0160] ASIC: Application Specific Integrated Circuit CD: Compact Disc CD-R: Recordable CD CD-RW: Rewritable CD DVD: Digital Versatile Disc and / or Digital Video Disc HTTP: Hypertext Transfer Protocol LAN: Local Area Network RAM: Random Access Memory DRAM: Dynamic RAM SRAM: Synchronous RAM ROM: Read-Only Memory PROM: Programmable ROM EPROM: Erasable PROM EEPROM: Electrically Erasable PROM USB: Universal Serial Bus XML: Extensible Markup Language WAN: Wide Area Network As used herein, "data structure" is the organization of data within a computing system stored in memory, a storage device, or other computerized systems. A data structure may be, for example, any one of a data field, a data file, a data array, a data record, a database, a data table, a graph, a tree, a linked list, etc. A data structure can be formed from and can contain many other data structures (e.g., a database contains many data records). According to other embodiments, other examples of data structures are similarly feasible.
[0161] As used herein, "computer-readable medium" or "computer storage medium" refers to a non-transitory medium that stores instructions and / or data configured to execute one or more of the disclosed functions when executed. The data can function as instructions in some embodiments. The computer-readable medium can take forms including, but not limited to, non-volatile media and volatile media. Non-volatile media can include, for example, optical disks, magnetic disks, etc. Volatile media can include, for example, semiconductor memory, dynamic memory, etc. Common forms of computer-readable media include floppy (registered trademark) disks, flexible disks, hard disks, magnetic tapes, other magnetic media, application specific integrated circuit (ASIC), programmable logic devices, compact disk (CD), other optical media, random access memory (RAM), read only memory (ROM), memory chips or cards, memory sticks, solid state storage device (SSD), flash drives, and other media that a computer, processor or other electronic device can function with, but are not limited thereto. When selected for implementation in an embodiment, each type of media can include stored instructions of an algorithm configured to execute one or more of the disclosed functions and / or the claimed functions.
[0162] As used herein, "logic" represents components implemented by computer or electrical hardware, a non-transitory medium storing instructions of executable applications or program modules, and / or combinations thereof, for performing any of the functions or operations disclosed herein and / or for causing a computer or electrical hardware to perform functions or operations from another disclosed logic, method, and / or system. Equivalent logic may include firmware, a microprocessor programmed with an algorithm, discrete logic (e.g., ASIC), at least one circuit, an analog circuit, a digital circuit, a programmed logic device, a memory device containing instructions of an algorithm, etc., any of which may be configured to perform one or more of the disclosed functions. In one embodiment, the logic may include one or more gates, combinations of gates, or other circuit components configured to perform one or more of the disclosed functions. When multiple logics are described, it may be possible to incorporate the multiple logics into one logic. Similarly, when a single logic is described, it may be possible to distribute the single logic among multiple logics. In one embodiment, one or more of these logics are corresponding structures related to performing the disclosed and / or claimed functions. The choice of which type of logic to implement may be based on desired system conditions or specifications. For example, if higher speed is a consideration, hardware may be selected to implement the function. If lower cost is a consideration, stored instructions / executable applications may be selected to implement the function. If lower cost is a consideration, stored instructions / executable applications may be selected to implement the function.
[0163] An "operable connection", or a connection by which an entity is "operably connected", is a connection through which signals, physical communication, and / or logical communication can be sent and / or received. An operable connection can include a physical interface, an electrical interface, and / or a data interface. An operable connection can include various combinations of interfaces and / or connections sufficient to enable operable control. For example, two entities can be operably connected to communicate signals directly with each other or through one or more intermediate entities (e.g., a processor, an operating system, logic, a non-transitory computer-readable medium). An operable connection can be created using a logical communication channel and / or a physical communication channel.
[0164] As used herein, "user" includes, but is not limited to, one or more persons, computers or other devices, or combinations thereof.
[0165] Although the disclosed embodiments have been illustrated and described in considerable detail, it is not intended to limit the appended claims to such detail or in any way. Of course, it is impossible to describe every conceivable combination of components or methodologies for purposes of explaining the various aspects of the subject matter. Accordingly, the disclosure is not limited to the specific details or specific examples illustrated and described. For this reason, the disclosure is intended to cover modifications, variations, and equivalents that fall within the scope of the appended claims.
[0166] To the extent that the term "includes" or "including" is used in the detailed description or claims it is intended to be inclusive in the same manner as the term "comprising" as interpreted when used as a transitional word in a claim.
[0167] As long as the word "or" is used in the detailed description or claims (e.g., A or B), it is intended to mean "A or B or both". If the applicant intends to indicate that it is "only A or B, not both", the expression "only A or B, not both" will be used. Therefore, the use of the word "or" in this specification is an inclusive use, not an exclusive use.
Claims
1. 1. A computer-implemented method for auditing results of a machine learning model, comprising: retrieving a set of state estimates for original time series data values from a database under audit, each of said state estimates being generated by a state estimation calculation for one of said time series data values, said method further comprising: inverting the state estimation calculation for each of the state estimates to generate a reconstructed time series data value for each of the state estimates; Retrieving said original time series data values from said database under audit; comparing the original time series data values with the reconstructed time series data values in a pairwise manner to determine whether the original and reconstructed time series match; generating a signal that the database under audit (i) has not been modified if the original time series and the reconstructed time series match, and (ii) has been modified if the original time series and the reconstructed time series do not match.
2. training a multivariate state estimation model with a set of training values selected from the original time series values; decomposing a matrix associated with the multivariate state estimation module into a set of eigenvectors; selecting a subset of principal eigenvectors from the set of eigenvectors; and generating a restricted multivariate state estimation model from the subset of principal eigenvectors. The method of claim 1 , wherein inverting the state estimate calculation further comprises generating an inverse of a calculation in which the constrained multivariate state estimation model forms a state estimate.
3. excluding training data values from the original time series data values from the database under audit; pre-processing the remaining original time series data values to mitigate the effects of sensor impairments and generating a model that records changes to the remaining original time series data; and performing state estimation on the preprocessed remaining original time series data values using the restricted multivariate state estimation model to create a compressed time series database.
4. pre-processing the original time series values to mitigate the effect of sensor impairments on the quality of the original time series values; generating a data preprocessing model that records one or more changes to the original time series values during the preprocessing; The inversion of the state estimate calculation further comprises: deriving the data preprocessing model; and reversing the pre-processing described by the data pre-processing model for each of the state estimates.
5. generating an electronic data report data structure, the electronic data report data structure comprising: a preprocessing model that records one or more changes to the original time series values during preprocessing to mitigate the effects of sensor impairments on the quality of the original time series values; and a compressed time series database, the compressed time series database comprising: performing state estimation with a restricted multivariate state estimation model; and excluding values from the compressed time series database that are not parameters of the restricted multivariate state estimation model. and the electronic data report data structure is further generated by:
2. The method of claim 1 , comprising one or more of: (i) a set of training values selected from the original time series values and used to train the restricted multivariate state estimation model; and (ii) the restricted multivariate state estimation model.
6. The inversion of the state estimate calculation further comprises: identifying an inverse state estimate calculation that undoes the steps performed by the state estimate calculation to form the state estimate from the original time series data values; generating a set of inverted state estimates for the original time series data from the set of state estimates, each of the inverted state estimates being generated by performing the inverted state estimation on one of the set of state estimates; The method of claim 1 , wherein the reconstructed time series data values for each of the state estimates are based on the inverted state estimate.
7. (i) in response to a signal that the database under audit has not been modified, generating an electronic verification report message indicating that the database under audit has been certified as uncorrupted and unaltered; (ii) in response to a signal that the database under audit has been modified, generating an electronic verification report message indicating that the database under audit has been corrupted or tampered with; 2. The method of claim 1, further comprising the step of: transmitting the generated electronic validation report to a computing device and causing the validation report message to be stored by the computing device or displayed by the computing device.
8. 1. A non-transitory computer-readable medium storing computer-executable instructions for auditing results of a machine learning model, the computer-executable instructions, when executed by at least a processor of a computer, causing the computer to: retrieving a set of state estimates for the original time series data values from the database under audit, said state estimates having been generated by the state estimation calculation for each of said time series data values; and inverting the state estimation calculation for each of the state estimates to generate a reconstructed time series data value for each of the state estimates; retrieving said original time series data values from said database under audit; comparing the original time series data values to the reconstructed time series data values in a pairwise manner to determine whether the original and reconstructed time series match; and (ii) generating a signal that the database under audit has not been modified if the original time series and the reconstructed time series do not match, and (iii) that the database under audit has been modified if the original time series and the reconstructed time series do not match.
9. The instructions further cause the computer to: training a multivariate state estimation model with a set of training values selected from the original time series values; decomposing a matrix associated with the multivariate state estimation module into a set of eigenvectors; selecting a subset of eigenvectors from the set of eigenvectors having maximum eigenvalues; generating a restricted multivariate state estimation model from the subset of eigenvectors; Pre-processing the original time series values to determine the effect of sensor impairments on the quality of the original time series values Reduce the creating a data preprocessing model that records one or more changes to the original time series values during the preprocessing; performing state estimation on the original time series data values using the restricted multivariate state estimation model to generate a compressed time series database; generating a report including the preprocessed model, the compressed time series database, and one or more of (i) the set of training values, and (ii) the restricted multivariate state estimation model; 9. The non-transitory computer-readable medium of claim 8, wherein the set of state estimates is the set of state estimates included in the compressed time series database in the report, and the reconstructed time series data is generated based on the report.
10. 1. A computing system for auditing results of a machine learning model, comprising: A processor; a memory operatively connected to the processor; a sensor interface operatively connected to the processor and the memory; and a non-transitory computer-readable medium operatively connected to the processor and the memory and storing computer-executable instructions, the computer-executable instructions, when executed by at least a processor of a computer, causing the computer to: retrieving from a database under audit a set of state estimates for the original time series data values received via said sensor interface, said state estimates having been generated by a state estimation calculation for each of said time series data values; and inverting the state estimation calculation for each of the state estimates to generate a reconstructed time series data value for each of the state estimates; retrieving said original time series data values from said database under audit; comparing the original time series data values to the reconstructed time series data values in a pairwise manner to determine whether the original and reconstructed time series match; and generating a signal that the database under audit (i) has not been modified if the original time series and the reconstructed time series match, and (ii) has been modified if the original time series and the reconstructed time series do not match.
11. The non-transitory computer readable medium further includes instructions that, when executed by at least the processor, cause the computing system to: training a multivariate state estimation model with a set of training values selected from the original time series values; decomposing a matrix associated with the multivariate state estimation module into a set of eigenvectors; selecting a subset of principal eigenvectors from the set of eigenvectors; generating a restricted multivariate state estimation model from the subset of principal eigenvectors; The computing system of claim 10 , wherein inverting the state estimation calculation further comprises generating an inverse of a calculation in which the constrained multivariate state estimation model forms a state estimate.
12. The non-transitory computer readable medium further includes instructions that, when executed by at least the processor, cause the computing system to: removing training data values from said original time series data values from said database under audit; The remaining original time series data values are preprocessed to reduce the effect of sensor failures, and the remaining original Generate a model that records changes to the time series data, 11. The computing system of claim 10, further comprising: performing state estimation on the preprocessed remaining original time series data values using the restricted multivariate state estimation model to create a compressed time series database.
13. The instructions for reversing the state estimation calculation further cause the computing system to generate an electronic data reporting data structure, the electronic data reporting data structure comprising: a preprocessing model that records one or more changes to the original time series values during preprocessing to mitigate the effects of sensor impairments on the quality of the original time series values; and a compressed time series database, the compressed time series database being generated by performing state estimation with a restricted multivariate state estimation model and excluding values from the compressed time series database that are not parameters of the restricted multivariate state estimation model, the electronic data report data structure further comprising: (i) a set of training values selected from the original time series values and used to train the restricted multivariate state estimation model; and (ii) the restricted multivariate state estimation model; The instructions for inverting the state estimate calculation further include causing the computer to: canceling the steps performed by said state estimate calculation using an inverse state estimate calculation; The computing system of claim 10 , further comprising: inverting the pre-processing described by the pre-processing model for each of the state estimates.
14. The non-transitory computer readable medium further includes instructions that, when executed by at least the processor, cause the computing system to: identifying an inverse state estimate computation that undoes steps performed by the state estimate computation to form the state estimates from the original time series data values; generating a set of inverted state estimates for the original time series data from the set of state estimates, each of the inverted state estimates being generated by performing the inverted state estimation on one of the set of state estimates; The computing system of claim 10 , wherein the reconstructed time series data values for each of the state estimates are based on the inverted state estimate.
15. The non-transitory computer readable medium further includes instructions that, when executed by at least the processor, cause the computing system to: (i) in response to a signal that the database under audit has not been modified, generating an electronic verification report message indicating that the database under audit has been certified as uncorrupted and unaltered; (ii) in response to a signal that the database under audit has been modified, generating an electronic verification report message indicating that the database under audit has been corrupted or tampered with; The computing system of claim 10 , further comprising: transmitting the generated electronic validation report to a computing device, where the validation report message is stored by the computing device or displayed by the computing device.
Citation Information
Patent Citations
Power system monitor
JP1993161265A
Management server, monitoring terminal, control method of management server and control program thereof
JP2019117428A
Machine learning system
JP2019212121A
Telemetry data analysis using multivariate sequential probability ratio test
US20100292959A1
MSET-based process for certifying provenance of time-series data in a time-series database
US20190197145A1