Computer system, computer program, and computer-implemented method (data quality-based reliability calculation derived from time-series data)
The system addresses data errors in time series data by calculating confidence values for corrected data, improving reliability and accuracy in analysis and decision-making.
Patent Information
- Application Number
- JP2021183697
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-25
- Filing Date
- 2021-11-10
- Publication Date
- 2025-10-09
- Estimated Expiration
- 2041-11-10
AI Technical Summary
Time series data collected from IoT devices and smart home devices often contains errors due to device malfunctions or system issues, leading to inaccurate analysis and decision-making, such as erroneous power shutoffs.
A system and method to calculate confidence values for corrected data by identifying potentially erroneous data instances, determining predicted replacement values, and generating explanatory evidence for resolution.
Enhances data reliability by correcting errors in time series data, ensuring accurate analysis and decision-making processes.
Smart Images

Figure 0007751937000001 
Figure 0007751937000002 
Figure 0007751937000003
Abstract
Description
[Background technology]
[0001] The present disclosure relates to calculating confidence values for corrected data in time series data, and more particularly, to providing replacement data confidence values for data having issues indicating one or more errors, wherein the data issues, replacement data, and confidence values are related to one or more key performance indicators (KPIs).
[0002] Many known entities, including businesses and residential entities, include systems that collect time series data from various sources, such as Internet of Things (IoT) devices, smart home devices, human activity, appliance activity, etc. The collected data can be analyzed to facilitate energy conservation, occupancy allocation, etc. Summary of the Invention [Problem to be solved by the invention]
[0003] Sometimes, some of the collected time series data may be in error due to various reasons such as malfunction of the controlled device, malfunction of the respective sensing device, problems with the data collection system, data storage system or data transmission system. [Means for solving the problem]
[0004] A system, computer program product, and method are provided for calculating a confidence value for corrected data in time series data.
[0005] In one aspect, a computer system is provided for calculating confidence values for corrected data in time-series data. The system includes one or more processing devices and at least one memory device operably coupled to the one or more processing devices. The one or more processing devices are configured to identify one or more potentially erroneous data instances in the time-series data stream and determine one or more predicted replacement values for the one or more potentially erroneous data instances. The one or more processing devices are also configured to determine a confidence value for each predicted replacement value of the one or more predicted values and resolve the one or more potentially erroneous data instances using one of the one or more predicted replacement values. The one or more processing devices are further configured to generate explanatory basis for the resolution of the one or more potentially erroneous data instances.
[0006] In another aspect, a computer program product is provided for calculating confidence values for corrected data in time-series data. The computer program product includes one or more computer-readable storage media and program instructions collectively stored on the one or more computer storage media. The product also includes program instructions for identifying one or more potentially erroneous data instances in the time-series data stream. The product further includes program instructions for determining one or more predicted replacement values for the one or more potentially erroneous data instances. The product also includes program instructions for determining a confidence value for each predicted replacement value of the one or more predicted values. The product further includes program instructions for resolving the one or more potentially erroneous data instances using a predicted replacement value of the one or more predicted replacement values. The product also includes program instructions for generating explanatory evidence for the resolving the one or more potentially erroneous data instances.
[0007] In yet another aspect, a computer-implemented method is provided for calculating confidence values for corrected data in time-series data. The method includes identifying one or more potentially erroneous data instances in a time-series data stream. The method also includes determining one or more predicted replacement values for the one or more potentially erroneous data instances. The method further includes determining a confidence value for each predicted replacement value of the one or more predicted replacement values. The method also includes resolving the one or more potentially erroneous data instances using a predicted replacement value of one of the one or more predicted replacement values. The method further includes generating explanatory evidence for the resolving the one or more potentially erroneous data instances.
[0008] This summary is not intended to describe each aspect, every implementation, or every embodiment of the present disclosure, or combinations thereof. These and other features and advantages will become apparent from the following detailed description of the present embodiments taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0009] The drawings included herein are incorporated into and form a part of the specification. They illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the disclosure. The drawings illustrate particular embodiments and are not intended to limit the disclosure.
[0010] [Figure 1] FIG. 1 is a schematic diagram illustrating a cloud computing environment according to some embodiments of the present disclosure.
[0011] [Figure 2] FIG. 1 is a block diagram illustrating a set of functional abstraction model layers provided by a cloud computing environment, according to some embodiments of the present disclosure.
[0012] [Figure 3]FIG. 1 is a block diagram illustrating a computer system / server that may be used as a cloud-based support system for implementing the processes described herein, according to some embodiments of the present disclosure.
[0013] [Figure 4] FIG. 1 is a schematic diagram illustrating a system for calculating a confidence value for corrected data in time series data according to some embodiments of the present disclosure.
[0014] [Figure 5A] 1 is a flowchart illustrating a process for calculating a confidence value for corrected data in time series data according to some embodiments of the present disclosure.
[0015] [Figure 5B] 5B is a continuation of the flowchart shown in FIG. 5A illustrating a process for calculating a confidence value for corrected data in time-series data according to some embodiments of the present disclosure.
[0016] [Figure 5C] 5A and 5B, illustrating a process for calculating a confidence value for corrected data in time-series data according to some embodiments of the present disclosure.
[0017] [Figure 6] FIG. 1 is a text diagram illustrating an algorithm for identifying related issues according to some embodiments of the present disclosure.
[0018] [Figure 7] FIG. 1 is a text diagram illustrating an algorithm for observable box KPI analysis according to some embodiments of the present disclosure.
[0019] [Figure 8] FIG. 1 is a text diagram illustrating an algorithm for unobservable box KPI analysis according to some embodiments of the present disclosure.
[0020] [Figure 9] FIG. 1 is a schematic diagram illustrating a portion of a process for snapshot simulation according to some embodiments of the present disclosure.
[0021] [Figure 10] FIG. 1 is a schematic diagram illustrating a process for method-based simulation according to some embodiments of the present disclosure.
[0022] [Figure 11] FIG. 1 is a schematic diagram illustrating a process for point-based simulation according to some embodiments of the present disclosure.
[0023] [Figure 12] FIG. 1 is a text diagram illustrating an algorithm for a snapshot optimizer according to some embodiments of the present disclosure.
[0024] [Figure 13] FIG. 1 is a graphical diagram illustrating KPI value estimation according to some embodiments of the present disclosure.
[0025] [Figure 14] FIG. 1 is a graphical / textual diagram illustrating the generation of a confidence measure according to some embodiments of the present disclosure.
[0026] [Figure 15] FIG. 1 is a graphical diagram illustrating a confidence measure according to some embodiments of the present disclosure.
[0027] [Figure 16] FIG. 1 is a text diagram illustrating a description of a confidence measure according to some embodiments of the present disclosure.
[0028] While the present disclosure is susceptible to various modifications and alternative forms, specific features thereof have been shown by way of example in the drawings and are described in detail below. It should be understood, however, that there is no intention to limit the disclosure to the particular embodiments described. On the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0029] It will be readily understood that the components of the present embodiments, as generally described and illustrated in the Figures herein, could be arranged and designed in a wide variety of different configurations. Thus, the following detailed description of the present apparatus, system, method and computer program product embodiments, as illustrated in the Figures, is not intended to limit the scope of the embodiments as claimed, but is merely representative of selected embodiments. Moreover, while specific embodiments have been described herein for purposes of illustration, it will be understood that various modifications may be made without departing from the spirit and scope of the embodiments.
[0030] Throughout this specification, references to "selected embodiments," "at least one embodiment," "one embodiment," "another embodiment," "other embodiments," or "embodiments" and similar phrases mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. Thus, appearances of the phrases "selected embodiments," "at least one embodiment," "in one embodiment," "another embodiment," "other embodiments," or "embodiments" in various places throughout this specification do not necessarily refer to the same embodiment.
[0031] The exemplary embodiments are best understood by referring to the drawings, in which like parts are designated by like numerals throughout. The following description is intended for purposes of example only and merely illustrates certain selected embodiments of devices, systems and processes consistent with the embodiments as claimed herein.
[0032] Although this disclosure includes specific references to cloud computing, it should be understood that implementation of the techniques recited herein is not limited to cloud computing environments. Rather, embodiments of the present disclosure can be implemented in conjunction with any other type of computing environment not now known or later developed.
[0033] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0034] The characteristics are as follows:
[0035] On-Demand Self-Service: Cloud consumers can unilaterally provision computing capabilities such as server time and network storage automatically as needed without requiring human interaction with the provider of the service.
[0036] Wide network access: Capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin-client and thick-client platforms (eg, cell phones, laptops, and PDAs).
[0037] Resource Pooling: Provider computing resources are pooled and serve multiple consumers using a multi-tenant model with different physical and virtual resources dynamically allocated and reallocated according to demand. Consumers generally have no control or knowledge of the exact location of the resources provided, although there is an implication of location independence in that they may be able to specify location at a higher level of abstraction (e.g., country, state, or data center).
[0038] Rapid Elasticity: Capacity can be provisioned quickly and elastically, in some cases automatically, to quickly scale out, and can be rapidly released to quickly scale in. To the consumer, the capacity available for provisioning often appears unlimited, and any amount can be purchased at any time.
[0039] Metered Services: Cloud systems automatically control and optimize resource usage using metering capabilities at several levels of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services used.
[0040] The service model is as follows:
[0041] Software as a Service (SaaS): The consumer is offered the ability to use a provider's applications running on a cloud infrastructure. The applications are accessible from a variety of client devices through a thin-client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.
[0042] Platform as a Service (PaaS): The ability offered to consumers is to deploy applications they create or acquire using programming languages and tools supported by the provider onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the application hosting environment configuration.
[0043] Infrastructure as a Service (IaaS): The ability provided to consumers is to provision processing, storage, network, and other underlying computing resources, on which the consumer can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but does have control over the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).
[0044] The deployment model is as follows:
[0045] Private Cloud: Cloud infrastructure is operated solely for an organization. The cloud infrastructure may be managed by that organization or a third party and may reside on-premise or off-premise.
[0046] Community Cloud: Cloud infrastructure is shared by several organizations to support a specific community of shared interests (e.g., mission, security requirements, policies, and compliance considerations). The cloud infrastructure may be managed by the organizations or a third party and may reside on-premises or off-premises.
[0047] Public Cloud: Cloud infrastructure is made available to the general public or large industry groups and is owned by organizations that sell cloud services.
[0048] Hybrid Cloud: A cloud infrastructure is a composition of two or more clouds (private, community, or public) that remain distinct entities but are tied together by standardized or proprietary technologies that allow for data and application portability (e.g., cloud bursting to balance load between clouds).
[0049] Cloud computing environments are service-oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0050] Referring now to FIG. 1 , an exemplary cloud computing environment 50 is shown. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10, with which local computing devices used by cloud consumers, such as, for example, a personal digital assistant (PDA) or cellular phone 54A, a desktop computer 54B, a laptop computer 54C, or a vehicle computer system 54N, or combinations thereof, may communicate. The nodes 10 may also communicate with each other. The nodes 10 may be physically or virtually grouped together in one or more networks (not shown), such as the private, community, public, or hybrid clouds described above, or combinations thereof. This enables the cloud computing environment 50 to provide infrastructure-as-a-service, platform-as-a-service, or software-as-a-service services, or combinations thereof, without the need for cloud consumers to maintain resources on their local computing devices. It will be understood that the types of computing devices 54A-N shown in FIG. 1 are intended to be exemplary only, and that computing node 10 and cloud computing environment 50 can communicate with any type of computerized device via any type of network or network-addressable connection (e.g., using a web browser), or both.
[0051] Referring now to Figure 2, a set of functional abstraction layers provided by cloud computing environment 50 (Figure 1) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 2 are intended to be exemplary only, and embodiments of the present disclosure are not limited thereto. As depicted, the following layers and corresponding functions are provided:
[0052] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (reduced instruction set computer) architecture-based server 62, server 63, blade server 64, storage device 65, and network and network components 66. In some embodiments, software components include network application server software 67 and database software 68.
[0053] The virtualization layer 70 provides an abstraction layer in which the following examples of virtual entities may be provided: virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75.
[0054] In one example, management layer 80 may provide the functions described below. Resource provisioning 81 provides dynamic procurement of computing and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 82 provides cost tracking as resources are used within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides protection for cloud consumer identities and tasks, as well as data and other resources. User portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 84 provides cloud computing resource allocation and management so that required service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides advance arrangements and procurement for cloud computing resources where future requirements are anticipated according to SLAs.
[0055] The workload layer 90 provides examples of functionality for which a cloud computing environment may be utilized. Examples of workloads and functionality that may be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom instruction delivery 93, data analytics processing 94, transaction processing 95, and calculation of confidence values for time series data 96.
[0056] 3, a block diagram of an exemplary data processing system is provided, referred to herein as computer system 100. System 100 may be embodied in a computer system / server at a single location, or, in at least one embodiment, may be configured in a cloud-based system that shares computing resources. For example, and without limitation, computer system 100 may be used as cloud computing node 10.
[0057] Aspects of computer system 100 may be embodied in a computer system / server at a single location, or, in at least one embodiment, configured in a cloud-based system sharing computing resources as a cloud-based support system implementing the systems, tools, and processes described herein. Computer system 100 is operable with numerous other general-purpose or special-purpose computer system environments or configurations. Examples of well-known computer systems, environments, or configurations, or combinations thereof, that may be suitable for use with computer system 100 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and file systems (e.g., distributed storage environments and distributed cloud computing environments) that include any of the above systems, devices, and their equivalents.
[0058] Computer system 100 may be described in the general context of computer system-executable instructions, such as program modules, executed by computer system 100. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. Computer system 100 may be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.
[0059] As shown in FIG. 3, computer system 100 is illustrated in the form of a general-purpose computing device. Components of computer system 100 may include, but are not limited to, one or more processors or processing devices 104 (which may also be referred to as processors and processing units), such as a hardware processor; system memory 106 (which may also be referred to as a memory device); and a communication bus 102 that couples various system components, including system memory 106, to processing device 104. Communication bus 102 may represent one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures include an Industry Standard Architecture (ISA) bus, a MicroChannel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus. Computer system 100 typically includes a variety of computer system-readable media. Such media may be any available media that is accessible by computer system 100, and includes both volatile and nonvolatile media, removable and non-removable media. Additionally, computer system 100 may include one or more persistent storage devices 108, communications units 110, input / output (I / O) units 112, and displays 114.
[0060] Processing device 104 serves to execute instructions for software that may be loaded into system memory 106. Processing device 104 may be multiple processors, a multi-core processor, or another type of processor, depending on the particular implementation. As used herein with respect to an item, a number may mean one or more of the item. Furthermore, processing device 104 may be implemented using a heterogeneous processor system in which a main processor resides on a single chip along with secondary processors. As another example, processing device 104 may be a symmetric multiprocessor system including multiple processors of the same type.
[0061] System memory 106 and persistent storage 108 are examples of storage devices 116. A storage device may be any piece of hardware capable of storing information, such as, without limitation, data, program code in a functional form, and / or other suitable information on either a temporary and / or permanent basis. System memory 106, in these examples, may be, for example, a random access memory or any other suitable volatile or non-volatile storage device. System memory 106 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) or cache memory, or both.
[0062] Persistent storage 108 may take various forms, depending on the particular implementation. For example, persistent storage 108 may include one or more components or devices. For example, without limitation, persistent storage 108 may be provided for reading from and writing to a non-removable, non-volatile magnetic medium (not shown, typically referred to as a “hard drive”). Although not shown, a magnetic disk drive may be provided for reading from and writing to a removable, non-volatile magnetic disk (e.g., a “floppy disk”), and an optical disk drive may be provided for reading from or writing to a removable, non-volatile optical disk, such as a CD-ROM, DVD-ROM, or other optical media. In such examples, each may be connected to communication bus 102 by one or more data media interfaces.
[0063] Communications unit 110, in these examples, may provide for communications with other computer systems or devices. In these examples, communications unit 110 is a network interface card. Communications unit 110 may provide for communications through the use of either or both physical and wireless communications links.
[0064] The input / output unit 112 may enable the input and output of data with other devices that may be connected to the computer system 100. For example, the input / output unit 112 may provide a connection for user input through a keyboard, a mouse, or some other suitable input device, or a combination thereof. Additionally, the input / output unit 112 may send output to a printer. The display 114 may provide a mechanism for displaying information to a user. Examples of the input / output unit 112 that facilitate establishing communications between various devices within the computer system 100 include, without limitation, a network card, a modem, and an input / output interface card. Furthermore, the computer system 100 may communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, via a network adapter (not shown in FIG. 3 ). It should be understood that other hardware and / or software components, not shown, may be used in conjunction with the computer system 100. Examples of such components include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.
[0065] Instructions for the operating system, applications, or programs, or combinations thereof, may be located in storage device 116, which communicates with processing device 104 over communication bus 102. In these examples, the instructions are in a functional form on persistent storage 108. These instructions may be loaded into system memory 106 for execution by processing device 104. The processes of the different embodiments may be performed by processing device 104 using computer-implemented instructions, which may be located in a memory, such as system memory 106. These instructions may be referred to as program code, computer-usable program code, or computer-readable program code, which may be read and executed by a processor in processing device 104. The program code in the different embodiments may be embodied on different physical or tangible computer-readable media, such as system memory 106 or persistent storage 108.
[0066] Program code 118 may be located in a functional form on computer-readable medium 120, which is selectively removable, and may be loaded onto or transferred to computer system 100 for execution by processing device 104. In these examples, program code 118 and computer-readable medium 120 may form computer program product 122. In one example, computer-readable medium 120 may be computer-readable storage medium 124 or computer-readable signal medium 126. Computer-readable storage medium 124 may include, for example, an optical or magnetic disk inserted into or placed in a drive or other device that is part of persistent storage 108 for transfer to a storage device, such as a hard drive that is part of persistent storage 108. Computer-readable storage medium 124 may also take the form of persistent storage, such as a hard drive, thumb drive, or flash memory that is connected to computer system 100. In some examples, computer-readable storage medium 124 may not be removable from computer system 100.
[0067] Alternatively, program code 118 may be transferred to computer system 100 using computer-readable signal medium 126. Computer-readable signal medium 126 may be, for example, a propagated data signal containing program code 118. For example, computer-readable signal medium 126 may be an electromagnetic signal, an optical signal, or any other suitable type of signal, or a combination thereof. These signals may be transmitted over communications links, such as wireless communications links, fiber optic cable, coaxial cable, a wire, or any other suitable type of communications link, or a combination thereof. That is, in the examples, the communications links and / or connections may be physical or wireless.
[0068] In some demonstrative embodiments, program code 118 may be downloaded to persistent storage 108 via computer-readable signal medium 126 from another device or computer system over a network and used within computer system 100. For example, program code stored in a computer-readable storage medium in a server computer system may be downloaded from the server over a network to computer system 100. The computer system providing program code 118 may be a server computer, a client computer, or some other device capable of storing and transmitting program code 118.
[0069] Program code 118 may include, by way of example and not limitation, one or more program modules (not shown in FIG. 3 ) that may be stored in system memory 106, as well as an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, may comprise an implementation of a network environment. The program modules of program code 118 generally perform the functions and / or methodologies of embodiments as described herein.
[0070] The different components illustrated for computer system 100 are not meant to provide architectural limitations on the manner in which different embodiments may be implemented. Different illustrative embodiments may be implemented in computer systems including components in addition to or instead of those illustrated for computer system 100.
[0071] The present disclosure may be a system, method, or computer program product, or combination thereof, at any conceivable level of technical detail of integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions for causing a processor to perform aspects of the present disclosure.
[0072] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction-execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or groove-embossed structures having instructions recorded thereon, and any suitable combination of the foregoing. Computer-readable storage medium, as used herein, is not to be construed as a transitory signal per se, such as, for example, radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted over a wire.
[0073] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium into each computing / processing device, or may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in the respective computing / processing device.
[0074] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium into each computing / processing device, or may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in the respective computing / processing device.
[0075] The computer-readable program instructions for carrying out the operations of the present disclosure may be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for an integrated circuit, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk® or C++, and procedural programming languages such as the “C” programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet Service Provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer readable program instructions using state information of the computer readable program instructions to personalize the electronic circuit to carry out aspects of the present disclosure.
[0076] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0077] These computer-readable program instructions may be provided to a computer processor or other programmable data processing apparatus to produce a machine, which, when executed by the computer processor or other programmable data processing apparatus, forms means for implementing the functions / acts specified in a block or blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored in a computer-readable storage medium, which may instruct a computer, programmable data processing apparatus and / or other device to function in a particular manner, such that a computer-readable storage medium having instructions stored thereon comprises a product including instructions that implement aspects of the functions / acts specified in a block or blocks of the flowcharts and / or block diagrams.
[0078] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device and executed on the computer, other programmable apparatus, or other device to produce a series of operational steps to generate a computer-implemented process, such that the instructions executing on the computer, other programmable apparatus, or other device implement the function / acts specified in a block or blocks of the flowchart and / or block diagram.
[0079] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and processing of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may actually be executed as a single step, may be executed concurrently or substantially concurrently, in a partially or fully overlapping manner, or in some cases, the blocks may be executed in the reverse order depending on the functionality involved. It should also be noted that each block of a block diagram or flowchart diagram or combination thereof, and combinations of blocks in block diagrams or flowchart diagrams or combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can execute a combination of dedicated hardware and computer instructions.
[0080] Many known entities, including businesses and residential entities, include systems that collect time series data from various sources, such as Internet of Things (IoT) devices, smart home devices, human activity, appliance activity, etc. The collected data may be analyzed to facilitate energy conservation, occupancy allocation, etc. Occasionally, portions of the collected time series data may be erroneous due to various reasons, such as malfunction of the controlled device, malfunction of the respective sensing device, issues with the data collection system, data storage system, or data transmission system, etc. For example, in one embodiment, an occupancy management system may evaluate power usage relative to a peak load value and, in response to an erroneous occupancy data value, erroneously initiate power shutoff of predetermined devices in the associated space to avoid peak usage charges.
[0081] Additionally, many known entities have one or more key performance indicators (KPIs); as used herein, KPI refers to one or more measurable indicators associated with one or more key objectives. KPIs facilitate achieving key objectives through measurement of success in meeting those key objectives. KPIs are scalable in that enterprise-wide KPIs may be used, as well as lower-level, sub-organization-specific KPIs, e.g., sales, marketing, HR, IT support, and maintenance KPIs. In some embodiments, KPIs are explicitly identified and described in one or more documents; in some embodiments, KPIs become apparent upon analysis of collected data; "hidden" KPIs may be "discovered" and previously defined KPIs may be validated.
[0082] Systems, computer program products, and methods are disclosed and described herein for collecting time series data from one or more sensor devices. In some embodiments, the system includes a data quality-KPI prediction confidence engine. The collected time series data is referred to herein as “original data” and “original time series data stream.” Unless otherwise indicated, each time series data stream originates from a single sensing device or from multiple sensors, as described herein for various embodiments, and the stream is a combination, e.g., either the output from each sensor or an aggregate of the output from one sensor among multiple sensors that have been auctioned or otherwise selected. Also, unless otherwise indicated, the system is configured to simultaneously analyze multiple data streams, depending on the number of data streams, without limitation, as described herein. Thus, the system is configured to analyze individual data streams as described herein.
[0083] In one or more embodiments, the quality of the data embedded within each data stream is analyzed, and a determination is made in response to one or more KPIs associated with each data stream through a two-stage process. First, the quality of the original data, such as data packets, is transmitted from the sensor to a data inspection module, where the data packets are inspected by a data inspection sub-module embedded within the data inspection module. In some cases, one or more data packets may contain issues that identify the respective data packets as containing potentially defective data. One such issue may be associated with sampling frequency. For example, and without limitation, the data inspection sub-module may check the sampling frequency of the data sensor to determine whether there are multiple sampling frequencies in the data, such as occasional perturbations in the sampling frequency and consecutive changes in the sampling frequency. Also, for example, and without limitation, the data inspection sub-module may check the timestamps of the data to determine whether the data is missing timestamps, whether data is missing for consecutive extended durations, and whether timestamps have changed formats. Also, by way of example and without limitation, the Data Inspection sub-module may check for syntactic value issues to determine whether data that is presumed to be numeric contains extensive durations of data that is "not a number (NaN)" and improper numeric rounding and truncation. Further, by way of example and without limitation, the Data Inspection sub-module may check for semantic value issues to determine whether any of the data contains anomalous events and noisy data. Thus, the Data Inspection sub-module may examine the data in the stream to determine whether the data is within predetermined tolerances and whether there are any suspected errors in the data.
[0084] In some embodiments, there are two modalities: processing the raw data that the system intends to manipulate to determine data quality (as described above), and determining one or more KPI formulations to be applied to the data. As used herein, KPI formulation includes one or more KPI characteristics, which also include, without limitation, formulation details, such as, without limitation, one or more data issues, and formulation algorithms refer to the algorithm itself and any parameters and definitions of each KPI. In some embodiments, both modalities are performed through the data inspection module, i.e., data quality is assessed through the data inspection submodule, and KPI-characterized formulation evaluation is performed through the KPI characteristic determination submodule, which is operably coupled to the data inspection submodule. In some embodiments, the KPI characteristic determination submodule is a separate module operably coupled to the data inspection module. Thus, the data inspection characteristics and the determination of relevant KPI formulation characteristics are tightly integrated.
[0085] In at least some embodiments, at least a portion of such KPI development characteristics are typically implemented as algorithms that manipulate the incoming data streams to provide the user with the output data and functionality necessary to support each KPI. In some embodiments, KPI development is easily located within a KPI development submodule embedded within a KPI characteristic determination submodule. Thus, as previously described, data is first checked to verify that it is within certain tolerances, and then a determination is made as to whether there is any correlation between one or more specific KPIs and potentially erroneous data. In one or more embodiments, at least a portion of the collected raw data is not associated with any KPIs, and therefore, such erroneous raw data does not have a strong impact on a given KPI. Therefore, a simple KPI relevance test is performed to perform an initial identification of related issues. For example, and without limitation, if one or more specific KPIs use average-based development and potentially erroneous data in the respective data streams includes unordered timestamps, it may be determined that the unordered timestamps have no impact on the respective one or more KPIs. Similarly, if one or more particular KPIs are median or mode-based formulations, the presence of outliers in the respective data streams will have no impact on the respective KPIs. Thus, some erroneous data characteristics may have no impact on a particular KPI, and such data is not relevant to the KPI-related analyses described further herein.
[0086] In some embodiments, one further mechanism for determining KPI relevance that may be employed is to pass at least a portion of the original time-series data stream having known error-free data, and in some embodiments, suspected error data, to one or more respective KPI development submodules in the KPI development submodule to generate numerical values therefrom, i.e., original KPI test values. Specifically, data without erroneous values may be manipulated to change at least one value to a known erroneous value, thereby generating imputed error-bearing data that is also passed to each of the one or more KPIs to generate imputed KPI test values. In some embodiments, the injected errors may include, without limitation, random selection of some of the data in the original data stream and removal of such random data to determine whether missing data issues are relevant, and random selection of known error-free data and injected values known to extend beyond established acceptable ranges to determine whether outlier issues are relevant. The imputed KPI test values are compared to the original KPI test values, and if there is a sufficient similarity between them, the original data, i.e., the issues associated with the original data, are classified as relevant to the respective KPI. If there is not a sufficient similarity between the imputed KPI test values and the original KPI test values, the original data including said issues is classified as unrelated to the respective KPI. Thus, to determine whether there is any relevant relationship between suspected or otherwise identified data errors in the original data stream, data having predetermined errors embedded therein is used to determine whether there is any relevance and significant impact of the erroneous data on the development of the respective KPI.
[0087] In at least some embodiments, KPI characterization is performed. The basis for each KPI characterization, sometimes referred to as a KPI characterization, includes one or more KPIs, e.g., one or more business-specific KPIs for a business and one or more residence-specific KPIs for a private residence. In some embodiments, the KPIs are predetermined and described as explicit measures of success, or lack thereof, aimed at achieving specific business goals, for example. In some embodiments, the KPIs are developed in response to the collection and analysis of operational data to determine otherwise unidentified measures for achieving the business goals, thereby facilitating the identification of one or more additional KPIs. Thus, regardless of origin, KPIs can be used to match relevant unique characteristics within the KPIs to each problem discovered in the original data, in some instances facilitating the identification of the relevant problem.
[0088] In one or more embodiments, a KPI characterization operation is performed on-the-fly as the original data is transmitted to the data inspection module. Furthermore, because the nature of the issues in the original data generated in real time is not known in advance, data inspection and KPI characterization are performed dynamically in real time. Thus, the determination of each KPI using the respective characteristics embedded within each KPI's respective formulation is performed in conjunction with the determination of issues affecting the incoming original data. At least a portion of the KPI characterization includes determining the nature of each KPI associated with the original data. In some embodiments, some of the incoming original data is not associated with any KPIs, and this data is not further manipulated in the context of the present disclosure; any embedded issues are ignored, and the data is either processed as is or a notification of the issue is transmitted to the user in one or more ways. In other embodiments, a relationship between the incoming original data and the associated KPI formulation is further determined.
[0089] In embodiments, KPI formulations are grouped into one of two types of formulations: “observable box” and “unobservable box” formulations. Observable box KPI formulations are available for inspection, i.e., the details are observed, and the KPI characterization sub-module includes an observable box sub-module. Unobservable box KPI formulations are opaque depending on the operations and algorithms contained therein; for example, without limitation, each unobservable box algorithm and operation may be proprietary in nature, and each user may require some level of confidentiality and secrecy over their contents. The KPI characterization sub-module includes an unobservable box sub-module. In some embodiments, for both observable box and unobservable box KPI formulations, the associated algorithms investigate whether a suitable KPI characterization includes one or more analyses of surrounding original data values for the data using questions including, without limitation, one or more of maximum value determination, minimum value determination, average value determination, median determination, and other statistical determinations, for example, without limitation, standard deviation analysis. As previously mentioned, if there is no relationship between the respective original data and the KPI development characteristics, no further action is taken on the data with issues per this disclosure. Thus, for those issues associated with original data that have a relationship with a KPI (both supplied by the user), the nature of the KPI development, i.e., characteristics, of an observable box or an unobservable box, is determined so that the associated data quality issues that may adversely affect the associated KPI can be properly classified and subsequent optimization can be performed.
[0090] In one or more embodiments, a snapshot generation module receives output from the data inspection module, including erroneous data with known embedded issues and respective KPI development characteristics. The snapshot generation module is configured to generate simulated data snapshots by simulating each data value through one or more models deployed in the product that facilitates simulation of the original data. In some embodiments, method-based simulation and point-based simulation are used. Either of the simulations may be used regardless of the nature of the issues in the erroneous data, including both simultaneously. While in some embodiments, the selection of the two simulations is based on the nature of the quality issues in the original data, and in some embodiments, the selection may be based on predetermined instructions generated by a user. However, in general, method-based simulations are well-suited to handle missing value issues, and point-based simulations are well-suited to handle outlier issues.
[0091] For example, in some embodiments, previous attempts by a user may have indicated that data is missing for a continuous extended duration, or that missing data can be determined based on whether a syntactic value issue exists (i.e., data presumed to be numeric includes extensive durations of data that are determined to be NaN or improper numeric rounding or truncation). Therefore, method-based simulation may provide a better analysis for the aforementioned conditions. If there are semantic value issues (i.e., if some of the data contains anomalous events or consistent or patterned noisy data), an outlier issue may be determined. Therefore, point-based simulation may provide a better analysis for the aforementioned conditions. Similarly, if a user determines that it may be uncertain whether method-based simulation or point-based simulation will provide a better simulation for a specified condition, both simulation methods may be used, as described above, for those conditions where either will provide a better simulation.
[0092] The snapshot generation module is configured to analyze one or more repair methods using method-based simulation, where each repair method may include, for example, without limitation, an algorithm for determining a mean, median, etc. Furthermore, the method-based simulation submodule may be used regardless of whether the KPI development characteristic is an unobservable box or an observable box. Each repair method includes generating one or more imputed values that are included in each simulation snapshot as potential solutions or replacements for the erroneous value if that particular repair method is used. In particular, the imputed values may or may not be potential replacement values. Because there is no predetermined notion of which repair method will result in the best or most correct replacement value for a particular existing condition, multiple models are used, each model being used to implement a respective repair method. In some embodiments, the method-based simulation submodule is communicatively coupled to the KPI development submodule. Additionally, the portion of the error-free data used to calculate the imputed value for the erroneous data depends on the particular repair technique. For example, if it is determined that a missing value should be replaced with the average of all values, then a substantially complete set of the respective data is used in the repair module. Alternatively, if only three surrounding values are used to calculate the missing value, then only these surrounding values are used by the repair module. Thus, method-based simulation is used to generate one or more simulated snapshots of the error-free original data and an imputed value for each erroneous original data value, each simulated value representing what the data value would look like when a particular repair method is used, thereby generating multiple imputed values, each resulting from a different repair method.
[0093] In at least some embodiments, data collection involves the use of heuristic-based features that facilitate the determination of patterns within data points as they are collected. As used herein, the terms “data point” and “data element” are used interchangeably. Under some conditions, one or more instances of original data may appear to be erroneous due to each data point exceeding a threshold based on the probability that the data point's value should conform to an established data pattern. For example, without limitation, apparent data excursions, i.e., data spikes or drops, may be generated either through erroneous data packets or in response to accurate depictions occurring in real time. Accordingly, the snapshot generation module is further configured to analyze errors using point-based simulations to determine whether apparent erroneous data is actually erroneous data.
[0094] In one or more embodiments, data includes known error-free original data and suspected potentially erroneous data points, and is combined in various configurations to begin determining the probability that a potentially erroneous data value is correct or erroneous. Each potentially erroneous data value is individually estimated as either a discrete "correct" or a discrete "incorrect," and the potentially erroneous data values are referred to as "estimated data points" to distinguish them from known error-free original data. As such, estimated data points have their original data value as transmitted and an estimated label as either correct or erroneous. The remaining analysis focuses narrowly on the estimated data points. Specifically, all possible combinations of estimated data points collected in the aforementioned simulation snapshots are evaluated. The generation of all possible combinations of discrete "correct" and discrete "incorrect" labels, and subsequent aggregation of such, facilitates further determination of whether the "best" action is to correct the erroneous data or accept the accurate data. These operations take into account the allowable error associated with the original data, which may or may not be erroneous, through determining one or more probabilities of the suspected potentially erroneous data value being "correct" or "incorrect." For example, and without limitation, for an example where there is an erroneous data point, 2 3 Eight combinations are generated by point-based simulation. For each combination, some of the erroneous values are expected to be incorrectly identified and some to be correctly identified as erroneous. Therefore, for each combination, the erroneous values are replaced with imputed values based on a predetermined correction method. Thus, each combination has a different set of correct and incorrect data points and requires different imputed values based on the predetermined correction technique.
[0095] The total number of possible combinations of discrete "correct" and "incorrect" estimated data points grows exponentially with the number of estimated data points (i.e., 2 x (where x = the number of estimated data values), generating all possible combinations and processing them can be time- and resource-intensive. Each combination of estimated data points is a potential simulation, and processing each combination as a potential simulation simply increases processing overhead. Thus, while the described possible combinations of estimated data points are further considered, the possible combinations of estimated data points are "pruned" so that only a subset of all possible combinations are further considered. Thus, the point-based simulation submodule is operatively coupled to the snapshot optimization submodule. In such an embodiment, snapshot optimization features are employed through the use of KPI development characteristics, as described above, that are determined regardless of whether the KPI development characteristics are unobservable or observable boxes. For example, without limitation, KPI development characteristics for maximum, minimum, average, and median analysis can be used to filter the simulations of estimated data points. Thus, the snapshot optimization module is communicatively coupled to the KPI development submodule. Generally, only those combinations of imputed values and estimated data points that successfully pass through the pruning process survive to generate respective simulations of suspected point values through the model and generate respective simulation snapshots using the clean original data and the imputed values for the identified erroneous data; some of the suspected erroneous point values may in fact be non-erroneous and do not require their replacement.
[0096] In at least some embodiments, simulation snapshots, whether method-based or point-based, are created by a snapshot generation module and transmitted to a KPI value estimation module. As described above, each simulation snapshot includes error-free original data and imputed values for the errored data. Each imputed value and associated original data is submitted to a respective KPI development to generate a predicted replacement value, i.e., an estimated snapshot value, for each imputed value in each simulation snapshot. Each estimated snapshot value is based at least in part on the respective KPI development in the context of the error-free original data on the time-series data stream. Thus, for each simulation snapshot transmitted to the KPI value estimation module, one or more predicted replacement values, i.e., estimated snapshot values, are generated.
[0097] In some embodiments, the estimated snapshot values are transmitted to a confidence measurement module, which generates an analytical score in the form of a confidence value (described further below) for each estimated snapshot value. For each respective scored estimated snapshot value for erroneous data, the highest confidence value is selected, and the respective estimated snapshot value is promoted to a KPI value selected to replace the erroneous data, where the selected KPI value is referred to as the estimated KPI value. Thus, the estimated KPI value is a value (i.e., an estimated snapshot value) selected from one or more predicted replacement values to resolve a potentially erroneous data instance.
[0098] In one or more embodiments, a confidence measurement module further receives the respective information to facilitate the selection of estimated KPI values and additional information to generate an explanation for the selected estimated KPI values. Typically, the confidence measurement module compares the estimated snapshot values generated through one or more of the aforementioned simulations with the respective original erroneous data. At least one result of the comparison is a respective confidence value in the form of a numerical value for each estimated snapshot value. The respective confidence values, as applied to the respective snapshots of data, indicate a predicted level of confidence that the respective estimated snapshot value is correct. A relatively low confidence value indicates that the respective estimated snapshot value, including the estimated KPI value, should either not be used or should be used with caution. A relatively high confidence value indicates that the respective estimated snapshot value, including the estimated KPI value, should be used. A threshold for the associated confidence value may be established by a user and may be used to train one or more models, both of which facilitate fully automated selection. Further, subsequent actions may be automated. For example, without limitation, for confidence values below a predetermined threshold, the respective estimated snapshot value is not passed on for further processing within the native application using the original data stream. Similarly, for confidence values above a predetermined threshold, the respective selected estimated KPI value is passed on for further processing within the native application using the original data. Thus, the systems and methods described herein automatically correct problems using erroneous data in the original data stream in a manner that prevents inadvertent actions or initiates appropriate actions as dictated by the conditions and accurate data.
[0099] Additionally, because the confidence value for the estimated KPI value may not be 100%, the confidence measurement module includes an explanation sub-module that provides explanatory evidence for the resolution of one or more potentially erroneous data instances through details about the selection of specific simulated snapshots and estimated KPI values. The explanation sub-module provides such details, including, without limitation, the types of issues detected in the dataset, the number and nature of simulations generated, statistical characteristics of the scores obtained from various simulations, and comparisons of the scores. Thus, the confidence measurement module generates information to help a user understand the nature of the variance in values to further provide clarity about the selection of various simulated snapshots and their respective estimated KPI values from the KPI value estimation module.
[0100] In some embodiments, the confidence measurement module also includes several additional sub-modules that facilitate generating the aforementioned confidence values and details and evidence to support such values. In some of these embodiments, there are three confidence measurement sub-modules: Size a spread-based reliability measurement sub-module; and a spread-based reliability measurement sub-module. Size and a spread-based confidence measurement sub-module.
[0101] SizeThe spread-based confidence measurement submodule is configured to take into account the magnitude of the value obtained from the KPI value estimation module and generate relevant confidence measurement information, for example, whether the magnitude of the KPI value is 50 or 1050, and the confidence in the resulting KPI value may differ in view of additional data and conditions. The spread-based confidence measurement submodule considers a range within which the simulated values lie and generates relevant confidence measurement information, i.e., instead of the absolute magnitude of the KPI value, the spread-based confidence measurement uses statistical properties such as the mean, minimum, maximum and standard deviation of the KPI value, and is therefore substantially unaffected by the magnitude. Size and a spread-based confidence measurement sub-module that considers the magnitude and range of values and generates associated confidence measurement information.
[0102] Referring to FIG. 4, a schematic diagram illustrating a system, namely, a time-series data substitution reliability system 400, configured to calculate reliability values for corrected data in time-series data is provided. The time-series data substitution reliability system 400 is referred to herein as "system 400" with respect to anything other than the identified time-series data substitution reliability system 400. The system 400 includes one or more processing devices 404 (only one shown) communicatively and operably coupled to one or more memory devices 406 (only one shown). The system 400 also includes a data storage system 408 communicatively coupled to the processing devices 404 and memory devices 406 through a communication bus 402. In one or more embodiments, the communication bus 402, the processing devices 404, the memory devices 406, and the data storage system 408 are similar to those shown in FIG. 3, namely, the communication bus 102, the processing devices 104, the system memory 106, and the persistent storage device 108, respectively.
[0103] In one or more embodiments, the system 400 includes a process control system 410 configured to operate any process that enables operation of the system 400, including, without limitation, electrical processes (e.g., energy management systems), mechanical processes (machine management systems), electromechanical processes (industrial manufacturing systems), and financial processes, as described herein. In some embodiments, the process control system 410 is an external system that is communicatively coupled to the system 400. As shown and described herein, the processing device 404, memory device 406, and data storage system 408, in some embodiments, are communicatively coupled to the process control system 410 through the input / output unit 112 (shown in FIG. 3 ).
[0104] The process control system 410 includes one or more processing devices 412 in conjunction with a respective one or more processes, which execute device / process control commands 414 generated through the interaction of associated program instructions through the processing device 404 and memory device 406. The process control system 410 also includes a sensor group 416 including sensors used to monitor the processing devices 412 and their respective processes and generate feedback 418 to the processing device 412 (e.g., without limitation, "sensor normal" and "sensor malfunction" signals), and one or more original time-series data streams 420 including data packets, hereinafter referred to as original data 422, representative of the processed measurement output of the sensor group 416.
[0105] The memory device 406 includes a process control algorithm and logic engine 430 configured to receive the original time-series data stream 420 and generate device / process control commands 414. In some embodiments, the memory device 406 also includes a data quality-KPI prediction confidence engine 440. In one or more embodiments, the data quality-KPI prediction confidence engine 440 includes one or more models 442 embedded therein. The system 400 also includes one or more output devices 450 communicatively coupled to the communication bus 402 to receive the output 444 of the data quality-KPI prediction confidence engine 440. The modules and sub-modules of the data quality-KPI prediction confidence engine 440 are described in relation to FIG. 5 .
[0106] Data storage system 408 stores data quality-KPI predicted confidence data 460, including, without limitation, original time series data 462 (supplemented through original time series data stream 420) and confidence values and explanations 464. Data storage system 408 also stores business KPIs 466, including formulations 468, properties and characteristics 470 (used interchangeably herein), and respective measurements 472, where formulations 468 include characteristics 470 and measurements 472.
[0107] 5A, a flow chart is provided illustrating a process 500 for calculating confidence values for corrected data in time series data. Also referring to FIG. 4, at least some of the modules and sub-modules of the data quality-KPI prediction confidence engine 440 are also shown and described in connection with FIG. 5A.
[0108] In one or more embodiments, the quality of the original data 504 (substantially similar to the original data 422) embedded within each original time-series data stream 420 is analyzed, and a determination is made as a function of each one or more KPIs associated with each original data 504 through a two-stage process. First, the quality of the original data 504 as a data packet is transmitted from each sensor of the sensor group 502 (substantially similar to the sensor group 416) to a data inspection module 510 (residing within the data quality-KPI prediction confidence engine 440), and the data packet is inspected by a data inspection sub-module 512 embedded within the data inspection module 510. In at least some embodiments, as described further, the data inspection module 510 also includes an integrated KPI characterization feature, thereby avoiding redundancy in the data inspection sub-module 512.
[0109] In some embodiments, one or more data packets of the original data 504 may include issues that identify the respective data packets as containing potentially defective data. One such issue may be associated with sampling frequency. For example, and without limitation, the data inspection sub-module 512 may check the sampling frequency of the sensor constellation 502 to determine whether multiple sampling frequencies are present in the original data 504, such as whether there are occasional perturbations in the sampling frequency and whether there are consecutive changes in the sampling frequency. Also, for example, and without limitation, the data inspection sub-module 512 may check the timestamps of the original data 504 to determine whether there are missing timestamps in the original data 504, whether the original data 504 is missing for consecutive extended durations, and whether there are timestamps with altered formats. Also, for example, and without limitation, the data inspection sub-module 512 may check for syntactic value issues to determine whether data that is supposed to be numeric contains "Not a Number (NaN)" values and extended durations of improper numeric rounding and truncation in the data 504. Further, for example and without limitation, the data inspection sub-module 512 checks for semantic value issues and determines whether any of the original data 504 contains anomalous events and noisy data. Accordingly, the data inspection sub-module 512 examines the original data 504 in each original time series data stream 420 to determine whether the original data 504 is within a predetermined tolerance range and whether there are any suspected errors in the original data 504 and the nature of the errors.
[0110] In some embodiments, there are two modalities: processing of raw data 504 that the system 400 intends to manipulate to determine data quality (as described above), and determination of one or more KPI formulations 468 to be applied to the raw data 504. As used herein, KPI formulation 468 includes one or more KPI characteristics 470, which also include, without limitation, details of the formulation 468, such as, without limitation, one or more data issues, and formulation 468 algorithms cover the algorithms themselves and any parameters and definitions of each KPI 466. KPIs 466, including formulations 468, characteristics 470, and measures 472, are stored in the data storage system 408. In some embodiments, both modalities are performed through the data inspection module 510: data quality is assessed through the data inspection sub-module 512, and KPI formulation characterization is performed through the KPI characterization determination sub-module 514, which is operatively coupled to the data inspection sub-module 512. In some embodiments, the KPI characteristic determination sub-module 514 is a separate module operatively coupled to the data inspection module 510. Thus, the determination of the data inspection characteristics and the relevant KPI development characteristics 470 are tightly integrated.
[0111] In at least some embodiments, at least a portion of such KPI development characteristics 470 are typically implemented as algorithms that manipulate the incoming raw time-series data stream 420 to provide the user with the output data and functionality necessary to support each KPI 466. Also, in some embodiments, the KPI development 468 is easily located within the KPI development sub-module 516 embedded within the KPI characteristic determination sub-module 514, and such KPI developments 468 can be imported from the data storage system 408. Thus, as previously described, the raw data 504 is first checked to verify that it is within certain tolerances, and then a determination is made as to whether there is any relevance between one or more particular KPIs 466 and any potentially erroneous data. In one or more embodiments, at least a portion of the collected raw data 504 is not associated with any KPIs 466, and therefore, such erroneous raw data 504 will not have a strong impact on a given KPI 466. Therefore, a simple KPI relevance test is performed to perform an initial identification of related issues. For example, and without limitation, if one or more particular KPIs 466 use a mean-based formulation and the potentially erroneous data 504 in the respective original time series data streams 420 includes unordered timestamps, the unordered timestamps may be determined to have no impact on the respective one or more KPIs 466. Similarly, if one or more particular KPIs 466 use a median- or mode-based formulation 468, the presence of outliers in the respective original time series data streams 420 may have no impact on the respective KPIs 466. Thus, some erroneous data characteristics may have no impact on a particular KPI 466, and such data is not relevant to the KPI-related analyses described further herein.
[0112] In some embodiments, one further mechanism for determining KPI relevance that may be employed is to pass at least a portion of the original time-series data stream 420 having known error-free data and, in some embodiments, suspected error data, to one or more respective KPI developments 468 in the KPI development sub-module 516 to generate numerical values therefrom, i.e., original KPI test values therefrom. Specifically, data without erroneous values may be manipulated to change at least one value to a known erroneous value, thereby generating imputed error-bearing data that is also passed to the respective one or more KPI developments 468 to generate imputed KPI test values therefrom.
[0113] Referring to FIG. 6, a text diagram illustrating an exemplary algorithm 600 for identifying relevant issues is presented. Also referring to FIGS. 4 and 5A, the algorithm 600 resides within the KPI development sub-module 516. The algorithm 600 includes an issue list operation 602, in which a predetermined set of potential data error issues is listed for selection within the algorithm 600, with each potential data error issue including one or more corresponding models 442. A data identification operation 604 is performed to identify which portions of the original data 504 in the original time series data stream 420 will be analyzed for potential errors, potential data substitutions, and a confidence determination of the substitutions. In some embodiments, the data quality-KPI prediction confidence engine 440 is scalable to simultaneously examine multiple streams of the original time series data stream 420, including, without limitation, a small portion of the original data 504, and scaling up to all of the original data 504 in all of the original time series data stream 420. KPI formulations 468, as developed by the user, are identified and retrieved in Identify and retrieve KPI formulations operation 606, and the selected raw data 504 to be analyzed is passed through each KPI formulation 468 in Raw Data-KPI Development operation 608. The impacting issues from Issue List operation 602 are cycled through Impacting Issues Analysis Selection Algorithm 610, either one at a time or simultaneously in parallel.
[0114] In one or more embodiments, the data-issues sub-algorithm 612 is executed, which includes injecting at least a portion of the original data 504 with imputed erroneous data via an imputed data injection operation 614. In some embodiments, the injected error may include, without limitation, a random selection of some of the original data 504 in the original time-series data stream 420 and removal of such random data to determine whether a missing data issue is relevant. Additionally, the injected error may include, without limitation, a random selection of known error-free original data 504 and injected values known to extend beyond established tolerances to determine whether an outlier issue is relevant. The imputed data is transmitted through KPI development 468 to determine imputed KPI test values via a KPI test value generation operation 616. The imputed KPI test values from operation 616 are compared to the original KPI test values from operation 608 via a KPI value comparison operation 618, and an issue determination operation 620 is executed in response to the comparison operation 618. In some embodiments, if there is sufficient similarity between the imputed KPI test values and the original KPI test values, the original data 504, i.e., the issues associated with the original data 504, are classified as relevant to the respective KPIs 466 through their KPI formulations 468. If there is not sufficient similarity between the imputed KPI test values and the original KPI test values, the original data 504, including such issues, are classified as unrelated to the respective KPIs 466 through their KPI formulations 468. Upon executing the sub-algorithm 612 through the depletion of issues from the issue list operation 602, the sub-algorithm 612 ends 622 and the algorithm 600 ends 624.Therefore, to determine if there is any association between suspected or otherwise identified data errors in the original time series data stream 420, data having predetermined errors embedded therein is used to determine if there is any relevance and significant impact of the erroneous data on the respective KPI formulation 468.
[0115] 4 and 5A, in at least some embodiments, KPI characterization is performed. The basis for each KPI characterization may be referred to as a KPI characterization and includes one or more KPIs 466. For example, without limitation, the basis for a business may be one or more business-specific KPIs 466, and for a private residence, the basis may be one or more residence-specific KPIs 466. In some embodiments, any entity-based KPI that enables the time series data replacement reliability system 400 disclosed herein may be used.
[0116] In some embodiments, KPIs 466 are predetermined and described, for example, as explicit measurements of success, or lack thereof, aimed at achieving specific business objectives. In some embodiments, KPIs 466 are developed in response to the collection and analysis of operational data to determine otherwise unidentified measurements for achieving business objectives, thereby facilitating the identification of one or more additional KPIs 466. Thus, for each issue 504 discovered in the original data, regardless of origin, a KPI 466 is available to match relevant unique characteristics within each KPI formulation 468, in some instances to facilitate the identification of the relevant issue.
[0117] In one or more embodiments, the KPI characterization operation is performed when the raw data 504 is transmitted to the data inspection module 512. Additionally, because the nature of the issues in the raw data being generated in real time is not known in advance, the data inspection and KPI characterization are performed dynamically in real time. Thus, the characterization of each KPI 466 using the respective characteristics 470 embedded within each formulation 468 of each KPI 466 is performed in conjunction with the determination of issues affecting the arriving raw data 422. At least a portion of the KPI characterization includes determining the nature of each KPI 466 associated with the raw data 504.
[0118] In some embodiments, some of the incoming raw data 504 has no association with KPIs 466, and this data is not further manipulated in terms of the present disclosure, and either any embedded issues are ignored and the data is processed as is, or a notification of the issue is transmitted to a user in one or more ways, for example, without limitation, through one or more of output devices 450. In other embodiments, a relationship between the incoming raw data and the associated KPI development 468 is further determined.
[0119] In embodiments, KPI formulations 468 are grouped into one of two types of formulations: "observable box" and "unobservable box" formulations. In some embodiments of both observable box and unobservable box KPI formulations, the associated algorithms investigate whether the appropriate KPI characteristic formulation includes one or more analyses of surrounding original data values for the problematic data, including, without limitation, one or more of a maximum determination, a minimum determination, an average determination, a median determination, and other statistical determinations, such as, without limitation, a mode determination, and a standard deviation analysis.
[0120] In at least some embodiments, the observable box KPI formulations 468 are available for inspection, i.e., the details are observed, and the KPI characteristic determination sub-module 514 includes an observable box sub-module 518. Referring to FIG. 7, a text diagram illustrating an example algorithm 700 for observable box KPI analysis is provided. Also referring to FIGS. 4 and 5A, the algorithm 700 resides within the observable box sub-module 518. The algorithm 700 includes a present KPI formulation operation 702, where the characteristics of each KPI formulation 468 are clearly presented to the user and the system 400 as described herein. The algorithm also includes a parse tree operation 704, where the KPI characteristics 470 are converted into an abstract syntax tree (AST) to generate the KPI characteristics 470 as an AST representing the source code in the respective programming language, such that the details of the KPIs 466 can be parsed and understood as nodes in the AST when various code blocks are available. As shown in FIG. 7, algorithm 700 includes a first sub-algorithm, namely, a function analysis operation 706 configured to determine whether a particular node in the AST is a function, such as, without limitation, a mathematical operation as further described.
[0121] 7, a second sub-algorithm, i.e., a median determination operation 708, is performed on those KPI development characteristics 470 defining a median determination of the original data 504, such that a KPI characteristic assignment operation 710 is performed, in this case the assigned KPI characteristic 470 is the "median" for subsequent portions of the process 500. The median determination operation 708 then terminates 712. In some embodiments, the algorithm includes one or more additional portions of the first sub-algorithm for other types of KPI characteristics, such as, without limitation, maximum value determination, minimum value determination, average value determination, and other statistical determinations, such as, without limitation, mode value determination and standard deviation analysis. In the exemplary embodiment of FIG. 7, a third sub-algorithm, i.e., a mean value determination operation 714, is performed on those KPI development characteristics 470 defining a mean value determination of the original data 504, such that a KPI characteristic assignment operation 716 is performed, in this case the assigned KPI characteristic 470 is the "mean" for subsequent portions of the process 500. The mean value determination operation 714 then ends 718. Any remaining possible KPI developing characteristics 470 are similarly determined, as described above. Once the function analysis operation 706 is complete, it ends 720.
[0122] Further, in one or more embodiments, as shown in FIG. 7 , algorithm 700 includes a fourth sub-algorithm, i.e., a binary operation analysis operation 722 configured to determine whether a particular node in the AST is a binary operation, for example, without limitation, a mathematical operation using two elements or operands to create another element. In the embodiment shown in FIG. 7 , a fifth sub-algorithm, i.e., a division sub-algorithm 724, is performed on these KPI development characteristics 470, which define a division operation on the original data 504. The division operation includes a sixth sub-algorithm, i.e., a combined addition and length operand, or combined mean value algorithm 726, where the length operand or operation provides the number of items to be added so that a KPI characteristic assignment operation 728 is performed, where the assigned KPI characteristic 470 is the “average value” for subsequent portions of process 500. The combined mean value algorithm 726 ends 730, the division sub-algorithm 724 ends 732, and the binary operation sub-algorithm 722 ends 734. An open sub-algorithm 736 is further indicated if further operations beyond functions and binary operations are required by the user. The parse tree operation 704 ends 738 when all of the respective observable box operations associated with each KPI 466 have been identified.
[0123] In at least some embodiments, the non-observable box KPI formulation 468 is opaque depending on the operations and algorithms contained therein; for example, without limitation, each non-observable box algorithm and operation may be proprietary in nature, with each user requiring some level of confidentiality and secrecy over their contents. In some embodiments, such non-observable box formulations may take the form of an application programming interface (API). Accordingly, one mechanism for determining KPI formulation characteristics 470 in the non-observable box KPI formulation 468 includes repeated sampling of the original data 504 to test the original data 504 through simulation of the formulation. Accordingly, the KPI characteristic determination sub-module 514 includes a non-observable box sub-module 520.
[0124] Referring to FIG. 8, a text diagram illustrating an example algorithm 800 for unobservable box KPI analysis is provided. Also referring to FIGS. 4 and 5A, the algorithm 800 resides within the unobservable box sub-module 520. In at least some embodiments, the algorithm 800 includes a data subset generation operation 802 in which the original data 504 is divided into K subsets of data, each subset having M data points therein, where M is a predetermined constant. For example, without limitation, a string of 100 data points may be divided into five subsets of 20 points each. Generating such subsets facilitates determining whether a particular error is recurring or a single-instance error, i.e., a one-off error. The algorithm 800 also includes a KPI developing characteristics list operation 804 configured to identify all of the potential KPI developing characteristics 470 that may be used within the unobservable box calculation. As previously described herein, such KPI development characteristics 470 include, without limitation, one or more of an average value determination ("average"), a median value determination ("median"), a mode value determination ("mode"), a maximum value determination ("max"), a minimum value determination ("min"), and other statistical determinations, such as, without limitation, a standard deviation analysis. Each of these KPI development characteristics 470 is investigated through one or more unobservable box model-based simulations to identify potential issues with erroneous data, which unobservable box model-based simulations are not directly related to the simulation modeling described further herein with respect to snapshot generation.
[0125] In one or more embodiments, an original KPI evaluation operation 806 is performed, in which each data element of each data subset is processed through a respective unobservable box model, such model not yet determined. As used herein, the terms “data point” and “data element” are used interchangeably. Thus, in an embodiment of 100 data points or data elements of original data 504, there are 100 respective KPI values, i.e., 20 KPI values for each of the five subsets of original data 504. Thus, the 100 processed data elements may be processed through unobservable box designs, whatever they may be, to generate 100 original KPI values through actual unobservable box designs. Also, in some embodiments, a correlation operation 808 is performed, which includes a simulation / correlation sub-algorithm 810. Specifically, in one or more embodiments, a simulated KPI evaluation operation 812 is performed, in which each data element of original data 504 is analyzed through a respective model of each KPI development characteristic 470 identified in KPI development characteristic list operation 804. An original KPI value-simulated KPI value correlation operation 814 is performed, comparing each original KPI value to each respective simulated KPI value generated through each model of the KPI development characteristics 470 identified from the KPI development characteristics list operation 804. Thus, for an embodiment using 100 data elements, there will be 100 correlations for each KPI development characteristic 470 identified from the KPI development characteristics list operation 804. In some embodiments, a statistical evaluation of each set of correlated data elements is performed to determine the strength of the correlation, for example, without limitation, a weak correlation and a strong correlation, and the definition of each correlation can be established by a user. A strong correlation indicates that the simulated KPI development follows the actual unobservable box KPI development 468.A weak correlation indicates that the simulated KPI formulation is not aligned with the actual unobservable box KPI formulation 468. Once processing through the correlation is complete, the simulation / correlation sub-algorithm 810 ends 816. The algorithm for unobservable box KPI analysis 800 includes a KPI formulation characteristic selection operation 818, where the characteristic with the highest correlation coefficient is selected. Once the unobservable box KPI formulation is determined, the algorithm 800 ends 820.
[0126] The output 522 of the data inspection module 510 includes an analysis of the original data 504 to determine whether there are any data errors therein and, if any, whether any KPI-developing characteristics 470 are affected. If there are no errors, the respective data is no longer processed through the process 500, the operations in the KPI characteristic determination submodule 514 are not invoked, and there is no output 522. If there are data errors in the original data 504, the output 522 is transmitted to a determination operation 524, which determines 524 whether the data issues are relevant to the identified KPI based on the analysis of the KPI characteristic determination submodule 514. As previously mentioned, if there is no relationship between the respective original data 504 and the KPI-developing characteristics 470, a "no" determination is generated and no further action is taken on the problematic data per this disclosure. The user may choose to take other action on the data errors if desired. For a "Yes" determination, i.e., for issues of these data errors that are associated through their respective properties with the original data 504 that have a relationship to the KPIs (both supplied by the user), the output 526 of determination operation 524 is transmitted for further processing, output 526 being substantially similar to output 522. Thus, if the characteristics 470 of KPI formulation 468 for erroneous data, whether observable or unobservable boxes, are determined to adversely affect the associated KPI, the error is further analyzed so that it can be properly classified and subsequent optimization can be performed.
[0127] 5B, a continuation of the flowchart shown in FIG. 5A is provided, further illustrating a process 500 for calculating a confidence value for corrected data in time-series data. Also referring to FIG. 4, in one or more embodiments, the process 500 further includes sending an output 526 to a snapshot generation module 530. The snapshot generation module 530 receives the output 526 of the data inspection module 510, which includes the erroneous data having known embedded issues and an identification of each KPI-developing characteristic 470. The snapshot generation module 530 is configured to generate a snapshot of simulated data through simulation of each data value through one or more models deployed in production to facilitate simulation of the original data 504.
[0128] 9, a schematic diagram is provided illustrating a portion of a process 900 for snapshot simulation using a snapshot generation module 904 substantially similar to snapshot generation module 530. Also referring to FIG. 5B, original data 902, substantially similar to original data 504, transmitted via output 526 to snapshot generation module 904 is further evaluated. The original data 902 (with erroneous data issues embedded therein) is processed through multiple models 532 (substantially similar to model 442 shown in FIG. 4) to generate multiple simulation snapshots 906 including simulated data, as further described herein. The simulated data snapshots 906 are later used for KPI interface 908 and confidence measurement 910 shown in FIG. 9 for context only.
[0129] 4 and 5B, in some embodiments, method-based simulation and point-based simulation are used. While either of the simulations may be used regardless of the nature of the issues in the erroneous data, including both at the same time, in some embodiments, the selection of the two simulations is based on the nature of the quality issues in the original data, and in some embodiments, the selection may be based on predetermined instructions generated by the user. However, in general, method-based simulations are well-suited to handle missing value issues, and point-based simulations are well-suited to handle outlier issues. For example, in some embodiments, previous attempts by a user may have indicated that data is missing for a continuous extended duration, or that missing data can be determined based on whether a syntactic value issue exists (i.e., data presumed to be numeric includes extensive durations of data that are determined to be NaN or improper numeric rounding or truncation). Therefore, method-based simulation may provide a better analysis for the aforementioned conditions. If there are semantic value issues (i.e., if some of the data contains anomalous events or consistent or patterned noisy data), an outlier issue may be determined. Therefore, point-based simulation may provide a better analysis for the aforementioned conditions. Similarly, if a user determines that it may be uncertain whether method-based simulation or point-based simulation will provide a better simulation for a specified condition, both simulation methods may be used, as described above, for those conditions where either may provide a better simulation.
[0130] In one or more embodiments, snapshot generation module 530 is configured to analyze one or more repair methods using method-based simulation, where each repair method may include, for example, without limitation, an algorithm for determining, at least in part, a mean, median, etc., depending on a respective KPI 466 affected by a data error. However, the repair methods are not necessarily limited to KPI development characteristics 470. Snapshot generation module 530 includes method-based simulation submodule 534, which may be used regardless of whether KPI development characteristics 470 are unobservable boxes or observable boxes.
[0131] Referring to FIG. 10, a schematic diagram illustrating a process 1000 for generating a methodology-based simulation is presented. Also referring to FIGS. 4 and 5B, a methodology-based simulation is generated through the methodology-based simulation submodule 534. A portion of the output 526 of the data inspection module 510, including erroneous data with embedded issues and identification of respective KPI-developing characteristics 470, is shown as a snippet 1002 having ten instances of error-free data 1004 and three instances of erroneous data 1006. The data snippet 1002 is transmitted to a plurality of repair methods 1010, including repair methods M1, M2, M3, and M4, each associated with a different respective model 532, the number 4 being non-limiting. Each repair method M1-M4 involves generating one or more imputed values that, if that particular repair method is used, are included in a respective simulation snapshot as potential solutions or replacements for the erroneous values. Because there is no predetermined notion as to which repair method M1-M4 will result in the best or most correct potential replacement value for the particular currently erroneous data 1006, multiple models 532 are used, each model 532 being used to implement a respective repair method M1-M4. In some embodiments, the method-based simulation sub-module 534 is communicatively coupled to the KPI development sub-module 516 for provision for accessing the KPI development 468 present therein.
[0132] In at least some embodiments, multiple simulated data snapshots 1020 are generated. For example, in an exemplary embodiment, repair method M1 utilizes each model 532 to calculate imputed values 1024 in the simulated data snapshots 1022. In some embodiments, the portion of error-free data 1004 used to calculate the imputed values 1024 for erroneous data 1006 depends on the particular repair technique associated with each repair method M1. For example, if it is determined to replace missing values with the average of all values, then a substantially complete set of each error-free data 1004 is used in each repair method M1. Alternatively, if only three values surrounding the error-free data 1004 are used to calculate the missing value, i.e., the erroneous data 1006, then only these surrounding values of the error-free data 1004 are used by each repair method M1. Similarly, simulated data snapshots 1032, 1042, and 1052 are generated through respective repair methods M2-M4 and include respective imputed values 1034, 1044, and 1054. Because repair methods M1-M4 are different, respective imputed values 1024, 1034, 1044, and 1054 should also be different. Referring to Figures 4 and 5B, simulated data snapshots 1022, 1032, 1042, and 1052 are shown as outputs 536 from method-based simulation submodule 534, which, in some embodiments, are transmitted to a data simulation snapshot storage module 538 residing within data storage system 408.
[0133] In at least one embodiment, such as an example embodiment, the three instances of erroneous data 1006 are substantially identical. In at least one embodiment, each of the instances of erroneous data 1006 is different. Thus, multiple models 532 and repair methods M1-M4 are used for all of the erroneous data 1006, facilitating the generation of multiple respective imputed values 1024, 1034, 1044, and 1054 for each different error. Thus, method-based simulations in the form of repair methods M1-M4 are used to generate one or more simulation snapshots 1022, 1032, 1042, and 1052 of the error-free original data 1004 and imputed values 1024, 1034, 1044, and 1054 for each erroneous original data value 1006, each of which indicates what the data value would look like if a particular repair method M1-M4 were used, and each of which imputed values 1024, 1034, 1044, and 1054 is the product of a different repair method M1-M4.
[0134] In at least some embodiments, the collection of the original time-series data stream 420 via the sensors 416 includes the use of heuristic-based features that facilitate the determination of patterns within the data elements as they are collected. Under some conditions, one or more instances of the original data 422 may appear to be erroneous due to a respective data element exceeding a threshold based on the probability that the value of the data element should conform to the established data pattern. For example, without limitation, apparent data excursions, i.e., data spikes or drops, may be generated either through erroneous data packets or in response to accurate depictions occurring in real time. Accordingly, the snapshot generation module 530 is further configured to analyze errors using point-based simulation to determine whether apparent erroneous data is actually erroneous data; i.e., the snapshot generation module 530 includes a point-based simulation submodule 540.
[0135] Referring to Figure 11, a schematic diagram illustrating a process 1100 for point-based simulation is provided. Also referring to Figures 4 and 5B, point-based simulations are generated through the point-based simulation submodule 540. A portion of the output 526 of the data inspection module 510, including erroneous data with embedded issues and identification of the respective KPI development characteristics 470, is shown as data snippet 1102 having ten instances of non-erroneous data points 1104 and three instances of suspected, potentially erroneous data points 1106. The three instances of suspected, potentially erroneous data points 1106 are individually referred to as 1106A, 1106B, and 1106C, and collectively referred to as 1106. In one or more embodiments, a data snippet 1102 containing known non-erroneous original data, i.e., non-erroneous data points 1104, and suspected, potentially erroneous data points 1106, is combined into a matrix 1110 of structures to begin determining the probability that the value of the suspected, potentially erroneous data points 1106 is correct or erroneous. As shown, the matrix 1110 includes three suspected, potentially erroneous data points 1106, i.e., two 3 or based on eight possible combinations of three suspected, potentially erroneous data points 1106. Matrix 1110 is made up of three columns 1112, 1114, and 1116, one for each of the suspected, potentially erroneous data points 1106A, 1106B, and 1106C, respectively. The resulting eight rows, individually referred to as D1 through D8 and collectively referred to as 1120, contain the available combinations of three suspected, potentially erroneous data points 1106.
[0136] Each of the three suspected potentially erroneous data points 1106 is individually estimated as either a discrete "correct" or a discrete "incorrect," and the potentially erroneous data values are referred to as "estimated data points" to distinguish them from the known, error-free original data, i.e., the error-free data points 1104. As shown in FIG. 11 , the estimated erroneous data points are collectively referred to as 1130. These estimated erroneous data points 1130 associated with suspected potentially erroneous data point 1106A are individually shown and referred to as 1122, 1132, 1162, and 1182 in column 1112. Additionally, these estimated erroneous data points 1130 associated with suspected potentially erroneous data point 1106B are individually shown and referred to as 1124, 1144, 1164, and 1174 in column 1114. Additionally, these suspected erroneous data points 1130 associated with suspected potentially erroneous data point 1106C are shown individually and are designated 1126, 1146, 1176, and 1186 in column 1116.
[0137] 11 , the estimated correct data points are collectively referred to as 1140. Those estimated correct data points 1140 associated with suspected, potentially erroneous data point 1106A are individually shown and referred to as 1142, 1152, 1172, and 1192 in column 1112. Those estimated correct data points 1140 associated with suspected, potentially erroneous data point 1106B are individually shown and referred to as 1134, 1154, 1184, and 1194 in column 1114. Those estimated correct data points 1140 associated with suspected, potentially erroneous data point 1106C are individually shown and referred to as 1136, 1146, 1166, and 1196 in column 1116. A simulation snapshot 542 of matrix 1120 is performed.
[0138] Thus, the first row D1 represents all three suspected, potentially erroneous data points 1106 as presumed erroneous data points 1130. Similarly, the eighth row D8 represents all three suspected, potentially erroneous data points 1106 as presumed correct data points 1140. The second, third, and fourth rows D2, D3, and D4 represent only one of the three suspected, potentially erroneous data points 1106 as the presumed erroneous data point 1130 and two of the three suspected, potentially erroneous data points 1106 as the presumed correct data point 1140, respectively. The fifth, sixth and seventh rows D5, D6 and D7 respectively represent two of the three suspected potentially erroneous data points 1106 as estimated erroneous data points 1130 and only one of the three suspected potentially erroneous data points 1106 as estimated correct data point 1140.
[0139] As such, the estimated erroneous data points 1130 and estimated correct data points 1140 have their original data values as transmitted and discrete estimated labels as either correct or erroneous. The remaining analysis focuses narrowly on the estimated erroneous data points 1130 and estimated correct data points 1140. Specifically, all of the possible combinations of estimated data points 1130 and 1140, as shown as D1 through D8, are collected in the aforementioned simulation snapshot 542 for further evaluation. The generation of all possible combinations of discrete "correct" labels, i.e., estimated correct data points 1140, and discrete "error" labels, i.e., estimated erroneous data points 1130, and subsequent aggregation of such, facilitates further determination of the "best" action and whether the "best" action is to correct the erroneous data or accept the correct data. These operations consider the allowable error associated with the original data in the data snippet 1102, which may or may not be erroneous, through determining the probability of one or more of the suspected potentially erroneous data values being "correct" or "incorrect." In each of the combinations D1 through D8, it is assumed that some of the suspected potentially erroneous values 1106 will be incorrectly identified as erroneous and some will be correctly identified as erroneous. Therefore, for each of the combinations D1 through D8, the erroneous values are replaced with imputed values based on a predetermined correction methodology similar to, but not limited to, that described with respect to FIG. 10. Thus, each combination D1 through D8 has a different set of correct and incorrect data points and requires different imputed values based on the predetermined correction technique.
[0140] As mentioned above, point-based simulations are well-suited to address the issue of outliers, which will be used to further illustrate the exemplary embodiment in FIG. 11 . As mentioned above, patterns may be identified in the original data 504, including the data snippet 1102 and the probability that the value of each data element will conform to the established data pattern. Thus, a discrete "false" estimated data point 1130 has a probability of being misidentified as erroneous with a percentage guarantee assigned to it. The probability of each of the three suspected erroneous values 1106 is used to determine whether the value 1106 is erroneous or not. As the various eight combinations D1 through D8 are evaluated, the probability that each of D1 through D8 is true is determined, and those rows D1 through D8 with the highest probability of true are submitted for further analysis. The overall probability of D1 through D8 is 100%. For example, and without limitation, after considering a heuristic analysis of each point 1122, 1124, and 1126 in D1 and their associated combined probabilities, it may be determined that all three points in D1 that are erroneous have a relatively low probability, such as row D8 (where all three values are correct). These two rows D1 and D8 are not considered further. In particular, for those embodiments where row D8, which has no erroneous values, has the highest probability of being correct, no further analysis needs to be performed, and value 1106 is not corrected through downstream processing, as will be further described. Thus, the value combinations with a higher probability of being true are processed further.
[0141] In general, the total number of possible combinations of discrete "correct" and "incorrect" estimated data points 1130 and 1140 grows exponentially with the number of estimated data points (i.e., 2 x(where x=the number of estimated data value generations), generating all possible combinations and processing them can be time- and resource-intensive. Each combination of estimated data points is a potential simulation, and processing each combination as a potential simulation simply increases processing overhead. Thus, although the illustrated possible combinations D1 through D8 of estimated data points 1130 and 1140 are further considered, however, the possible combinations of estimated data points 1130 and 1140 are "pruned" so that only a subset of all possible combinations are further considered. As described above, initial pruning occurs when low-probability combinations of potentially erroneous values are excluded from further processing.
[0142] The point-based simulation submodule 540 is operatively coupled to the snapshot optimization submodule 544. In such an embodiment, snapshot optimization features are utilized through the use of KPI development characteristics 470, which are determined whether the KPI development characteristics 470 are unobservable boxes or observable boxes, as described above. For example, without limitation, KPI development characteristics 470 for maximum, minimum, average, and median analysis may be used to filter the simulation of estimated data points 1130 and 1140. Accordingly, the snapshot optimization module 544 is communicatively coupled to the KPI development submodule 516. Generally, only those combinations of estimated data points that successfully pass through the pruning process survive to generate respective simulations of suspect point values through the model and generate respective simulation snapshots using the error-free original data and values imputed for the identified erroneous data; some of the suspected erroneous point values may, in fact, be non-erroneous and do not require their replacement.
[0143] Referring to FIG. 12, a text diagram illustrating an exemplary algorithm 1200 for a snapshot optimizer configured for execution within the snapshot optimization submodule 544 (as shown in FIG. 5B) is provided. Referring to FIGS. 4, 5A, 5B, and 11, the algorithm 1200 includes an operation 1202 that determines KPI development characteristics 470 as previously determined by the KPI characteristic determination submodule 514 and as described with respect to FIGS. 6-8. The data, as represented in the exemplary embodiment as matrix 1120, i.e., the data embedded in the remaining rows that have been excluded due to low probability as described above, is further analyzed to generate a pruning effect as described herein through a data presentation operation 1204. As described above, the exemplary embedding includes analyzing outliers. In one or more embodiments, the first sub-algorithm, i.e., the "best" sub-algorithm 1206, is considered for execution. In the event that the previously determined KPI development characteristics 470 are the best characteristics, a modified data operation 1208 is executed through one or more of the models 532. The corrected data operation 1208 involves determining whether the suspected, potentially erroneous data 1106 is an outlier within an upward peak of the data snippet 1102 of the original data 504. If the data snippet 1102 does not exhibit an upward trend, thereby ruling out any opportunity for an upward peak, the algorithm 1200 proceeds to the next set of operations. If the data snippet 1102 exhibits an upward trend, the affected outlier is replaced with a value that provides a smoothing effect on the upward trend per the corrected data operation 1208, using the previously described probability values that provide some level of certainty that the suspected erroneous data was, in fact, erroneous. These data points are selected for simulation through one or more models 532. Once the data replacement identification or "correction" has been performed, the best sub-algorithm 1206 ends 1210.
[0144] The second sub-algorithm, i.e., the "worst" sub-algorithm 1212, is considered for execution. In the event that the previously determined KPI development characteristic 470 is the worst characteristic, a modified data operation 1214 is run through one or more of the models 532. The modified data operation 1214 includes determining whether the suspected potentially erroneous data 1106 is an outlier within a downward peak of a data snippet 1102 of the original data 504. If the data snippet 1102 does not exhibit a downward trend, thereby eliminating any opportunity for a downward peak, the algorithm 1200 proceeds to the next set of operations. If the data snippet 1102 exhibits a downward trend, the affected outlier is replaced with a value that provides a smoothing effect on the downward trend per modified data operation 1214, using the previously described probability value that provides some level of certainty that the suspected erroneous data was, in fact, erroneous. These data points are selected for simulation through one or more models 532. Once the data repair or "fix" has been performed, the lowest sub-algorithm 1212 ends 1216.
[0145] The third sub-algorithm, i.e., the "Mean Value" sub-algorithm 1218, is considered for execution. In the event that the previously determined KPI development characteristic 470 is an average characteristic, a modified data operation 1220 is run through one or more models 532. The modified data operation 1220 includes determining whether the suspected potentially erroneous data 1106 is an outlier through considering all of the issues, i.e., the affected suspected potentially erroneous data 1106, and all of the respective probabilities described above, and grouping them into one or more clusters of potentially erroneous data values based on the proximity of their respective values to one another. In some embodiments, there may be multiple clusters of potentially erroneous data values exhibiting average characteristics that are used as the basis for clustering. A cluster consideration operation 1222 is performed, and a representative point, for example, without limitation, a collection of the mean values from each cluster, is considered as the representative point for the simulation. Once the data selection for the simulation has been run through one or more models 532, the mean value sub-algorithm 1218 ends 1224.
[0146] The fourth sub-algorithm, i.e., "Median" sub-algorithm 1226, is considered for execution. In the event that the previously determined KPI development characteristic 470 is a median characteristic, a corrected data operation 1228 is executed through one or more of the models 532. The corrected data operation 1228 includes determining whether the suspected potentially erroneous data 1106 is an outlier through consideration of all of the issues, i.e., the affected suspected potentially erroneous data 1106 and all of the respective probabilities described above. If the suspected potentially erroneous data 1106 is indeed an outlier, no further action is taken on the data because the median-based KPI is not affected by the value perturbation, and the median sub-algorithm 1226 ends 1230. In some embodiments, the sub-algorithms 1206, 1212, 1218, and 1226 may be executed simultaneously in parallel. The output of snapshot optimization module 544, shown as optimized simulated data snapshot 546, in some embodiments is transmitted to data simulation snapshot storage module 538, which resides within data storage system 408. Thus, multiple simulation snapshots 536 and 546 are generated for further processing, and simulation snapshots 536 and 546 are generated in a manner that significantly reduces the number of other imputed values.
[0147] 4, 5B, 10, and 11, in at least some embodiments, simulation snapshots 536 and 546 are created by a snapshot generation module and transmitted, whether method-based or point-based, to a KPI value estimation module 550. As described above, each simulation snapshot of simulation snapshots 536 and 546 includes error-free original data (e.g., 1004 and 1104) and imputed values for the established errored data (e.g., 1006 and 1106). Each of the imputed values and associated original data is submitted to a respective KPI development 468 to generate predicted replacement values, i.e., estimated snapshot values, for each of the imputed values in each simulation snapshot 536 and 546. To that end, original data 504 is also transmitted to the KPI value estimation module 550.
[0148] Referring to FIG. 13, a graphical illustration illustrating at least a portion of a KPI value estimation process 1300 is presented. Also referring to FIGS. 4 and 5B, the estimated snapshot values for simulation snapshots 536 and 546 are based on the respective KPI formulations 468 and are in the context of the original, error-free data on the time-series data stream. Thus, for each simulation snapshot 536 and 546 transmitted to KPI value estimation module 550, a predicted replacement value, i.e., an estimated snapshot value, is generated. FIG. 13 illustrates a horizontal axis (Y-axis) 1302 and a vertical axis (X-axis) 1304. Y-axis 1302 is shown to span from 41.8 to 42.6, with the values being unitless. X-axis 1304 is shown as unitless and unitless. The nature of the values is not important; however, process 1300 illustrates some of the values determined for simulation snapshots 536 and 546 depending on the presented KPI formulation characteristics 470. The original KPI values 1306, i.e., the values generated by processing the suspected erroneous data through the respective KPI formulations 468, are presented as a reference, each having a value of 42.177. The simulated KPI maximum snapshot 1308 presents an estimated snapshot value of 42.548, the simulated KPI average snapshot 1310 presents an estimated snapshot value of 42.091, and the simulated KPI minimum snapshot 1312 presents an estimated snapshot value of 41.805. These estimated snapshot values will be used in the discussion of subsequent portions of the process 500.
[0149] Referring to FIG. 5C , a continuation of the flowchart shown in FIGS. 5A and 5B is provided, illustrating a process 500 for calculating confidence values for corrected data in time-series data. Also referring to FIG. 5B , the output of KPI value estimation module 550 includes estimated point-based snapshot values 552, estimated method-based snapshot values 554, and original data 504, which are transmitted to a confidence measurement module 570 communicatively coupled to KPI value estimation module 550. Generally, for each estimated snapshot value for erroneous data generated from simulation snapshots in KPI value estimation module 550 within confidence measurement module 570, the estimated snapshot value is scored individually. Each scoring method includes generating scored estimated point-based snapshot values 562, i.e., estimated point-based snapshot values 552, with respective confidence values. Furthermore, each scoring method generates scored estimated method-based snapshot values 564, i.e., estimated method-based snapshot values 554, with respective confidence values. The generation of the confidence value is further described below. The highest analysis score is selected and each estimated snapshot value is promoted to a KPI value selected to replace the erroneous data, where the selected KPI value is referred to as estimated KPI value 566. Thus, estimated KPI value 566 is a value selected from one or more predicted replacement values (i.e., scored estimated snapshot values 562 and 564) to resolve the potentially erroneous data instance.
[0150] In some embodiments, the confidence measurement module 570 includes several additional sub-modules that facilitate the generation of confidence values and details and evidence to support such values. In some of these embodiments, there are three confidence measurement sub-modules: Size a spread-based confidence measurement submodule 572, a spread-based confidence measurement submodule 574, and Sizeand spread-based confidence measurement sub-module 576 is used.
[0151] Size The spread-based confidence measurement submodule 572 is configured to take into account the magnitude of the values obtained from the KPI value estimation module 550 and generate associated confidence measurement information including a respective confidence score. For example, without limitation, the confidence in the resulting KPI value, whether the magnitude of the KPI value is 50 or 1050, may differ in light of additional data and conditions. The spread-based confidence measurement submodule 574 considers the range of simulated values and generates associated confidence measurement information including a respective confidence score. Instead of the absolute magnitude of the KPI value, the spread-based confidence measurement uses statistical characteristics such as the mean, minimum, maximum, and standard deviation of the KPI value and is therefore substantially unaffected by the magnitude. Size and spread-based confidence measurement sub-module 576 considers a range of magnitude-order values to generate associated confidence measurement information, including a respective confidence score. In some embodiments, all three of sub-modules 572, 574, and 576 are used in parallel, with the results of each considered and evaluated for selection. In some embodiments, only one or two of sub-modules 572, 574, and 576 are selected based on the nature of the arrived estimated KPI values 566 and other data 568 (described further below).
[0152] Referring to FIG. 14 , a graphic / text diagram is provided illustrating the generation of a numerical confidence measure 1400. Also referring to FIGS. 5B and 5C , confidence values are generated for the estimated point-based snapshot value 552 and the estimated method-based snapshot value 554. A linear graphical representation 1410 is presented along with the four values shown in FIG. 13 . Specifically, other data 568 (shown in FIG. 5C ) are shown, such as, without limitation, a KPI minimum snapshot value 1412 having an estimated snapshot value of 41.805, a simulated KPI average snapshot value 1414 having an estimated snapshot value of 42.091, an original KPI value 1416 of 42.117, and a KPI maximum snapshot value 1418 having an estimated snapshot value of 42.548. Also presented in FIG. 14 is a first set of confidence measure evaluation algorithms, namely, a maximum deviation confidence measure algorithm 1430. The confidence measure 1A algorithm determines the relationship between the maximum variance of the estimated snapshot values 1412, 1414, and 1418 depending on the original KPI value 1416. The confidence measure 1B algorithm determines the relationship between the maximum variance of the estimated snapshot values 1412, 1414, and 1418 depending on the simulated KPI average snapshot value 1414. Also presented in FIG. 14 is a second set of confidence measure evaluation algorithms, namely, mean deviation confidence measure algorithms 1440. The confidence measure 2A algorithm determines the relationship between the variance between the original KPI value 1416 and the simulated KPI average snapshot value 1414 depending on the original KPI value 1416. The confidence measure 2B algorithm determines the relationship between the variance between the original KPI value 1416 and the simulated KPI average snapshot value 1414 depending on the simulated KPI average snapshot value 1414.14 also presents a spread-based measurement algorithm 1450, i.e., an algorithm for confidence measure 3, which evaluates the deviation 1452 between the original KPI value 1416 and the simulated KPI average snapshot value 1414 depending on the spread 1454 between the simulated KPI maximum value 1418 and the simulated KPI minimum value 1412. A maximum deviation confidence measure algorithm 1430 for confidence measures 1A and 1B and an average deviation confidence measure algorithm 1440 for confidence measures 2A and 2B are: Size Based on the reliability measurement submodule 572 and Size and spread-based confidence measurement sub-module 576. Similarly, the confidence measurement 3 algorithm of spread-based measurement algorithm 1450 resides in spread-based confidence measurement sub-module 574 and Size and in the spread-based confidence measurement sub-module 576.
[0153] 15, a graphical diagram, i.e., column chart 1500, is provided illustrating confidence measures using values calculated from the algorithm and the values provided in FIG. 14 with comparisons therebetween. Column chart 1500 includes a vertical axis (Y-axis 1502) representing the values of the calculated confidence values ranging between 0% and 100%. Column chart 1500 also includes a horizontal axis (X-axis) 1504 identifying confidence measures 1A, 1B, 2A, 2B, and 3. The confidence value of confidence measures 2A and 2B provides the highest value of 99.8. Thus, simulated KPI average snapshot value 1414 provides the highest confidence value for the erroneous data. In at least some embodiments, simulated KPI average snapshot value 1414 is the estimated KPI value 566 for this example.
[0154] Generally, the confidence measurement module 570 compares the estimated snapshot values 552 and 554 generated through one or more of the aforementioned simulations with the respective original erroneous data. At least one result of the comparison is a confidence value in the form of a numerical value for each of the estimated snapshot values 552 and 554 to be applied to the respective snapshot of data indicating a level of confidence that the estimated snapshot values 552 and 554 are appropriate replacements for the erroneous data. A relatively low confidence value indicates that the respective estimated snapshot values 552 and 554, including the resulting estimated KPI value 566, should either not be used or should be used with caution. A relatively high confidence value indicates that the respective estimated snapshot values 552 and 554, including the resulting estimated KPI value 566, should not be used. The associated confidence value thresholds may be established by a user and may be used to train one or more models, both conditions facilitating fully automated selection. Furthermore, subsequent actions may be automated. For example, and without limitation, for confidence values below a predetermined threshold, the estimated KPI value 566 is not passed on for further processing within the native application, e.g., the process control system 410, using the original data stream 420. Similarly, for confidence values above a predetermined threshold, the estimated KPI value 566 is passed on for further processing within the native application, e.g., the process control system 410, using the original data stream 422. Thus, the systems and methods as described herein prevent inadvertent actions or automatically correct issues using erroneous data in the original data stream 420 in a manner that initiates appropriate actions as dictated by the conditions and accurate data.
[0155] 5C , confidence measurement module 570 includes an explanation sub-module 578 configured to receive confidence-based data 580 from confidence measurement sub-modules 572, 574, and 576. The confidence-based data 580 includes, without limitation, the estimated KPI value 566 and its associated confidence value, respective information associated with the selection of the estimated KPI value 566, and additional information, and generates an explanation for the selected estimated KPI value 566 including other data 568 including, without limitation, all of the estimated snapshot values 552 and 554, including their respective confidence values. Furthermore, because the prediction, i.e., the confidence value for the estimated KPI value 566, may not be 100%, the explanation sub-module 578 provides explanatory basis for resolving one or more potentially erroneous data instances through providing details and evidence for the selection of the particular scored estimated snapshot values 562 and 564 as the estimated KPI value 566. The explanation sub-module 578 provides such details including, without limitation, the types of issues detected in the dataset, the number and nature of simulations generated, statistical properties of the scores obtained from various simulations, and comparisons of the scores. Accordingly, the confidence measure module 570 generates various confidence measures for the scored estimated snapshot values 562 and 564 and information to facilitate a user's understanding of the nature of the variance in the scored estimated snapshot values 562 and 564, and further to provide clarity for the selection of each estimated KPI value 566, thereby generating confidence scores and explanations 582 as outputs of the process 500.
[0156] 16, a text diagram is provided showing a confidence measure description 1600. The data provided in the confidence measure description 1600 is essentially self-explanatory.
[0157] As disclosed herein, systems, computer program products, and methods help overcome the drawbacks and limitations of inadvertently processing erroneous time-series data and potentially encountering unexpected results. For example, for a given business KPI as each data is generated, the automated systems and methods described herein determine whether data quality issues have a significant impact on the respective business KPI. Furthermore, the systems and methods described herein identify the nature (or characteristics) of the relevant business KPIs so that relevant data issues can be identified, and can perform optimization regardless of whether the correct KPI formulation is explicitly visible, i.e., whether the formulation is actually an observable or unobservable box. The systems and methods described herein also resolve the identified data issues by selecting a scored prediction of a replacement value for the erroneous data. Furthermore, the systems and methods described herein optimize the selection of possible replacement values to efficiently use system resources. The scored predictions are accompanied by a numerical confidence value, along with an explanation of the confidence value for the estimated confidence measure and the reason for the value. Thus, as described herein, data quality issues are filtered based on the analysis of a given KPI, and the data is modified to mitigate the quality issues, considering various scenarios, to calculate their impact on the measurement of the given KPI, and to further measure the confidence of the predicted replacement value.
[0158] Furthermore, the features of the systems, computer program products, and methods disclosed herein may extend beyond implementation in strictly business-based embodiments. Non-business implementations are also contemplated to overcome similar drawbacks and limitations of inadvertently processing erroneous time-series data and potentially encountering unexpected results. Specifically, any computer-implemented process that relies on time-series data to properly perform its respective function may be improved through implementation of the features of the present disclosure. For example, without limitation, the use of any time-series data collected from IoT devices, including residential and vehicle users, may avoid inadvertent and unnecessary automated actions through the replacement of missing data values with the highest degree of confidence. Specifically, for residential users, erroneous data indicating erroneously low voltage from their respective electric utility may be prevented from inadvertently and unnecessarily activating low-voltage protection circuits that would otherwise disrupt the supply of sufficient power to the respective residence. In such an implementation, one respective KPI may be maintaining continuity of power to residential users. Also, specifically, for a vehicle user, erroneous data falsely indicating excessive propulsion mechanism temperature may be prevented from inadvertently and unnecessarily activating an automated emergency engine shutdown. In such an implementation, one respective KPI may be to maintain continuity of propulsion for the vehicle user.
[0159] Accordingly, the embodiments disclosed herein provide improvements to computer technology by efficiently, effectively, and automatically identifying issues associated with erroneous time series data; determining whether data quality issues are impactful to a given business KPI through identifying characteristics of the business KPI such that relevant data issues can be identified and optimization can be performed on whether the precise KPI characteristics are straightforwardly defined, i.e., whether the KPI formulation is truly observable or unobservable; and providing a mechanism for resolving the identified data issues while presenting a confidence analysis of the potential solutions investigated.
[0160] The description of various embodiments of the present disclosure has been provided for illustrative purposes and is not intended to be exhaustive or to be limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best explain the principles of the embodiments, practical applications of, or technical improvements to, the technology found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. 1. A computer system comprising: one or more processing devices; and at least one memory device operably coupled to the one or more processing devices, wherein the one or more processing devices: Identifying one or more potentially erroneous data instances within the time series data stream; determining one or more predicted replacement values for the one or more potentially erroneous data instances; determining a confidence value for each predicted replacement value of the one or more predicted replacement values; resolving the one or more potentially erroneous data instances with a predicted replacement value of the one or more predicted replacement values; generating explanatory evidence regarding said resolution of said one or more potentially erroneous data instances; The system is configured as follows:
2. The one or more processing devices further include: The system of claim 1 , configured to identify one or more key performance indicators (KPIs) that are affected through the one or more potentially erroneous data instances.
3. The one or more processing devices further include:
3. The system of claim 2, configured to determine that one or more KPI formulation characteristics are associated with the one or more potentially erroneous data instances, each KPI of the one or more KPIs including one or more formulations thereof, and each formulation of the one or more formulations including one or more characteristics thereof.
4. The one or more processing devices further include: Analyze observable KPI formulation, Analyzing unobservable KPIs The system of claim 3 , configured to:
5. The one or more processing devices further include:
5. The system of claim 3 or 4, configured to generate one or more simulation snapshots, each simulation snapshot of the one or more simulation snapshots including one or more imputed values, and wherein each predicted replacement value of the one or more predicted replacement values is based at least in part on the one or more imputed values and the one or more KPI development characteristics.
6. The one or more processing devices further include: generating a plurality of estimated data points from the one or more potentially erroneous data instances, the method comprising alternately assigning one of a discrete correct label and a discrete erroneous label to each potentially erroneous data instance of the one or more potentially erroneous data instances; generating a set of all possible combinations of the plurality of estimated data points; determining the probability that the plurality of estimated data points are in fact erroneous; and generating a plurality of point-based simulation snapshots for only a subset of the set of all possible combinations of the plurality of estimated data points, each point-based simulation snapshot of the plurality of point-based simulation snapshots including the one or more imputed values; generating the plurality of simulation snapshots through a point-based simulation having generating the one or more imputed values for each potentially erroneous data instance, each imputed value of the one or more imputed values being generated through a respective repair operation; generating multiple simulation snapshots through a method-based simulation having The system of claim 5 , configured to:
7. The one or more processing devices further include:
7. The system of claim 6, configured to generate the subset of the set of all possible combinations of the plurality of estimated data points using a snapshot optimization feature through the use of the one or more KPI development characteristics.
8. The one or more processing devices further include:
8. The system of claim 1 , configured to resolve the one or more potentially erroneous data instances and generate explanatory evidence for the resolution of the one or more potentially erroneous data instances through one or more of a magnitude-based confidence measure and a spread-based confidence measure.
9. The processor identifying one or more potentially erroneous data instances in a time series data stream; determining one or more predicted replacement values for the one or more potentially erroneous data instances; determining a confidence value for each predicted replacement value of the one or more predicted replacement values; resolving the one or more potentially erroneous data instances using a predicted replacement value of the one or more predicted replacement values; generating an explanatory basis for said resolution of said one or more potentially erroneous data instances; A computer program for executing
10. the processor, 10. The computer program of claim 9, further comprising identifying one or more key performance indicators (KPIs) that are affected through the one or more potentially erroneous data instances.
11. the processor, 11. The computer program product of claim 10, further comprising: a step of determining that one or more KPI formulation characteristics are associated with the one or more potentially erroneous data instances, wherein each KPI of the one or more KPIs includes one or more formulations thereof, and each formulation of the one or more formulations includes one or more characteristics thereof.
12. the processor, 12. The computer program product of claim 11, further comprising: generating one or more simulation snapshots, each simulation snapshot of the one or more simulation snapshots including the one or more imputed values, and wherein each predicted replacement value of the one or more predicted replacement values is based at least in part on the one or more imputed values and the one or more KPI development characteristics.
13. 1. A computer-implemented method comprising: identifying one or more potentially erroneous data instances within the time series data stream; determining one or more predicted replacement values for the one or more potentially erroneous data instances; determining a confidence value for each predicted replacement value of the one or more predicted replacement values; resolving the one or more potentially erroneous data instances using a predicted value of the one or more predicted values; generating an explanatory basis for said resolution of said one or more potentially erroneous data instances; A method for providing
14. The method of claim 13 , further comprising identifying one or more key performance indicators (KPIs) that are affected through the one or more potentially erroneous data instances.
15. Each KPI of the one or more KPIs includes one or more formulations thereof, and each formulation of the one or more formulations includes one or more characteristics thereof, and the method further comprises: The method of claim 14 , further comprising determining that one or more of the one or more KPI development characteristics are associated with the one or more potentially erroneous data instances.
16. determining the one or more KPI development characteristics; A stage of analyzing observable KPI formulation; The stage of analyzing unobservable KPI formulation and 16. The method of claim 15, comprising:
17. Determining the one or more predicted replacement values comprises:
17. The method of claim 15 or 16, comprising generating one or more simulation snapshots, each simulation snapshot of the one or more simulation snapshots including the one or more imputed values, and wherein a predicted replacement value for each of the one or more predicted replacement values is based at least in part on the one or more imputed values and the one or more KPI development characteristics.
18. generating a plurality of estimated data points from the one or more potentially erroneous data instances, the method comprising alternately assigning one of a discrete correct label and a discrete erroneous label to each potentially erroneous data instance of the one or more potentially erroneous data instances; generating a set of all possible combinations of the plurality of estimated data points; determining the probability that the plurality of estimated data points are in fact erroneous; and generating a plurality of point-based simulation snapshots for only a subset of the set of all possible combinations of the plurality of estimated data points, each point-based simulation snapshot of the plurality of point-based simulation snapshots including the one or more imputed values; generating the plurality of simulation snapshots through a point-based simulation having: generating the one or more imputed values for each potentially erroneous data instance, each imputed value of the one or more imputed values being generated through a respective repair operation; generating a plurality of simulation snapshots through a method-based simulation having the method; 20. The method of claim 17, further comprising one or more of:
19. 20. The method of claim 18, further comprising generating the subset of the set of all possible combinations of the plurality of estimated data points with utilizing a snapshot optimization feature through the use of the one or more KPI development characteristics.
20. Resolving the one or more potentially erroneous data instances and generating the explanation basis may include:
20. The method of any one of claims 13 to 19, comprising using one or more of a magnitude-based confidence measure and a spread-based confidence measure.
Citation Information
Patent Citations
Management method for time series data, information disclosing device and recording means
JP2002358397A
Systems and methods for detecting, correcting, and validating bad data in data streams
JP2015082843A
System for maintenance recommendation based on maintenance effectiveness estimation
JP2017194967A
Systems and methods for detecting, correcting, and validating bad data in data streams
US20150121160A1
Data processing device, method and program-stored medium
WO2019159602A1