Using Machine Learning to Identify Non-Technical Losses
By applying machine learning algorithms to generate a classifier model, identifying signals related to energy usage conditions, solving the problem of difficulty in efficiently identifying non-technical losses in the prior art, and achieving a more efficient and accurate identification process.
Patent Information
- Application Number
- CN202110829239.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2014-09-24
- Filing Date
- 2015-09-24
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2035-09-24
AI Technical Summary
The difficulty in detecting and identifying non-technical losses (NTLs) efficiently in the prior art has led to challenges in identifying and resolving these losses.
By configuring the system to select signals related to multiple energy usage conditions, machine learning algorithms are applied to generate multiple N-dimensional representations and generate a classifier model for identifying non-technical losses.
It realizes efficient identification of non-technical losses, reduces the need for manpower analysis, and improves the accuracy and efficiency of identification.
Smart Images

Figure CN113918139B_ABST
Abstract
Description
[0001] This application is a divisional application of a Chinese patent application with an application date of September 24, 2015, an application number of 201580050923.0, and an invention title of "Using Machine Learning to Identify Non-Technical Losses" (the corresponding PCT application has an application date of September 24, 2015 and an application number of PCT / US2015 / 052048). Field of the Invention
[0002] The present invention technology relates to the field of energy management. More specifically, the present invention technology provides techniques for using machine learning to identify non-technical losses (NTLs). Background
[0003] Conventional energy management tools are designed to help companies track energy usage. For example, such tools can collect certain types of energy-related information, including billing statements and energy meter readings. The information collected can be used to understand or analyze energy usage. Such tools can also generate detailed reports of energy-related information and usage.
[0004] In some cases, energy providers (e.g., utility companies) may face challenges related to the losses of the energy provided, such as technical losses and non-technical losses (NTLs). Technical losses can include energy losses during normal use due to expected or natural limitations (such as power losses due to the resistance of cables, wires, power lines, etc.). Non-technical losses can include one or more losses that are not due to such limitations. Non-technical losses may be related to non-compliant (or undesired) energy usage, such as losses in the form of energy theft or malfunctions in the energy distribution system.
[0005] For energy providers, non-technical losses can be costly. Conventional methods for detecting non-technical losses typically require a large amount of manpower. In addition, conventional methods may also be inaccurate, inefficient, or ineffective. As a result, cases of non-technical losses are often overlooked, undetected, misdiagnosed, or otherwise not adequately addressed. These and other issues can pose challenges to energy providers and their customers.
[0006] Summary
[0007] Various embodiments of the present disclosure can include systems, methods, and non-transitory computer-readable media configured to select a set of signals related to multiple energy usage conditions. Signal values of the set of signals can be determined. Machine learning can be applied to the signal values to identify energy usage conditions associated with non-technical losses.
[0008] In an embodiment, multiple N-dimensional representations can be generated for multiple energy usage conditions. The multiple N-dimensional representations can be generated based on signal values. Applications of machine learning can include applying at least one machine learning algorithm to the multiple N-dimensional representations to produce a classifier model for identifying non-technical losses.
[0009] In an embodiment, at least a first portion of the multiple N-dimensional representations can be pre-identified as corresponding to non-technical losses. At least a second portion of the multiple N-dimensional representations can be pre-identified as corresponding to normal energy usage.
[0010] In an embodiment, at least one machine learning algorithm can include a regulatory process that classifies at least a third portion within an allowable N-dimensional neighborhood range of the first portion in the multiple N-dimensional representations as corresponding non-technical losses. The regulatory process can also classify at least a fourth portion within an allowable N-dimensional neighborhood range of the second portion in the multiple N-dimensional representations as corresponding to normal energy usage.
[0011] In an embodiment, new signal values of the set of signals can be received. The new signal values can be associated with a specific energy usage condition. A new N-dimensional representation can be generated for the specific energy usage condition based on the new signal values. The new N-dimensional representation can be classified based on the classifier model.
[0012] In an embodiment, at least one machine learning algorithm can be applied to the new N-dimensional representation to modify the classifier model.
[0013] In an embodiment, the new N-dimensional representation can be identified as corresponding to a non-technical loss. The non-technical loss can be reported to an energy provider associated with the specific energy usage condition.
[0014] In an embodiment, at least one of a confirmation or non-confirmation that a specific energy usage condition is associated with a non-technical loss can be obtained from one or more entities.
[0015] In an embodiment, the classifier model can be modified based on at least one of the confirmation or non-confirmation.
[0016] In an embodiment, at least one machine learning algorithm can be associated with at least one of a support vector machine, boosted decision tree, classification tree, regression tree, bagging tree, random forest, neural network, or rotation forest.
[0017] In an embodiment, multiple utility meters having a likelihood associated with a non-technical loss can be identified. The multiple utility meters can be ranked based on the likelihood associated with the non-technical loss.
[0018] In an embodiment, it may be determined that at least some of the plurality of meters meet a specified ranking threshold criterion. At least some of the plurality of meters may be identified as candidates for investigation.
[0019] In an embodiment, one or more signals of the set of signals may be associated with at least one of the following: an account attribute signal category, an abnormal load signal category, a computed status signal category, a current analysis signal category, a missing data signal category, a disconnection signal category, a meter event signal category, a monthly meter abnormal load signal category, an inactive monthly meter consumption signal category, an interruption signal category, a stolen meter signal category, an abnormal production signal category, a work order signal category, or a zero read signal category.
[0020] In an embodiment, a set of formulas for the set of signals may be obtained. Each formula in the set of formulas may correspond to a respective signal in the set of signals. The signal values of the set of signals may be calculated based on the set of formulas.
[0021] In an embodiment, at least some of the signal values may be derived from data obtained from a plurality of meters associated with a plurality of energy usage conditions.
[0022] In an embodiment, a first signal of the set of signals may be generated based on a modification of a second signal of the set of signals.
[0023] In an embodiment, at least one signal related to an energy usage condition and not included in the set of signals may be received from an energy provider to identify non-technical losses.
[0024] In an embodiment, at least one machine learning algorithm may include an unsupervised process. In some instances, the unsupervised process may utilize unclassified data to identify non-technical losses.
[0025] Many other features and embodiments of the disclosed technology will be apparent from the accompanying drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 An example environment of an energy management platform in accordance with an embodiment of the present disclosure is shown.
[0027] Figure 2 An example energy management platform in accordance with an embodiment of the present disclosure is shown.
[0028] Figure 3 An example application server of an energy management platform in accordance with an embodiment of the present disclosure is shown.
[0029] Figure 4Shows an example non - technical loss (NTL) identification module configured to utilize machine learning to identify non - technical losses according to an embodiment of the present disclosure.
[0030] Figure 5 Shows an example table of example signal values including a set of example signals according to an embodiment of the present disclosure.
[0031] Figure 6 Shows an example graph including an example N - dimensional representation generated based on example signal values according to an embodiment of the present disclosure.
[0032] Figure 7 Shows an example method for utilizing machine learning to identify non - technical losses according to an embodiment of the present disclosure.
[0033] Figure 8 Shows an example machine in which a set of instructions for causing the machine to perform one or more embodiments described herein can be executed according to an embodiment of the present disclosure.
[0034] The drawings depict various embodiments of the present disclosure for illustrative purposes only, where the drawings use the same reference numerals to identify the same elements. Those skilled in the art will readily recognize from the following discussion that alternative embodiments of the structures and methods shown in the figures can be employed without departing from the principles of the disclosed technology described herein. Detailed Description
[0035] Energy is consumed or used daily for various purposes. In one example, a consumer can use gas to power various appliances in a home, and an enterprise can use gas to operate various machines. In another example, consumers and enterprises can use electricity to power various electronic devices and other electrical devices and components. Energy consumption is facilitated by an energy provider that supplies the energy to meet the demand.
[0036] An energy provider, such as a utility company, can provide one or more forms of energy, such as gas and electricity. The energy provider can utilize an energy distribution system to provide or deliver energy to its intended customers or users. In some cases, there may be energy losses during delivery. For example, even in normal use, there may be resistance in power lines, cables, and / or wires, etc., such that electrical energy is lost during delivery through these channels. This energy loss is attributed to expected or natural causes and can be referred to as a technical loss. However, in some cases, there may be additional energy losses in addition to technical losses. Energy may be lost due to non-compliant or undesirable energy usage. For example, energy may be lost due to theft and / or malfunctions in the energy distribution system and its distribution nodes (e.g., malfunctioning utility meters). This energy loss can be referred to as non-technical loss (NTL).
[0037] Energy losses, such as non-technical loss (NTL), can be expensive for energy providers. However, conventional methods for attempting to detect, prevent, and reduce non-technical losses are problematic. Typically, conventional methods require a large amount of manpower to analyze information in attempts to detect non-technical losses, such as due to theft or utility meter malfunctions. Additionally, conventional methods typically only consider a limited amount of information. Worse still, conventional methods often rely on manual estimates and approximations, which can lead to inaccuracies and miscalculations. Therefore, improved methods for detecting, preventing, and reducing non-technical losses may be advantageous.
[0038] Various embodiments of the present disclosure are designed to consider all types of comprehensive information, such as information associated with an energy provider, energy customers, utility meters, and other components of an energy distribution or management system. The information can be analyzed (such as by utilizing machine learning techniques) to determine attributes or characteristics that may be associated with non-technical losses. Situations of energy usage with similar properties or characteristics can be classified as potentially corresponding to non-technical losses. Such situations of energy usage can be identified and reported to help prevent or reduce further non-technical losses. It is also anticipated that many variations are possible.
[0039] Figure 1 An example environment 100 for energy management in accordance with an embodiment of the present disclosure is shown. Environment 100 includes an energy management platform 102, external data sources 1041-n, an enterprise 106, and a network 108. The energy management platform 102, discussed in more detail herein, provides functionality that allows the enterprise 106 to track, analyze, and optimize the energy usage of the enterprise 106. The energy management platform 102 can constitute an analysis platform. The analysis platform can handle data management, multi-layer analysis, and data visualization capabilities for all applications of the energy management platform 102. The analysis platform can be specifically designed to process and analyze large amounts of frequently updated data while maintaining a high level of performance.
[0040] The energy management platform 102 can communicate with the enterprise 106 through a user interface (UI) presented by the energy management platform 102 for the enterprise 106. The UI can provide information to the enterprise 106 and receive information from the enterprise 106. The energy management platform 102 can communicate with external data sources 104 1-n through APIs and other communication interfaces. The communication involving the energy management platform 102, the external data sources 104 1-n, and the enterprise 106 is discussed in more detail herein.
[0041] The energy management platform 102 can be implemented as a computer system, such as a server or a series of servers and other hardware (e.g., application servers, analytical computing servers, database servers, data integrator servers, network infrastructure (e.g., firewalls, routers, communication nodes)). The servers can be arranged as a server farm or cluster. Embodiments of the present disclosure can be implemented on the server side, on the client side, or a combination of both. For example, embodiments of the present disclosure can be implemented by one or more servers of the energy management platform 102. As another example, embodiments of the present disclosure can be implemented by a combination of the servers of the energy management platform 102 and the computer systems of the enterprise 106.
[0042] The external data sources 104 1-n can represent numerous possible data sources related to energy management analysis. The external data sources 104 1-n can include, for example, power grids and utility operation systems, meter data management (MDM) systems, customer information systems (CIS), billing systems, utility customer systems, utility enterprise systems, utility energy conservation measures and discount databases. The external data sources 104 1-n can also include, for example, building feature systems, weather data sources, third-party property management systems, and industrial standard benchmark databases.
[0043] The enterprise 106 can represent a user (e.g., a customer) of the energy management platform 102. The enterprise 106 can include any private or public business, such as large companies, small and medium-sized enterprises, households, individuals, management groups, government agencies, non-governmental organizations, non-profit organizations, etc. The enterprise 106 can include energy providers and suppliers (e.g., utilities), energy service companies (ESCOs), and energy consumers. The enterprise 106 can be associated with one or many facilities distributed across many geographical locations. The enterprise 106 can be associated with any purpose, industry, or other type of profile.
[0044] Network 108 may use standard communication technologies and protocols. Accordingly, network 108 may include links using technologies such as Ethernet, 802.11, Worldwide Interoperability for Microwave Access (WiMAX), 3G, 4G, CDMA, GSM, LTE, Digital Subscriber Line (DSL), and so on. Similarly, network protocols used on network 108 may include Multiprotocol Label Switching (MPLS), Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Simple Mail Transfer Protocol (SMTP), File Transfer Protocol (FTP), etc. Data exchanged over network 108 may be represented using technologies and / or formats including Hypertext Markup Language (HTML) and Extensible Markup Language (XML). Additionally, conventional encryption technologies such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), and Internet Protocol Security (IPsec) may be used to encrypt all or some of the links.
[0045] In an embodiment, each of energy management platform 102, external data sources 1041-n, and enterprise 106 may be implemented as a computer system. The computer system may include one or more machines, each of which may be implemented as machine 800 described in further detail herein Figure 8 below.
[0046] Figure 2 An example energy management platform 202 in accordance with an embodiment of the present disclosure is shown. In certain embodiments, example energy management platform 202 may be implemented as Figure 1 energy management platform 102. In an embodiment, energy management platform 202 may include a data management module 210, an application server 212, a relational database 214, and a key / value store 216.
[0047] Data management module 210 may support the ability to automatically and dynamically scale the network of computing resources of energy management platform 202 according to demands on energy management platform 202. The dynamic scaling supported by data management module 210 may include the ability to provide additional computing resources (or nodes) to accommodate increased computing demands. Similarly, data management module 210 may include the ability to release computing resources to accommodate decreased computing demands. Data management module 210 may include one or more actions 218, queues 220, schedulers 222, resource managers 224, and cluster managers 226.
[0048] Action 218 may represent a task performed in response to a request provided to the energy management platform 202. Each action 218 may represent a unit of work performed by the application server 212. Action 218 may be associated with a data type and bound to an engine (or module). The request may pertain to any task supported by the energy management platform 202. For example, the request may pertain to, e.g., analytical processing, loading energy-related data, retrieving Energy Star readings, retrieving benchmark data, etc. Action 218 is provided to the action queue 220.
[0049] The action queue 220 may receive each action 218. The action queue 220 may be a distributed task queue and represents work to be routed to appropriate computing resources and then executed.
[0050] The scheduler 222 may associate the queued actions to the engine that will execute the action and switch the queued actions to the engine for that action. The scheduler 222 may control the routing of each queued action to a particular one of the application servers 212 based on load balancing and other optimization considerations. When the current computing resources are at or above the threshold capacity, the scheduler 222 may receive instructions from the resource manager 224 for provisioning new nodes. When the current computing resources are at or below the threshold capacity, the scheduler 222 may receive instructions from the resource manager for releasing nodes. The scheduler 222 may thus instruct the cluster manager 226 to dynamically provision new nodes or release existing nodes based on the demand for computing resources. Nodes may be computing nodes or storage nodes connected to the application server 212, the relational database 214, and the key / value store 216.
[0051] The resource manager 224 may monitor the action queue 220. The resource manager 224 may also monitor the current load on the application server 212 to determine the availability of resources for executing the queued actions. Based on the monitoring, the resource manager may communicate with the cluster manager 226 via the scheduler 222 to request the dynamic allocation and deallocation of nodes.
[0052] The cluster manager 226 may be a distributed entity that manages all nodes of the application server 212. The cluster manager 226 may dynamically provision new nodes or release existing nodes based on the demand for computing resources. The cluster manager 226 may implement a group membership service protocol. The cluster manager 226 may also perform a task monitoring function. The task monitoring function may involve tracking resource usage, such as CPU utilization, amount of data read / written, storage size, etc.
[0053] The application server 212 may execute processes that manage or host analytical server execution, data requests, etc. Engines provided by the energy management platform 202 (such as engines that perform data services, batch processing, stream services) may be hosted within the application server 212. Engines are discussed in more detail herein.
[0054] In an embodiment, the application server 212 can be part of a computer cluster of multiple loosely or tightly connected computers that are coordinated to work as a system when executing the services and applications of the energy management platform 202. Nodes of the cluster (e.g., servers) can be connected to each other via a fast local area network (“LAN”), where each node runs an instance of its own operating system. The application server 212 can be implemented as a computer cluster to increase its performance and availability to levels higher than that of a single computer, while generally being more cost-effective than a single computer with comparable speed or availability. The application server 212 can be a combination of software, hardware, or both.
[0055] The relational database 214 can maintain various data that supports the energy management platform 202. In an embodiment, as discussed in more detail herein, non-time series data can be stored in the relational database 214.
[0056] The key / value store 216 can maintain various data that supports the energy management platform 202. In an embodiment, as discussed in more detail herein, time series data (e.g., meter readings, meter events, etc.) can be stored in the key / value store. In an embodiment, the key / value store 216 can be implemented with Apache Cassandra, where Apache Cassandra is an open-source distributed database management system designed to handle large amounts of data on numerous commodity servers. In an embodiment, other database management systems for key / value stores can be used.
[0057] In an embodiment, one or more of the application server 212, the relational database 214, and the key / value store 216 can be implemented by an entity that owns, maintains, or controls the energy management platform 202.
[0058] In an embodiment, one or more of the application server 212, the relational database 214, and the key / value store 216 can be implemented by a third party that can provide a computing environment for lease to an entity that owns, maintains, or controls the energy management platform 202. In an embodiment, the application server 212, the relational database 214, and the key / value store 216 implemented by the third party can communicate with the energy management platform 202 via a network such as network 208.
[0059] The computing environment provided by a third party for the entity that owns, maintains, or controls the energy management platform 202 can be a cloud computing platform that allows the entity that owns, maintains, or controls the energy management platform 202 to lease virtual computers and run its own computer applications on the virtual computers. Such applications can include, for example, the applications executed by the application server 200, as discussed in more detail herein. In an embodiment, the computing environment can allow for scalable deployment of applications by providing a web service through which the entity that owns, maintains, or controls the energy management platform 202 can initiate virtual appliances for creating virtual machines containing any desired software. In an embodiment, the entity that owns, maintains, or controls the energy management platform 202 can create, start, and terminate server instances as needed and pay based on time of use, data usage, or any combination of these or other factors. The ability to provide and release computing resources in this manner supports the ability of the energy management platform 202 to scale dynamically based on demand on the energy management platform 202.
[0060] Figure 3 An example application server 300 of an energy management platform according to an embodiment of the present disclosure is shown. In an embodiment, Figure 2 one or more of the application servers 212 can be implemented using Figure 3 the application server 300. The application server 300 includes a data integrator (data loading) module 302, an integration service module 304, a data service module 306, a computing service module 308, a stream analysis service module 310, a batch parallel processing analysis service module 312, a normalization module 314, an analysis container 316, a data model 318, and a user interface (UI) service module 324. In certain embodiments, the application server 300 can also include a non-technical loss (NTL) identification module 330.
[0061] The analysis platform supported by the application server 300 includes multiple services, each handling a specific data management or analysis capability. The services include the data integrator module 302, the integration service module 304, the data service module 306, the computing service module 308, the stream analysis service module 310, the batch parallel processing analysis service module 312, and the UI service module 324. All or some of the services within the analysis platform can be modular and are thus specifically constructed to perform their respective capabilities for large data volumes and high speeds. The services can be optimized in software for high-performance distributed computing on a computer cluster including the application servers 212.
[0062] Figure 3The modules and components of the application server 300 in [description] and all the drawings in this document are merely exemplary and can be combined differently into fewer modules and components or separated into additional modules and components. The functions of the described modules and components can be performed by other modules and components.
[0063] The data integrator module 302 is a tool for automatically importing data maintained in software systems or databases of external data sources 1041 - n into Figure 1 the energy management platform 102. The imported data can be used for various applications of the energy management platform 102 or the application server 300. The data integrator module 302 accepts data from a wide range of data sources, including the power grid and operating systems (such as MDM, CIS, and billing systems) as well as third - party data sources (such as weather databases, building databases (e.g., city planning commission databases), third - party property management systems, and external benchmark databases). The imported data can include, for example, meter data (e.g., electricity consumption, water consumption, natural gas consumption) provided at a minimum daily or other time intervals (e.g., 15 - minute intervals), weather data (e.g., temperature, humidity) at daily or other time intervals (e.g., hourly intervals), building data (e.g., square footage, occupancy, age, building type, number of floors, air - conditioned square footage), cluster definitions (hierarchies) (e.g., building to which a meter belongs, city block to which a building belongs, area identification of a building), and asset data (e.g., quantity and type of HVAC assets, quantity and type of production units (for factories)).
[0064] The data integrator module 302 also has the ability to import information from flat files (such as Excel spreadsheets) and the ability to capture information directly input into applications of the energy management platform 102. By combining data from a wide array of sources, the application server 300 is able to perform complex and detailed analyses, achieving greater business insights.
[0065] The data integrator module 302 provides a set of standardized specification object definitions (standardized interface definitions) that can be used to load data into applications of the application server 300. The specification objects of the data integrator module 302 can be based on current or emerging utility industry standards (such as the Common Information Model (CIM), Green Button, and Open Automated Data Exchange) or on the specifications of the application server 300. The application server 300 can support these and other standards to ensure that a wide range of utility data sources will be able to easily connect to the energy management platform 102. Specification objects can include, for example:
[0066]
[0067]
[0068] Once the data in canonical form is received, the data integrator module 302 can transform the data into individual data entities according to the data model 318 so that the data can be loaded into the database schema for storage, processing, and analysis.
[0069] The data integrator module 302 is capable of processing very large amounts of data (e.g., "big data"). For example, the data integrator module 302 can frequently process interval data from millions of digital meters. To receive data, the application server 300 can provide a consistent secure web service API (e.g., REST). The integration can be performed in asynchronous batch processing or real-time mode. The data integrator module 302 can combine real-time and batch data from, for example, utility customer systems, building feature systems, industrial standard benchmark systems, utility energy conservation measures and discount databases, utility enterprise systems, MDM, and utility operation systems. When an external data source does not have an API or computerized device for extracting data, the application server 300 can directly extract data from a web page associated with the external data source (e.g., by using web scraping).
[0070] The data integrator module 302 can also perform initial data validation. The data integrator module 302 can check the structure of the input data to ensure that the required fields are present and the data has the correct data type. For example, the data integrator module 302 can identify that the format of the provided data does not match the expected format (e.g., a numeric value is incorrectly provided as a text format), prevent the mismatched data from being loaded, and record the issue for review and investigation. In this way, the data integrator module 302 can serve as the first line of defense to ensure that the input data meets the requirements for accurate analysis.
[0071] The integration service module 304 serves as the second layer of data validation or proofreading to ensure that the data is error-free before being loaded into the database or storage. The integration service module 304 receives data from the data integrator module 302, monitors the data as it flows in, performs a second round of data checking, and passes the data to the data service module 306 for storage.
[0072] The integration service module 304 can provide various data management functions. The integration service module 304 can perform duplicate processing. The integration service module 304 can identify cases of data duplication to ensure that analysis is accurately performed on a single data set. The integration service module 304 can be configured to handle duplicates according to user-specified business requirements (e.g., treating two duplicate records as the same or taking the average of duplicate records). This flexibility allows the application server 300 to conform to customer standards for data processing.
[0073] The integration service module 304 can perform data validation. The integration service module 304 can detect data gaps and data anomalies (e.g., statistical anomalies), identify outliers, and perform referential integrity checks. The referential integrity checks ensure that the data has the correct association network for analysis and aggregation, such as ensuring that the loaded meter data is associated with the facility, or conversely, ensuring that the facility has associated meters. The integration service module 304 solves the data validation problem according to the user-specified business requirements. For example, if there are data gaps, linear interpolation can be used to fill in the missing data, or the gaps can be left as they are.
[0074] The integration service module 304 can perform data monitoring. The integration service module 304 can provide end-to-end visibility throughout the data loading process. The user can monitor the data integration process from duplicate detection to data storage. This monitoring helps ensure that the data is loaded properly and there are no duplicates and validation errors.
[0075] The data service module 306 is responsible for persisting (storing) large and growing amounts of data while also making the data easily available for analytical calculations. The data service module 306 splits the data into relational and non-relational (key / value store) databases and also performs operations on the stored data. These operations include creating, reading, updating, and deleting data. The data engine of the data service module 306 can persist data for stream processing. The data engine of the data service module 306 can also identify data sets to be processed in conjunction with batch jobs for batch parallel processing.
[0076] The data service module 306 can perform data partitioning. The data service module 306 takes advantage of relational data stores and non-relational data stores (such as Figure 2 the relational database 214 and the key / value store 216). By "partitioning" the data into two separate data stores, the relational database 214 and the key / value store 216, the application server 300 ensures that its applications can effectively process and analyze large amounts of data, such as interval data from meters and grid sensors. The data in the relational database 214 and the key / value store 216 is stored according to the data model 318 of the energy management platform 102.
[0077] The relational database 214 is designed to manage structured and slowly changing data. Examples of such data include organization (e.g., customer) and facility data. Relational databases (like the relational database 214) are designed for random access updates.
[0078] The key / value store 216 is designed to manage very large amounts of interval (time series) data, such as meter and grid sensor data. Key / value stores (like key / value store 216) are designed for large "append-only" data streams that are read in a specific order. "Append-only" means that new data is only added to the end of the associated file. By using a dedicated key / value store 216 for interval data, the application server 300 ensures that this type of data is stored efficiently and can be accessed quickly.
[0079] The data service module 306 can perform distributed data management. The data service module 306 can include an event queue that schedules notifications to perform stream processing and batch parallel processing. With respect to batch parallel processing, the scheduling can be based on rules that consider the availability of processing resources in the associated clusters in the energy management platform 102. As the data volume grows, the data service module 306 automatically adds nodes to the cluster to accommodate (e.g., store and process) the new data. As nodes are added, the data service module 306 automatically rebalances and partitions the data across all nodes, ensuring continued high performance and reliability.
[0080] The compute service module 308 is a library of analytical functions that are called by the stream analysis service module 310 and the batch parallel processing analysis service module 312 to perform business analytics. The functions can be executed individually or in combination to form complex analyses. The services provided by the compute service module 308 can be modular (i.e., dedicated to a single task), such that the compute service module 308 can process a large number of computations simultaneously and quickly in parallel, which allows for significant computational scalability.
[0081] The compute service module 308 can also leverage distributed processing to create even greater scalability. For example, if a user is interested in calculating the average annual electricity consumption of hundreds of thousands of meters, the energy management platform 102 can quickly respond by distributing the request across multiple servers.
[0082] The stream analysis service module 310 performs complex analysis on real-time and near real-time data streams. The streams can represent, for example, feeds of high-volume data from meters, sub-meters, or grid sensors. In an implementation, the streams can be data supervisory control and data acquisition (SCADA) feeds. The stream analysis service module 310 can be called to analyze the data when the analysis needs to be performed soon after the data is generated.
[0083] The flow analysis service module 310 may include a flow processor to convert a flow into data according to the data model 318. The flow analysis service module 310 may also include flow processing logic, which may be provided by a user of the energy management platform 102. The flow processing logic may provide calculation results that can be retained and used for subsequent analysis. The flow processing logic may also provide an alert based on the calculation results. For example, when there is an unexpected and significant drop or spike in the load, a utility may want to receive an alert and immediate analysis. Such load changes may be caused by a faulty part of a device or a sudden breakdown of a device, and may represent a great risk to the distribution system or end users. Data regarding unexpected load changes can be quickly identified, analyzed, and used to send the necessary alerts. The flow processing logic may also provide a new flow for another purpose or for an application of the energy management platform 102 based on the processed original flow after processing the original flow.
[0084] The flow analysis service module 310 may perform near-real-time continuous processing. Since the processing performed by the flow analysis service module 310 occurs very quickly after the data arrives, the time-sensitive high-priority analysis provided by the energy management platform 102 is relevant and feasible.
[0085] The flow analysis service module 310 may provide horizontal scalability. To manage a large amount of data simultaneously, the processing performed by the flow analysis service module 310 may be distributed across an entire server cluster (a group of computers working together).
[0086] The flow analysis service module 310 may provide fault tolerance. The flow may be retained. If a processing failure occurs on one node (e.g., a computer in the cluster), the workload will be distributed to other nodes within the cluster without data loss. After the processing performed on the flow is completed, the flow may be discarded.
[0087] A non-limiting example is provided to illustrate the performance of the flow analysis service module 310. Assume a flow of recently generated power consumption and demand data. The flow may be provided to an event queue associated with the data service module 306. When the data arrives at the event queue, an automatic analysis process is triggered. Multiple analysis processes or analyses may run on the same data set. The analysis processes may be executed in parallel. Parallel processing of the same data set enables multiple analyses to be processed faster. The outputs of these analysis processes may be alerts and calculations, which are then stored in a database and made available as analysis results for a specified end user. The analysis processes and processing tasks may be distributed across multiple servers supporting the flow analysis service module 310. In this way, a large amount of data can be quickly processed by the flow analysis service module 310.
[0088] The batch parallel processing analysis service module 312 can perform most of the analyses required by the users of the energy management platform 102. The batch parallel processing analysis service module 312 can analyze large data sets consisting of current and historical data to create reports and analyses, such as periodic key performance indicator (KPI) reports, historical power usage analysis, predictions, outlier analysis, analysis of the financial impact of energy efficiency projects, and so on. In an implementation, the batch parallel processing analysis service module 312 can be based on MapReduce (a programming model for processing large data sets and distributing computations across one or more computer clusters). The batch parallel processing analysis service module 312 automatically performs tasks of parallelization, fault tolerance, and load balancing, thus improving the performance and reliability of processing-intensive tasks.
[0089] Non-limiting examples are provided to illustrate the performance of the batch parallel processing analysis service module 312. As an example, a benchmark analysis of energy intensity, a performance summary for key performance indicators, and an analysis of unbilled energy due to non-technical losses can be jobs processed by the batch parallel processing analysis service module 312. When a batch job is invoked in the energy management platform 102, the input reader associated with the batch parallel processing analysis service module 312 decomposes the processing job into multiple smaller batches. This decomposition reduces the complexity and processing time of the job. Then, each batch is passed to a worker process to perform its assigned task (e.g., calculation or estimation). Then the results are "shuffled", which refers to the rearrangement of the data set so that the next set of worker processes can efficiently complete the calculation (or estimation) and quickly write the results to the database through an output writer.
[0090] The batch parallel processing analysis service module 312 can distribute worker processes across multiple servers. This distributed processing is used to fully utilize the computing power of the cluster and ensure that the calculations are completed quickly and efficiently. In this way, the batch parallel processing analysis service module 312 provides scalability and high performance.
[0091] The normalization module 314 can normalize the meter data to be maintained in the key / value store 216. For example, the normalization of meter data can involve filling in blanks in the data and resolving outliers in the data. For example, if it is expected that the meter data is at consistent intervals, but the data actually provided to the energy management platform 102 does not have meter data at some intervals, the normalization module 314 can apply certain algorithms (e.g., interpolation) to provide the missing data. As another example, abnormal values of energy usage can be detected and resolved by the normalization module 314. In an implementation, the normalization performed by the normalization module 314 can be configurable. For example, the algorithms (e.g., linear, non-linear) used by the normalization module 314 can be specified by the administrator or user of the energy management platform 102. The normalized data can be provided to the key / value store 216.
[0092] The UI service module 324 provides a graphical framework for all applications of the energy management platform 102. The UI service module provides visualization of analysis results, enabling end users to receive clear and actionable insights. After the analysis is completed by the flow analysis service module 310 or the batch parallel processing analysis service module 312, they can be graphically presented by the UI service module 324, provided to the appropriate applications of the energy management platform 102, and ultimately presented on the user's computer system (e.g., machine). This provides data insights to the user in an intuitive and easy-to-understand format.
[0093] The UI service module 324 provides many features. The UI service module 324 can provide a library of chart types and a library of page layouts. All variants of chart types and page layouts are maintained by the UI service module 324. The UI service module 324 can also provide page layout customization. Users (such as administrators) can add, rename, and group fields. For example, the energy management platform 102 allows a utility administrator to group energy intensity, energy consumption, and energy demand together on a page for easy viewing. The UI service module 324 can provide role-based access control. An administrator can determine which parts of an application will be visible to certain types of users. Using these features, the UI service module 324 ensures that end users enjoy a consistent visual experience, access capabilities and data relevant to their roles, and can interact with charts and reports that provide clear business insights.
[0094] In addition, in some implementations, the application server 300 includes a non-technical loss (NTL) identification module 330, as Figure 3 shown. The non-technical loss identification module 330 can be configured to identify non-technical losses using machine learning. In certain embodiments, the non-technical loss identification module 330 can be implemented as hardware, software, and / or a combination thereof. It is also contemplated that in some instances, one or more parts or components of the non-technical loss identification module 330 can be implemented using Figure 1 one or more other modules, engines, and / or components of the energy management platform 102.
[0095] In one example, the non-technical loss identification module 330 can be configured to obtain or determine signal values of a set of signals indicative of the presence of non-technical losses (e.g., NTL). The set of signals indicative of the presence of non-technical losses can directly or indirectly reflect various conditions of energy usage. Such energy usage conditions can relate to, for example, the type of energy usage, the state of energy usage, the amount of energy usage, the energy usage reading of a meter, the operating state of a meter, the state of a customer account of an energy provider, and any other considerations directly or indirectly reflecting energy supply, usage, availability, and payment. Each signal from the set of signals can reflect a specific energy usage condition. The signal values of the signals from the set of signals can be numerical, boolean, binary, or qualitative values describing the magnitude, type, or presence (or absence) of the energy usage condition associated with the signal. For example, the energy usage condition can refer to various situations during which energy is being used or consumed, including situations of zero consumption or non-use. In some cases, the energy usage condition can represent the state (e.g., current state) of energy usage measured by an energy or utility meter (e.g., gas meter, electric meter, water meter, etc.). In some cases, a specific energy usage condition can be associated with the use of a specific type of energy by a specific energy consumer or customer at a specific geographical location at a specific time or interval at a specific premise. Thus, the energy usage condition can be associated not only with the meter measuring the usage, but also with customer information, location information, premise type, date, and time, etc.
[0096] The set of signals can correspond to a selected set of analytics or features generated based on the acquired data (such as data received from Figure 1 the external data sources 1041-n). In certain embodiments, the set of signals can be selected, chosen, or determined based on research, development, observation, machine learning, and / or experimentation, etc. For example, based on empirical analysis, certain signals can be determined to be more useful for indicating non-technical losses (NTL), and thus these signals are selected or prioritized over other signals that cannot or are less likely to indicate non-technical losses. The data received from the data sources can include but are not limited to AMI system data (meter data management and head-end data), customer information data, customer consumption data, billing information, contract information, meter event information, outage management system (OMS) data, generator generation, work order management (WO) data, verified theft and fault data, weather, and geolocation. The data sources can include but are not limited to power grids and utility operating systems, meter data management (MDM) systems, customer information systems (CIS), billing systems, utility customer systems, utility enterprise systems, utility energy conservation measures, discount databases, building feature systems, weather data sources, third-party property management systems, industry standard benchmark databases, etc.
[0097] By means of a large variety of signals in a large number of signal categories and the corresponding signal values of these signals, a better understanding of energy usage can be achieved. Each signal of a category from various signal categories can be generated, and its corresponding signal value can be calculated based on at least a part of the acquired data. In some cases, there can be dozens of signal categories, and within each signal category, there are hundreds of signals or more. This disclosure will only discuss a few examples. It should be understood that many signal categories and their signals other than those explicitly discussed herein can also be utilized. In some implementations, the signal value can be a numerical value, a value between 0 and 1, a binary value, etc.
[0098] An example signal category is the "account attribute" signal category. The "account signal" category can include various signals. For example, the first signal of the "account signal" category can be called the "seasonal meter" signal. The "seasonal meter" signal can indicate whether a house (or customer) is recorded as seasonal, such as for a vacation home. Data from CIS (such as customer information and customer consumption data) can indicate that the house is seasonal, and a signal value can be set for the "seasonal meter" signal to indicate that the house is seasonal.
[0099] As another example, the second signal of the "account attribute" signal category can be called the "service disconnect" signal. The "service disconnect" signal can indicate whether a house has a service point that has been terminated or disconnected at the relevant analysis time (e.g., the time of data acquisition). If the service point has been disconnected, the signal value of the "service disconnect" signal will indicate that the service point has been disconnected. If the service point has not been disconnected, the signal value will indicate that the service point has not been disconnected.
[0100] Another example signal category is the "abnormal load" signal category. The "abnormal load" signal category can include the "active power vs. reactive power curve analysis" signal, which involves analyzing active and reactive power data and identifying abnormal patterns indicating theft and / or faults. For example, the signal value of the "active power vs. reactive power curve analysis" signal can characterize the irregular changes in the annual consumption pattern of a given customer, which can indicate the likelihood of theft and / or faults. The "abnormal load" signal category can also include the "number of days with decreasing annual consumption" signal, which involves recording the number of days with decreasing usage year over year. In addition, the "abnormal load" signal category can include the "year-over-year change (every quarter hour)" signal, which involves calculating the maximum difference in consumption during the same month from one year to the previous year. In addition, the "abnormal load" signal category can include the "consumption decrease" signal related to tracking the consumption curve and recording when the 15-day rolling average consumption of the meter drops by more than 20%.
[0101] Another example signal category is the "Computation Status" signal category, which may include signals that facilitate cross-checking the status of a meter, such as by checking whether the meter status is set to active or whether the meter reports a communication problem. The "Meter Location Indoor" signal in this category may indicate that the meter is indoors. The "Meter Location Outdoor" signal in this category may indicate that the meter is outdoors. The "Service Inactive Consumption (Electricity)" signal in this category may indicate that the service is inactive, but there is still electricity consumption on the meter.
[0102] Another example signal category is the "Inactive Consumption" signal category. The "Inactive Consumption" signal category may include "Inactive Consumption" signals, which relate to detecting customers with non-zero consumption who have service accounts disconnected by a utility company. The "Inactive Consumption (Gas)" signal in the "Inactive Consumption" signal category may also relate to a situation where there is no service agreement activity, but there is gas consumption on the meter.
[0103] Another example signal category is the "Current Analysis" signal category. Signals in this category may be associated with analyzing historical current (amperage) curves to evaluate any inconsistencies among load harmonics, real and reactive power measurements, and potential outages. This category may include the "CT > 0.5 Amps" signal indicating intervals where the current transformer (CT) is greater than 0.5 amps and the "CT < 0.05 Amps" signal indicating intervals where the current transformer (CT) is less than 0.05 amps.
[0104] Another example signal category is the "Missing Data" signal category, which includes signals related to missing data. The "Missing Data" signal in this signal category relates to identifying whether the meter has missing consumption data.
[0105] Another example signal category is the "Disconnection" signal category, which includes signals associated with evaluating whether a meter has been disconnected from the communication network. The "Electric Disconnection Unreachable" signal in this category may indicate the number of days after a remotely disconnected Advanced Metering Infrastructure (AMI) meter has become unreachable. The "Communication after Hard Disconnection" signal in this category may indicate detecting a Network Interface Controller (NIC) power recovery event after a service point has been disconnected at a pole or service head. The "Days Disconnected before Unreachable" signal in this category may indicate the number of days the meter was disconnected before becoming unreachable.
[0106] Another example signal category is the "Meter Event" signal category, which includes signals that track various meter events (e.g., meter tampering events, meter failure events, meter last gasp events, etc.) and filter out any noise (e.g., due to the large number of meter events reported by meters, many of which are false positives). The "Failure Event" signal in this category can identify meters with failure events and can calculate the number of failure events that have been triggered. The "Failure and Shutdown Event Count" signal in this category can identify meters with failure events and can calculate the readings of failure and shutdown events. The "Damage Event Count" signal in this category can evaluate the number of recorded meter damage events. The "Damage of Fault Combinations Combined with Shutdown Meter Events" in this signal category can identify meters with combined meter events, including damage events, failure events, and shutdown events.
[0107] Another example signal category is the "Monthly Meter" signal category, which includes signals associated with meters that report data at monthly intervals. These signals can provide insights into monthly reporting meters or, more generally, can facilitate predicting patterns with less available data. The "Maximum Monthly Consumption Drop" signal in this signal category can record the maximum monthly-to-month consumption drop. The "Year-to-Year Change (Monthly, Seasonal)" signal in this category can calculate the maximum difference in consumption for non-seasonal meters during the same month period from one year to the previous year. The "Inactive Meter Consumption (Monthly)" signal can identify that the meter contract has ended and record non-zero (monthly) consumption after the contract termination date.
[0108] Another example signal category is the "Interruption" signal category, which includes signals that can track interruptions, disturbances, and can be related to consumption curves to provide more insights into whether a meter is damaged or whether a meter has experienced an interruption. The "Line Interruption Event" signal in this category can identify whether a line interruption event has been recorded for a meter. The "Interruption Related to Consumption Drop" signal in this category can track interruption data and set a flag when there is an interruption related to a drop in the consumption curve. The "Partial Line Interruption Event" signal in this category can track whether a partial line interruption event has been detected.
[0109] Another example signal category is the "Stolen Meter" signal category. The "Interruption and Stolen Meter" signal in this category relates to whether a meter has been stolen and whether it occurred during a short interruption. The "Stolen Meter Distance" signal in this category relates to whether the meter is > 300 feet from the expected installation location.
[0110] Another example signal category is the "abnormal production" signal category, which includes signals that can track network metering customers producing electricity (e.g., solar electricity) and can detect that the production data is abnormal. The "production after dark" signal in this category can identify whether production (reverse consumption) is detected during dark times. The "electricity production after dark" signal in this category can indicate the electricity generated during dark times.
[0111] Another example signal category is the "work order" signal category, which includes signals that track work orders to gain insights into whether a customer has been reported for theft or has an unpaid history on his or her account, etc. Signals in the "work order" category can be powerful in insights related to consumption patterns and theft patterns. The "cancel work order" signal in this category can identify the cancellation of services for customers who miss payments. The "contract change" signal in this category can identify whether a service change to a contract has been registered. The "meter change" signal in this category can generate results for each work order corresponding to a meter change.
[0112] Another example signal category is the "zero reading" signal category, which includes signals that track zero readings on meters to detect zero consumption patterns that do not match those of the nearest neighbors or peer account clusters. The "intermittent zero readings" signal in this category can identify meter zero readings that persist for a specified number of sequential meter readings (e.g., within a specified time period). The "continuous zero readings related to an interruption (non-seasonal)" signal in this category can track continuous zero readings related to an interrupted (non-seasonal) meter (e.g., for more than 7 days). The "intermittent zero" signal in this category can indicate a zero reading period that persists for a specified time period (e.g., at least 6 hours).
[0113] Similarly, the signals and signal categories described herein are exemplary and for illustrative purposes. Other suitable signals and signal categories may be employed additionally or alternatively. Many variations are also contemplated as possible. In some cases, there may be more (or fewer) signals than those described herein. In certain embodiments, the first signal in a set of signals may be generated based on a modification of a second signal in the set. In one example, the first signal may be generated based on a permutation of the second signal. In another example, the first signal may be generated based on a combination of a second signal and a third signal.
[0114] In some instances, there may be more (or fewer) signal categories than those described herein. For example, in certain embodiments, one or more signals in the set of signals may be associated with at least one of the following: account attribute signal category, abnormal load signal category, computed status signal category, inactive consumption signal category, current analysis signal category, missing data signal category, disconnect signal category, meter event signal category, monthly meter abnormal load signal category, inactive monthly meter consumption signal category, interruption signal category, stolen meter signal category, abnormal production signal category, work order signal category, or zero read signal category.
[0115] After determining a set of selected signals from the selected signal categories, the signal values of the signals can be determined based on the data received from the data source. In some embodiments, determining the signal values may include determining a set of formulas for the set of signals. Each formula in the set of formulas may correspond to a respective signal in the set of signals. Then, the signal values of the set of signals can be calculated based on the set of formulas. By way of illustration, the signal value of the "consumption decline" signal may correspond to the amount of digital consumption decline of the meter compared to the average consumption of the meter. It should be understood that many other formulas can be obtained or developed for various other signals. Additionally, in some embodiments, the signal values can be normalized across the set of signals.
[0116] After determining the signal values for the set of signals, the non-technical loss identification module 330 can generate multiple N-dimensional representations (e.g., points in an N-dimensional space) for multiple energy usage conditions based on the signal values, where N represents the number of signals (i.e., the signal volume) in a set of signals indicating the presence of non-technical losses. For example, if there are 150 signals, the N-dimensional representation can have 150 dimensions. Each dimension can correspond to a respective signal. A particular energy usage condition among the multiple energy usage conditions can be represented as a point in the N-dimensional space with coordinates based on the signal values.
[0117] The non-technical loss identification module 330 can also apply at least one machine learning algorithm to the multiple N-dimensional representations to generate a classifier model for identifying non-technical losses. The classifier model can be used to identify energy usage conditions that may involve non-technical losses in the form of, for example, theft or malfunction.
[0118] Figure 4 An example non-technical loss (NTL) identification module 400 configured to utilize machine learning to identify non-technical losses in accordance with an embodiment of the present disclosure is shown. The example non-technical loss identification module 400 can be implemented as Figure 3 the non-technical loss identification module 330. As described above, in certain embodiments, various parts of the non-technical loss identification module 400 can be implemented as Figure 2One or more components of the energy management platform 202. For example, in certain embodiments, at least some portions of the non-technical loss identification module 400 may be implemented as Figure 3 One or more components of the application server 300.
[0119] Such as Figure 4 As shown, the non-technical loss identification module 400 may include a signal data acquisition module 402, an N-dimensional representation module 404, a machine learning module 406, and a result processing module 408. The signal data acquisition module 402 may be configured to determine a set of signals of the set of signals and associated signal values. The signal values may be associated with multiple energy usage conditions. In certain embodiments, the signal data acquisition module 402 may be implemented as residing in Figure 3 The data integrator module 302, and / or operate in conjunction with Figure 3 The data integrator module 302. Data may be received from external data sources 1041-n, and the set of signals may be generated based on such received data. The signal data acquisition module 402 may determine the signal values of the set of signals, such as by applying a set of formulas to the set of signals. Each formula in the set of formulas may correspond to a respective signal in the set of signals. In some cases, the set of formulas may be derived or developed from research, analysis, observation, experimentation, etc. The signal data acquisition module 402 may be configured to calculate the signal values of the set of signals based on the set of formulas. In some cases, each condition of energy usage may be represented by one or more corresponding signal values. For example, a particular set of signal values may be associated with the current state of a particular utility meter of a particular customer at a particular location and premises.
[0120] The N-dimensional representation module 404 may be configured to generate multiple N-dimensional representations of multiple energy usage conditions. The multiple N-dimensional representations may be generated based on the signal values. Each N-dimensional representation may be generated based on the signal values associated with the respective condition of energy usage. Each N-dimensional representation may have N dimensions corresponding to the amount of signals of the set of signals. In one example, each energy usage condition may be represented as a point in an N-dimensional space and may have coordinates corresponding to its respective signal values. In another example, each energy usage condition may be represented as an N-dimensional vector having vector values corresponding to its respective signal values. Other N-dimensional representations may also be used.
[0121] The machine learning module 406 may be configured to apply at least one machine learning algorithm to the multiple N-dimensional representations. A classifier model for identifying non-technical losses may be produced, developed, or generated based on applying at least one machine learning algorithm to the multiple N-dimensional representations.
[0122] In some embodiments, at least one machine learning algorithm can be associated with a regulatory process. In one example, at least a first portion of a plurality of N-dimensional representations can have been pre-identified or verified as corresponding to non-technical losses. At least a second portion of the plurality of N-dimensional representations can have been pre-identified or verified as corresponding to normal energy usage. Machine learning module 406 can classify a new signal value associated with a new energy usage condition as normal or associated with NTL based on the proximity of the new signal value to N-dimensional representations that have been verified as normal or associated with NTL. Machine learning module 406 can be configured to determine one or more N-dimensional representations that are close to or cluster with the first portion. Machine learning module 406 can classify these one or more N-dimensional representations that are close to or cluster with the first portion as corresponding to non-technical losses because they have attributes (e.g., signal values) similar to those of the first portion. In some cases, a first representation is adjacent (or clusters, is close, etc.) to a second representation when the first and second representations are within an allowable N-dimensional proximity (or threshold) of each other. For example, machine learning module 406 can classify at least a third portion of the plurality of N-dimensional representations (which is within the allowable N-dimensional proximity of the first portion) as corresponding to non-technical losses.
[0123] Similarly, machine learning module 406 can classify one or more N-dimensional representations that are close to or cluster with the second portion as corresponding to normal energy usage because they have attributes similar to those of the second portion (e.g., signal values). For example, machine learning module 406 can classify at least a fourth portion of the plurality of N-dimensional representations (within the allowable N-dimensional proximity of the second portion) as corresponding to normal energy usage.
[0124] In addition, the machine learning module 406 can be configured to receive or obtain new signal values of the set of signals. The new signal values can be associated with an environment of a change regarding a new energy usage condition. For example, new data can be received from a specific utility meter, and the new signal values can be calculated based on the received new data. The machine learning module 406 can generate a new N-dimensional representation for the new energy usage condition based on the new signal values. For example, the new signal values can be used to generate new points in the N-dimensional space. Since the signal values and the N-dimensional representation are new, they have not been classified yet. The machine learning module 406 can classify the new N-dimensional representation based on the classifier model. For example, if the classifier model indicates that the new N-dimensional representation is similar to (or close enough in N-dimensional proximity, adjacent, clustered, etc.) another representation cluster that has been classified as corresponding to non-technical losses, the new N-dimensional representation can also be classified as corresponding to non-technical losses. Thus, at least one machine learning algorithm can contribute to mapping at least some N-dimensional representations to non-technical losses based on the signal values. On the other hand, if the classifier model determines that the new representation is similar to another representation classified as normal energy usage, the new representation can be classified as normal energy usage.
[0125] In some instances, at least one machine learning algorithm includes an unsupervised process. Thus, unclassified data (e.g., new signal values) can be used to detect new patterns, trends, attributes, and / or characteristics for identifying non-technical losses. For example, it can be assumed that the N-dimensional representations of high-density clusters correspond to normal usage. The unsupervised process can attempt to classify small clusters of N-dimensional representations that are outside or substantially separated from the high-density clusters. If one of the representations in the small cluster is verified as corresponding to non-technical losses, the entire small cluster can be classified as corresponding to non-technical losses. In some cases, human review or confirmation can facilitate the unsupervised process.
[0126] In some cases, one or more new signal values associated with the new energy usage condition can be obtained and analyzed to continuously or periodically train the classifier model. Through a supervised process or an unsupervised process, the new signal values can be analyzed to provide an improved understanding for more accurately identifying energy usage conditions that may be associated with non-technical losses and those that may be normal. As the machine learning module 406 receives new signal values indicating non-technical losses and new signal values indicating normal energy usage, at least one machine learning algorithm can modify the classifier model to account for the new signal values. Thus, the classifier model can learn, change, and improve over time. In certain embodiments, the classifier model can determine that certain signals used for classifying energy usage conditions may not be particularly relevant or important for the determination of non-technical losses based on their signal values. Thus, the energy usage identification module 400 can selectively consider eliminating certain signals in the identification of non-technical losses.
[0127] In some embodiments, signals can be selected to maximize output. In this context, output can refer to the figure of correctly identified leads relative to the total leads associated with potential instances of non-technical losses. Signals can also be selected to minimize false positives. False positives can refer to instances of non-technical losses that are not correctly identified, which can result in associated costs and delays.
[0128] In some embodiments, at least one machine learning algorithm can be associated with at least one of a support vector machine, boosted decision tree, classification tree, regression tree, bagging tree, random forest, neural network, or rotation forest. It should be understood that many other variations, methods, techniques, and / or processes can be utilized.
[0129] The result processing module 408 can be configured to facilitate the processing of data, such as applying at least one machine learning algorithm to data generated from multiple N-dimensional representations. In some embodiments, the result processing module 408 can be configured to identify multiple utility meters having a likelihood associated with non-technical losses, such as gas meters, power meters, and water meters. For example, the identified meters can be associated with energy usage conditions represented by certain N-dimensional representations that have been classified as corresponding to non-technical losses.
[0130] In addition, the result processing module 408 can rank the multiple identified utility meters based on the likelihood associated with the non-technical losses. For example, the result processing module 408 can generate a ranking or score for the identified meters based on their respective likelihoods of being associated with non-technical losses. In some implementations, the likelihood of an identified meter being associated with a particular energy usage condition can depend on the N-dimensional proximity between a representation associated with one energy usage condition and another representation verified to correspond to a non-technical loss. A smaller N-dimensional proximity can indicate a higher likelihood.
[0131] The result processing module 408 can also determine that at least some of the multiple meters meet specified ranking threshold criteria and can provide at least some of the multiple utility meters as candidates for investigation regarding potential non-technical losses. In one example, the ranking threshold criteria can specify a minimum percentage amount of likelihood. In another example, the ranking threshold criteria can specify an amount of the highest likelihood. The meters ranked that meet the ranking threshold criteria can be the meters most likely to encounter non-technical losses, such as due to theft or malfunction.
[0132] In addition, as previously discussed, the new N-dimensional representation can be recognized as corresponding to non-technical losses. The result processing module 408 can report non-technical losses to one or more entities associated with specific energy usage conditions. For example, the meters determined to be most likely to encounter non-technical losses can be presented to one or more energy providers or suppliers (e.g., utility companies). The energy provider or supplier can then investigate and resolve any issues.
[0133] In some cases, the result processing module 408 can obtain at least one of a confirmation or non-confirmation that specific energy usage conditions are associated with non-technical losses from one or more entities such as energy providers. For example, one or more entities can conduct on-site investigations or perform other processes to confirm the presence or absence of non-technical losses. The entity can report its findings back to the non-technical loss identification module 400. Additionally, in some instances, the classifier model can be modified, improved, or enhanced based on at least one of the confirmation or non-confirmation.
[0134] Figure 5 An example table 500 including a set of example signal values of example signals according to an embodiment of the present disclosure is shown. As Figure 5 shown, the example table 500 can show a set of three example signals (Signal A, Signal B, and Signal N). Thus, the number of signals for this set of example signals is three. Many variations can be expected.
[0135] In Figure 5 the example, Signal A is a "consumption decline" signal. For example, the signal value of Signal A is calculated to be 0.82. In this example, Signal B can correspond to a "line outage event" signal and can have a signal value of 0.74. For example, Signal N can be a "cancel work order" signal with a signal value of 0.91. These signal values can be associated with specific energy usage conditions. For example, these signal values can be associated with a specific utility meter at a specific time. Based on these signal values, an N-dimensional representation can be generated, which will be discussed in more detail with reference to Figure 6 more detail.
[0136] Figure 6 An example graph 600 including an example N-dimensional representation generated based on example signal values according to an embodiment of the present disclosure is shown. The example graph 600 can show an N-dimensional representation (e.g., points) 610 generated based on the signal values of the set of signals shown in the example table 500 of Figure 5 .
[0137] Since Figure 5 the number of signals in the set in Figure 6Each dimension in the N-dimensional space is associated with an axis and can correspond to Figure 5 the corresponding signal in. Thus, dimension A602 can correspond to Figure 5 signal A, dimension B 604 can correspond to signal B, and dimension N 606 can correspond to signal N. Thus, the N-dimensional representation 610 has coordinates (A = 0.82, B = 0.74, N = 0.91) and is accordingly presented in the example diagram 600.
[0138] As Figure 6 shown in the example of, the representation 610 is within a cluster 612 that includes other N-dimensional representations, which can represent other energy usage conditions related to, for example, other meters. In one example, if the representation 610 is within an allowable distance from the cluster 612, the representation can be classified according to the cluster 612. For example, if the cluster 612 has been verified to be associated with NTL (or alternatively, normal energy usage), then the representation 610 will be similarly classified as being associated with NTL (or, alternatively, normal energy usage) when located within the distance allowed from the cluster 612.
[0139] In another example, if it is verified that the representation 610 corresponds to a non-technical loss, then the entire cluster 612 to which the representation 610 belongs can be classified as corresponding to a non-technical loss (and vice versa for normal energy). If another representation in the cluster 612 is verified to correspond to a non-technical loss and if the representation 610 has not yet been classified, then the representation 610 (and the entire cluster 612) can be classified as corresponding to a non-technical loss (and vice versa for normal energy usage). The other clusters in the example diagram 600 can be classified in a similar manner.
[0140] Furthermore, it should be understood that the example diagram 600 of Figure 6 is provided for illustrative purposes. In some embodiments, the N-dimensional representation need not be presented in a graphical or visual form.
[0141] Figure 7 An example method 700 for using machine learning to identify non-technical losses in accordance with an embodiment of the present disclosure is shown. It should be understood that additional, fewer, or alternative steps may be performed in a similar or alternative order (or in parallel) within the scope of the various embodiments unless otherwise stated.
[0142] At block 702, example method 700 may select a set of signals related to multiple energy usage conditions. In some cases, the set of signals may be associated with multiple energy usage conditions. In some embodiments, the set of signals may be determined, in whole or in part, by an operator of the energy management platform 102. The set of signals may be stored in a library, either within or outside of the energy management platform 102. In some instances, the set of signals may grow, shrink, and / or change over time. For example, the number of signals in the set may be modified based on a machine learning algorithm used to classify energy usage conditions. In some implementations, an energy provider, such as a utility company, may create their own signals and provide those signals to the energy management platform 102 for use in addition to or in place of the set of signals determined by the operator of the energy management platform 102.
[0143] At block 704, example method 700 may determine signal values for the set of signals. In some instances, multiple N-dimensional representations of the multiple energy usage conditions may be generated based on the signal values. Additionally, each N-dimensional representation may have N dimensions corresponding to the number of signals in the set.
[0144] At block 706, example method 700 may apply machine learning to the signal values to identify energy usage conditions associated with non-technical losses. In some instances, the application of machine learning to the signal values may involve applying at least one machine learning algorithm to the multiple N-dimensional representations to produce a classifier model for identifying non-technical losses. In some embodiments, the classifier model may be modified, improved, and / or refined over time. Additional details of example method 700 are discussed above and will not be repeated here.
[0145] It is further contemplated that there may be many other uses, applications, and / or variations associated with the various embodiments of the present disclosure.
[0146] Figure 8 An example machine 800 is shown in accordance with an embodiment of the present disclosure, in which a set of instructions may be executed to cause the machine to perform one or more embodiments described herein. The machine may be connected (e.g., networked) to other machines. In a networked deployment, the machine may operate in the capacity of a server or a client machine in a client-server network environment, or as a peer machine in a peer-to-peer (or distributed) network environment.
[0147] Machine 800 includes a processor 802 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), a main memory 804, and a non-volatile memory 806 (e.g., volatile RAM and non-volatile RAM), which communicate with each other via a bus 808. For example, in some embodiments, machine 800 can be a desktop computer, a laptop computer, a personal digital assistant (PDA), or a mobile phone. In one embodiment, machine 800 further includes a video display 810, an alphanumeric input device 812 (e.g., a keyboard), a cursor control device 814 (e.g., a mouse), a disk drive unit 816, a signal generation device 818 (e.g., a speaker), and a network interface device 820.
[0148] In one embodiment, video display 810 includes a touch-sensitive screen for user input. In one embodiment, the touch-sensitive screen is used instead of a keyboard and a mouse. Disk drive unit 816 includes a machine-readable medium 822, on which is stored one or more sets of instructions 824 (e.g., software) embodying any one or more of the methods or functions described herein. Instructions 824 may also reside, completely or at least partially, within main memory 804 and / or within processor 802 during execution by computer system 800. Instructions 824 may also be sent or received via network interface device 820 over network 840. In some embodiments, machine-readable medium 822 further includes a database 825.
[0149] Volatile RAM can be implemented as dynamic RAM (DRAM), which continuously requires power to refresh or maintain the data in the memory. Non-volatile memory is typically a magnetic hard disk drive, a magnetic optical drive, an optical drive (e.g., DVD RAM), or other types of memory systems that retain data even after power is removed from the system. Non-volatile memory can also be random access memory. Non-volatile memory can be a local device directly coupled to the rest of the components in the data processing system. Non-volatile memory located far from the system can also be used, such as a network storage device coupled to any of the computer systems described herein via a network interface (such as a modem or an Ethernet interface).
[0150] Although the machine-readable medium 822 is shown as a single medium in the exemplary embodiment, the term "machine-readable medium" should be considered to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store a set or multiple sets of instructions. The term "machine-readable medium" will also be considered to include any medium that is capable of storing, encoding, or carrying a set of instructions for execution by a machine and that causes the machine to perform any one or more of the methods of the present disclosure. Thus, the term "machine-readable medium" should be considered to include, but not be limited to, solid-state memories, optical and magnetic media, and carrier signals. As used herein, the term "storage module" may be implemented using a machine-readable medium.
[0151] In general, routines executed to implement the embodiments of the present disclosure may be implemented as part of an operating system or a specific application, component, program, object, module, or sequence of instructions referred to as a "program" or "application". For example, one or more programs or applications may be used to execute the specific processes described herein. A program or application typically includes one or more instructions set at various times in various memories and storage devices in a machine, and when read and executed by one or more processors, it causes the machine to perform operations to execute elements related to various aspects of the embodiments described herein.
[0152] Executable routines and data may be stored in various places, including, for example, ROM, volatile RAM, non-volatile memory, and / or caches. Portions of these routines and / or data may be stored in any of these storage devices. In addition, the routines and data may be obtained from a centralized server or a peer-to-peer network. Different portions of the routines and data may be obtained from different centralized servers and / or peer-to-peer networks at different times and in different communication sessions or in the same communication session. The routines and data may be obtained in their entirety before the execution of the application. Alternatively, portions of the routines and data may be obtained dynamically in a timely manner when needed. Thus, it is not required that the routines and data be present in their entirety on a machine-readable medium at a particular time instance.
[0153] Although the embodiments have been described fully in the context of a machine, those skilled in the art will understand that the various embodiments can be distributed in various forms as a program product, and the embodiments described herein are equally applicable regardless of the particular type of machine-readable medium or computer-readable medium used for the actual implementation of the distribution. Examples of machine-readable media include, but are not limited to, recordable media such as volatile and non-volatile memory devices, floppy disks and other removable disks, hard disk drives, optical disks (e.g., compact disk read-only memories (CD ROMs), digital versatile disks (DVDs), etc.), and transmission-type media such as digital and analog communication links.
[0154] Alternatively or in combination, dedicated circuits with or without software instructions, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs), may be used to implement the embodiments described herein. Embodiments may be implemented using hardwired circuitry without software instructions or in combination with software instructions. Accordingly, the techniques are not limited to any specific combination of hardware circuitry and software, nor to any specific source of the instructions executed by a data processing system.
[0155] For purposes of illustration, numerous specific details are set forth in order to provide a thorough understanding of this specification. However, it will be apparent to one of ordinary skill in the art that the embodiments of the present disclosure may be practiced without these specific details. In some instances, modules, structures, processes, features, and devices are shown in block diagram form in order to avoid obscuring the description. In other instances, functional block diagrams and flowcharts are shown to represent data and logical flows. The components of the block diagrams and flowcharts (e.g., modules, engines, blocks, structures, devices, features, etc.) may be combined, separated, removed, reordered, and replaced in different ways than explicitly described and depicted herein.
[0156] References in this specification to "one embodiment", "an embodiment", "other embodiments", "another embodiment", etc., mean that a particular feature, design, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. For example, the phrases "according to an embodiment", "in one embodiment", "in an embodiment", or "in another embodiment" that appear throughout the specification are not necessarily all referring to the same embodiment, nor to separate or alternative embodiments that are mutually exclusive of other embodiments. Moreover, various features are described, which may be combined differently and included in some embodiments but variously omitted in other embodiments, whether or not there is an explicit reference to "embodiments" and the like. Similarly, various features are described, which may be preferred or required for some embodiments but not for others.
[0157] While embodiments have been described with reference to specific exemplary embodiments, it will be apparent that various modifications and changes can be made to these embodiments. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. The foregoing specification provides a description with reference to specific exemplary embodiments. Obviously, various modifications can be made without departing from the broader spirit and scope set forth in the appended claims. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.
[0158] Although some of the figures show multiple operations or method steps in a particular order, steps that are not order-dependent may be reordered and other steps may be combined or omitted. While some reorderings or other groupings are specifically mentioned, others will be apparent to those of ordinary skill in the art and thus no exhaustive list of alternatives is provided. Additionally, it should be recognized that these stages may be implemented in hardware, firmware, software, or any combination thereof.
[0159] It should also be understood that various changes may be made without departing from the essence of the disclosure. Such changes are also implicitly included in the description. They still fall within the scope of the disclosure. It should be understood that the disclosure is intended to yield patents covering multiple aspects of the disclosed technology, whether independently or as an overall system, and in both method and apparatus modes.
[0160] Furthermore, each of the various elements of the present disclosure and the claims may also be implemented in various ways. The present disclosure should be understood to include each such variation, being a variation of any apparatus embodiment, method or process embodiment, or even merely a variation of any element thereof.
Claims
1. A computer-implemented method, comprising: Obtain a set of signals related to energy and / or customer-related data from multiple data sources, the multiple data sources including a meter data management system and a customer information system; Determine the signal values of the set of signals by correlating data from the multiple data sources; Generate a representation for one or more energy usage or related conditions at least in part based on the signal values, wherein the representation includes a multi-dimensional representation of the energy usage conditions, and wherein each representation corresponds to an energy usage condition; Apply at least one machine learning algorithm to the representation to produce a classifier model for identifying non-technical losses, wherein applying the at least one machine learning algorithm includes performing an unsupervised machine learning process on the multi-dimensional representation to determine multiple clusters of the representation, and wherein at least one of the multiple clusters is identified as corresponding to non-technical losses based on the corresponding energy usage condition of at least one representation of the at least one of the multiple clusters; and Train the classifier model by: (1) modifying the classifier model to consider more relevant new signal values and new energy usage conditions, and / or (2) selectively eliminating one or more less relevant signals in the identification of the non-technical losses.
2. The computer-implemented method according to claim 1, wherein the unsupervised machine learning process is configured to process unclassified data to identify the non-technical losses.
3. The computer-implemented method according to claim 1, wherein at least a first part of the representation has been pre-identified as corresponding to the non-technical losses, and at least a second part of the representation has been pre-identified as corresponding to normal energy usage.
4. The computer-implemented method according to claim 1, further comprising: Determine the new signal values of a new set of input signals, the new signal values being associated with the one or more energy usage or related conditions; Generate a new representation for the one or more energy usage or related conditions based on the new signal values; and Classify the new representation using the classifier model.
5. The computer-implemented method according to claim 4, further comprising: Identify one or more of the new representations as corresponding to the non-technical losses; and Report the non-technical losses to an energy provider for investigation.
6. The computer-implemented method according to claim 5, further comprising: Obtain a confirmation signal or a non-confirmation signal from the energy provider related to the one or more energy usage or related conditions and the non-technical losses.
7. The computer-implemented method according to claim 6, further comprising: Modify the classifier model at least based on the confirmation signal or the non-confirmation signal.
8. The computer-implemented method according to claim 1, wherein the at least one machine learning algorithm includes a support vector machine, an enhanced decision tree, a classification tree, a regression tree, a bagging tree, a random forest, a neural network, or a rotation forest.
9. The computer-implemented method according to claim 1, further comprising: Identify multiple meters having a likelihood associated with the non-technical losses; and Rank the multiple meters based on the likelihood associated with the non-technical losses.
10. The computer-implemented method according to claim 9, further comprising: Determine that at least some of the multiple meters meet a ranking threshold criterion; and Identify the at least some of the multiple meters as candidates for investigation.
11. The computer-implemented method according to claim 1, wherein one or more signals in the set of signals are associated with at least one of the following: account attribute signal category, abnormal load signal category, computed status signal category, current analysis signal category, missing data signal category, disconnection signal category, meter event signal category, monthly meter abnormal load signal category, inactive monthly meter consumption signal category, interruption signal category, stolen meter signal category, abnormal production signal category, work order signal category, or zero reading signal category.
12. The computer-implemented method according to claim 1, further comprising: Obtain a set of formulas for the set of signals, each formula in the set of formulas corresponding to a respective signal in the set of signals; and Calculate the signal values of the set of signals based on the set of formulas.
13. The computer-implemented method according to claim 1, wherein at least some of the signal values are derived from data obtained from a plurality of meters associated with the one or more energy usages or related conditions.
14. The computer-implemented method according to claim 1, wherein a first signal in the set of signals is generated at least in part based on a modification of a second signal in the set of signals.
15. The computer-implemented method according to claim 1, further comprising receiving, from an energy provider, at least one signal related to the one or more energy usages or related conditions that is not included in the set of signals to identify the non-technical losses.
16. The computer-implemented method according to claim 1, wherein two or more signal values of two or more different signals are associated with a specific condition from the one or more energy usages or related conditions.
17. The computer-implemented method according to claim 16, wherein the two or more different signals include a first signal associated with a consumption drop, a second signal associated with a line interruption event, and / or a third signal associated with a cancelled work order.
18. A system for identifying non-technical losses using machine learning, the system comprising: A server communicating with multiple data sources; and A memory storing instructions which, when executed by the server, cause the server to perform operations including: Obtain a set of signals related to energy and / or customer-related data from multiple data sources, the multiple data sources including a meter data management system and a customer information system; Determine the signal values of the set of signals by correlating data from the multiple data sources; Generate a representation for one or more energy usages or related conditions, at least in part based on the signal values, wherein the representation includes a multi-dimensional representation of the energy usage conditions, and wherein each representation corresponds to an energy usage condition; Apply at least one machine learning algorithm to the representation to produce a classifier model for identifying non-technical losses, wherein applying the at least one machine learning algorithm includes performing an unsupervised machine learning process on the multi-dimensional representation to determine a plurality of clusters of the representation, and wherein at least one of the plurality of clusters is identified as corresponding to a non-technical loss based on the corresponding energy usage condition of at least one representation of the at least one of the plurality of clusters; and Train the classifier model by: (1) modifying the classifier model to account for more relevant new signal values and new energy usage conditions, and / or (2) selectively eliminating one or more less relevant signals in the identification of the non-technical losses.
19. The system according to claim 18, wherein the unsupervised machine learning process is configured to process unclassified data to identify the non-technical losses.
20. The system according to claim 18, wherein at least a first portion of the representation has been pre-identified as corresponding to the non-technical losses, and at least a second portion of the representation has been pre-identified as corresponding to normal energy usage.
21. The system according to claim 18, wherein the operation further comprises: Determine new signal values for a new set of input signals, the new signal values being associated with the one or more energy usages or related conditions; Generate a new representation for the one or more energy usages or related conditions based on the new signal values; and Classify the new representation using the classifier model.
22. The system according to claim 21, wherein the operation further comprises: Identify one or more of the new representations as corresponding to the non-technical losses; and Report the non-technical losses to an energy provider for investigation.
23. The system according to claim 22, wherein the operation further comprises: Obtain a confirmation signal or a non-confirmation signal from the energy provider related to the one or more energy usages or related conditions associated with the non-technical losses.
24. The system according to claim 23, wherein the operation further comprises: Modify the classifier model based at least on the confirmation signal or the non-confirmation signal.
25. The system according to claim 18, wherein the at least one machine learning algorithm comprises a support vector machine, boosted decision tree, classification tree, regression tree, bagging tree, random forest, neural network, or rotation forest.
26. The system according to claim 18, wherein the operation further comprises: Identify a plurality of meters having a likelihood associated with the non-technical losses; and Rank the plurality of meters based on the likelihood associated with the non-technical losses.
27. The system according to claim 26, wherein the operation further comprises: Determine that at least some of the plurality of meters meet a ranking threshold criterion; and Identify at least some of the plurality of meters as candidates for investigation.
28. The system according to claim 18, wherein one or more signals of the set of signals are associated with at least one of the following: account attribute signal category, abnormal load signal category, computed status signal category, current analysis signal category, missing data signal category, disconnect signal category, meter event signal category, monthly meter abnormal load signal category, inactive monthly meter consumption signal category, interruption signal category, stolen meter signal category, abnormal production signal category, work order signal category, or zero read signal category.
29. The system according to claim 18, wherein the operation further comprises: Obtain a set of formulas for the set of signals, each formula in the set of formulas corresponding to a respective signal in the set of signals; and Calculate the signal values of the set of signals based on the set of formulas.
30. The system according to claim 18, wherein at least some of the signal values are derived from data obtained from a plurality of meters associated with the one or more energy usages or associated conditions.
31. The system according to claim 18, wherein the first signal in the set of signals is generated at least in part based on a modification of a second signal in the set of signals.
32. The system according to claim 18, wherein the operation further comprises: Receive from an energy provider at least one signal related to the one or more energy usages or related conditions that is not included in the set of signals to identify the non-technical losses.
33. The system according to claim 18, wherein the signal values of two or more different signals are associated with a specific condition from the one or more energy usages or associated conditions.
34. The system according to claim 33, wherein the two or more different signals include a first signal associated with a consumption decrease, a second signal associated with a line interruption event, and / or a third signal associated with a cancelled work order.
35. A non - transitory computer - readable storage medium comprising instructions that, when executed by at least one processor of a computing system, cause the computing system to perform operations comprising: Obtain a set of signals related to energy and / or customer - related data from a plurality of data sources, the plurality of data sources including a meter data management system and a customer information system; Determine the signal values of the set of signals by correlating data from the plurality of data sources; Generate a representation for one or more energy usages or related conditions, at least in part based on the signal values, wherein the representation includes a multi-dimensional representation of the energy usage conditions, and wherein each representation corresponds to an energy usage condition; Apply at least one machine learning algorithm to the representation to generate a classifier model for identifying non-technical losses, wherein applying the at least one machine learning algorithm includes performing an unsupervised machine learning process on the multi-dimensional representation to determine a plurality of clusters of the representation, and wherein at least one of the plurality of clusters is identified as corresponding to a non-technical loss based on a corresponding energy usage condition of at least one representation of the at least one of the plurality of clusters; and Train the classifier model by: (1) modifying the classifier model to account for more relevant new signal values and new energy usage conditions, and / or (2) selectively eliminating one or more less relevant signals in the identification of the non-technical losses.
36. The non - transitory computer - readable storage medium according to claim 35, wherein the unsupervised machine - learning process is configured to process unclassified data to identify the non - technical losses.
37. The non - transitory computer - readable storage medium according to claim 35, wherein at least a first portion of the representation has been pre - identified as corresponding to the non - technical losses and at least a second portion of the representation has been pre - identified as corresponding to normal energy usage.
38. The non - transitory computer - readable storage medium according to claim 35, wherein the operation further comprises: Determine new signal values for a new set of input signals, the new signal values being associated with the one or more energy usages or related conditions; Generate a new representation for the one or more energy usages or related conditions based on the new signal values;and Classify the new representation using the classifier model.
39. The non - transitory computer - readable storage medium according to claim 38, wherein the operation further comprises: Identify one or more of the new representations as corresponding to the non-technical losses; and Report the non-technical losses to an energy provider for investigation.
40. The non - transitory computer - readable storage medium according to claim 39, wherein the operation further comprises: Obtain from the energy provider a confirmation signal or a non-confirmation signal related to the non-technical losses for the one or more energy usages or related conditions.
41. The non - transitory computer - readable storage medium according to claim 40, wherein the operation further comprises: Modify the classifier model based at least on the confirmation signal or the non-confirmation signal.
42. The non - transitory computer - readable storage medium according to claim 35, wherein the at least one machine - learning algorithm comprises a support vector machine, a boosted decision tree, a classification tree, a regression tree, a bagging tree, a random forest, a neural network, or a rotation forest.
43. The non - transitory computer - readable storage medium according to claim 35, wherein the operation further comprises: Identify a plurality of meters having a likelihood associated with the non-technical losses; and Rank the plurality of meters based on the likelihood associated with the non-technical losses.
44. The non - transitory computer - readable storage medium according to claim 43, wherein the operation further comprises: Determine that at least some of the plurality of meters meet a ranking threshold criterion; and Identify the at least some of the plurality of meters as candidates for investigation.
45. The non - transitory computer - readable storage medium according to claim 35, wherein one or more signals in the set of signals are associated with at least one of the following: account attribute signal category, anomaly load signal category, computed status signal category, current analysis signal category, missing data signal category, disconnect signal category, meter event signal category, monthly meter anomaly load signal category, inactive monthly meter consumption signal category, interruption signal category, stolen meter signal category, abnormal production signal category, work order signal category, or zero - read signal category.
46. The non - transitory computer - readable storage medium according to claim 35, wherein the operation further comprises: Obtain a set of formulas for the set of signals, each formula in the set of formulas corresponding to a respective signal in the set of signals; and Calculate the signal values of the set of signals based on the set of formulas.
47. The non-transitory computer-readable storage medium according to claim 35, wherein at least some of the signal values are derived from data obtained from a plurality of meters associated with the one or more energy usages or related conditions.
48. The non-transitory computer-readable storage medium according to claim 35, wherein a first signal in the set of signals is generated at least in part based on a modification of a second signal in the set of signals.
49. The non-transitory computer-readable storage medium according to claim 35, wherein the operation further comprises: Receive from an energy provider at least one signal related to the one or more energy usages or related conditions that is not included in the set of signals to identify the non-technical losses.
50. The non-transitory computer-readable storage medium according to claim 35, wherein signal values of two or more different signals are associated with a specific condition from the one or more energy usages or related conditions.
51. The non-transitory computer-readable storage medium according to claim 50, wherein the two or more different signals include a first signal associated with a consumption decrease, a second signal associated with a line interruption event, and / or a third signal associated with a work order cancellation.