Automated progressive adjustment of machine learning models for fault location in overhead transmission lines

Through the automated gradual adjustment method, the architecture and parameters of the machine learning model are jointly adjusted, which solves the problem of real-time execution of large models on devices with limited computing capabilities, and generates a fault location model with strong adaptability and low computing resource requirements.

CN120344977APending Publication Date: 2025-07-18HITACHI ENERGY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380085194.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-30
Filing Date
2023-12-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to execute large machine learning models in real time on devices with limited computing power, especially in power system fault location. Traditional adjustment methods are costly and slow and cannot effectively adapt to environmental changes.

Method used

Through an automated progressive adjustment method, the architecture and parameters of the machine learning model are adjusted in combination, and retrained layer by layer until the stop condition is met, and the model is pruned to accommodate devices with limited computing power.

Benefits of technology

It realizes rapid adaptation to environmental changes on devices with limited computing capabilities, generates smaller intensive machine learning models, meets real-time fault location requirements, reduces computing resource requirements, and improves model adaptability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120344977A_ABST
    Figure CN120344977A_ABST
Patent Text Reader

Abstract

Conventional methods for reducing the size of machine learning models generally involve expensive and slow searching for new architectures, or pruning model parameters to produce sparse machine learning models. However, some environments, such as regression tasks performed on time series data in real time to be embedded in hardware, require pruned machine learning models, are dense. Thus, automated progressive adjustments are disclosed to jointly adjust both the architecture and parameters of a machine learning model initially trained for a first environment to deploy in a second environment. In embodiments, during adjustment, constraints are applied to the dimensions of the machine learning model to prune the machine learning model to a smaller dense machine learning model suitable for embedding in hardware and other potential environments.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technical Field

[0002] The embodiments described herein generally relate to machine learning, and more particularly, to the automated progressive adjustment of machine learning models for, e.g., fault location in overhead transmission lines. Background Art

[0003] Artificial neural networks such as convolutional neural networks (CNNs) and deep neural networks (DNNs) can achieve satisfactory performance levels in various applications. However, the best-performing neural networks are typically large and complex. This size can impede implementation in applications that require stringent real-time execution with limited computational capabilities, such as when the neural network must be embedded in hardware.

[0004] One such application is power system protection, which must disconnect the faulty part of the power system in real time to limit physical damage to the rest of the power system. For example, a line protection system can be used to protect an overhead transmission line by tripping a circuit breaker as quickly as possible when a fault is detected within a predefined range setting. One task of such a line protection system is to estimate the location of the fault. A machine learning regression model can be trained to perform such fault location. However, the trained model must be installed in a relay, which has limited computational capabilities and thus cannot accommodate large machine learning models.

[0005] Relearning or "tuning" of machine learning models can be employed to quickly adapt a machine learning model to a new dataset. For example, in transfer learning, an existing machine learning model that has been previously trained for one task is retrained for a different but generally related task. Tuning a pre-trained model has several advantages compared to training from scratch. These advantages include faster learning, improved performance, and prevention of overfitting with respect to smaller datasets. However, reusing the same architecture as an existing machine learning model limits the ability of the machine learning model to relearn.

[0006] This limitation also exists in traditional pruning methods. For example, Pytortch TM and TensorFlow TM include model pruning libraries that perform local (i.e., layer-by-layer) and global pruning of machine learning models by setting and fixing weights and biases to zero. The weights and biases are selected randomly or based on importance, such as determined according to a metric (e.g., using L1 regularization). The resulting machine learning model is sparse, but there is not much performance loss. However, while such pruning can reduce the computational requirements when executing the machine learning model, it does not reduce the dimension of the pruned part of the machine learning model.

[0007] Yong-Deok et al., "Compression of Deep Convolutional Neural Networks for Fast and Low Power Mobile Applications", International Conference on Learning Representations (ICLR) 2016 is hereby incorporated by reference herein (as if set forth in full), which applies one-shot full-network compression, decomposing a D×R×S×T-dimensional tensor into slices of D×S1, R×S s , S×S3, and T×S4, and then checking the accuracy of the decomposed kernels with reduced sizes. This method has three main steps: (i) using global analytic variational Bayesian matrix factorization to analyze the principal subspaces of the modulo-3 and modulo-4 matrixizations of the kernel tensors of each layer; (ii) applying Tucker decomposition (also known as high-order singular value decomposition, SVD) to the kernel tensors of each layer with a previously determined rank; and (iii) using standard backpropagation to fine-tune the entire network.

[0008] Han et al., "Once-for-All: Train One Network and Specialize It for Efficient Deployment", ICLR 2020 is hereby incorporated by reference herein (as if set forth in full), which proposes to train large machine learning models and then select sub-networks. In this method, the training phase is decoupled from the specialization phase. In the training phase, the focus is on improving the accuracy of all sub-networks, which are obtained by selecting different parts of the once-for-all network. In the specialization phase, a subset of sub-networks is sampled to train the accuracy predictor and the latency predictor. Based on the target hardware and constraints, predictor-guided architecture search is performed to obtain a dedicated sub-network.

[0009] Yihui et al., "AMC: AutoML for Model Compression and Acceleration on Mobile Devices", arXiv 2019 is hereby incorporated by reference herein (as if set forth in full), which proposes a model compression method that uses reinforcement learning to provide model compression policies. First, an iterative pruning method is proposed, in which a certain configuration is selected from a candidate configuration set, and then the configuration is scheduled in training. The candidate configuration set is gradually pruned by removing low-performance configurations using a comparison between the estimated confidence bounds associated with various configurations. Second, for a given layer of a machine learning model, compression parameters are calculated based on the pruning loss of each hidden layer in the machine learning model. Third, the pruning method clusters the input weights according to their patterns and prunes each cluster to achieve a predetermined sparsity. Fourth, the Sparse Distillation Framework (SDF) distills knowledge from a computationally intensive teacher model while pruning the student model in a single training.

[0010] Broadly speaking, these publications and other publications can be classified into several types of adaptations: (i) selecting the best sub-model using expensive and slow processes (e.g., neural architecture search, reinforcement learning, etc.) through a supermodel; (ii) sparsifying a machine learning model; or (iii) scheduling the best candidate configurations for pruning based on performance estimates.

[0011] What is needed is an adaptation process that can obtain a dense machine learning model with reduced dimensions, which can be executed in real time on a device with limited computing power (such as a relay in a line protection system) to perform, for example, a regression task using time series data. SUMMARY OF THE INVENTION

[0012] Accordingly, systems, methods, and non-transitory computer-readable media for automated progressive adaptation of machine learning models are disclosed. The aim of certain embodiments is to rapidly and jointly adapt both the architecture and parameters of a machine learning model to environmental changes (e.g., environmental changes in terms of input-output data distribution and / or computing resources). Another aim of certain embodiments is to iteratively retrain the layers of a hierarchical machine learning model (such as a neural network) in a defined order until a stop condition is met. Another aim of certain embodiments is to prune a machine learning model until a specific reduction metric is achieved to produce a smaller dense machine learning model. Another aim of certain embodiments is to change the learning rates of different layers, for example, in the case of pruning, by decaying the learning rate after each iteration.

[0013] In an embodiment, a method includes using at least one hardware processor to: receive a machine learning model having an architecture and parameters that are trained to operate in a first environment; jointly adjust both the architecture and parameters of the machine learning model using a training dataset of a second environment that is different from the first environment; and deploy the adjusted machine learning model to the second environment.

[0014] The machine learning model may include a neural network. For example, the neural network may be a convolutional neural network.

[0015] Jointly adjusting both the architecture and parameters of the machine learning model may include: for each of one or more layers of the neural network, retraining the layer until a stopping condition is met. The one or more layers may be a plurality of layers. The plurality of layers may be trained in an order from a layer closest to the output of the machine learning model among the plurality of layers to another layer closest to the input of the machine learning model among the plurality of layers.

[0016] Jointly adjusting both the architecture and parameters of the machine learning model may include: in each of one or more iterations: selecting one layer from a plurality of layers of the neural network that has not been retrained in any previous iteration of the one or more iterations; retraining the selected layer; determining whether the stopping condition is met; when the stopping condition is met, stopping the joint adjustment; and when the stopping condition is not met, adding a next iteration to the one or more iterations. The one or more iterations may be a plurality of iterations, wherein each selected layer is retrained according to a learning rate, and wherein jointly adjusting both the architecture and parameters of the machine learning model further includes: in at least one iteration of the plurality of iterations, when the stopping condition is not met, changing the learning rate before the next iteration.

[0017] Retraining each selected layer may include: pruning the selected layer. The stopping condition may include a threshold indicating a measure of reduction in the size of the machine learning model. Each selected layer may be retrained according to a learning rate, wherein jointly adjusting both the architecture and parameters of the machine learning model further includes: in each of the one or more iterations, when the stopping condition is not met, reducing the learning rate before the next iteration.

[0018] Retraining the selected layer may include: selecting a plurality of alternative layers, each alternative layer having a different size from the selected layer; training the plurality of alternative layers using the training dataset to minimize the cross-entropy loss of the machine learning model; and selecting one alternative layer having the lowest error metric among the plurality of alternative layers as the retrained layer. The cross-entropy loss may include negative log-likelihood loss, wherein the error metric includes mean squared error. The size of each alternative layer among the plurality of alternative layers may be smaller than the selected layer. Each of the plurality of alternative layers may be selected to have a size different from the selected layer and the size of any other layer among the plurality of alternative layers within a restricted search space around the size of the selected layer. The size of each alternative layer among the plurality of alternative layers may be smaller than the selected layer, wherein the neural network is a convolutional neural network, and wherein the restricted search space is defined based on the number of filters to be pruned from the selected layer. The method may further include: using the at least one hardware processor to determine the number of filters to be pruned using principal component analysis.

[0019] The adjusted machine learning model may estimate the location of a fault on a power line based on one or more measurement parameters, wherein the training dataset includes labeled feature vectors, and wherein each of the labeled feature vectors includes a value of each of the one or more measurement parameters and is labeled with the fault location, wherein the joint adjustment includes pruning the machine learning model, wherein the second environment is a relay configured to trip a circuit breaker on the power line, and wherein deploying the adjusted machine learning model includes installing the adjusted machine learning model in a controller of the relay.

[0020] It should be understood that any feature in the above methods may be implemented alone or in any combination with any subset of other features. Thus, even though the appended claims indicate specific dependencies between features, the disclosed embodiments are not limited to these specific dependencies. Instead, any feature described herein may be combined with any other feature described herein or implemented in any combination of features without any one or more of the other features described herein. Additionally, any method described above and elsewhere herein may be embodied, alone or in any combination, in executable software modules in a processor-based system (such as a server) and / or in executable instructions stored in a non-transitory computer-readable medium. Brief Description of the Drawings

[0021] By studying the accompanying drawings, details of both the structure and operation of the present invention can be partially gathered, in which like reference numerals refer to like parts, and in the drawings:

[0022] Figure 1 illustrates an example infrastructure in which one or more processes described herein can be implemented, according to an embodiment;

[0023] Figure 2 illustrates an example processing system that can be used to perform one or more processes described herein, according to an embodiment;

[0024] Figure 3 illustrates an environment of a line protection system, according to an embodiment;

[0025] Figure 4 illustrates a process for automated progressive adjustment, according to an embodiment;

[0026] Figure 5 illustrates an example of a process for adjusting a hierarchical machine learning model, according to an embodiment;

[0027] Figure 6 illustrates an example of a retraining process, according to an embodiment;

[0028] Figure 7 illustrates the operations of stop conditions and decreasing learning rates during automated progressive pruning, according to an embodiment;

[0029] Figure 8 illustrates a process for making a tripping decision, according to an embodiment;

[0030] Figures 9A to 9D illustrates blind spots through various example implementations of a machine learning model for fault location, according to an experiment; and

[0031] Figure 10 is a graph of the convergence time of embodiments of automated progressive pruning and brute-force methods, according to an experiment. Detailed Description

[0032] In an embodiment, systems, methods, and non-transitory computer-readable media for automated progressive adjustment of a machine learning model are disclosed. After reading this specification, those skilled in the art will be clear how to implement the present invention in various alternative embodiments and alternative applications. However, while various embodiments of the present invention will be described herein, it should be understood that these embodiments are presented by way of example and illustration only, and not by way of limitation. Therefore, this detailed description of the various embodiments should not be construed as limiting the scope or breadth of the present invention as set forth in the appended claims.

[0033] 1. Example infrastructure

[0034] Figure 1 Illustrates an example infrastructure in which one or more of the disclosed processes may be implemented. The infrastructure may include a platform 110 (e.g., one or more servers) that hosts and / or executes one or more of the various functions, processes, methods, and / or software modules described herein. Platform 110 may include dedicated servers or, alternatively, may be implemented in a computing cloud where resources of one or more servers are dynamically and elastically allocated to multiple tenants based on demand. In either case, the servers may be collocated and / or geographically distributed. Platform 110 may also include or be communicatively coupled to software 112 and / or one or more databases 114. Additionally, platform 110 may be communicatively coupled to one or more user systems 130 via one or more networks 120. Platform 110 may also be communicatively connected to one or more external systems 140 via one or more networks 120.

[0035] (Multiple) networks 120 may include the Internet, and platform 110 may communicate with (multiple) user systems 130 via the Internet using standard transport protocols as well as proprietary protocols, such as the Hypertext Transfer Protocol (HTTP), HTTP Secure (HTTPS), File Transfer Protocol (FTP), FTP Secure (FTPS), Secure Shell FTP (SFTP), etc. Although platform 110 is illustrated as being connected to various systems via a single set of (multiple) networks 120, it should be understood that platform 110 may be connected to various systems via different sets of one or more networks. For example, platform 110 may be connected to a subset of user systems 130 and / or external systems 140 via the Internet, but may also be connected to one or more other user systems 130 and / or external systems 140 via an intranet. Additionally, although only a few user systems 130 and external systems 140, a single set of software 112, and a single set of (multiple) databases 114 are illustrated, it should be understood that the infrastructure may include any number of user systems, external systems, software applications, and databases.

[0036] (Multiple) user systems 130 may include any one or more types of computing devices capable of wired and / or wireless communication, including but not limited to desktop computers, laptop computers, tablet computers, smartphones or other mobile phones, servers, game consoles, televisions, set-top boxes, electronic self-service terminals, and / or point-of-sale terminals, etc. However, it is generally envisioned that user system 130 will include a personal computer or workstation of an agent (such as an operator of a power system or a developer of a line protection system of a power system) of an organization responsible for designing machine learning models. Each user system 130 may include or be communicatively connected to a client application 132 and / or one or more local databases 134.

[0037] In an embodiment, (multiple) external systems 140 may include one or more devices to which a machine learning model (such as a neural network (e.g., CNN, DNN, etc.)) is deployed. For example, external system 130 may be a line protection system including relays for overhead transmission lines, or other protection systems. Although such systems may be any type of system performing any type of task, the disclosed embodiments may be particularly beneficial to real-time systems with limited computing capabilities (such as limited processing resources, limited volatile and / or non-volatile memory, etc.).

[0038] Platform 110 may include a web server hosting one or more websites and / or web services. In embodiments where a website is provided, the website may include a graphical user interface including, for example, one or more screens (e.g., web pages) generated in Hypertext Markup Language (HTML) or other languages. Platform 110 transmits or provides one or more screens of the graphical user interface in response to requests from user system(s) 130. In some embodiments, these screens may be provided in the form of a wizard, in which case two or more screens may be provided in a sequential manner, and one or more sequential screens may depend on the interaction of the user or user system 130 with one or more previous screens. Requests to platform 110 and responses from platform 110 (including screens of the graphical user interface) may both be transmitted via network(s) 120 using standard communication protocols (e.g., HTTP, HTTPS, etc.), which may include the Internet. These screens (e.g., web pages) may include a combination of content and elements such as text, images, video, animation, references (e.g., hyperlinks), frames, inputs (e.g., text boxes, text areas, check boxes, radio buttons, drop-down menus, buttons, tables, etc.), scripts (e.g., JavaScript), etc., including elements formed or derived from data stored in one or more databases (e.g., database(s) 114) accessible locally and / or remotely to platform 110. It should be understood that platform 110 may also respond to other requests from user system(s) 130.

[0039] Platform 110 may include one or more databases 114, be communicatively coupled with, or otherwise access, the one or more databases. For example, platform 110 may include one or more database servers managing one or more databases 114. Software 112 executing on platform 110 and / or client application 132 executing on user system 130 may submit data (e.g., user data, table data, etc.) to be stored in database(s) 114, and / or request access to data stored in database(s) 114. Any suitable database may be utilized, including but not limited to MySQL TM 、Oracle TM 、IBM TM 、Microsoft SQL TM 、Access TM 、PostgreSQL TM 、MongoDB TMetc., including cloud-based databases and proprietary databases. For example, data can be sent to platform 110 using well-known POST requests supported by HTTP, via FTP, etc. This data and other requests can be processed, for example, by server-side network technologies executed by platform 110, such as server applets or other software modules (e.g., included in software 112).

[0040] In embodiments that provide network services, platform 110 can receive requests from external system(s) 140 and provide responses in Extensible Markup Language (XML), JavaScript Object Notation (JSON), and / or any other suitable or desired format. In such embodiments, platform 110 can provide an application programming interface (API) that defines the manner in which user system(s) 130 and / or external system(s) 140 can interact with the network service. Thus, user system(s) 130 and / or external system(s) 140 (which can themselves be servers) can define their own user interfaces and rely on the network service to implement or otherwise provide the backend processes, methods, functions, and / or storage, etc., described herein. For example, in such embodiments, a client application 132 executing on one or more user systems 130 can interact with software 112 executing on platform 110 to perform one or more or a portion of one or more of the various functions, processes, methods, and / or software modules described herein.

[0041] The client application 132 can be a "thin" client application, in which case the processing is mainly performed on the server side by the software 112 on the platform 110. A basic example of a thin client application 132 is a browser application that simply requests, receives, and presents web pages at the (multiple) user systems 130, while the software 112 on the platform 110 is responsible for generating the web pages and managing database functions. Alternatively, the client application can be a "thick" client application, in which case the processing is mainly performed on the client side by the (multiple) user systems 130. It should be understood that depending on the design goals of a particular implementation, the client application 132 can perform the processing volume relative to the software 112 on the platform 110 at any point within the range between such "thin" and "thick". In any case, the software described herein can reside entirely on the platform 110 (e.g., in which case the software 112 performs all processing) or on the (multiple) user systems 130 (e.g., in which case the client application 132 performs all processing) or be distributed between the platform 110 and the (multiple) user systems 130 (e.g., in which case both the software 112 and the client application 132 perform processing), and can include one or more executable software modules that include instructions implementing one or more of the processes, methods, or functions described herein.

[0042] 2. Example processing device

[0043] Figure 2 FIG. is a block diagram illustrating an example wired or wireless system 200 that can be used in conjunction with various embodiments described herein. For example, the system 200 can be used as or in conjunction with one or more of the functions, processes, or methods described herein (e.g., for storing and / or executing software), and can represent components of the platform 110, the (multiple) user systems 130, the (multiple) external systems 140 (e.g., line protection systems), and / or other processing devices described herein. The system 200 can be a server or any conventional personal computer, or any other processor-enabled device capable of wired or wireless data communication. It will be apparent to those skilled in the art that other computer systems and / or architectures can also be used.

[0044] System 200 preferably includes one or more processors 210. The (multiple) processors 210 may include a central processing unit (CPU). Additional processors may be provided, such as a graphics processing unit (GPU), an auxiliary processor for managing input / output, an auxiliary processor for performing floating-point mathematical operations, a dedicated microprocessor having an architecture suitable for quickly executing signal processing algorithms (e.g., a digital signal processor), a slave processor subordinate to the main processing system (e.g., a backend processor), an additional microprocessor or controller and / or a coprocessor for a dual-processor or multi-processor system. Such auxiliary processors may be discrete processors or may be integrated with the processor 210. Examples of processors that may be used with the system 200 include, but are not limited to, any processor provided by Intel Corporation of Santa Clara, California (e.g., Pentium TM , Core i7 TM , Xeon TM , etc.), any processor provided by Advanced Micro Devices, Inc. (AMD) of Santa Clara, California, any processor provided by Apple Inc. of Cupertino (e.g., A series, M series, etc.), any processor provided by Samsung Electronics Co., Ltd. of Seoul, Korea (e.g., Exynos TM ), any processor provided by NXP Semiconductors N.V. of Eindhoven, the Netherlands, and so on.

[0045] Processor 210 is preferably connected to a communication bus 205. The communication bus 205 may include data channels for facilitating the transfer of information between the storage devices of the system 200 and other peripheral components. In addition, the communication bus 205 may provide a set of signals for communicating with the processor 210, including a data bus, an address bus, and / or a control bus (not shown). The communication bus 205 may include any standard or non-standard bus architecture, such as an Industry Standard Architecture (ISA), an Extended Industry Standard Architecture (EISA), a Micro Channel Architecture (MCA), a Peripheral Component Interconnect (PCI) local bus, a bus architecture that complies with standards published by the Institute of Electrical and Electronics Engineers (IEEE), including the IEEE 488 General-Purpose Interface Bus (GPIB), IEEE 696 / S-100, etc.

[0046] System 200 preferably includes a main memory 215 and may also include an auxiliary memory 220. The main memory 215 provides instruction and data storage for programs (such as any software discussed herein) executed on the processor 210. It should be understood that the programs stored in the memory and executed by the processor 210 can be written and / or compiled in any suitable language, including but not limited to C / C++, Java, JavaScript, Perl, Visual Basic, NET, etc. The main memory 215 is typically a semiconductor-based memory, such as dynamic random access memory (DRAM) and / or static random access memory (SRAM). Other semiconductor-based memory types include, for example, synchronous dynamic random access memory (SDRAM), Rambus dynamic random access memory (RDRAM), ferroelectric random access memory (FRAM), etc., including read-only memory (ROM).

[0047] The auxiliary memory 220 is a non-transitory computer-readable medium on which computer-executable code (e.g., any software disclosed herein) and / or other data are stored. The computer software or data stored on the auxiliary memory 220 is read into the main memory 215 for execution by the processor 210. The auxiliary memory 220 may include, for example, semiconductor-based memories, such as programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable read-only memory (EEPROM), and flash memory (a block-oriented memory similar to EEPROM).

[0048] The auxiliary memory 220 may optionally include an internal medium 225 and / or a removable medium 230. Reading and / or writing to the removable medium 230 is performed in any well-known manner. The removable storage medium 230 may be, for example, a tape drive, a compact disc (CD) drive, a digital versatile disc (DVD) drive, other optical drives, a flash memory drive, etc.

[0049] In an alternative embodiment, the auxiliary memory 220 may include other similar devices for allowing a computer program or other data or instructions to be loaded into the system 200. Such devices may include, for example, a communication interface 240 that allows software and data to be transferred from an external storage medium 245 to the system 200. Examples of the external storage medium 245 include an external hard drive, an external optical drive, an external magneto-optical drive, etc.

[0050] As mentioned above, system 200 may include a communication interface 240. The communication interface 240 allows software and data to be transferred between system 200 and external devices (such as printers), networks, or other information sources. For example, computer software or executable code may be transferred from a network server (such as platform 110) to system 200 via the communication interface 240. Examples of the communication interface 240 include built-in network adapters, network interface cards (NICs), Personal Computer Memory Card International Association (PCMCIA) network cards, CardBus network adapters, wireless network adapters, Universal Serial Bus (USB) network adapters, modems, wireless data cards, communication ports, infrared interfaces, IEEE 1394 FireWire, and any other device capable of coupling system 200 to a network (such as (multiple) networks 120) or another computing device. The communication interface 240 preferably implements industry-released protocol standards, such as Ethernet IEEE 802 standards, Fibre Channel, Digital Subscriber Line (DSL), Asynchronous Digital Subscriber Line (ADSL), Frame Relay, Asynchronous Transfer Mode (ATM), Integrated Services Digital Network (ISDN), Personal Communication Service (PCS), Transmission Control Protocol / Internet Protocol (TCP / IP), Serial Line Internet Protocol / Point-to-Point Protocol (SLIP / PPP), etc., but may also implement custom or non-standard interface protocols.

[0051] The software and data transferred via the communication interface 240 typically take the form of electrical communication signals 255. These signals 255 may be provided to the communication interface 240 via a communication channel 250. In an embodiment, the communication channel 250 may be a wired or wireless network (such as (multiple) networks 120), or any other type of communication link. The communication channel 250 carries the signals 255 and may be implemented using various wired or wireless communication means, including wires or cables, optical fibers, conventional telephone lines, cellular phone links, wireless data communication links, radio frequency (“RF”) links, or infrared links, to name a few.

[0052] Computer-executable code (such as a computer program, such as the disclosed software) is stored in the main memory 215 and / or the secondary memory 220. The computer-executable code may also be received via the communication interface 240 and stored in the main memory 215 and / or the secondary memory 220. Such a computer program, when executed, enables system 200 to perform the various functions of the disclosed embodiments described elsewhere herein.

[0053] In this specification, the term "computer-readable medium" is used to refer to any non-transitory computer-readable storage medium for providing computer-executable code and / or other data to or within system 200. Examples of such media include main memory 215, secondary memory 220 (including internal memory 225 and / or removable media 230), external storage media 245, and any peripheral device (including a network information server or other network device) communicatively coupled to communication interface 240. These non-transitory computer-readable media are means for providing software and / or other data to system 200.

[0054] In software-implemented embodiments, the software can be stored on a computer-readable medium and loaded into system 200 via removable media 230, I / O interface 235, or communication interface 240. In such embodiments, the software is loaded into system 200 in the form of an electrical communication signal 255. The software, when executed by processor 210, preferably causes processor 210 to perform one or more of the processes and functions described elsewhere herein.

[0055] In an embodiment, I / O interface 235 provides an interface between one or more components of system 200 and one or more input and / or output devices. Example input devices include, but are not limited to, sensors, keyboards, touchscreens or other touch-sensitive devices, cameras, biometric sensing devices, computer mice, trackballs, pen-based pointing devices, etc. Examples of output devices include, but are not limited to, other processing devices, cathode ray tubes (CRTs), plasma displays, light emitting diode (LED) displays, liquid crystal displays (LCDs), printers, vacuum fluorescent displays (VFDs), surface conduction electron emission displays (SEDs), field emission displays (FEDs), etc. In some cases, the input and output devices can be combined, such as in the case of a touch panel display (e.g., in a smartphone, tablet computer, or other mobile device).

[0056] System 200 may also include optional wireless communication components that facilitate wireless communication via a voice network and / or a data network (e.g., in the case of user system 130). The wireless communication components include antenna system 270, radio system 265, and baseband system 260. In system 200, radio frequency (RF) signals are transmitted and received through the air by antenna system 270 under the management of radio system 265.

[0057] In an embodiment, the antenna system 270 may include one or more antennas and one or more multiplexers (not shown) that perform a switching function to provide transmit and receive signal paths to the antenna system 270. In the receive path, the received RF signal may be coupled from the multiplexer to a low noise amplifier (not shown), which amplifies the received RF signal and sends the amplified signal to the radio system 265.

[0058] In an alternative embodiment, the radio system 265 may include one or more radios configured to communicate at various frequencies. In an embodiment, the radio system 265 may combine a demodulator (not shown) and a modulator (not shown) in one integrated circuit (IC). The demodulator and modulator may also be separate components. In the incoming path, the demodulator removes the RF carrier signal, leaving a baseband received audio signal that is sent from the radio system 265 to the baseband system 260.

[0059] If the received signal contains audio information, the baseband system 260 decodes the signal and converts it to an analog signal. The signal is then amplified and sent to the speaker. The baseband system 260 also receives analog audio signals from the microphone. These analog audio signals are converted to digital signals and encoded by the baseband system 260. The baseband system 260 also encodes the digital signals for transmission and generates a baseband transmitted audio signal that is routed to the modulator section of the radio system 265. The modulator mixes the baseband transmitted audio signal with the RF carrier signal, thereby generating an RF transmitted signal that is routed to the antenna system 270 and may pass through a power amplifier (not shown). The power amplifier amplifies the RF transmitted signal and routes it to the antenna system 270, where the signal is switched to an antenna port for transmission.

[0060] The baseband system 260 is also communicatively coupled to the processor(s) 210. The processor(s) 210 may access data storage areas 215 and 220. The processor(s) 210 is preferably configured to execute instructions (i.e., computer programs such as the disclosed software) that may be stored in the main memory 215 or the secondary memory 220. The computer program may also be received from the baseband processor 260 and stored in the main memory 210 or the secondary memory 220, or executed upon receipt. Such a computer program, when executed, may enable the system 200 to perform the various functions of the disclosed embodiments.

[0061] 3. Example line protection system

[0062] Figure 3Illustrates the operation of a line protection system 330 according to an embodiment. The line protection system 330 is an example of an external system 140 to which a machine learning model developed on the platform 110 can be deployed. A power line 310 can be provided between two substations 320A and 320B. One or more voltage measurement units 312 can measure the voltage on the power line 310 and output the voltage measurement results to the line protection system 330. Additionally, one or more current measurement units 314 can measure the current on the power line 310 and output the current measurement results to the line protection system 330.

[0063] The line protection system 330 can include a controller 332, which can implement the system 200 or a subset thereof. For example, the controller 332 can include one or more processors 210 and store a machine learning model 334 in a memory (e.g., main memory 215 and / or secondary memory 220). The controller 332 can receive voltage measurement results from the (multiple) voltage measurement units 312, current measurement results from the (multiple) current measurement units 314, and / or other measurement results from one or more other sensors or devices associated with the power line 310. The controller 332 can also derive one or more additional measurement results from the received measurement results.

[0064] The controller 332 can apply the machine learning model 334 to the measurement results as part of a time-domain protection scheme. For example, the measurement results can be input into the machine learning model 334 to estimate the location of a fault 340 or output a decision on whether to trip (i.e., open) a circuit breaker 316 on the power line 310. If the location of the fault 340 is within the protection zone defined by a range setting 336 of the line protection system 330, the controller 332 can determine to trip the circuit breaker 316 (i.e., switch the circuit breaker 316 to the open state). Otherwise, the controller 332 can determine not to trip the circuit breaker 316 (i.e., maintain the circuit breaker 316 in the closed state). When the controller 332 determines to trip the circuit breaker 316, the controller 332 can send a control signal to the circuit breaker 316 (e.g., directly or via a relay to control the circuit breaker 316) to open the circuit breaker 316, and thereby isolate the fault 340 from the substation 320A. It should be understood that a similar line protection system 330 may exist at the other end of the power line 110, which performs the same function to isolate the fault 340 from the substation 320B if the location of the fault 340 is within its corresponding protection zone.

[0065] 4. Automated progressive adjustment

[0066] A method for Automated Progressive Tuning (APT) will now be described in detail. In an embodiment, Automated Progressive Tuning utilizes joint learning of the architecture and parameters of a machine learning model to train the machine learning model. In the case of a hierarchical machine learning model, such as a neural network, this joint training can be iterated over the various layers of the machine learning model until a stopping condition is met. It should be understood that in the context of "joint tuning", "joint learning", "joint training", etc., the term "joint" refers to simultaneously tuning, learning, training, etc. the architecture of the machine learning model and the parameters of the machine learning model. This is in contrast to, for example, learning the parameters with a fixed architecture or learning the architecture with fixed parameters. One purpose of Automated Progressive Tuning is to address the limitations of restricted architectures faced by traditional tuning methods. The disclosed Automated Progressive Tuning can be used for environment-specific tuning, which is less costly and faster than traditional tuning methods, such as Neural Architecture Search (NAS), which searches for the optimal architecture for general tasks and is known to be generally very expensive and slow.

[0067] Figure 4 FIG. illustrates a process 400 for Automated Progressive Tuning according to an embodiment. It should be understood that process 400 can be embodied in one or more software modules executed by one or more hardware processors (e.g., processor 210), e.g., as a software application (e.g., software 112, client application 132, and / or a distributed application including both software 112 and client application 132), which can be executed entirely by the (multiple) processors of platform 110, entirely by the (multiple) processors of (multiple) user systems 130, or can be distributed across platform 110 and (multiple) user systems 130 such that some parts or modules of the software application are executed by platform 110 and other parts or modules of the software application are executed by the (multiple) user systems 130. The described process can be implemented as instructions represented in source code, object code, and / or machine code. These instructions can be executed directly by the (multiple) hardware processors 210 or, alternatively, can be executed by a virtual machine running between the object code and the (multiple) hardware processors 210. Additionally, the disclosed software can be built on or interface with one or more existing systems.

[0068] Alternatively, the described processes may be implemented as hardware components (e.g., general purpose processors, integrated circuits (ICs), application specific integrated circuits (ASICs), digital signal processors (DSPs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, etc.), combinations of hardware components, or combinations of hardware and software components. For clarity of explanation of the interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps are generally described herein in terms of functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be construed as causing a departure from the scope of the present invention. Additionally, the functional grouping within a component, block, module, circuit, or step is for ease of description. Without departing from the present invention, a particular function or step may be moved from one component, block, module, circuit, or step to another.

[0069] Furthermore, although process 400 is illustrated by a certain arrangement and ordering of sub - processes, process 400 may be implemented with fewer, more, or different sub - processes and different arrangements and / or orderings of sub - processes. Additionally, it should be understood that any sub - process that does not depend on the completion of another sub - process may be executed before, after, or in parallel with that other independent sub - process, even if the sub - processes are described or illustrated in a particular order.

[0070] Initially, in sub - process 410, a machine learning model having an architecture and parameters that have been trained to operate in a first environment is received. The machine learning model may be an artificial neural network, which may be referred to herein simply as a “neural network”. A neural network is any collection of connected units or nodes (referred to as “neurons”) that loosely mimics the neurons in a biological brain in that these neurons send signals to other neurons via connections (referred to as “edges”). As an example, the neural network may be a convolutional neural network (CNN) and / or a deep neural network (DNN). A convolutional neural network is a specialized type of neural network that includes an input layer, one or more hidden layers, and an output layer, and uses convolutions instead of general matrix multiplication in at least one of these layers. A deep neural network is any neural network that includes multiple hidden layers. Thus, a convolutional neural network may be a deep neural network and vice versa. However, it should be understood that the machine learning model may include other types of neural networks (such as recurrent neural networks (RNNs)) and / or other types of models (such as generalized linear models, spline interpolation models, etc.).

[0071] In sub - process 420, the training dataset of the second environment is used to jointly adjust both the architecture and parameters of the machine - learning model received in sub - process 410. Although not required, it is generally envisioned that the second environment will be different from the first environment. Even if the environments are different, the tasks performed in each environment can be the same. For example, the first environment can be the first input - output data distribution of a task, while the second environment can be the second input - output data distribution of the same task, but the second input - output data distribution is different from the first input - output data distribution. As another example, the first environment can be a computationally relatively rich environment (e.g., a server - based platform such as platform 110), while the second environment can be a computationally relatively poor environment with computational power significantly lower than that of the first environment (e.g., an embedded system (such as line protection system 330), a mobile system (such as a smart phone), an Internet of Things (IoT) device, etc.). In such a case, sub - process 420 can include performing pruning to reduce the size of the machine - learning model, as discussed in more detail elsewhere in this document. The differences in computational power can include differences in processing resources, memory resources, available power (e.g., plugged - in vs. battery - powered), and / or network bandwidth, etc.

[0072] In sub - process 430, the adjusted machine - learning model is deployed to the second environment. Sub - process 430 can include installing the adjusted machine - learning model on one or more external systems 140. For example, the adjusted machine - learning model can be trained to determine the fault location on a power line based on measurements of the power line (e.g., voltage and / or current), and the external system(s) 140 can include one or more line protection systems 330. In such a case, the adjusted and preferably pruned machine - learning model can be installed in the memory (e.g., main memory 215 and / or secondary memory 220) of the controller 332 of the line protection system 330. It should be understood that the machine - learning model can be trained, adjusted, and / or deployed to any other type of external system 140 in a similar manner for any other task.

[0073] Figure 5 An example of sub - process 420 of a hierarchical machine - learning model according to an embodiment is illustrated. In the illustrated example, the joint adjustment of both the architecture and parameters of the machine - learning model is progressive because the process adjusts the machine - learning model layer by layer until a stopping condition is met. Similarly, examples of hierarchical machine - learning models include, but are not limited to, convolutional neural networks, deep neural networks, recurrent neural networks, and many other types of neural networks.

[0074] In sub - process 422, it is determined whether a stop condition is met. The stop condition can include one or more stop criteria, and the one or more stop criteria can be defined according to the purpose(s) of a particular tuning process. For example, in the case where the tuning includes pruning, the stop condition can be that a measure of the size reduction of the machine - learning model meets a threshold. As an alternative example, the stop condition can include that another metric of the tuned machine - learning model meets a threshold, the number of iterations exceeds a predefined number of cycles, a user action is received (e.g., terminating the tuning via a graphical user interface provided by software 112), and / or a timer expires, etc. In any case, if the stop condition is not met (i.e., "no" in sub - process 422), then sub - process 420 proceeds to sub - process 424. Otherwise, if the stop condition is met (i.e., "yes" in sub - process 422), then sub - process 420 can end (e.g., and proceed to sub - process 430).

[0075] In sub - process 424, a new layer is selected for tuning. In particular, one layer can be selected from among multiple layers in a machine - learning model (e.g., a neural network) that have not been retrained in any previous iteration of sub - process 426. In an embodiment, the multiple layers can be trained in a defined order from the layer closest to the output of the machine - learning model (i.e., farthest from the input), which can be referred to as the "top - most" layer, to the layer closest to the input of the machine - learning model (i.e., farthest from the output), which can be referred to as the "bottom - most" layer. In an alternative embodiment, the multiple layers can be trained in order from the bottom - most layer to the top - most layer. In an alternative embodiment, the multiple layers can be trained from a certain higher layer (e.g., other than the top - most layer) to a certain lower layer (e.g., other than the bottom - most layer) or from a certain lower layer to a certain higher layer. In yet another alternative, the multiple layers can be trained in another order (e.g., in the order of every Nth layer from higher layer to lower layer or from lower layer to higher layer according to the priority assigned to each layer, etc.) or in a random order. As yet another alternative, layers can be selected in each iteration of sub - process 424 according to one or more selection criteria. In this case, the selection criteria can include that a measure of the importance or other property of the layer is the minimum or maximum value relative to all layers that have not been retrained in any previous iteration of sub - process 426.

[0076] In sub - process 426, the layer selected in sub - process 424 is retrained. The retraining in sub - process 426 includes jointly retraining both the architecture and the parameters of the layer. In other words, both the architecture and the parameters of the layer can be changed during retraining. In an embodiment, such retraining can include pruning, in which case the size of the layer is reduced in at least one dimension. Once the layer has been retrained, sub - process 420 returns to sub - process 422 to determine whether the stop condition has been met after retraining.

[0077] Each layer can be retrained in sub - process 426 according to a learning rate. In an embodiment, when a layer has been retrained in sub - process 426 and the stop condition has not been met yet in sub - process 422, the learning rate can be changed before the next iteration of sub - process 424 or 426. For example, in the case where the retraining includes pruning, the learning rate can be decreased according to a decay function before the next iteration. Alternatively, the learning rate can be increased before the next iteration. As another alternative, the learning rate can be kept constant for all iterations. As yet another alternative, the learning rate can be varied up and down or kept the same for each iteration according to one or more criteria or patterns.

[0078] The following algorithm represents an example of pseudo - code for an embodiment of sub - process 420 for automated progressive adjustment:

[0079]

[0080]

[0081] Figure 6 An example of retraining in sub - process 426 according to an embodiment is illustrated. In the illustrated example, layer i of neural network 600 is retrained during an iteration of sub - process 426. Although neural network 600 is illustrated as including at least a specific number of layers, it should be understood that neural network 600 can include any number of layers, including fewer or more layers than those illustrated. Additionally, the layer being retrained can be any layer of neural network 600. For example, while a particular implementation may retrain the convolutional layers of a convolutional neural network, the same retraining can be performed on other layers of a convolutional neural network or other neural networks, such as gated recurrent units (GRUs) (e.g., in a recurrent neural network), fully - connected (FC) layers, etc.

[0082] Layer i of the existing neural network 600 has a specific size, which can be defined as the number of units or filters in the context of a convolutional neural network. In an embodiment, adjusting the architecture of layer i includes changing the number of input units and output units, and the number of input units of the subsequent layer i+1. To this end, sub-process 426 can select a plurality of alternative layers, each alternative layer having a size different from that of the existing layer i within a restricted search space around the size of the existing layer i, and the sizes of these alternative layers are different from each other. In the illustrated example, the plurality of alternative layers are represented as layer i+11, layer i+12,... layer i+1 n , where each layer has a different size. The number n of alternative layers can be a predefined number greater than one. Each of these plurality of alternative layers represents a different trainable path from layer i-1 to layer i+1, where each trainable path has a different set of input-output structures that can be learned simultaneously.

[0083] In an embodiment where sub-process 426 is pruning layer i, the restricted search space can be defined based on the number of units (e.g., the number of filters) to be pruned from layer i. In this case, the sizes of all alternative layers will be smaller than the existing layer i. In an embodiment, principal component analysis (PCA) can be used to determine the number of units to be pruned. Principal component analysis can be used to calculate the principal components of layer i, and the search space can be defined to include the number of units starting from (and including) the number of principal components up to (but not including) the number of units in the existing layer i.

[0084] In sub-process 426, the training dataset of the second environment can be used to train the plurality of alternative layers to minimize the cross-entropy loss of the machine learning model (e.g., neural network 600). To jointly adjust the architecture and parameters of the machine learning model, the cross-entropy loss can be added to the mean squared error (MSE) loss of the task of the entire machine learning model (e.g., a regression task for fault location). To this end, the following joint optimization problem can be solved:

[0085] Equation (1):

[0086]

[0087] where, represents the loss function, X is the input data (e.g., represented as a matrix), y is the target label (e.g., represented as a vector), f ω,θ is the learnable machine learning model, ω is the parameter representing the architecture, and θ is the parameter representing the model parameters of the architecture ω.

[0088] The training loss can be extended as follows:

[0089] Equation (2):

[0090]

[0091] Where NLL(·) represents the negative log-likelihood loss, and MSE(·) represents the mean squared error. β1 and β2 are hyperparameters, each having a value greater than zero, representing the relative importance of the learning architecture and the learning parameters. The values of β1 and β2 can vary according to the needs of the application. For example, β1 = 1.0 and β2 = 1.5 will focus the optimization problem on reducing the estimated average MSE value.

[0092] The architecture parameter ω refers to the number of learnable alternative layers. Thus, represents the average MSE value from all alternative layers.

[0093] The target y of the NLL classifier class can be obtained as follows:

[0094] Equation (3):

[0095] y class = argmin ω (MSE(y, f ω,θ (X)))

[0096] It should be noted that instead of the continuous parameterization ω of the architecture, in the embodiment, a set of discrete sizes of alternative layers is used as the architecture parameter ω. Thus, the training loss is used to learn the architecture as a classification loss. In particular, the negative log-likelihood loss of the MSE loss obtained through the alternative layers (i.e., different architectures) is used as the overall training loss. During the retraining of layer i in subprocess 426, the other layers receive more updates regarding the alternative layer presenting the lowest MSE among all the multiple alternative layers. One layer presenting the lowest cross-entropy loss among the multiple alternative layers can be selected as the new architecture of layer i.

[0097] During the iterations of sub - process 426, the learning rate can vary or remain constant. In an embodiment, the retraining in sub - process 420 proceeds from the layers near the top (e.g., the top - most layer, the last convolutional layer in a convolutional neural network, or other higher layers) towards the layers near the bottom (e.g., the bottom - most layer, the first convolutional layer in a convolutional neural network, or other lower layers), where the learning rate gradually decreases. In other words, the higher layers are adjusted more aggressively than the lower layers. The learning rate can decay linearly at each iteration. Alternatively, the learning rate can decay according to other suitable curves (including non - linear curves). As another alternative, the learning rate can gradually increase linearly or non - linearly. As yet another alternative, the learning rate can remain constant during all the iterations of sub - process 426. As even yet another alternative, the learning rate can decrease, increase, or remain constant iteratively according to one or more criteria.

[0098] In some applications (e.g., when the existing machine - learning model is very large), it may be sufficient to stop the adjustment before reaching the bottom - most layer in the sequence of layers to be trained, for example, if the size of the adjusted machine - learning model at that point satisfies a certain pruning objective. In such a case, continuing to retrain the layers is wasteful. Thus, sub - process 422 can be used to define a stopping condition such that once one or more stopping criteria are met, the adjustment stops even if only a portion of the layers have been retrained.

[0099] 5. Automated progressive pruning

[0100] A direct way to produce a smaller dense machine - learning model using traditional methods is to define the architecture as a hyper - parameter and use a method for tuning the hyper - parameter to adjust the architecture. However, when significantly reducing the size of a machine - learning model (e.g., when reducing from a very large existing machine - learning model to a very small machine - learning model), such a method is expensive and wasteful because it retrains all layers for all candidate architectures.

[0101] Thus, in an embodiment, the automated progressive adjustment of process 400 is used to prune an existing machine - learning model by adding constraints to at least one dimension of the existing machine - learning model. In particular, at least one dimension of the adjusted machine - learning model is constrained to be strictly smaller than that dimension in the existing machine - learning model. As an example of a convolutional neural network, the dimension can be the kernel and filter sizes or the number of filters in a layer of the convolutional neural network. Advantageously, the pruning performed by process 400 can produce a smaller dense machine - learning model without the expensive and wasteful retraining of the above - mentioned direct method. The pruning performed by the automated progressive adjustment of process 400 is referred to herein as automated progressive pruning (APP).

[0102] In an embodiment of automated progressive pruning, the stopping condition in sub - process 422 can be that the size reduction metric of at least one dimension from an existing machine - learning model to the adjusted machine - learning model meets a threshold. The reduction metric can be an overall size reduction metric, such as the ratio of the number of parameters of the adjusted machine - learning model to the number of parameters of the existing machine - learning model. In this case, when the ratio drops below the threshold, the stopping condition can be met. Alternatively, the size reduction metric can be the radius of a circular neighborhood around the desired number of parameters in the finally adjusted machine - learning model. In this case, when the radius drops below the threshold, the stopping condition can be met.

[0103] The following algorithm represents example pseudocode for an embodiment of automated progressive pruning:

[0104]

[0105]

[0106] Compared with traditional methods, embodiments of automated progressive pruning start from an existing machine - learning model and jointly learn the architecture and model parameters in an iterative or progressive manner. Automated progressive pruning changes at least one dimension of the existing architecture while optimizing the parameters, rather than decoupling the learning of the architecture from the optimization of the parameters and searching for the architecture (which can be expensive and slow). Thus, automated progressive pruning addresses the limitations of traditional methods that reuse the same architecture of an existing machine - learning model.

[0107] In an embodiment, the adjustment in sub - process 420 starts from higher layers (i.e., closer to the output of the machine - learning model) and moves towards lower layers (i.e., closer to the input of the machine - learning model) with each iteration. After each iteration, the learning rate can be decreased before the next iteration. The learning rate can be gradually decreased according to some algorithm. This gradual decrease in the learning rate is because the inventors recognize that for the task of implementing a machine - learning model, higher layers (i.e., closer to the output of the machine - learning model) are more likely to be redundant than lower layers (i.e., closer to the input of the machine - learning model). Therefore, higher layers can be pruned more aggressively, while the role of lower layers is mainly to embed input data, which requires dense connectivity for effective embedding.

[0108] Figure 7 Illustrates the operations of the stopping condition and the gradually decreasing learning rate during automated progressive pruning according to an embodiment. In this illustration, it is assumed that sub - process 420 iterates from higher layers to lower layers. In iteration j of sub - process 420, a layer (e.g., Figure 6 layer i inFigure 6 The architectures and parameters of layers i and i+1) in are jointly retrained, where α is the initial learning rate. After the retraining in sub-process 426, the reduction metric for the next iteration of sub-process 422 is calculated In sub-process 422, the stopping condition can be satisfied when the following equation is achieved:

[0109] Equation (4):

[0110]

[0111] where r is the target value of the reduction metric and ε is the tolerance. The definition of the reduction metric may vary depending on the requirements of a specific application. Generally, the reduction metric can be defined according to the machine learning model as:

[0112] Equation (5):

[0113] r := g(f ω,θ (X))

[0114] where g(·) is a function that depends on the specific application (e.g., specified by the user via the graphical user interface provided by software 112). An example of the reduction metric is the ratio of the number of parameters of the adjusted machine learning model to the number of parameters of the existing machine learning model. In this case, r represents the target value of the ratio, and represents the actual value of the ratio after iteration j.

[0115] Assume that after iteration j is not less than the tolerance ε, then the stopping condition is not satisfied in sub-process 422. Therefore, for iteration j+1, the learning rate is reduced to α / (j+1). Then, in iteration j+1, a new layer (e.g., Figure 6 layer i-1 in ) is selected in sub-process 424, and the architectures and parameters of two layers (e.g., Figure 6 layer i-1 and layer i in ) are jointly retrained in sub-process 426. At the end of the joint retraining in iteration j+1, the reduction metric is calculated and the value of is compared with the tolerance ε in another iteration of sub-process 422. This process continues until or there are no remaining layers to be retrained (i.e., "yes" in sub-process 422).

[0116] The overall training can be regarded as a two-layer optimization, given by the following equation:

[0117]

[0118] where,

[0119]

[0120] wherein, is the trained machine learning model obtained by equation (1).

[0121] 6. Example use case

[0122] In an embodiment, the adjusted machine learning model generated by sub-process 420 can be deployed as machine learning model 334 in the controller 332 of the line protection system 330. In this case, the machine learning model 334 can be a convolutional neural network, which is generally sufficient for fault location. The convolutional neural network can efficiently embed time series data to successfully learn labels such as the fault location (e.g., fault distance), and provides enhanced interpretability relative to other machine learning models by means of the filtering operations performed by the one-dimensional convolutional layer. The process 400 is faster and more transparent than the NAS method.

[0123] Figure 8 FIG. illustrates a process 800 for making a tripping decision according to an embodiment, wherein the machine learning model for a computationally richer environment is pruned into the machine learning model 334 by the process 400 and deployed to the computationally poorer environment of the line protection system 330 to estimate the location of a fault. In sub-process 810, the controller 332 of the line protection system 330, which acts as a protection device, can receive measurement signals from a plurality of sensors connected to the power line 310. For example, these measurement signals can include or be composed of voltage values measured by the voltage measurement unit 312 and current values measured by the current measurement unit 314. In sub-process 820, the controller 332 can apply the adjusted machine learning model 334 to the measurement signals to estimate the location of the fault 340 on the power line 310. Then, in sub-process 830, the controller 332 can compare the estimated fault location with a predetermined threshold representing the range setting 336. For example, the controller 332 can determine whether the estimated fault location is less than or equal to the range setting 336. When it is determined that the estimated fault location meets the threshold (i.e., "yes" in sub-process 830), indicating that the fault 340 is within the protection area defined by the range setting 336, the controller 332 can trip the circuit breaker 316 in sub-process 840 to electrically isolate the fault location on the power line 310. Otherwise, when it is determined that the estimated fault location does not meet the threshold (i.e., "no" in sub-process 830), the controller 332 determines not to trip the circuit breaker 316 in sub-process 850.

[0124] In a fault location task, appropriate metrics are necessarily required when evaluating the performance of a machine learning model 334. The metrics must reliably capture the ability of the machine learning model 334 to make correct decisions as quickly as possible. Since the problem is inherently causal and temporal, there is an asymmetry in the decisions at each time sample because the machine learning model 334 can only use past trajectories (e.g., a single threshold violation at the fault location) to make a tripping decision. The correctness of the tripping decision depends on the time sample considered and affects the extent of misoperations (i.e., false negatives) and maloperations (i.e., false positives) in the resulting decision. Additionally, the speed of true positives is a separate performance metric. The trade-off between decision accuracy and decision speed makes the process of quantifying performance a meaningful task.

[0125] Since the machine learning model 334 outputs a fault location, the idea is to compare the fault location with a threshold that defines a protection zone (e.g., sub-process 830). This enables the machine learning model 334 to be tested for a set of different range settings 336. For each range setting 336, the machine learning model 334 is tested against a specific safety counter. It is expected that if the actual fault location is close to the range setting 336, the decision may not be perfect. Therefore, the machine learning model 334 is evaluated by computing the distance from the range setting 336 at which the machine learning model 334 makes a perfect decision. This can be represented by a dead zone for a specific range setting 336, in which at least one decision fails as a misoperation or a maloperation. The dead zone captures the total width of the zone, which contains the error estimates generated by false positive (type I) and false negative (type II) errors. A machine learning model 334 with a good estimate will produce a narrow dead zone, while a machine learning model 334 with a poor estimate will produce a wide dead zone. A dead zone equal to zero would mean that the machine learning model 334 makes perfect decisions for all cases.

[0126] The machine learning model 334 for estimating fault location pruned using automated progressive pruning as disclosed herein was tested in a transmission system in the presence of renewable generation interfaced through converters. The transmission system includes two lines, one of which is optionally open.

[0127] Figures 9A to 9D Illustrates the dead zones achieved by various example implementations of the machine learning model for fault location during testing. In Figures 9A to 9DIn each of them, the x-axis represents different range settings 336, which correspond to the thresholds applied to the fault location estimates output by the respective machine learning models 334 (e.g., in subprocess 820) (e.g., in subprocess 830). The y-axis represents the fault location and the corresponding blind zone range. The central solid line represents the case where the fault location is exactly equal to the range setting 336. The cases above the central line are those where the fault location is greater than the range setting 336, such that the line protection system 330 should decide not to trip (i.e., inhibit), as represented by subprocess 850. The cases below the central line are those where the fault location is less than the range setting 336, such that the line protection system 330 should decide to trip, as represented by subprocess 840. The dashed lines represent the user-specified error estimate boundaries. The narrower the blind zone outside the dashed lines, the better the machine learning model. The shaded areas together represent the blind zone, where the shaded area above the central line represents the area containing at least one false negative (i.e., incorrect inhibition) for a given fault location, and the shaded area below the central line represents the area containing at least one false positive (i.e., incorrect trip) for a given fault location. For testing, the average speed (including the safety counter delay) is calculated for all correct trip decisions at all range settings 336. Typically, 70% of the range settings 336 are used in the line protection system 330. Therefore, the blind zone width of 70% of the range settings 336 (represented as 0.7 on the x-axis) has been highlighted to show the performance of the corresponding machine learning model at this range setting 336. The safety and dependability of the machine learning model at 70% of the range settings 336 are also calculated.

[0128] Figure 9A Represents the blind zone of the existing machine learning model before pruning. At 70% of the range settings 336, the blind zone is 17.9%, the calculated safety is 99.3%, the calculated dependability is 98.4%, and the average speed is 24.8 milliseconds.

[0129] Figure 9B Represents the blind zone of the machine learning model after pruning the existing machine learning model using a brute-force method, where a given layer of the existing machine learning model is manually designed. At 70% of the range settings 336, the blind zone is 19.4%, the calculated safety is 98.7%, the calculated dependability is 98.8%, and the average speed is 14.8 milliseconds.

[0130] Figure 9C Represents the blind zone of the machine learning model after automated progressive pruning according to the first embodiment, where all layers in the new architecture remain learnable such that fine-tuning affects all layers. At 70% of the range settings 336, the blind zone is 18.4%, the calculated safety is 98.8%, the calculated dependability is 98.3%, and the average speed is 14.8 milliseconds.

[0131] Figure 9D Represents the blind spot of the machine learning model after automated progressive pruning according to the second embodiment, where all layers except the two trained layers are frozen during the iterations of sub-process 426, and several epochs using Adam are used to train the two learnable layers. In this second embodiment, after training the two learnable layers in each iteration of sub-process 426, the entire architecture is fine-tuned. At a range setting 336 of 70%, the blind spot is 18.8%, the calculated security is 98.7%, the calculated dependability is 98.5%, and the average speed is 14.8 milliseconds. This second embodiment results in faster pruning, but there is a slight performance degradation compared to the first embodiment.

[0132] In standard practice, a blind spot of 20% or less at a range setting 336 of 70% is considered satisfactory. Thus, both the first and second embodiments of automated progressive pruning achieve satisfactory results. Additionally, the two embodiments of automated progressive pruning produce a smaller blind spot than the brute-force method. Furthermore, compared to the brute-force method depicted in the table below, the two embodiments of automated progressive pruning achieve a smaller pruned machine learning model:

[0133] Pruning method # Number of parameters before pruning # Number of parameters after pruning % Pruning APP - Learnable 466,673 416,253 10.8 APP - Frozen 466,673 428,858 8.1 Brute force 466,673 440,144 5.7

[0134] It is worth noting that when compared to the time required for the brute-force method to prune and fine-tune the machine learning model with a given layer (i.e., one iteration), as the number of iterations in automated progressive pruning increases, the time required for pruning and fine-tuning increases sub-linearly. Figure 10 Is a graph of the convergence times of the two embodiments of automated progressive pruning and the brute-force method. The times corresponding to the embodiments of automated progressive pruning are for architectures with two to four alternative layers and two pruning iterations, while the times corresponding to the brute-force method correspond to one alternative layer and one pruning iteration.

[0135] 7. Example embodiment

[0136] The disclosed embodiments jointly adjust both the architecture and parameters of a machine learning model, such as a neural network. The neural network can be a convolutional neural network, a deep neural network, a recurrent neural network, or other types of artificial neural networks. In an embodiment, the adjustment is performed with constraints on the dimensions of the machine learning model to prune the machine learning model. Pruning reduces the overall size of the machine learning model such that it is suitable for execution (e.g., in real time) on, for example, a system with limited computing power, such as a line protection system 330 for a power line or other embedded system. Different from traditional pruning methods that produce sparse machine learning models, the disclosed pruning produces a dimension-reduced dense machine learning model that satisfactorily handles regression tasks (e.g., fault location) regarding time series data. The dense machine learning model is particularly beneficial for deployment in hardware that cannot accommodate the computational complexity associated with sparse matrix operations. The disclosed automated progressive adjustment is also less costly and easier to interpret than traditional methods.

[0137] The disclosed automated progressive adjustment can be used to adjust a machine learning model for any application that requires adapting the machine learning model from a first environment to a second environment. The first environment and the second environment can be different in some characteristics. The differences can be any change in the data, such as a change in the input-output distribution of the data. For example, both the first environment and the second environment may need to perform the same task (e.g., fault location), but have different source impedance ratios or distributions of line impedances. More generally, the first environment and the second environment can involve performing the same task with different data distributions. Alternatively, the first environment and the second environment can involve the execution of different but related tasks.

[0138] In an embodiment, the disclosed automated progressive adjustment can be automated progressive pruning. The disclosed automated progression is particularly beneficial in environments that require a pruned dense machine learning model. One such environment is hardware embedding, in which case functions are written in a low-level programming language (e.g., C language) for computational efficiency, and sparse matrix calculations are expensive. More generally, automated progressive pruning can be used to reduce the size of a machine learning model trained for a computationally richer first environment (e.g., in which more computational resources are available) for deployment in a computationally poorer second environment (e.g., in which computational resource availability is lower).

[0139] A description and test example of a computationally lean second environment is the line protection system 330 for protecting overhead lines. In particular, the fault location task in time domain protection was tested to demonstrate the performance of the machine learning model generated by automated progressive pruning. However, it should be understood that this is only an example, and the first environment and / or the second environment may involve other tasks of other applications. Examples of such tasks include, but are not limited to, power system state estimation, topology estimation, parameter estimation (e.g., grid inertia), power flow estimation, load prediction, etc. It is worth noting that all these specific examples benefit from a final dense machine learning model small enough to be embedded in hardware. Thus, a full model can be trained to achieve a satisfactory performance level in a computationally richer first environment, and then the disclosed automated progressive pruning can be used for pruning to obtain a smaller dense model for deployment in a computationally leaner second environment without much performance loss.

[0140] Example 1: A method includes using at least one hardware processor to: receive a machine learning model having an architecture and parameters trained to operate in a first environment; jointly adjust both the architecture and parameters of the machine learning model using a training dataset of a second environment different from the first environment; and deploy the adjusted machine learning model to the second environment.

[0141] Example 2: The method according to Example 1, wherein the machine learning model includes a neural network.

[0142] Example 3: The method according to Example 2, wherein the neural network is a convolutional neural network.

[0143] Example 4: The method according to Example 2 or 3, wherein jointly adjusting both the architecture and parameters of the machine learning model includes: for each layer in one or more layers of the neural network, retraining the layer until a stopping condition is met.

[0144] Example 5: The method according to Example 4, wherein the one or more layers are multiple layers.

[0145] Example 6: The method according to Example 5, wherein the multiple layers are trained in an order from a layer closest to the output of the machine learning model in the multiple layers to another layer closest to the input of the machine learning model in the multiple layers.

[0146] Example 7: The method according to any one of Examples 2 to 6, wherein jointly adjusting both the architecture and the parameters of the machine learning model comprises: in each of one or more iterations: selecting one of a plurality of layers in the neural network that has not been retrained in any previous iteration of the one or more iterations; retraining the selected layer; determining whether a stopping condition is met; when the stopping condition is met, stopping the joint adjustment; and when the stopping condition is not met, adding a next iteration to the one or more iterations.

[0147] Example 8: The method according to Example 7, wherein the one or more iterations are multiple iterations, wherein each selected layer is retrained according to a learning rate, and wherein jointly adjusting both the architecture and the parameters of the machine learning model further comprises: in at least one of the multiple iterations, when the stopping condition is not met, changing the learning rate before the next iteration.

[0148] Example 9: The method according to Example 7 or 8, wherein retraining each selected layer comprises: pruning the selected layer.

[0149] Example 10: The method according to Example 9, wherein the stopping condition comprises a threshold indicative of a size reduction metric of the machine learning model.

[0150] Example 11: The method according to Example 9 or 10, wherein each selected layer is retrained according to a learning rate, and wherein jointly adjusting both the architecture and the parameters of the machine learning model further comprises: in each of the one or more iterations, when the stopping condition is not met, reducing the learning rate before the next iteration.

[0151] Example 12: The method according to any one of Examples 7 to 11, wherein retraining the selected layer comprises: selecting a plurality of alternative layers, each alternative layer having a different size from the selected layer; training the plurality of alternative layers using the training dataset to minimize the cross-entropy loss of the machine learning model; and selecting one alternative layer having the lowest error metric among the plurality of alternative layers as the retrained layer.

[0152] Example 13: The method according to Example 12, wherein the cross-entropy loss comprises a negative log-likelihood loss, and wherein the error metric comprises a mean squared error.

[0153] Example 14: The method according to Example 12 or 13, wherein the size of each alternative layer among the plurality of alternative layers is smaller than the selected layer.

[0154] Example 15: The method according to any one of Examples 12 to 14, wherein each of the plurality of alternative layers is selected to have a size different from the size of the selected layer and the size of any other one of the plurality of alternative layers within a restricted search space around the size of the selected layer.

[0155] Example 16: The method according to Example 15, wherein each of the plurality of alternative layers is smaller in size than the selected layer, wherein the neural network is a convolutional neural network, and wherein the restricted search space is defined based on the number of filters to be pruned from the selected layer.

[0156] Example 17: The method according to Example 16, wherein the method further comprises: using the at least one hardware processor to determine the number of filters to be pruned using principal component analysis.

[0157] Example 18: The method according to any one of the foregoing examples, wherein the adjusted machine learning model estimates the location of a fault on a power line based on one or more measured parameters, wherein the training data set includes labeled feature vectors, and wherein each of the labeled feature vectors includes a value of each of the one or more measured parameters and is labeled with a fault location, wherein the joint adjustment includes pruning the machine learning model, wherein the second environment is a relay configured to trip a circuit breaker on the power line, and wherein deploying the adjusted machine learning model includes installing the adjusted machine learning model in a controller of the relay.

[0158] Example 19: A system comprising: at least one hardware processor; and software configured to perform the method according to any one of Examples 1 to 18 when executed by the at least one hardware processor.

[0159] Example 20: A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform the method according to any one of Examples 1 to 18.

[0160] The foregoing description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles described herein may be applied to other embodiments without departing from the spirit or scope of the invention. Accordingly, it is to be understood that the description and drawings presented herein represent the presently preferred embodiments of the invention and thus represent the broad subject matter contemplated by the invention. It should be further understood that the scope of the invention fully encompasses other embodiments that may be obvious to those skilled in the art and that, accordingly, the scope of the invention is not limited.

[0161] As used herein, the terms "comprising," "comprise," and "comprises" are open-ended. For example, "A comprises B" means that A can include any of the following scenarios: (i) only B; or (ii) a combination of B with one or more (and potentially any number of) other elements. In contrast, the terms "consisting of," "consist of," and "consists of" are closed-ended. For example, "A consists of B" means that A includes only B and no other elements in the same context.

[0162] Combinations such as "at least one of A, B, or C," "one or more of A, B, or C," "at least one of A, B, and C," "one or more of A, B, and C," and "any combination of A, B, C, or thereof" as described herein include any combination of A, B, and / or C and may include multiples of A, multiples of B, or multiples of C. Specifically, combinations such as "at least one of A, B, or C," "one or more of A, B, or C," "at least one of A, B, and C," "one or more of A, B, and C," and "any combination of A, B, C, or thereof" can be only A, only B, only C, A and B, A and C, B and C, or A and B and C, and any such combination can include one or more members of its components A, B, and / or C. For example, the combination of A and B can include one A and multiple Bs, multiple As and one B, or multiple As and multiple Bs.

Claims

1. A method includes using at least one hardware processor to: Receive a machine learning model having an architecture and parameters that are trained to operate in a first environment; Jointly adjust both the architecture and parameters of the machine learning model using a training dataset of a second environment that is different from the first environment; and Deploy the adjusted machine learning model to the second environment.

2. The method according to claim 1, wherein, The machine learning model includes a neural network.

3. The method according to claim 2, wherein, The neural network is a convolutional neural network.

4. The method according to claim 2, wherein, Jointly adjusting both the architecture and parameters of the machine learning model includes: for each layer in one or more layers of the neural network, retraining the layer until a stopping condition is met.

5. The method according to claim 4, wherein, The one or more layers are multiple layers.

6. The method according to claim 5, wherein, The multiple layers are trained in an order from a layer closest to the output of the machine learning model among the multiple layers to another layer closest to the input of the machine learning model.

7. The method according to claim 2, wherein Jointly adjusting both the architecture and parameters of the machine learning model includes: in each of one or more iterations: Select a layer from the multiple layers of the neural network that has not been retrained in any previous iteration of the one or more iterations; Retrain the selected layer; Determine whether the stopping condition is met; When the stopping condition is met, stop the joint adjustment; and When the stopping condition is not met, add the next iteration to the one or more iterations.

8. The method according to claim 7, wherein, The one or more iterations are multiple iterations, wherein each selected layer is retrained according to a learning rate, and wherein jointly adjusting both the architecture and parameters of the machine learning model further includes: in at least one of the multiple iterations, when the stopping condition is not met, changing the learning rate before the next iteration.

9. The method according to claim 7, wherein, Retraining each selected layer includes: pruning the selected layer.

10. The method according to claim 9, wherein, The stopping condition includes a threshold indicating a size reduction metric of the machine learning model.

11. The method according to claim 9, wherein, Each selected layer is retrained according to a learning rate, and wherein jointly adjusting both the architecture and parameters of the machine learning model further includes: in each of the one or more iterations, when the stopping condition is not met, reducing the learning rate before the next iteration.

12. The method according to claim 7, wherein, Retraining the selected layer includes: Selecting multiple alternative layers, each alternative layer having a different size from the selected layer; Training the multiple alternative layers using the training dataset to minimize the cross-entropy loss of the machine learning model; and Selecting one alternative layer having the lowest error metric among the multiple alternative layers as the retrained layer.

13. The method according to claim 12, wherein, The cross-entropy loss includes negative log-likelihood loss, and wherein the error metric includes mean squared error.

14. The method according to claim 12, wherein, The size of each alternative layer among the multiple alternative layers is smaller than the selected layer.

15. The method according to claim 12, wherein, Each of the plurality of alternative layers is selected to have a size different from the size of the selected layer and the size of any other one of the plurality of alternative layers within a restricted search space around the size of the selected layer.

16. The method according to claim 15, wherein, Each of the plurality of alternative layers has a smaller size than the selected layer, where the neural network is a convolutional neural network, and where the restricted search space is defined based on the number of filters to be pruned from the selected layer.

17. The method according to claim 16, wherein, The method further includes: using the at least one hardware processor to determine the number of filters to be pruned using principal component analysis.

18. The method according to claim 1, wherein, The adjusted machine learning model estimates the location of a fault on a power line based on one or more measured parameters, where the training data set includes labeled feature vectors, and where each of the labeled feature vectors includes a value of each of the one or more measured parameters and is labeled with a fault location, where the joint adjustment includes pruning the machine learning model, where the second environment is a relay configured to trip a circuit breaker on the power line, and where deploying the adjusted machine learning model includes installing the adjusted machine learning model in a controller of the relay.

19. A system, comprising: at least one hardware processor; and software configured to, when executed by the at least one hardware processor, receive a machine learning model having an architecture and parameters trained to operate in a first environment, jointly adjust both the architecture and parameters of the machine learning model using a training data set of a second environment different from the first environment, and deploy the adjusted machine learning model to the second environment.

20. A non-transitory computer-readable medium storing instructions, wherein, The instructions, when executed by a processor, cause the processor to: receive a machine learning model having an architecture and parameters trained to operate in a first environment; jointly adjust both the architecture and parameters of the machine learning model using a training data set of a second environment different from the first environment; and deploy the adjusted machine learning model to the second environment.