Overhead line power outage prediction neural network structure search method and device
By building a basic model and searching the neural network structure, sampling and training the subnet to determine the optimal model, the problem of overhead power outage prediction under extreme weather in the prior art is solved, and efficient power outage prediction and deployment are achieved.
Patent Information
- Application Number
- CN202210577252.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-25
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-05-25
AI Technical Summary
The prior art is difficult to effectively predict power outages of 10kV overhead distribution lines in extreme weather, and the deep learning model is difficult to apply due to the long time spent in actual deployment.
By constructing a basic model and performing structural searches based on time-consuming requirements and network parameter-time-consuming mapping relationships, the first subnet is sampled and trained to convergence, and then the second subnet is sampled and the performance is evaluated to determine the optimal subnet as the power outage prediction model.
It realizes that the performance of the overhead line power outage prediction model is improved while meeting the time-consuming requirements, balances the algorithm performance and model scale, and ensures the operational efficiency and model performance in deployment.
Smart Images

Figure CN115034449B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power distribution network big data, and in particular to a method and device for searching a neural network structure for predicting power outages of overhead lines. Background Art
[0002] With the development of the economy and society, the coverage of substations and distribution networks in power grid systems has been continuously expanding. The intrusion of disastrous weather such as strong winds and lightning on the safe operation of the power grid has increased accordingly. Under the influence of extreme weather, 10kV overhead distribution lines are more likely to trip, causing regional power grid line failures and shutdowns, affecting the normal production and life of residents.
[0003] With the increasing number of severe weather such as ice, the normal operation of power grid equipment has been greatly affected. There is an urgent need to conduct in-depth analysis of meteorological disasters in the power grid and seek preventive measures. Current research mainly focuses on weather such as ice and snow, strong winds, and lacks research on the prediction of overhead line power outages in severe convective weather. Moreover, it is mainly based on large-scale data training of machine learning or deep learning models to obtain power outage prediction models, and then deploy them to the central processing cluster to realize overhead line power outage prediction. Considering the computing power limitations of the central cluster, the actual scale of the overhead line power outage prediction model cannot be too large, otherwise the central cluster will find it difficult to calculate the model results within a certain time range. In current related research, the scale of the neural network is not limited, resulting in the algorithm being too time-consuming during actual deployment and difficult to apply in actual environments. However, if the deep learning model is too small, it is easy to underfit and difficult to mine big data information. Therefore, how to balance the algorithm performance and model scale is crucial for the deployment of the overhead line power outage prediction model in a production environment.
[0004] In the related art, the invention patent application with application number 202010109054.7 discloses a neural network structure search method and a neural network structure search device, which trains a sampling model given a search space and resource constraints of a target device, so that the sampling model samples a neural network structure that satisfies the resource constraints from the given search space, and then uses the sampling model to search the target neural network structure from the candidate search space sampled by the sampling model from the given search space. This technical solution can ensure that a neural network that satisfies the resource constraints of the target device is searched, thereby improving the search efficiency of the neural network structure. However, this technical solution is to obtain the network structure by directly searching for the hyperparameters such as the operators and number of layers of the neural network and training it. This direct search method for the operators, depth, and width of the lightweight network is inefficient, and because the lightweight network does not have good initialization network parameters, the performance of the lightweight neural network finally trained is poor. Summary of the invention
[0005] The technical problem to be solved by the present invention is how to improve the performance and scale of the balance model and realize the prediction of overhead line power outage.
[0006] The present invention solves the above technical problems through the following technical means:
[0007] In a first aspect, an embodiment of the present invention adopts a neural network structure search method for overhead line power outage prediction, the method comprising:
[0008] Based on the time consumption requirements and the network parameter-time consumption mapping relationship, a base model is constructed;
[0009] According to the target network scale, sampling a first sub-network from the base model, training the first sub-network using a training data set, and evaluating the first sub-network using an evaluation data set, so that the performance of the base model converges in the overhead line power outage prediction task;
[0010] Sampling a second subnetwork from the converged base model to obtain a plurality of second subnetworks that meet the time consumption requirement, and evaluating the performance of the second subnetwork based on the evaluation data set and the evaluation criteria to determine the optimal subnetwork;
[0011] The optimal subnetwork is used as a power outage prediction model to implement the overhead line power outage prediction.
[0012] By searching for appropriate model parameters, the target network can achieve optimal performance while meeting the time requirements for neural network deployment. The above algorithm can better balance the algorithm performance and model scale, ensuring both the actual deployment running time and the optimal performance of the model under a certain capacity.
[0013] Furthermore, the process of constructing the network parameter-time consumption mapping relationship includes:
[0014] Setting the parameter space, including the width and depth of the neural network, the number of neurons, and the type of activation function;
[0015] Traversing the parameter space to obtain k neural networks;
[0016] The time consumption of the k neural networks is measured, and the network parameter-time consumption mapping relationship is determined as follows:
[0017] time=T(flops,activation)=T(g(N,M,D),actionvation)
[0018] Among them, time is the time consumed by the neural network on the actual computing chip obtained through machine calibration, flops is a function of the number of input neurons and output neurons in each layer of the neural network, activation is the neuron activation function, N is the number of input neurons in each layer of the neural network, M is the number of output neurons, and D is the depth of the neural network.
[0019] Furthermore, the base model is constructed based on the time consumption requirement and the network parameter-time consumption mapping relationship, including:
[0020] Based on the network parameter-time consumption mapping relationship, possible network parameters under the time consumption requirement are solved as proposed parameters;
[0021] According to the proposed parameters, the base model is constructed and expressed as follows:
[0022] net=Network(L(N,activation),D)
[0023] Among them, N is the number of input neurons in each layer of the neural network, D is the depth of the neural network, and activation is the neuron activation function.
[0024] Furthermore, sampling a first subnetwork from the base model according to the target network scale, training the first subnetwork using a training data set, and evaluating the first subnetwork using an evaluation data set, so that the performance of the base model converges in the overhead line power outage prediction task, includes:
[0025] According to the target network scale, the representation of the first subnetwork is set as: subnet(subnet_N, D, activation), wherein the depth D and the neuron activation function activation of the first subnetwork are consistent with the base model, the width of the first subnetwork is smaller than the width of the neural network in the base model, and the width subnet_N of the first subnetwork is obtained by the network parameter-time consumption mapping relationship;
[0026] Sampling subnet_N neurons from each layer of the neural network of the base model, and training the first sub-network using a training data set, and evaluating the first sub-network using an evaluation data set to obtain model parameters;
[0027] The sampling is repeated s times, and the model parameters obtained by the current sampling training overwrite the model parameters obtained by the previous sampling training, until the performance of the base model converges in the overhead line power outage prediction task.
[0028] Further, sampling a second subnetwork from the converged base model to obtain a plurality of second subnetworks that meet the time consumption requirement, and evaluating the performance of the second subnetwork based on the evaluation data set and the evaluation criteria to determine the optimal subnetwork includes:
[0029] Using uniform sampling, a second subnetwork is sampled from each layer of the neural network in the converged base model to obtain c second subnetworks that meet the time consumption requirement.
[0030] On the evaluation data set, calculating the index value of the second sub-network according to the evaluation criteria;
[0031] The second sub-network corresponding to the maximum indicator value is used as the optimal sub-network.
[0032] Furthermore, the evaluation criteria are:
[0033] acc = number of correctly predicted samples / total number of samples.
[0034] Furthermore, the time consumption of the base model is m times the time consumption of the target network scale, where m>1.
[0035] In a second aspect, an embodiment of the present invention adopts a neural network structure search device for predicting power outages of overhead lines, the device comprising:
[0036] A model building module is used to build a base model based on time consumption requirements and a network parameter-time consumption mapping relationship;
[0037] A training module, used to sample a first sub-network from the base model according to the target network scale, train the first sub-network using a training data set, and evaluate the first sub-network using an evaluation data set, so that the performance of the base model converges in the overhead line power outage prediction task;
[0038] A search module, configured to sample a second sub-network from the converged base model to obtain a plurality of second sub-networks that meet the time-consuming requirement, and to evaluate the performance of the second sub-network based on the evaluation data set and the evaluation criteria to determine the optimal sub-network;
[0039] A prediction module is used to use the optimal sub-network as a power outage prediction model to realize the overhead line power outage prediction.
[0040] Furthermore, the network parameter-time consumption mapping relationship is:
[0041] time=T(flops,activation)=T(g(N,M,D),actionvation)
[0042] Among them, time is the time consumed by the neural network on the actual computing chip obtained through machine calibration, flops is a function of the number of input neurons and output neurons in each layer of the neural network, activation is the neuron activation function, N is the number of input neurons, M is the number of output neurons, and D is the depth of the neural network.
[0043] Furthermore, the training module includes:
[0044] According to the target network scale, the representation of the first subnetwork is set as: subnet(subnet_N, D, activation), wherein the depth D and the neuron activation function activation of the first subnetwork are consistent with the base model, the width of the first subnetwork is smaller than the width of the neural network in the base model, and the width subnet N of the first subnetwork is obtained by the network parameter-time consumption mapping relationship;
[0045] A first sampling unit is used to sample N neurons of a subnet from each layer of the neural network of the base model, and train the first subnet using a training data set, and evaluate the first subnet using an evaluation data set to obtain model parameters;
[0046] A training unit, used for repeating sampling s times, and overwriting the model parameters obtained by the previous sampling training with the model parameters obtained by the current sampling training, until the performance of the base model converges in the overhead line power outage prediction task;
[0047] The search module includes:
[0048] The second sampling unit is used to sample the second sub-network from each layer of the neural network in the converged base model in a uniform sampling manner to obtain c second sub-networks that meet the time consumption requirement.
[0049] An evaluation unit, configured to calculate an indicator value of the second sub-network based on the evaluation data set according to an evaluation standard;
[0050] A determination unit is used to use the second sub-network corresponding to the maximum indicator value as the optimal sub-network.
[0051] The advantages of the present invention are:
[0052] (1) The present invention designs the base model network structure according to the network scale limitation, then uniformly samples the network path to generate the corresponding first sub-network, trains the first sub-network to achieve the best performance in the overhead line power outage prediction task, and finally makes the base model performance converge in the overhead line power outage prediction task; then samples the second sub-network from the base model, and evaluates the performance of the second sub-network in the overhead line power outage prediction task, and takes the network with the best performance as the overhead line power outage prediction model that can be actually deployed. By searching for appropriate model parameters, the target network can achieve the best performance while meeting the time requirements for neural network deployment. According to the above algorithm, the algorithm performance and model scale can be better balanced, which not only ensures the actual deployment operation time, but also ensures the optimal performance of the model under a certain capacity.
[0053] (2) Since the base model has a large capacity difference compared to the target model, we use sampling to train it. That is, we sample a sub-network of the size of the target network at a time, train the sub-network on the training data set traindata, and evaluate it on evaluationdata to achieve the optimal performance as the converged model.
[0054] (3) Through sub-network sampling and shared weight training, the effect of one-time base model training and multiple sub-network searches is achieved, which greatly reduces the search time.
[0055] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is a flow chart of a neural network structure search method for predicting power outages of overhead lines in Embodiment 1 of the present invention;
[0057] Figure 2 is a schematic diagram of a neural network structure in Embodiment 1 of the present invention;
[0058] Figure 3 It is a schematic diagram showing the principle of the basic model in the first embodiment of the present invention;
[0059] Figure 4 This is a schematic diagram of the sampling principle of the base model in the first embodiment of the present invention;
[0060] Figure 5 It is a structural diagram of the neural network structure search device for predicting overhead line power outages in the second embodiment of the present invention. DETAILED DESCRIPTION
[0061] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in combination with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0062] like Figure 1 As shown, the first embodiment of the present invention also discloses a neural network structure search method for overhead line power outage prediction, the method comprising the following steps:
[0063] S10, constructing a base model based on the time consumption requirement and the network parameter-time consumption mapping relationship;
[0064] S20, sampling a first subnetwork from the base model according to the target network scale, training the first subnetwork using a training data set, and evaluating the first subnetwork using an evaluation data set, so that the performance of the base model converges in the overhead line power outage prediction task;
[0065] It should be noted that in models with extremely large scale, there are more parameters, the greater the training difficulty, and the slower the convergence speed. Therefore, this embodiment adopts a step-by-step training mode. According to the target network scale, some neurons are sampled from the neurons in each layer of the base model to obtain the first sub-network, and the first sub-network is evaluated using an evaluation data set to reduce model training parameters, thereby improving training effect and training speed.
[0066] S30, sampling a second sub-network from the converged base model to obtain a plurality of second sub-networks that meet the time consumption requirement, and evaluating the performance of the second sub-network based on the evaluation data set and the evaluation criteria to determine the optimal sub-network;
[0067] It should be noted that the time requirement refers to the number of neurons N in the network and the depth D of the network. The base model is used to sample the depth and number of neurons in each layer, and the sampled second sub-network is evaluated to obtain the sub-network with the best network structure performance. The sub-network is trained using the evaluation data set, and the parameters of the second sub-network are fine-tuned to finally obtain the optimal sub-network.
[0068] S40: Using the optimal sub-network as a power outage prediction model to implement the overhead line power outage prediction.
[0069] It should be noted that the present invention designs the base model network structure according to the network scale limitation, then uniformly samples the path of the network to generate the corresponding first sub-network, trains the first sub-network to achieve the best in the overhead line power outage prediction task, and finally makes the base model performance converge in the overhead line power outage prediction task; then samples the second sub-network from the base model, and evaluates the performance of the second sub-network in the overhead line power outage prediction task, and takes the network with the best performance as the overhead line power outage prediction model that can be actually deployed. By searching for appropriate model parameters, the target network can achieve the best performance while meeting the time-consuming requirements for neural network deployment. According to the above algorithm, the algorithm performance and model scale can be better balanced, which not only ensures the actual deployment operation time consumption, but also ensures the optimal performance of the model under a certain capacity.
[0070] like Figure 2 As shown, the neural network in this embodiment includes an input layer facture, a backbone network backbone and a prediction network head. The backbone network includes a multi-layer neural network, and each layer of the neural network includes a plurality of neurons.
[0071] In some embodiments, the process of constructing the network parameter-time consumption mapping relationship includes:
[0072] Setting the parameter space, including the width and depth of the neural network, the number of neurons, and the type of activation function;
[0073] Traversing the parameter space to obtain k neural networks;
[0074] The time consumption of the k neural networks is measured, and the network parameter-time consumption mapping relationship is determined as follows:
[0075] time=T(flops,activation)=T(g(N,M,D),actionvation)
[0076] Among them, time is the time consumed by the neural network on the actual computing chip obtained through machine calibration, flops is a function of the number of input neurons and output neurons in each layer of the neural network, activation is the neuron activation function, N is the number of input neurons in each layer, M is the number of output neurons, and D is the depth of the neural network.
[0077] It should be noted that for the convenience of calculation, the number of neurons in each layer of the model is assumed to be the same, but in actual applications, they may be different, such as N1, N2, N3..., N D Respectively represent the number of input neurons in each layer of the model.
[0078] It should be noted that the calibration algorithm designed in this embodiment solves the mapping relationship between the scale of the neural network and the runtime consumption.
[0079] Specifically, when the algorithm is actually deployed, there is a time limit for the algorithm. For example, there are 100 locations in Anhui Province that need to simultaneously predict the impact of severe convective weather on the status of overhead lines. The algorithm needs to respond to system requirements with a delay of seconds. Therefore, for a single prediction, assuming that parallel computing is not considered, the algorithm takes 1s / 100 times, which is a processing time limit of about 10ms. The calculation process of the model time is:
[0080] The time consumption of the overhead line power outage prediction neural network model during operation is mainly related to factors such as the depth, width, activation function category, and number of residual connections of the neural network. Generally speaking, the scale of neural network operations is measured by the sum of the number of multiplications and additions of the network, that is, Flops. Taking the calculation of one layer of neurons in the overhead line power outage prediction neural network as an example, the number of input neurons is N, the output is M, and the neuron calculation function is:
[0081]
[0082] Therefore, the sum of the number of multiplications and additions is:
[0083] Flops = N*M + (N-1)*M + M = 2N*M
[0084] Therefore, the flops of each layer is a function of the number of input neurons and output neurons. In addition, the neural network is composed of multiple layers of neurons, so the depth of the neural network D is also one of the factors affecting flops:
[0085] Flops = g(N,M,D)
[0086] Finally, the activation function of the neuron, such as the sigmoid function or the ReLU function, has different time consumption, so the network time consumption is also affected by the neuron activation category. In summary, the neural network time consumption is affected by the activation function and the number of multiplications and additions. Large number of multiplications and additions and high-time activation functions will increase the neural network time consumption.
[0087] Then, we first set the parameter space range, N, M, and D are within a certain range, and the action functions are sigmoid, ReLu, and ReLU6. Then, we traverse the parameter space as above to obtain the neural network {Net-1, Net-2, …, Net-k}, and calculate the time consumption of k neural networks. Finally, we get the time consumption mapping table based on the mapping relationship between parameters and time consumption, that is, we get the assumed mapping relationship between the neural network time consumption and the activation function and the number of multiplications and additions as T.
[0088] In some embodiments, step S10 includes:
[0089] Based on the network parameter-time consumption mapping relationship, possible network parameters under the time consumption requirement are solved as proposed parameters;
[0090] According to the proposed parameters, the base model is constructed and expressed as follows:
[0091] net=Network(L(N,activation),D)
[0092] Among them, N is the number of input neurons in each layer of the neural network, D is the depth of the neural network, and activation is the neuron activation function.
[0093] It should be noted that if Figure 3 As shown, the neural network is composed of multiple layers of neurons, and each neuron has its corresponding activation function. This embodiment decouples the network structure search space into two steps: neuron module design and neural network depth design. Neuron module design is to design the structure of neurons in each layer. The neurons in each layer can be expressed as:
[0094] layer = L(N, activation)
[0095] N is the neuron width of each layer of the overhead line power outage prediction neural network. Activation is the activation function of each layer of neurons.
[0096] After the neuron module is designed, it is repeated D times to obtain a complete neural network. Therefore, the neural network structure can be expressed as:
[0097] net=Network(layer,D)
[0098] This embodiment assumes that the width of neurons in all layers is consistent and the activation function of neurons in each layer is the same.
[0099] In this embodiment, according to the algorithm time requirement, the corresponding possible base models N, D, activation are solved by converting the algorithm time into a mapping relationship of the neural network scale, and then the base model Network (L (N, activation), D) is obtained.
[0100] Specifically, the application scenario of this embodiment is a single algorithm running time of 10ms. Considering the redundancy problem of the base model compared with the target model, the design time of the base model is 1.5 times that of the target network model, that is, the base model running time constraint is 15ms. Therefore, by converting the algorithm time into a mapping relationship of the neural network scale, the corresponding possible base models N, D, activation are solved, and then the base model Network (L (N, activation), D) is obtained.
[0101] In some embodiments, step S20 includes:
[0102] According to the target network scale, the representation of the first subnetwork is set as: subnet(subnet_N, D, activation), wherein the depth D and the neuron activation function activation of the first subnetwork are consistent with the base model, the width of the first subnetwork is smaller than the width of the neural network in the base model, and the width subnet_N of the first subnetwork is obtained by the network parameter-time consumption mapping relationship;
[0103] Sampling subnet_N neurons from each layer of the neural network of the base model, and training the first sub-network using a training data set, and evaluating the first sub-network using an evaluation data set to obtain model parameters;
[0104] The sampling is repeated s times, and the model parameters obtained by the current sampling training overwrite the model parameters obtained by the previous sampling training, until the performance of the base model converges in the overhead line power outage prediction task.
[0105] It should be noted that since the base model has a large capacity difference compared to the target network model, this embodiment uses sampling for training, that is, sampling the first sub-network of the target network size once, and training the first sub-network using the training data set traindata, and evaluating it on evaluationdata to achieve the optimal performance as the converged model. Repeat uniform sampling s times to obtain the final trained base model.
[0106] In this embodiment, each layer of neurons in the neural network is partially sampled, and the network parameters of the next sampling network are partially repeated with those obtained in the previous training. When the first sub-network obtained by the next sampling is trained, parameters different from those of the network obtained by the previous sampling are trained, and the parameters of the same part of the previous training are used as initialization parameters for the next sampling network training, and the parameters of the entire network model are gradually covered, so as to achieve better training results.
[0107] (1) According to the 10ms limit of the algorithm, the subnet expression is set as subnet(subnet_N, D, activation), and the subnet depth and activation are kept consistent with the base model. The algorithm time consumption is converted into a mapping relationship of the neural network scale to solve the corresponding possible subnet width. Since the algorithm time consumption is positively correlated with the neuron width, generally speaking, the width of the base model is larger than the subnet width, so the neurons of the subnet can be sampled from the neurons of the base model.
[0108] (2) Base model sampling: subnet_N neurons are sampled from each layer of N neurons in the base model, with a total of There are several possibilities, as follows Figure 4 As shown, this patent samples subnet_N neurons in each layer, as shown by gray neurons, to obtain a subnetwork that meets the time consumption requirements of the algorithm.
[0109] (3) Training the first sub-network: For the first sub-network sampled above, the first sub-network is trained using the training data set traindata, and is evaluated on the evaluationdata to achieve the optimal performance as a converged model.
[0110] (4) The sampling of the first sub-network is repeated s times. Each time the sampling training is performed, part of the parameters of the base model are obtained, and the parameters of the previous base model are used as initialization parameters. After obtaining the new sub-network, part of the existing parameters of the base model are overwritten. After repeating s, the complete base model parameters are obtained.
[0111] In some embodiments, the unqualified user S30 includes:
[0112] By uniform sampling, the second sub-network is sampled from each layer of the neural network in the converged base model to obtain c second sub-networks subnet_1, subnet_2, subnet_3 ... subnet_c that meet the time consumption requirements.
[0113] On the evaluation data set, calculating the index value of the second sub-network according to the evaluation criteria;
[0114] The second sub-network corresponding to the maximum indicator value is used as the optimal sub-network.
[0115] It should be noted that this embodiment performs sub-network performance evaluation based on evaluationdata, and the evaluation standard is acc, so the acc index values {acc_1, acc_2, ..., acc_c} of c sub-networks are calculated, and the sub-network index corresponding to the largest acc is found, which is the optimal sub-network.
[0116] It should be noted that this embodiment evaluates the overhead line power outage prediction model under severe convective weather based on EvaluationData. The prediction accuracy is used as the evaluation index, and the evaluation index is calculated as follows:
[0117] acc = number of correctly predicted samples / total number of samples.
[0118] This embodiment achieves the effect of one-time base model training and multiple sub-network searches through sub-network sampling and shared weight training, which greatly reduces the search time.
[0119] like Figure 5 As shown, the second embodiment of the present invention discloses a neural network structure search device for predicting power outages of overhead lines, the device comprising:
[0120] A model building module 10 is used to build a base model based on the time consumption requirement and the network parameter-time consumption mapping relationship;
[0121] A training module 20, for sampling a first sub-network from the base model according to a target network scale, training the first sub-network using a training data set, and evaluating the first sub-network using an evaluation data set, so that the performance of the base model converges in the overhead line power outage prediction task;
[0122] A search module 30 is used to sample a second sub-network from the converged base model to obtain a plurality of second sub-networks that meet the time consumption requirement, and to evaluate the performance of the second sub-network based on the evaluation data set and the evaluation criteria to determine the optimal sub-network;
[0123] The prediction module 40 is used to use the optimal sub-network as a power outage prediction model to realize the overhead line power outage prediction.
[0124] In some embodiments, the network parameter-time consumption mapping relationship is:
[0125] time=T(flops,activation)=T(g(N,M,D),actionvation)
[0126] Among them, time is the time consumed by the neural network on the actual computing chip obtained through machine calibration, flops is a function of the number of input neurons and output neurons in each layer of the neural network, activation is the neuron activation function, N is the number of input neurons, M is the number of output neurons, and D is the depth of the neural network.
[0127] In some embodiments, the training module includes:
[0128] According to the target network scale, the representation of the first subnetwork is set as: subnet(subnet_N, D, activation), wherein the depth D and the neuron activation function activation of the first subnetwork are consistent with the base model, the width of the first subnetwork is smaller than the width of the neural network in the base model, and the width subnet_N of the first subnetwork is obtained by the network parameter-time consumption mapping relationship;
[0129] A first sampling unit is used to sample subnet_N neurons from each layer of the neural network of the base model, and train the first sub-network using a training data set, and evaluate the first sub-network using an evaluation data set to obtain model parameters;
[0130] A training unit, used for repeating sampling s times, and overwriting the model parameters obtained by the previous sampling training with the model parameters obtained by the current sampling training, until the performance of the base model converges in the overhead line power outage prediction task;
[0131] The search module includes:
[0132] The second sampling unit is used to sample the second sub-network from each layer of the neural network in the converged base model in a uniform sampling manner to obtain c second sub-networks that meet the time consumption requirement.
[0133] An evaluation unit, configured to calculate an indicator value of the second sub-network based on the evaluation data set according to an evaluation standard;
[0134] A determination unit is used to use the second sub-network corresponding to the maximum indicator value as the optimal sub-network.
[0135] It should be noted that the overhead line power outage prediction neural network structure search device disclosed in this embodiment corresponds to the overhead line power outage prediction neural network structure search method disclosed in the above embodiment, and has corresponding technical features and technical effects, which will not be repeated here.
[0136] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0137] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0138] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0139] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0140] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. A neural network structure search method for predicting power outages in overhead lines, characterized in that: The method comprises: Based on the time consumption requirements and the network parameter-time consumption mapping relationship, a base model is constructed; According to the target network scale, a first subnetwork is sampled from the base model, the first subnetwork is trained using a training data set, and the first subnetwork is evaluated using an evaluation data set, so that the performance of the base model converges in the overhead line power outage prediction task, including setting the representation of the first subnetwork as: subnet(subnet_N, D, activation) according to the target network scale, wherein the depth D and the neuron activation function activation of the first subnetwork are consistent with the base model, the width of the first subnetwork is smaller than the width of the neural network in the base model, and the width subnet_N of the first subnetwork is obtained through the network parameter-time consumption mapping relationship; subnet_N neurons are sampled from each layer of the neural network of the base model, and the first subnetwork is trained using a training data set, and the first subnetwork is evaluated using an evaluation data set to obtain model parameters; the sampling is repeated s times, and the model parameters obtained by the current sampling training cover the model parameters obtained by the previous sampling training, until the performance of the base model converges in the overhead line power outage prediction task; Sampling a second subnetwork from the converged base model to obtain a plurality of second subnetworks that meet the time consumption requirement, and evaluating the performance of the second subnetwork based on the evaluation data set and the evaluation criteria to determine the optimal subnetwork, including sampling a second subnetwork from each layer of the neural network in the converged base model in a uniform sampling manner to obtain c second subnetworks that meet the time consumption requirement, On the evaluation data set, the index value of the second sub-network is calculated according to the evaluation criteria; the second sub-network corresponding to the maximum index value is taken as the optimal sub-network; The optimal subnetwork is used as a power outage prediction model to implement the overhead line power outage prediction.
2. The method for searching the neural network structure for predicting power outages of overhead lines according to claim 1, characterized in that: The process of constructing the network parameter-time consumption mapping relationship includes: Setting the parameter space, including the width and depth of the neural network, the number of neurons, and the type of activation function; Traversing the parameter space to obtain k neural networks; The time consumption of the k neural networks is measured, and the network parameter-time consumption mapping relationship is determined as follows: time=T(flops,activation)=T(g(N,N,D),actionvation) Among them, time is the time consumed by the neural network on the actual computing chip obtained through machine calibration, flops is a function of the number of input neurons and output neurons in each layer of the neural network, activation is the neuron activation function, N is the number of input neurons, M is the number of output neurons, and D is the depth of the neural network.
3. The method for searching the neural network structure for predicting power outages of overhead lines according to claim 2, characterized in that: The base model is constructed based on the time consumption requirement and the network parameter-time consumption mapping relationship, including: Based on the network parameter-time consumption mapping relationship, possible network parameters under the time consumption requirement are solved as proposed parameters; According to the proposed parameters, the base model is constructed and expressed as follows: net=Network(L(N,activation),D) Among them, N is the number of input neurons in each layer of the neural network, D is the depth of the neural network, and activation is the neuron activation function.
4. The method for searching the neural network structure for predicting power outages of overhead lines according to claim 1, characterized in that: The evaluation criteria are: acc = number of correctly predicted samples / total number of samples.
5. The method for searching the neural network structure for predicting power outages of overhead lines according to claim 1, characterized in that: The time consumption of the base model is m times the time consumption of the target network scale, where m>1.
6. A neural network structure search device for predicting power outages of overhead lines, characterized in that: The device comprises: A model building module is used to build a base model based on time consumption requirements and a network parameter-time consumption mapping relationship; A training module, used to sample a first sub-network from the base model according to the target network scale, train the first sub-network using a training data set, and evaluate the first sub-network using an evaluation data set, so that the performance of the base model converges in the overhead line power outage prediction task; A search module, configured to sample a second sub-network from the converged base model to obtain a plurality of second sub-networks that meet the time-consuming requirement, and to evaluate the performance of the second sub-network based on the evaluation data set and the evaluation criteria to determine the optimal sub-network; A prediction module, used for using the optimal sub-network as a power outage prediction model to implement the overhead line power outage prediction; The training module includes: According to the target network scale, the representation of the first subnetwork is set as: subnet(subnet_N, D, activation), wherein the depth D and the neuron activation function activation of the first subnetwork are consistent with the base model, the width of the first subnetwork is smaller than the width of the neural network in the base model, and the width subnet_N of the first subnetwork is obtained by the network parameter-time consumption mapping relationship; A first sampling unit is used to sample subnet_N neurons from each layer of the neural network of the base model, and train the first sub-network using a training data set, and evaluate the first sub-network using an evaluation data set to obtain model parameters; A training unit, used for repeating sampling s times, and overwriting the model parameters obtained by the previous sampling training with the model parameters obtained by the current sampling training, until the performance of the base model converges in the overhead line power outage prediction task; The search module includes: The second sampling unit is used to sample the second sub-network from each layer of the neural network in the converged base model in a uniform sampling manner to obtain c second sub-networks that meet the time consumption requirement. An evaluation unit, configured to calculate an indicator value of the second sub-network based on the evaluation data set according to an evaluation standard; A determination unit is used to use the second sub-network corresponding to the maximum indicator value as the optimal sub-network.
7. The overhead line power outage prediction neural network structure search device according to claim 6, characterized in that: The network parameter-time consumption mapping relationship is: time=T(flops,activation)=T(g(N,M,D),actionvation) Among them, time is the time consumed by the neural network on the actual computing chip obtained through machine calibration, flops is a function of the number of input neurons and output neurons in each layer of the neural network, activation is the neuron activation function, N is the number of input neurons in each layer of the neural network, M is the number of output neurons, and D is the depth of the neural network.
8. The overhead line power outage prediction neural network structure search device according to claim 7, characterized in that: The formula of the constructed base model is expressed as follows: net=Network(L(N,activation),D) Among them, N is the number of input neurons in each layer of the neural network, D is the depth of the neural network, and activation is the neuron activation function.
9. The overhead line power outage prediction neural network structure search device according to claim 6, characterized in that: The evaluation criteria are: acc = number of correctly predicted samples / total number of samples.
10. The overhead line power outage prediction neural network structure search device according to claim 6, characterized in that: The time consumption of the base model is m times the time consumption of the target network scale, where m>1.
Citation Information
Patent Citations
Neural network structure searching method and neural network structure searching device
CN111382868A
Super-network training method and device
CN111738418A
Model construction method and device for target detection and neural network architecture
CN112215269A