An inspection model training method and device based on a small-batch merchant gradient system
By adopting the training method of a small batch quotation gradient system in power system inspection, the problem that the model is prone to fall into local minimum points in the prior art is solved, the accuracy and efficiency of the model are improved, and fast and accurate obstacle recognition is achieved.
Patent Information
- Application Number
- CN202510322459.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-19
AI Technical Summary
Existing training methods are prone to falling into the problem of local minimum points or global minimum points, resulting in low model accuracy and difficult to achieve fast and accurate obstacle recognition.
The inspection model training method based on the small batch quotient gradient system is adopted. By constructing the transmission line obstacle patrol data set, the deep learning network is trained using the small batch data set, and a weight correction function is added to solve the overfitting and underfitting problems of unbalanced data sets.
The test accuracy of the model is improved, the memory limitation problem caused by excessive data volume is solved, and the overall performance of the unbalanced data set is improved through the weight correction function, achieving a better obstacle inspection model.
Smart Images

Figure CN119851095B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power inspection, and particularly relates to a method and device for training an inspection model based on a small-batch commercial gradient system. Background Art
[0002] Inspecting transmission lines is one of the daily tasks of the power department. Effective line inspection can promptly detect potential safety hazards of transmission lines and avoid major accidents. However, with the development of the power system, the total mileage of transmission lines has increased exponentially, making the power line inspection task more arduous and significant. It consumes a large amount of manpower and material resources and has become a burden on the power department. There is a need for a method that can automatically identify obstacles to replace manual inspection, reduce input costs, improve inspection efficiency, and achieve the goal of real-time monitoring of the entire network.
[0003] The gradient descent method is the most popular deep network optimization algorithm at present. Every state-of-the-art deep learning library contains various variants of the gradient descent algorithm. Among the many variants of the gradient descent algorithm, the stochastic gradient descent method (SGD) is one of the most widely used in machine learning. SGD obtains gradient estimates from individual training data samples, which can greatly reduce the computational cost of large redundant data sets, is easy to implement, and has a low algorithm complexity. However, because SGD cannot change the step size, the convergence speed of gradient estimation is very slow. Even when the objective function is a strongly convex function, SGD still cannot achieve linear convergence and will fall into local optimal points. The batch gradient descent method (BGD) uses the entire data set to determine the search direction, which can effectively reduce the number of iterations and converge more accurately towards the direction of the extreme value. However, when training on large sample data sets, calculating all samples is required for each iteration, resulting in a slow training process and easy memory overflow. Mini-batch gradient descent (MBGD) reduces the computational cost by an order of magnitude through matrix operations and the use of mini-batches, and can produce better solutions than online stochastic gradients. However, its training effect for unbalanced data is poor.
[0004] Since the gradient descent algorithm iteratively solves for the local or global minimum point of the network by taking the first-order series of the Taylor series, the error is relatively large. To address such problems, existing literature has proven the equivalence between the union of the regular stable equilibrium manifolds of the Quotient Gradient System (QGS) and the entire feasible region of the OPF problem, and further proven that QGS is completely stable and each trajectory converges to an equilibrium manifold. This provides a theoretical basis for the calculation of the optimal power flow solution in power systems. On this basis, existing literature has proposed a state estimation algorithm based on the dynamic system by constructing a quotient gradient dynamic system, making the degenerate stable equilibrium manifold of the system correspond to the solution of WLS, overcoming the limitations of the traditional WLS algorithm, having the advantages of strong robustness, strong convergence, and being able to solve state estimation problems with residual constraints, strictly satisfying the equality constraints and residual inequality constraints, and ensuring the rationality of the solution. Existing literature has given a description and visualization of the feasible region of the combined heat and power system through QGS. Its theoretical basis is the relationship between the entire feasible region of QGS and the union of non-degenerate stable equilibrium manifolds. And starting from any point, the trajectory will converge to an equilibrium manifold QGS, which demonstrates the effectiveness of the trajectory-based method. To solve the partial state estimation problem of the distribution network, existing literature has proposed a solution method for the partial state estimation problem of the distribution network based on QGS. An equivalent relationship has been established between the objective of WLS and the DSEM of QGS, with global convergence. Existing literature has proposed a method based on QGS, which can be used to calculate the feasible solution of the load demand. It does not require matrix inversion, thus avoiding the ill-conditioning of the Jacobian matrix along the trajectory. Summary of the Invention
[0005] The present invention aims to solve at least one of the technical problems in the above related technologies to a certain extent.
[0006] To this end, the purpose of the present invention is to provide an inspection model training method and device based on a mini-batch quotient gradient system, which can solve the problem that the existing training method is prone to falling into local or global minimum points, improve the model accuracy, and achieve fast and accurate obstacle recognition.
[0007] To solve the above technical problems, the present invention is implemented as follows:
[0008] An embodiment of the present invention provides an inspection model training method based on a mini-batch quotient gradient system, and the method includes:
[0009] S1. Construct a transmission line obstacle inspection data set;
[0010] S2. Based on the power transmission line obstacle inspection dataset, the deep learning network is trained using the method based on mini-batch quotient gradient. By reducing the computational complexity, the memory limitation problem when the data volume is too large is solved; by adding a weight correction function, the overfitting in the training of large categories and the underfitting in the training of small categories are solved, and the overall performance of the imbalanced dataset is improved to obtain a better power transmission line obstacle inspection model.
[0011] S3. Use the trained power transmission line obstacle inspection model to identify obstacles during the intelligent inspection of the power system transmission line.
[0012] In addition, according to the inspection model training method based on the mini-batch quotient gradient system of the present invention, the following additional technical features may also be included:
[0013] In some of the embodiments, the method based on mini-batch quotient gradient adopts the idea of training with mini-batch datasets and variable step-size integration to accelerate training, and fits the search direction in the programming manner of Limited memory.
[0014] In some of the embodiments, the solution of the deep learning network model is obtained by tracking the DSEM trajectory of the quotient gradient system.
[0015] In some of the embodiments, the deep learning network is a convolutional neural network, and its equality constraint function is a generalized equality constraint function considering the loss function of the mini-batch, and its formula is:
[0016] ;
[0017] In the formula, represents the convolutional network loss function, represents the measurement residual constraint, S represents the slack variable, z is the objective function, h () is the loss function, x is the neuron weight, H E (x) represents the equality constraint, H I ( x , S ) represents the inequality constraint.
[0018] In some of the embodiments, the solution of the generalized equality constraint function is related to the steady state of the mini-batch quotient gradient system; the way of the steady state connection is:
[0019] ;
[0020] Among them, is 's Jacobian matrix, is the weight correction function, x is the neuron weight, S represents the slack variable.
[0021] In some of these embodiments, the specific way to obtain the solution of the deep learning network model by tracking the DSEM trajectory of the quotient gradient system is as follows:
[0022] ;
[0023] In the formula, x is the neuron weight, is the Jacobian matrix of is the weight correction function, represents the penalty factor, S represents the slack variable.
[0024] In some of these embodiments, the pseudo-transient continuation method is used to quickly calculate the steady-state solution, and the specific content includes:
[0025] Referring to the two-norm results of the quotient gradient system calculated in the previous and subsequent times, the training speed can be accelerated following the correction of the search direction of the quotient gradient system.
[0026] In some of these embodiments, the content of step S1 includes: obtaining the original data of the power system inspection,
[0027] extracting features from the original data using the resnet-50 network, extracting two-dimensional features using the principal component analysis method, and then classifying the data according to the two-dimensional features to obtain the required inspection data set of transmission line obstacles.
[0028] In some of these embodiments, a weight correction function is added during data classification to solve the problem of large errors caused by data underfitting; the weight correction function is a weight correction function considering the loss value.
[0029] In some of these embodiments, when using the mini-batch quotient gradient system for model training, the mini-batch size is 128, and the number of iterations for each mini-batch is 20.
[0030] The embodiment of the present invention also provides an inspection model training device based on the mini-batch quotient gradient system, including a processor and a memory. A software program is stored on the memory, and when the processor runs the software program, it can implement the content of the inspection model training method based on the mini-batch quotient gradient system described above.
[0031] Compared with the prior art, the present invention has at least the following beneficial effects:
[0032] In the embodiments of the present invention, the provided inspection model training method based on the small batch quotient gradient system is trained using the model training method based on Minibatch QGS, which well solves the problem that when the data volume is large enough, due to memory limitations, QGS cannot normally train data, and can greatly improve the test accuracy of the trained model.
[0033] The inspection model training device based on the small batch quotient gradient system of the present invention can implement the content of the inspection model training method based on the small batch quotient gradient system, and thus at least has all the features and advantages of the inspection model training method based on the small batch quotient gradient system, which will not be elaborated here. The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a flowchart of the inspection model training method based on the small batch quotient gradient system disclosed in an embodiment of the present invention;
[0035] Figure 2 is disclosed in an embodiment of the present invention specific correction diagram of the function;
[0036] Figure 3 is a test error curve diagram of SGD and Minibatch QGS on the test set disclosed in an embodiment of the present invention;
[0037] Figure 4 is a two-dimensional feature distribution diagram of the power system inspection data set disclosed in an embodiment of the present invention;
[0038] Figure 5 is a diagram of the proportion of each category of data in the power system inspection data set disclosed in an embodiment of the present invention;
[0039] Figure 6 is a test error curve diagram of SGD and Minibatch QGS on the test set disclosed in another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] Next, the embodiments of the present invention will be described in detail through specific embodiments and their application scenarios in conjunction with the accompanying drawings.
[0042] Please refer to Figure 1 As shown, in some embodiments of the present invention, a training method for an inspection model based on a small-batch quotient gradient system is provided, including constructing a power transmission line obstacle inspection data set; based on the power transmission line obstacle inspection data set, using the inspection model training method based on the small-batch quotient gradient system to train a deep learning network, solving the memory limitation problem when the data volume is too large by reducing the computational complexity; by adding a weight correction function, solving the overfitting problem in the training of large categories and the underfitting problem in the training of small categories, improving the overall performance of the class-imbalanced data set to obtain a better power transmission line obstacle inspection model; using the trained power transmission line obstacle inspection model to identify obstacles in the intelligent inspection process of the power system transmission line; the inspection model training method based on the small-batch quotient gradient system adopts the idea of training with a small-batch data set and variable step-size integration to accelerate training, and fits the search direction in the programming manner of Limited memory.
[0043] Furthermore, QGS is applied to the training of deep networks and some improvement work has been done. The content includes:
[0044] 1) QGS is applied to the training of deep learning networks;
[0045] 2) Adopting the idea of mini-batch and variable step-size integration to accelerate training;
[0046] 3) Using the programming manner of Limited memory to fit the search direction and training large networks.
[0047] The network with 3.87 million parameters is trained by mini-batch QGS, and the accuracy of the MNIST data set is 99.7%, and the accuracy of the power system inspection data set reaches 96.3%.
[0048] The present invention first introduces the objective function of the convolutional neural network below; then introduces the concept of mini-batch QGS and proposes a deep network training method for imbalanced data sets; finally, tests the deep learning network based on mini-batch QGS.
[0049] (1) Explanation of the objective function of the convolutional neural network and the power system inspection data set
[0050] The convolutional neural network is a deep feedforward neural network model. Compared with other neural network structures, the convolutional neural network requires relatively fewer parameters and is widely used. The CNN has a multi-layer network structure, and its basic structure mainly includes an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The formula is as follows:
[0051] ;
[0052] In the formula, b is the offset, and represent the convolutional input and output in the l +1-th layer, generating a feature map, is the dimension of. represents the pixel at a certain point on the feature map, K is the number of channels of the feature map, f、 and p are the sizes of the convolutional kernels respectively.
[0053] The convolutional layer is also called the feature extraction layer and is the core of the CNN. Its main function is to perform convolutional calculation processing on the input data and complete the process of extracting image features through local calculations with different convolutional kernels. After the convolutional operation is completed, it enters the pooling layer. Pooling is achieved through a pooling kernel. Generally, a continuous range in the image is selected as the pooling area, and the features generated by the same hidden units are clustered and statistically analyzed. Therefore, these pooling units have translational invariance, which is also the most important feature of the pooling layer.
[0054] The optimized model of the CNN is as follows:
[0055] (4)
[0056] In the formula, represents the loss function of the convolutional network, and β represents the measurement residual constraint.
[0057] To convert the constraint formula into an unconstrained formula, a slack variable is added to the inequality in Equation (4), and the penalty factor is used as the measurement residual constraint:
[0058] ;
[0059] Among them, S represents the slack variable. Under the penalty factor , the measurement residuals with higher precision are still not satisfactory. To overcome this shortcoming, the QGS-based method will be adopted.
[0060] Explanation of the concept of mini-batch QGS and the training method of the deep network for unbalanced data sets
[0061] Some dynamic systems have many excellent characteristics, such as the absence of convergence and near-convergence problems. If a non-linear dynamic system can be constructed such that the local optimal solution of the problem to be solved corresponds to the stable equilibrium point of a suitable dynamic system, then the problem is transformed into the solution of the dynamic system. How to construct an effective dynamic system and solve it through its properties is the key to this method. The specific construction method and its properties will be introduced below.
[0062] By adding the mini-batch loss function to Equation (4), the following generalized equality constraint function is generated:
[0063] (5)
[0064] Where, represents the convolutional network loss function, represents the measurement residual constraint, S represents the slack variable, z is the objective function, h () is the loss function, x is the neuron weight, H E (x) represents the equality constraint, H I ( x , S ) represents the inequality constraint.
[0065] In the proposed method, a non-linear dynamic system based on the equality constraint (5) is derived, and the solution of (5) is related to the (stable) steady state of the QGS system:
[0066] (6)
[0067] Where, is 's Jacobian matrix, is the weight correction function, x is the neuron weight, S represents the slack variable.
[0068] Theorem 1: Local optimal
[0069] Any stable equilibrium manifold of the quotient gradient system (6), denoted as , is a local minimum point of the function .
[0070] Definition 3: Degenerate stable equilibrium manifold
[0071] For a stable equilibrium manifold of the quotient gradient system (6), if and , then It is called a degenerate stable equilibrium manifold.
[0072] According to Theorem 1 and Definition 3, a degenerate stable equilibrium manifold of the quotient gradient system is a non-zero local minimum point of the function .
[0073] Since:
[0074] (7)
[0075] is also a local minimum. Therefore, by tracking the DSEM trajectory of the QGS, the solution of the deep learning network optimization model can be obtained.
[0076] Therefore, Problem (4) can be solved by solving the QGS:
[0077] (8)
[0078] Balance analysis:
[0079] In the power system inspection dataset, the imbalance in the amount of data between different categories leads to overfitting in the training of large categories or underfitting in the training of small categories. In traditional methods, a smaller weight is often given to data with larger errors. However, in the training of deep networks, since the power system inspection training dataset is carefully labeled manually and the possibility of errors is extremely low, when there are data with larger errors, it is mainly due to underfitting, and a larger weight will be given to increase its contribution to the training gradient. The present invention adds a weight correction function as follows:
[0080] ;
[0081] The specific correction diagram of the function is as Figure 2 shown.
[0082] The present invention proposes a new loss weight correction function. Setting γ can reduce the relative loss of large sample examples for classification and focus more on small sample classification examples.
[0083] Intuitively, this weight function can automatically reduce the weight of large sample examples during the training process and quickly focus the model on small category examples.
[0084] The mini-batch QGS small batch training proposed by the present invention calculates the gradient on a small part of the dataset, which well solves the problem that when the amount of data is large enough, due to memory limitations, the QGS cannot normally train the data.
[0085] Doing so has two advantages:
[0086] 1) Reduce the variance of parameter updates, so that convergence can be more stable;
[0087] 2) Compared with the classical batch algorithm, the minibatch processing algorithm reduces the computational cost by an order of magnitude and produces better solutions than online learning.
[0088] In order to quickly calculate the steady-state solution, the present invention adopts a technique of a pseudo-transient continuation method (abbreviated as PTC).
[0089] The present invention corrects the step size according to the magnitude of the QGS two-norm. By referring to the QGS two-norm results of the previous and subsequent calculations, the training speed can be accelerated following the correction of the QGS search direction.
[0090] ;
[0091] The present invention adopts a method of restricting memory to find the iteration direction, and approximates the matrix B through the matrices s and y (V is an m*n matrix, m << n).
[0092] ;
[0093] Among them, represents the initial value at the k-th iteration, , represents the weight at the k-th iteration, , , .
[0094] (3) Testing of the deep learning network based on mini-batch QGS
[0095] When testing the performance of the deep learning network training algorithm, the datasets used are the standard MNIST dataset and the CIFAR-10 dataset. In order to further detect the performance of the training method based on the small-batch entropy gradient system in the field of power systems, the present invention creates a transmission line obstacle inspection dataset. This dataset contains 43,000 images, the image size is 32*32*3, and there are six classifications: crane, tower crane, construction machinery, foreign object on the wire, dump truck. Among them, there are 33,000 images in the training set and 10,000 images in the test set.
[0096] The experiment was carried out using 1 machine; 1 Intel CPU core (2.67 GHz) and 1 GPU of GeForce RTX2080TI. The experiment was carried out using python and its GPU plugin.
[0097] In the following experiments, standard metrics in machine learning are shown: the objective function evaluated over time based on test data (i.e., test error). The objective function evaluated during training shows a similar trend.
[0098] The computer is used to calculate SGD and Minibatch QGS to train a convolutional neural network. The convolutional neural network has two convolutional layers, two Pooling layers, and one fully connected layer. Therefore, the model has approximately 3.22*10^7 (i.e., 3.22 multiplied by 10 to the 7th power) parameters, which poses a challenge for general optimization methods.
[0099] For the Minibatch QGS algorithm, there are two optimization parameters: the mini-batch size and the number of iterations per mini-batch. The Minibatch was replaced after 20 iterations, and the size of the adopted minibatch is 128. Through a large number of experiments, it is proved that the above parameters are relatively effective.
[0100] For Minibatch QGS, the present invention uses the above fixed parameters for training; while for SGD, the present invention has tried 10 combinations of optimization parameters, including changing the mini-batch size among {1, 10, 100, 1000}, and taking the number of iteration steps as a reference, using values between 0.001 - 1 for variable step sizes with different strategies.
[0101] Comparing the test errors of SGD and Minibatch QGS on the test set, the summary of the best results is as Figure 3 shown. The results show that Minibatch converges faster than carefully tuned ordinary SGD. In particular, the performance of Minibatch QGS is better, with a maximum test accuracy of 99.7%.
[0102] The resnet-50 network is used to extract features from the transmission line obstacle inspection dataset, and the principal component analysis method is used to extract two-dimensional features, as Figure 4 shown: It can be seen from the figure that the dispersion and intersection of construction machinery, foreign objects on the wire, and fireworks are very high; counting the data proportions of each category in the transmission line obstacle inspection dataset, as Figure 5 shown: It can be seen from the figure that the data volumes of each category are very unbalanced, with the largest data proportion reaching 40%, and the smallest being only 2%, because samples such as wildfires and kites hanging on transmission lines are very difficult to collect. Based on the above analysis, this classification task is very difficult.
[0103] The computer is used to calculate SGD and Minibatch QGS to train a convolutional neural network. The convolutional neural network has two convolutional layers, two Pooling layers, and one fully connected layer. Therefore, the model has approximately 3.87*10^7 parameters to fit this dataset.
[0104] For the Minibatch QGS algorithm, the Minibatch was replaced after 20 iterations, and the size of the minibatch used was 128. Compare the test errors of SGD and Minibatch QGS on the test set and summarize the best results in Figure 6 . The results show that the Minibatch converges faster than the well-tuned ordinary SGD. In particular, the performance of Minibatch QGS is better, with a maximum test accuracy of 96.1%.
[0105] Table 1
[0106] Major hazard category code Test accuracy Crane 98.4% Tower crane 99.6% Construction machinery 97.3% Foreign object on conductor 94.7 Smoke and fire 92.2% Skip 93.5%
[0107] The results show that the Minibatch converges faster than the well-tuned ordinary SGD. In particular, the performance of Minibatch QGS is better, with a maximum test accuracy of 96.1%.
[0108] For the parts not described in detail in the present invention, reference can be made to the prior art or the well-known techniques to those skilled in the art.
[0109] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention, and all of them fall within the protection scope of the present invention.
Claims
1. A method for training an inspection model based on a small batch quotient gradient system, characterized in that: The method comprises: S1. Construct a transmission line obstacle inspection dataset; the inspection dataset is a picture dataset; S2. Based on the transmission line obstacle inspection dataset, a small batch quotient gradient-based method is used to train the deep learning network. By reducing the computational complexity, the memory limitation problem when the data volume is too large is solved. By adding a weight correction function, the overfitting problem of large category training and the underfitting problem of small category training are solved, and the overall performance of the class imbalanced dataset is improved to obtain a better transmission line obstacle inspection model. S3, using the trained transmission line obstacle inspection model to identify obstacles during the intelligent inspection of power system transmission lines; The method based on mini-batch quotient gradient adopts the idea of mini-batch data set training and variable step-size integral accelerated training, and fits the search direction in a limited memory programming way; By tracking the DSEM trajectory of the quotient gradient system, the solution of the deep learning network model is obtained; The deep learning network is a convolutional neural network, and its equality constraint function is a generalized equality constraint function that takes into account the loss function of a small batch, and its formula is: y=(x,S) In the formula, h zero (x) represents the convolutional network loss function, β represents the measurement residual constraint, s represents the slack variable, z is the objective function, h() is the loss function, x is the neuron weight, HE(x) represents the equality constraint, H I (x, S) represents an inequality constraint; The solution of the generalized equality constraint function is steadily connected to the small batch quotient gradient system; the steady state connection is in the form of: y=(x,S) Where DH(y) is the Jacobian matrix of H(y), W -1 is the weight correction function, x is the neuron weight, and s is the slack variable.
2. The inspection model training method based on a small batch quotient gradient system according to claim 1 is characterized in that: The specific method of obtaining the deep learning network model solution by tracking the DSEM trajectory of the quotient gradient system is: y=(x,S) Where x is the neuron weight, DH E (x) is H E The Jacobian matrix of (x), is the weight correction function, α is the penalty factor, and S is the slack variable.
3. The inspection model training method based on a small batch quotient gradient system according to claim 1 is characterized in that: The pseudo-transient continuous method is used to quickly calculate the steady-state solution, including: By referring to the second norm results of the quotient gradient system calculated twice before and after, the training speed can be accelerated by following the correction of the search direction of the quotient gradient system.
4. The inspection model training method based on a small batch quotient gradient system according to claim 1 is characterized in that: Step S1 includes: obtaining the original data of the power system inspection, The resnet-50 network is used to extract features from the original data, and the principal component analysis method is used to extract two-dimensional features. Then the data is classified according to the two-dimensional features to obtain the required transmission line obstacle inspection data set.
5. The inspection model training method based on a small batch quotient gradient system according to claim 1 is characterized in that: When performing data classification, a weight correction function is added to solve the problem of large errors caused by under-fitting of data; the weight correction function is a weight correction function that takes the loss value into consideration.
6. A patrol model training device based on a small batch quotient gradient system, comprising a processor and a memory, wherein a software program is stored in the memory, characterized in that: When the processor runs the software program, it can implement the content of the inspection model training method based on a small batch quotient gradient system as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Boundary determination method and device of steady-state security domain, equipment, medium and product
CN119294930A
Optimal sparse placement of phasor measurement units and state estimation of key buses in distribution networks
US20210083474A1