A fast load characteristics matching method based on AdaBoost model
Through the fast matching method of load characteristics based on the AdaBoost model, the problem of low accuracy of existing load recognition algorithms in complex scenarios is solved, and the fast, efficient and accurate matching of load characteristics is achieved, which improves training speed and accuracy.
Patent Information
- Application Number
- CN202011391968.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-01
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2040-12-01
AI Technical Summary
The existing load recognition algorithm has low accuracy when dealing with complex scenarios, low solution efficiency for mathematical optimization, and insufficient accuracy of load recognition algorithms based on unsupervised learning.
The fast matching method of load characteristics based on the AdaBoost model is adopted. By obtaining the graphical feature data of the load V-I curve trajectory, the weak recognizer weight distribution is initialized, data training is performed, and an enhanced strong recognizer is generated, which is used to quickly, efficiently and accurately match the load characteristics.
It speeds up the training speed, improves the training accuracy, and can quickly, efficiently and accurately match load characteristics, which has good economic benefits and practical value.
Smart Images

Figure CN112560906B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electric load monitoring, and in particular to a load characteristic fast matching method based on an AdaBoost model. Background Art
[0002] With the continuous progress of ICT technology, the intelligent and digital construction of distribution networks has been accelerated, and the development and application of the new generation of power collection systems have gained new driving forces. Intelligent power consumption and intelligent management of power networks have become the trend of future development. As the most basic link to realize intelligent power consumption and intelligent management in the context of information-based and digitalized power networks, the innovation and progress of load identification algorithms are of great value. At present, mathematical optimization and pattern recognition are the two major methods for solving load identification and decomposition problems, but the efficiency of mathematical optimization is low, and accurate identification requires a complete load feature library, which is often difficult to meet in practice; in pattern recognition, supervised learning-based load identification algorithms emerge in an endless stream, but they involve few types of loads and deal with simple scenarios. The accuracy of unsupervised learning-based load identification algorithms is low. Summary of the invention
[0003] The technical problem to be solved and the technical task proposed by the present invention are to improve and perfect the existing technical solutions and provide a load feature fast matching method based on the AdaBoost model to achieve the purpose of accelerating the training speed and improving the training accuracy. To this end, the present invention adopts the following technical solutions.
[0004] A method for fast matching of load characteristics based on an AdaBoost model, characterized in that it comprises the following steps:
[0005] 1) Obtain data, including training data, label data, and weak identifier G m (x) and the number of iterations M; the training data is the graphical feature data in the load VI curve trajectory; the training data and the label data are input into the training data set to form a training data set, wherein the training data set is:
[0006] T={(x 1 ,y 1 ),(x 2 ,y 2 ),...,(x n ,y n )} (1)
[0007] Where: T is the training data set; x i is the load characteristic; i is the label corresponding to the load feature. The label refers to the working status of the equipment. When it is -1, it means it is in the stopped state, and when it is 1, it means it is working. y i ∈Y=-1,1, the number of iterations is M
[0008] 2) Initialize the weight distribution of weak identifiers;
[0009] 3) Data training to obtain weak identifiers;
[0010] 4) Determine whether the recognition error rate is less than or equal to the set threshold, or greater than the number of iterations; if so, proceed to the next step, if not, return to step 3);
[0011] 5) Calculate the weight of the weak identifier in the strong identifier;
[0012] 6) Update the weight of weak identifier;
[0013] 7) Obtaining a strong identifier for identifying the working state of the corresponding load device according to the weight of the weak identifier;
[0014] 8) When identification is required, the acquired new identification sample data is input into the strong identifier, and the strong identifier identifies the identification sample data and outputs the load type of the corresponding equipment.
[0015] This technical solution uses the graphic features in the load VI curve trajectory as identification samples, defines identification labels for each load, selects appropriate weak identifiers, generates enhanced strong identifiers, and uses the enhanced identifiers to identify new sample data; after the entire process is completed, the existing feature library can be used to train an identifier F(x) that can identify different load devices, and after inputting the newly added sample data x, the identifier can output the corresponding device working status. The use of the AdaBoost identifier can extract important training data features, exclude some unnecessary training data features, and place them on key training data, reducing the training dimension, speeding up the training, and improving the training accuracy. It can quickly, efficiently and accurately match load features, and has good economic benefits and practical value.
[0016] As a preferred technical means: Step 2) Initialize the weight distribution of the weak identifier as:
[0017]
[0018] Where: D 1 is the weight distribution set for initializing training samples, w i,n is the nth training weight of the i-th weak identifier.
[0019] As a preferred technical means: in step 4), the recognition error rate of the weak identifier on the training data set is calculated:
[0020]
[0021] Where: e m is the recognition error rate, G m (x) is the mth weak identifier.
[0022] As a preferred technical means: in step 5), the weight of the weak identifier in the strong identifier is calculated:
[0023]
[0024] Where: α m is the weight of the mth weak identifier in the strong identifier.
[0025] As a preferred technical means: in step 6), update the weight distribution of the training data set:
[0026]
[0027]
[0028] Where: z m is a normalization factor used to make the probability distribution of samples sum to 1.
[0029] As a preferred technical means: in step 7), the strong identifier expression is:
[0030]
[0031] Where: F is a strong identifier, x is the newly added sample data.
[0032] As a preferred technical means: in step 1), the data collection of the graphical characteristic data in the load VI curve trajectory includes the following steps:
[0033] 101) Total electricity load data collection;
[0034] The voltage and current data of a single electrical appliance are collected through an oscilloscope, and the VI trajectory curve is drawn;
[0035] 102) Data preprocessing;
[0036] Data preprocessing includes current and voltage standardization; current and voltage standardization is achieved by dividing the current and voltage signals by their root mean square respectively; the current and voltage data after standardization are plotted with voltage as the horizontal axis and current as the vertical axis to draw a current and voltage VI trajectory curve;
[0037] The current and voltage standardization calculation formula is:
[0038]
[0039] v norm(n) =vn / V rms (1≤n≤N) (9)
[0040]
[0041] i norm(n) =i n / I rms (1≤n≤N) (11)
[0042] Where: N is the total number of points in the collected trajectory curve, V rms is the RMS voltage in the VI trace, V norm(n) The voltage v = v n The standard voltage at rms is the RMS current in the VI trajectory curve, i norm(n) For current i=i n The standard current when
[0043] 103) VI feature extraction;
[0044] After preprocessing the collected data, the VI trajectory area is made using the preprocessed voltage and current signals and the characteristic values of the trajectory area are extracted;
[0045] 104) Acquire a feature data set and store it to form a feature library to provide data support for user load identification;
[0046] The characteristic data set includes: closed area, average curve distortion, number of self-intersection points, slope of the middle section of the average curve, and circulation direction; the VI waveform characteristic data set is obtained as follows: X = {A norm , DMC, N ip , tanθ, K}.
[0047] For the same or similar power loads, the current and voltage trajectories of this technical solution are similar. Due to the difference in electrical power consumption, the VI trajectories will form different shapes. By extracting features from the VI trajectories and classifying them, user behavior perception can be achieved. This technical solution draws a VI trajectory curve based on the collected voltage and current data. Through the VI trajectory curve, combined with the feature extraction model proposed in the present invention, a feature data set is finally obtained to provide data support for user power behavior perception. The feature extraction speed is fast and the implementation is simple. Multiple features jointly participate in user behavior perception with high accuracy. It can provide data support for user load identification and has good economic benefits and practical value.
[0048] As a preferred technical means: in step 103), VI feature extraction includes:
[0049] 1) Closed area:
[0050]
[0051] Where: A norm is the area covered by the VI trajectory curve, V L is the minimum voltage in the curve, V R is the maximum voltage in the curve, v n is the voltage value at any point between the minimum and maximum voltages, i nu is the voltage v in the trajectory curve n The larger of the two corresponding current values, i nd is the voltage v in the trajectory curve n The smaller current value of the two corresponding current values;
[0052] 2) Distortion of the average curve:
[0053]
[0054] Where: m n v n The average current of the trajectory curve at;
[0055] Connect the first and last coordinates of the VI trajectory curve to obtain a straight line, whose equation is shown in formula (14):
[0056]
[0057] Where: I L V L The current value of the trajectory curve, I R V R The current value of the trajectory curve;
[0058] DMC=max(|m n -i n ′|) (15)
[0059] Where: DMC is the distortion of the average curve;
[0060] 3) Number of self-intersection points:
[0061] Get the number of self-intersection points N in the VI trajectory curve ip ;
[0062] 304) The slope of the middle section of the average curve:
[0063] The slope of the average curve near zero point in the middle section can characterize the power-electronic characteristics of the electrical equipment, which is used to represent the angle between the tangent of the average curve near zero point and the V axis, as shown in formula (9):
[0064]
[0065] Where: θ is the angle between the tangent line of the average curve near zero point of the VI trajectory curve and the V axis, The average curve is close to zero Current value is the point to the right of zero The current value;
[0066] 4) Circulation direction:
[0067] "Circulation direction" refers to the average curvature of the VI trajectory in counterclockwise or clockwise time. The clockwise curvature refers to the total capacitive load behavior, and the counterclockwise curvature refers to the total inductive load characteristics. The calculation model of the curvature is as follows:
[0068]
[0069]
[0070] Where: k n is the curvature of any point of the VI trajectory curve, i n 、i n+1 、i n+2 They are v=v n 、v=v n+1 、v=v n+2 The corresponding current value.
[0071] Beneficial effects: This technical solution uses the graphic features in the load VI curve trajectory as identification samples, defines identification labels for each load, selects appropriate weak identifiers, generates enhanced strong identifiers, and uses the enhanced identifiers to identify new sample data; after the entire process is completed, the existing feature library can be used to train an identifier F(x) that can identify different load devices, and after inputting the newly added sample data x, the identifier can output the corresponding equipment working status. The use of the AdaBoost identifier can extract important training data features, exclude some unnecessary training data features, and place them on key training data, reducing the training dimension, speeding up the training, and improving the training accuracy. It can quickly, efficiently and accurately match load features, and has good economic benefits and practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 It is a flow chart of the present invention.
[0073] Figure 2 It is a flowchart of constructing a strong identifier of the present invention.
[0074] Figure 3It is a flow chart of data collection of graphic characteristics in the load VI curve trajectory of the present invention. DETAILED DESCRIPTION
[0075] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings.
[0076] Embodiment 1:
[0077] like Figure 1 As shown, the present invention comprises the following steps:
[0078] 1) Obtaining data, including training data, label data, weak identifiers and iteration times; the training data is the graphical feature data in the load VI curve trajectory; inputting the training data and label data into a training data set to form a training data set, wherein the training data set is:
[0079] T={(x 1 ,y 1 ),(x 2 ,y 2 ),...,(x n ,y n )} (1)
[0080] Where: T is the training data set; x i is the load characteristic; i is the label corresponding to the load feature. The label refers to the working status of the equipment. When it is -1, it means it is in the stopped state, and when it is 1, it means it is working.
[0081] 2) Initialize the weight distribution of weak identifiers;
[0082]
[0083] Where: D 1 is the weight distribution set for initializing training samples, w i,n is the nth training weight of the i-th weak identifier.
[0084] 3) Data training to obtain weak identifiers;
[0085] 4) Determine whether the recognition error rate is less than or equal to the set threshold; if so, proceed to the next step, if not, return to step 3); the recognition error rate calculation formula is:
[0086]
[0087] Where: e m is the recognition error rate, G m (x) is the mth weak identifier.
[0088] 5) Calculate the weight of the weak identifier in the strong identifier;
[0089]
[0090] Where: α m is the weight of the mth weak identifier in the strong identifier.
[0091] 6) Update the weight of weak identifier;
[0092]
[0093]
[0094] Where: z m is the normalization factor (to make the probability distribution of the samples sum to 1).
[0095] 7) Obtaining a strong identifier for identifying the working state of the corresponding load device according to the weight of the weak identifier;
[0096]
[0097] Where: F is a strong identifier, x is the newly added sample data.
[0098] 8) When identification is required, the acquired new identification sample data is input into the strong identifier, and the strong identifier identifies the identification sample data and outputs the load type of the corresponding equipment.
[0099] The present invention can train an identifier F(x) that can identify different load devices through the existing feature library, and the identifier can output the corresponding device working status after inputting the newly added sample data x.
[0100] like Figure 2 The figure shows the generation process of the strong identifier. The AdaBoost identifier can be used to extract important training data features, exclude some unnecessary training data features, and put them on the key training data, which reduces the dimension of training, speeds up training, and improves training accuracy. It can quickly, efficiently and accurately match load characteristics, and has good economic benefits and practical value.
[0101] Embodiment 2: The same parts as the embodiment will not be repeated, the difference is that:
[0102] like Figure 3 As shown, the data collection of the graphical characteristic data in the load VI curve trajectory includes the following steps:
[0103] 101) Total electricity load data collection;
[0104] The voltage and current data of a single electrical appliance are collected through an oscilloscope, and the VI trajectory curve is drawn;
[0105] 102) Data preprocessing;
[0106] Data preprocessing includes current and voltage standardization; current and voltage standardization is achieved by dividing the current and voltage signals by their root mean square respectively; the current and voltage data after standardization are plotted with voltage as the horizontal axis and current as the vertical axis to draw a current and voltage VI trajectory curve;
[0107] The current and voltage standardization calculation formula is:
[0108]
[0109] v norm(n) =v n / V rms (1≤n≤N) (9)
[0110]
[0111] i norm(n) =i n / I rms (1≤n≤N) (11)
[0112] Where: N is the total number of points in the collected trajectory curve, V rms is the RMS voltage in the VI trace, V norm(n) The voltage v = v n The standard voltage at rms is the RMS current in the VI trajectory curve, i norm(n) For current i=i n The standard current when
[0113] 103)VI feature extraction;
[0114] After preprocessing the collected data, the VI trajectory area is made using the preprocessed voltage and current signals and the characteristic values of the trajectory area are extracted;
[0115] 104) Acquire a feature data set and store it to form a feature library to provide data support for user load identification;
[0116] The characteristic data set includes: closed area, average curve distortion, number of self-intersection points, slope of the middle section of the average curve, and circulation direction; the VI waveform characteristic data set is obtained as follows: X = {A norm , DMC, N ip , tanθ, K}.
[0117] Among them, VI feature extraction includes:
[0118] 1) Closed area:
[0119]
[0120] Where: A norm is the area covered by the VI trajectory curve, V L is the minimum voltage in the curve, V R is the maximum voltage in the curve, v n is the voltage value at any point between the minimum and maximum voltages, i nu is the voltage v in the trajectory curve n The larger of the two corresponding current values, i nd is the voltage v in the trajectory curve n The smaller current value of the two corresponding current values;
[0121] 2) Distortion of the average curve:
[0122]
[0123] Where: m n v n The average current of the trajectory curve at;
[0124] Connect the first and last coordinates of the VI trajectory curve to obtain a straight line, whose equation is shown in formula (14):
[0125]
[0126] Where: I L V L The current value of the trajectory curve, I R V R The current value of the trajectory curve;
[0127] DMC=max(|m n -i n ′|) (15)
[0128] Where: DMC is the distortion of the average curve;
[0129] 3) Number of self-intersection points:
[0130] Get the number of self-intersection points N in the VI trajectory curve ip ;
[0131] 304) The slope of the middle section of the average curve:
[0132] The slope of the average curve near zero point in the middle section can characterize the power-electronic characteristics of the electrical equipment, which is used to represent the angle between the tangent of the average curve near zero point and the V axis, as shown in formula (9):
[0133]
[0134] Where: θ is the angle between the tangent line of the average curve near zero point of the VI trajectory curve and the V axis, The average curve is close to zero Current value is the point to the right of zero The current value;
[0135] 4) Circulation direction:
[0136] "Circulation direction" refers to the average curvature of the VI trajectory in counterclockwise or clockwise time. The clockwise curvature refers to the total capacitive load behavior, and the counterclockwise curvature refers to the total inductive load characteristics. The calculation model of the curvature is as follows:
[0137]
[0138]
[0139] Where: k n is the curvature of any point of the VI trajectory curve, i n 、i n+1 、i n+2 They are v=v n 、v=v n+1 、v=v n+2 The corresponding current value.
[0140] In the process of collecting the graphic feature data in the load VI curve trajectory of this embodiment, the current-voltage trajectory curve is collected, and the closed area of the trajectory curve, the average curve distortion, the number of curve self-intersection points, the current-voltage trajectory curve, the slope of the near-zero point of the middle section of the average curve, and the clockwise or counterclockwise average curvature of the VI trajectory are considered. The features are extracted mathematically to obtain the feature data set for user behavior perception analysis. The feature extraction speed is fast and the implementation is simple. Multiple features participate in user behavior perception together with high accuracy. It can provide data support for user load identification and has good economic benefits and practical value.
[0141] above Figure 1 , 2 The load feature fast matching method based on the AdaBoost model shown in 3 is a specific embodiment of the present invention, which has embodied the essential characteristics and progress of the present invention. According to actual use needs and under the guidance of the present invention, equivalent modifications in shape, structure, etc. can be made to it, which are all within the protection scope of this scheme.
Claims
1. A fast load feature matching method based on AdaBoost model, Features: The following steps are involved: 1) Obtain data, including training data, label data, weak identifiers, and number of iterations; The training data is the graphical feature data in the load VI curve trajectory; the training data and label data are input into the training data set to form a training data set, where the training data set is: T={(x 1 ,and 1 ),(x 2 ,and 2 ),…,(x n ,and n )} (1) Where: T is the training data set; x i is the load characteristic; i is the label corresponding to the load feature. The label refers to the working status of the equipment. When it is -1, it means it is in the stopped state, and when it is 1, it means it is working. 2) Initialize the weight distribution of weak identifiers; 3) Data training to obtain weak identifiers; 4) Determine whether the recognition error rate is less than or equal to the set threshold, or greater than the number of iterations; if so, proceed to the next step, if not, return to step 3); 5) Calculate the weight of the weak identifier in the strong identifier; 6) Update the weight of weak identifier; 7) Obtaining a strong identifier for identifying the working state of the corresponding load device according to the weight of the weak identifier; 8) When identification is required, the new identification sample data is input into the strong identifier, and the strong identifier identifies the identification sample data and outputs the load type of the corresponding equipment; In step 1), data collection of graphical characteristic data in the load VI curve trajectory includes the following steps: 101) Total electricity load data collection; The voltage and current data of a single electrical appliance are collected through an oscilloscope, and the VI trajectory curve is drawn; 102) Data preprocessing; Data preprocessing includes current and voltage standardization; current and voltage standardization is achieved by dividing the current and voltage signals by their root mean square respectively; the current and voltage data after standardization are plotted with voltage as the horizontal axis and current as the vertical axis to draw a current and voltage VI trajectory curve; The current and voltage standardization calculation formula is: Where: N is the total number of points in the collected trajectory curve, V rms is the RMS voltage in the VI trace, V norm(n) The voltage v = v n The standard voltage at rms is the RMS current in the VI trajectory curve, i norm(n) For current i=i n The standard current when 103) VI feature extraction; After preprocessing the collected data, the VI trajectory area is made using the preprocessed voltage and current signals and the characteristic values of the trajectory area are extracted; 104) Obtaining a feature data set; The characteristic data set includes: closed area, average curve distortion, number of self-intersection points, slope of the middle section of the average curve, and circulation direction; the VI waveform characteristic data set is obtained as follows: X = {A norm , DMC, N ip ,tanθ,K}; In step 103), VI feature extraction includes: 1) Closed area: Where: A norm is the area covered by the VI trajectory curve, V L is the minimum voltage in the curve, V R is the maximum voltage in the curve, v n is the voltage value at any point between the minimum and maximum voltages, i nu is the voltage v in the trajectory curve n The larger of the two corresponding current values, i nd is the voltage v in the trajectory curve n The smaller current value of the two corresponding current values; 2) Distortion of the average curve: Where: m n v n The average current of the trajectory curve at; Connect the first and last coordinates of the VI trajectory curve to obtain a straight line, whose equation is shown in formula (14): Where: I L V L The current value of the trajectory curve, I R V R The current value of the trajectory curve; DMC=max(|m n -i n ′|) (15) Where: DMC is the distortion of the average curve; 3) Number of self-intersection points: Get the number of self-intersection points N in the VI trajectory curve ip ; 304) The slope of the middle section of the average curve: The slope of the average curve near zero point in the middle section can characterize the power-electronic characteristics of the electrical equipment, which is used to represent the angle between the tangent of the average curve near zero point and the V axis, as shown in formula (9): Where: θ is the angle between the tangent line of the average curve near zero point of the VI trajectory curve and the V axis, The average curve is close to zero Current value is the point to the right of zero The current value; 4) Circulation direction: "Circulation direction" refers to the average curvature of the VI trace in counterclockwise or clockwise time. The clockwise curvature refers to the total capacitive load behavior, and the counterclockwise curvature refers to the total inductive load characteristics. The calculation model of curvature is as follows: Where: k n is the curvature of any point of the VI trajectory curve, i n 、i n+1 、i n+2 They are v=v n 、v=v n+1 、v=v n+2 The corresponding current value.
2. According to claim 1, a method for rapid matching of load characteristics based on the AdaBoost model, Features: Step 2) Initialize the weight distribution of the weak identifier as: Where: D 1 is the weight distribution set for initializing training samples, w i,n is the nth training weight of the i-th weak identifier.
3. According to claim 2, a method for rapid matching of load characteristics based on the AdaBoost model, Features: In step 4), the recognition error rate of the weak recognizer on the training data set is calculated: Where: e m is the recognition error rate, G m (x) is the mth weak identifier.
4. According to claim 3, a method for rapid matching of load characteristics based on the AdaBoost model, Features: In step 5), the weight of the weak classifier in the strong classifier is calculated: Where: α m is the weight of the mth weak identifier in the strong identifier.
5. According to claim 4, a method for rapid matching of load characteristics based on the AdaBoost model, Features: In step 6), update the weight distribution of the training data set: where: z m is a normalization factor used to make the sum of the probability distributions of the samples equal to 1.
6. According to claim 5, a method for rapid matching of load characteristics based on the AdaBoost model, Features: In step 7), the strong identifier expression is: Where: F is a strong identifier, x is the newly added sample data.
Citation Information
Patent Citations
Load prediction method based on feature analysis and combinatorial learning
CN111563615A