A method for assisting in evaluating typical working conditions of enterprises based on Markov chain
Through the method based on the Markov chain, typical operating conditions of enterprises are constructed, and the working conditions deviation problem caused by inaccurate clustering in the existing technology is solved, and the higher-precision operating conditions construction and auxiliary evaluation are achieved.
Patent Information
- Application Number
- CN202310386278.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-12
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2043-04-12
AI Technical Summary
When the prior art is constructing typical operating conditions of enterprises, the clustering method is inaccurate when screening samples, resulting in a large deviation between the built operating conditions and the characteristic parameter values of the sample data.
The enterprise typical working condition assisted evaluation method based on the Markov chain is adopted, and the enterprise daily load curve data is collected through non-invasive detection methods, preprocessing and periodic analysis is performed, and the dynamic time planning algorithm is used for clustering, a state probability transfer matrix is created, and a corporate typical working condition curve cluster is obtained through Markov chain fragment synthesis.
It improves the accuracy of building typical working conditions, can intuitively characterize the variation laws of load curves, and enhances the auxiliary evaluation ability of enterprise operating conditions.
Smart Images

Figure CN116756592B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of typical operating conditions of loads, and particularly to an auxiliary evaluation method for typical operating conditions of enterprises based on Markov chains. Background Art
[0002] Establishing a determination method for typical operating conditions of enterprise production lines helps provide data support for enterprise departments, achieve the healthy operation of enterprises, improve ecological environment problems, and at the same time can also provide some practical suggestions for enterprise equipment management. Effective real-time classification of user daily load curves can provide a basis for aspects such as power grid system planning, load modeling, load forecasting, and demand-side management, and at the same time help power grid staff to identify load models.
[0003] In addition, deeply exploring user load characteristics is an urgent need for precise analysis of power users under the background of power big data to support more high-quality power grid-friendly interactive regulation. For the mining and analysis of power consumption load data, the load data collected by smart meters accurately records the power consumption of each power user within a fixed time interval. Representing the load data as a curve can vividly display its load and operating condition characteristics. The construction of enterprise standard operating conditions usually uses clustering to mine and analyze user power consumption data, and directly obtains the typical load curve or curve cluster of users. However, if only clustering is used, it is efficient in classifying and statistically analyzing characteristic parameters but inaccurate in sample screening, often resulting in a large deviation between the constructed operating conditions and the characteristic parameter values of sample data. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide an auxiliary evaluation method for typical operating conditions of enterprises based on Markov chains, which can intuitively characterize the change law of the load curve within a specified time period and can further improve the construction accuracy of typical operating conditions.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: An auxiliary evaluation method for typical operating conditions of enterprises based on Markov chains, comprising the following steps:
[0006] Step S1: At the entrance of the power enterprise, use non-invasive detection means to collect the active power with a time period of 15 minutes to obtain the enterprise daily load curve data with a total time length of 3 months, preprocess these data, and at the same time perform periodic analysis of the load curve;
[0007] Step S2: Divide the processed load curve data of the enterprise within a specified time unit into change points, use the dynamic time warping algorithm for clustering, and create the state probability transition matrices between and within each class;
[0008] Step S3: Use the inter-class and intra-class state probability transition matrices to synthesize Markov chain sub-fragments, thereby obtaining a cluster of typical operating condition curves for the enterprise;
[0009] Step S4: Combine the single-loop theorem in random matrix theory to complete the auxiliary assessment of the enterprise's operating conditions.
[0010] In a preferred embodiment, step S1 specifically includes the following steps:
[0011] Step S101: Select 2 known point data x q on each side of the missing value x, and use Newton interpolation method to complete the filling to replace the missing value x 1 , x 2 , x 3 , x 4 . The Newton divided-difference interpolation polynomial is expressed as: q
[0012] P n (x) = f(x 1 ) + a 2 (x - x 1 ) + a 2 (x - x 1 )(x - x 2 ) + a 3 (x - x 1 )(x - x 2 )(x - x 3 )(x - x k )
[0013] k where the coefficients are: a 1 = f[x 2 , x k , k = 0, 1,...;
[0014] Step S102: The processing method for outlier rejection is to define a custom time window ω. Starting from the first data of the load curve X m , take the subsequent ω data. Then continuously shift the window forward. Each time, calculate the standard deviation σ and the mean μ of the data within the window. If the residual error v i of a certain measurement value x i , 1 ≤ i ≤ n, satisfies the formula:
[0015] v i = |x - x i | > 3δ
[0016] then the data that appears outside the range of μ ± 3σ is regarded as abnormal data. Take the mean of the 6 data before and after the abnormal data as the replacement value to replace the abnormal data. Then continuously shift the window forward and replace the abnormal data each time until all the data in this feature is processed;
[0017] Step S103: The method for judging whether the enterprise production has periodicity is as follows:
[0018] Use the STL algorithm based on the decomposition model to perform time series periodicity test. If the time series has periodicity, calculate the sequence period T at the same time.
[0019] In a preferred embodiment, step S2 specifically includes the following steps:
[0020] Step S201: Since the dimension of this data set is relatively high and it takes a long time to analyze using the clustering algorithm, it is necessary to perform dimensionality reduction on this data set; assume there are M samples {X 1 , X 2 , …, X M}, and each sample has N-dimensional features is the average value of the feature dimension. Set multiple features to distinguish different segments after segmentation, and calculate the covariance coefficient between all different features as:
[0021]
[0022] Form the covariance matrix as:
[0023]
[0024] Secondly, calculate the eigenvalues and eigenvectors of the covariance matrix, sort the calculated eigenvalues, and retain the eigenvectors corresponding to the top N' largest eigenvalues so that the cumulative contribution rate exceeds 90%. Transform the original features into the new space constructed by the above N' eigenvectors. At this time, the data features are denoted as X ∈ R n×m ;
[0025] Step S202: Use the dynamic time warping algorithm to calculate and cluster the load curve within the specified time period of the enterprise to determine the state transition space dimension between different classes of the inter-class Markov chain. The specific method is as follows:
[0026] When the lengths of two time series are inconsistent, the ordinary Euclidean method becomes invalid when calculating the distance; at this time, it is necessary to consider the points on the load curve within the specified time period of an enterprise, which can match multiple points on the load curve within the specified time period of another enterprise, and the load curve within the specified time period of another enterprise can also match multiple points of this specified time period load curve sequence. Then the correct matching result of the elements of the load curve within the specified time period of the enterprise is to walk from the upper left corner to the lower right corner of the distance matrix, add up the elements passed by, and the minimum sum of the accumulation is the distance obtained by these two sequences using the dynamic programming algorithm, which is used as the similarity reference value; define the minimum cumulative distortion function as:
[0027]
[0028] Among them, the matching point pairs (i, j) and (i', j') respectively represent the point coordinates at different times on two compared curves; d(x i , y i ) represents the cumulative matching distance of the optimal path among all possible paths in front of the matching point pair (i, j); W n represents the weight of the matching path. If the search is optimized horizontally or vertically, its weight is 1. If the search is optimized diagonally, its weight is 2;
[0029] Take the data X ∈ R n×m , X = {X 1 , X 2 , …, X k} after dimensionality reduction at this time as the input for working condition clustering, and perform clustering based on dynamic time warping DTW. The number of clustering k is determined by the "elbow method"; furthermore, cluster the k centers in the space, classify the objects closest to them, and update the values of the cluster centers successively through an iterative method to minimize the squared error as:
[0030]
[0031] Among them, is the mean vector of cluster X i ; finally, obtain the clustering result M = [M 1 , M 2 , …, M n , where n is the number of clusters and is also the state transition space dimension of the inter-class Markov chain;
[0032] Step S203: Perform change point partitioning within the class. Segment and partition the load curve data of the enterprise within the specified time period after clustering processing, and determine the value of the within-class working condition state transition matrix. The specific method is as follows:
[0033] Partition the load curve at an appropriate time scale by considering every possible partition of the time series and segmenting it to the best position; for change point detection, use the sliding T-test. For a time series x with n sample sizes, artificially set a certain moment as the reference point. The sample sizes of the two subsequences Δ 1 and Δ 2 before and after the reference point are n 1 and n 2 respectively. The average values of the two subsequences are and The variances are and Define the statistic as:
[0034]
[0035] where s is:
[0036]
[0037] Then test these two subsections Δ 1 and Δ 2 , if it exceeds the significance level, it is considered that a mutation has occurred, and check whether their approximation errors are lower than the specified threshold θ th ; if not, recursively continue to divide the subsequence until the approximation errors of all segments are lower than the threshold θ th ; recursively divide the active power time series of the power enterprise until a certain stop condition is met;
[0038] For different M i ∈M, count the active power values of all working condition sub - segments in M i . Since the electricity consumption data of the enterprise satisfies randomness and lack of after - effect, regard the operation mode of the enterprise's electricity consumption as a stochastic process of a Markov chain, and define it as:
[0039]
[0040] where P{X n+1 =j|X n =i} is the probability that the working condition X is in state i at time n and in state j at time n + 1 (the next time) after a period of time;
[0041] First, convert the active power - time change into a state - time change, calculate the transition probability between two adjacent states, establish the state probability transition matrix of the working conditions within the class, and after statistics, establish the state transition probability matrix as:
[0042]
[0043]
[0044] In the formula, P ij is the probability of transferring from state i to j, m represents the number of states, F ij is the total number of transfers from state i to j within the class, F s is the number of all values in the segment, so as to determine the within - class transition probability value of the Markov chain;
[0045] Step S204: According to the working conditions obtained in the previous step, for different M i ∈M, count the active power values of all segments in M i ;
[0046] First, convert the active power-time variation into a state-time variation, calculate the transition probability between two adjacent states, and establish a state transition probability matrix after statistics as follows:
[0047]
[0048]
[0049] In the formula, P' ij is the probability of transitioning from state i to j, m represents the number of states, F' ij is the total number of transitions from state i to j between classes, F' s is the number of all values in the segment, and determine the inter-class transition probability value of the Markov chain.
[0050] In a preferred embodiment, step S3 specifically includes the following steps:
[0051] Step S301: The process of constructing the operating conditions based on the Markov chain is to first set the initial state, generate a random number λ, compare λ with the corresponding elements of the state transition matrix to determine the next state, and repeat the calculation until the cumulative time reaches the target length;
[0052] Step S302: Selection of the initial state; The first operating condition sub-segment of different clustering centers, that is, the starting process of the enterprise starting the equipment, is used as the initial segment of the candidate operating condition;
[0053] Step S303: Use the Gaussian distribution to construct the operating conditions to form a curve band, and its probability density function is:
[0054]
[0055] where μ is the mean of each clustering center, and σ is the variance of each clustering center; When constructing the operating conditions, the state value of the next moment j is determined by the state transition probability matrix. At each state jump, without repetition, according to the intra-class state transition matrix P 1 and the random number λ, select the next intra-class state; If the values before and after the time-varying point during splicing are different or there is a mutation, take the average value of the front and back time-varying points as the connection point for splicing, and at the same time, treat this point as an intermediate transition link; At the same time, set thresholds for the upper and lower limits, and if the threshold is exceeded, re-select and splice the sub-segments;
[0056] Step S304: Due to the uncertainty of the length and time of the enterprise operating condition sub-segment, one of the following three situations occurs at the end of the operating condition segment:
[0057] When ΔT = Δt = 0, the operating condition sub-segment is just spliced to complete a typical load operating condition of one cycle, and no additional processing is required;
[0058] When ΔT > Δt > 0 and the length of the working condition segment ΔT exceeds the length to be spliced Δt, the excess sequence length ΔT - Δt of the time series is used as the start segment of the load curve in the next specified time unit. At the same time, the next inter-class state is selected according to the inter-class state transition matrix P' and the random number λ.
[0059] When ΔT < Δt and Δt > 0, the length of the working condition segment ΔT is shorter than the length to be spliced Δt, and the splicing of the working condition segment continues.
[0060] Where ΔT is the length of the time series of the working condition required for splicing at the cut-off, and Δt is the length of the time series required for the last segment of the working condition to be spliced at the cut-off.
[0061] The working condition segments constructed in the above manner use the length of the typical working condition to be constructed as a constraint condition, select and combine segments from the segment library, and finally achieve that the time of the generated working condition is equal to the target working condition.
[0062] In a preferred embodiment, step S4 specifically includes the following steps:
[0063] Step S401: For the non-Hermitian random matrix X = {x i,j} ∈ C N×T , whose elements are independent and identically distributed random variables, satisfying the mean μ(x) = 0 and the variance σ 2 (x) = 1, perform singular value decomposition on the matrix X to obtain the singular value equivalent matrix X u ∈ C N×N ; for L different matrices X i (i = 1, 2,..., L), the product of the corresponding L singular value equivalent matrices is
[0064] When N, T → ∞ and c = N / T ∈ (0, 1], the empirical spectral distribution of the time series is almost surely distributed according to the single-ring law, and the probability density function is:
[0065]
[0066] According to the single-ring theorem, the eigenvalues of the high-dimensional non-Hermitian matrix will be distributed between the outer ring radius 1 and the inner ring radius (1 - c) L / 2 , which indicates that the system is in a stable state;
[0067] Step S402: Use the large-dimensional random matrix theory to identify working condition anomalies; if the newly identified working condition is within the single ring, the reasonable working condition is updated in real time and added to the working condition library; otherwise, it is determined as an abnormal working condition and reported or alarmed; the evaluation result reflects the abnormal situation of the typical working conditions of industrial large users.
[0068] Compared with the prior art, the present invention has the following beneficial effects:
[0069] 1. A method for constructing the operating conditions of large industrial users based on the Markov chain is proposed, which can construct a cluster of typical operating condition curves of the industrial load of the user;
[0070] 2. A method for auxiliary energy consumption assessment of enterprises based on the random matrix theory is proposed, which is convenient for analyzing the power consumption mode of large industrial users, finding weak links for targeted energy efficiency improvement;
[0071] 3. Considering the operating condition combinations between various devices and production lines, it is more in line with the actual production process flow. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 is a schematic flow chart of a preferred embodiment of the present invention;
[0073] Figure 2 is a schematic diagram of operating condition construction in a preferred embodiment of the present invention;
[0074] Figure 3 is a schematic diagram of the operation steps in a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0075] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0076] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0077] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0078] An auxiliary assessment method for typical operating conditions of enterprises based on the Markov chain, refer to Figures 1 to 3 , step S1: At the entrance of the power enterprise, use non-invasive detection means to collect the active power with a time period of 15 minutes, and the enterprise daily load curve data with a total time length of 3 months, and preprocess these data, and at the same time conduct a periodic analysis of the load curve.
[0079] Step S2: Divide the processed load curve data of the enterprise within the specified time unit into change points, use the dynamic time warping algorithm for clustering, and create the state probability transition matrices between and within each class.
[0080] Step S3: Use the state probability transition matrices between and within classes to synthesize Markov chain sub-fragments, thereby obtaining the typical operating condition curve clusters of the enterprise.
[0081] Step S4: Combine the single-loop theorem in the random matrix theory to complete the auxiliary assessment of the enterprise's operating conditions.
[0082] Further, the data preprocessing and periodic analysis described in Step S1 are specifically the following steps:
[0083] Step S101: Select 2 known point data x q on each side of the missing value x 1 , x 2 , x 3 , x 4 , and use Newton interpolation method to complete it to replace the missing value x q . The Newton divided-difference interpolation polynomial can be expressed as:
[0084] P n (x) = f(x 1 ) + a 2 (x - x 1 ) + a 2 (x - x 1 )(x - x 2 ) + a 3 (x - x 1 )(x - x 2 )(x - x 3 )
[0085] where the coefficients are: a k = f[x 1 , x 2 , …, x k , k = 0, 1, ….
[0086] Step S102: The processing method for outlier rejection is to define a custom time window ω. Starting from the first data of the load curve X m , take the subsequent ω data. Then, continuously shift the window forward. Each time, calculate the standard deviation σ and the mean μ of the data within the window. If the residual error v i of a certain measurement value x i (1 ≤ i ≤ n) satisfies the formula:
[0087] v i = |x - x i | > 3δ
[0088] It means that the data appearing outside the range of μ±3σ is regarded as abnormal data, and the average value of the 6 data before and after the abnormal data is taken as the replacement value to replace the abnormal data. Then the window is continuously shifted forward, and the abnormal data is replaced each time until all the data in this feature are processed.
[0089] Step S103: The method for judging whether the enterprise production has periodicity is as follows:
[0090] The periodic performance in time series data may be very complex. The daily cycle, weekly cycle, and annual cycle may exist intertwined and there are some special cases. Therefore, it cannot be simply analyzed with a specified time length as the time span unit. This patent uses the STL (Seasonal and Trend decomposition using Loess) algorithm based on the decomposition model for time series periodicity test. If the time series has periodicity, the period T of this sequence is calculated simultaneously.
[0091] Furthermore, the method for clustering based on dynamic time programming and creating the probability transition matrix between each class in step S2 is specifically as follows:
[0092] Step S201: Since the dimension of this data set is relatively high and it takes a long time to analyze using the clustering algorithm, it is necessary to perform dimensionality reduction on this data set. Suppose there are M samples {X 1 , X 2 ,..., X M}, and each sample has N-dimensional features is the average value of the feature dimension. Multiple features are set to distinguish different segments after segmentation, and the covariance coefficient between all different features is calculated as:
[0093]
[0094] The covariance matrix formed is:
[0095]
[0096] Secondly, calculate the eigenvalues and eigenvectors of the covariance matrix, sort the calculated eigenvalues, and retain the eigenvectors corresponding to the first N' largest eigenvalues so that the cumulative contribution rate exceeds 90%. The original features are transformed into the new space constructed by the above N' eigenvectors, thus realizing feature compression and laying a foundation for improving the efficiency of subsequent algorithms. At this time, the data features are denoted as X∈R n×m .
[0097] Step S202: Use the dynamic time warping algorithm to calculate the clustering of the load curves of an enterprise within a specified time period to determine the state transition space dimension between different classes of the inter-class Markov chain. The specific method is as follows:
[0098] When the lengths of two time series are inconsistent, the ordinary Euclidean method becomes invalid when calculating the distance. At this time, it is necessary to consider that a point on the load curve within a specified time period of an enterprise can match multiple points on the load curve within a specified time period of another enterprise, and the load curve within a specified time period of another enterprise can also match multiple points on the load curve sequence within this specified time period. Then, the correct matching result of the elements of the load curve within the specified time period of the enterprise is to walk from the upper left corner to the lower right corner of the distance matrix, accumulate the elements passed by, and the minimum sum of the accumulation is the distance obtained by using the dynamic programming algorithm for these two sequences, which is used as the similarity reference value. Define the minimum cumulative distortion function as:
[0099]
[0100] where the matching point pairs (i, j) and (i', j') represent the point coordinates at different times on two compared curves respectively. d(x i ,y i ) represents the cumulative matching distance of the best path among all possible paths up to the matching point pair (i, j). W n represents the weight of the matching path. If it is a horizontal or vertical search for the optimal path, its weight is 1. If it is a diagonal search for the optimal path, its weight is 2.
[0101] Take the data X ∈ R n×m , X = {X 1 , X 2 ,..., X k} after dimensionality reduction at this time as the input for working condition clustering, and perform clustering according to the dynamic time warping DTW. The number k of clusters is determined by the "elbow method". Furthermore, cluster the k centers in the space, classify the objects closest to them, and update the values of the cluster centers successively through an iterative method to minimize the squared error as:
[0102]
[0103] where, is the mean vector of the cluster X i . Finally, obtain the clustering result M = [M 1 , M 2 , …, M n , where n is the number of clusters and is also the state transition space dimension of the inter-class Markov chain.
[0104] Step S203: Conduct change point division within the class. Segment the load curve data of the enterprise within the specified time period after clustering processing, and determine the value of the in-class working condition state transition matrix. The specific method is as follows:
[0105] Divide the load curve at an appropriate time scale by considering every possible division of the time series and segmenting it at the optimal position. Change point detection uses a sliding T-test. For a time series x with n sample sizes, a certain moment is artificially set as the reference point, and the two subsequences Δ 1 and Δ 2 before and after the reference point have sample sizes of n 1 and n 2 respectively. The average values of the two subsequences are and The variances are and Define the statistic as:
[0106]
[0107] where s is:
[0108]
[0109] Then test the two subsections Δ 1 and Δ 2 . If it exceeds a certain significance level, it can be considered that there is a mutation. Check whether their approximation error is lower than the specified threshold θ th . If not, recursively continue to segment the subsequence until the approximation error of all segments is lower than the threshold θ th . Recursively segment the active power time series of the power enterprise until a certain stop condition is met.
[0110] For different M i ∈M, statistically analyze the active power values of all working condition sub-fragments in M i . Since the electricity consumption data of the enterprise satisfies randomness and lack of aftereffect, the operation mode of the enterprise's electricity consumption can be regarded as a stochastic process of a Markov chain, defined as:
[0111]
[0112] where P{X n+1 =j|X n =i} is the probability that the working condition X is in state i at time n and in state j at time n + 1 (the next moment) after a period of time.
[0113] First, convert the active power - time variation into a state - time variation, calculate the transition probability between two adjacent states, establish a state probability transition matrix for the in - class operating conditions, and after statistics, establish the state transition probability matrix as follows:
[0114]
[0115]
[0116] Where P ij is the probability of transitioning from state i to j, m represents the number of states, and F ij is the total number of transitions from state i to j within the class, and F s is the number of all values in the segment, thereby determining the in - class transition probability value of the Markov chain.
[0117] Step S204: According to the operating conditions obtained in the previous step, for different M i ∈M, count the active power values of all segments in M i .
[0118] First, convert the active power - time variation into a state - time variation, calculate the transition probability between two adjacent states, and after statistics, establish the state transition probability matrix as follows:
[0119]
[0120]
[0121] Where P' ij is the probability of transitioning from state i to j, m represents the number of states, and F' ij is the total number of transitions from state i to j between classes, and F' s is the number of all values in the segment, determining the between - class transition probability value of the Markov chain.
[0122] Furthermore, the specific steps for constructing the typical operating condition curve cluster of the enterprise in step S3 are as follows:
[0123] Step S301: The process of constructing the operating conditions based on the Markov chain is to first set the initial state, generate a random number λ, compare λ with the corresponding elements of the state transition matrix to determine the next state, and repeat the calculation in this way until the cumulative time reaches the target length.
[0124] Step S302: Selection of the initial state. The selection of the initial segment plays a crucial role in the splicing of subsequent segments and the deviation of the final result in the process of constructing the operating conditions by the Markov method. In this patent, the first operating condition sub - segment of different clustering centers (i.e., the starting process of the enterprise starting equipment) is used as the initial segment of the candidate operating condition.
[0125] Step S303: Use the Gaussian distribution to construct the operating conditions to form a curve band. Its probability density function is:
[0126]
[0127] where μ is the mean of each cluster center and σ is the variance of each cluster center. When constructing the operating conditions, the state value at the next moment j is determined by the state transition probability matrix. At each state jump, without repetition, according to the within-class state transition matrix P 1 and the random number λ, the next within-class state is selected. If the values before and after the time-varying point during splicing are different or there is a mutation, the average value of the front and rear change points is taken as the connection point for splicing, and this point is treated as an intermediate transition link. At the same time, thresholds are set for the upper and lower limits. If the threshold is exceeded, the selection and splicing of sub-fragments are redone.
[0128] Step S304: Due to the uncertainty in the length and time of the enterprise operating condition sub-fragments, the following three situations may occur at the end of the operating condition fragment:
[0129] When ΔT = Δt = 0, the operating condition sub-fragment is just spliced to complete a typical load operating condition for one cycle, and no additional processing is required;
[0130] When ΔT > Δt > 0, the length of the operating condition fragment ΔT exceeds the length Δt to be spliced. Then, the excess sequence length ΔT - Δt of the time series is used as the start segment of the load curve within the next specified time unit. At the same time, according to the between-class state transition matrix P' and the random number λ, the next between-class state is selected;
[0131] When ΔT < Δt and Δt > 0, the length of the operating condition fragment ΔT is shorter than the length Δt to be spliced, and the splicing of the operating condition fragment continues;
[0132] where ΔT is the length of the time series of the operating condition required for splicing at the end, and Δt is the length of the time series required for the last segment of the operating condition to be spliced at the end.
[0133] For the operating condition fragments constructed in the above manner, with the length of the typical operating condition to be constructed as the constraint condition, fragments are selected and combined from the fragment library, and finally, the time of the generated operating condition is made equal to that of the target operating condition.
[0134] Furthermore, the specific steps for using the single-ring theorem in random matrix theory to complete the auxiliary evaluation of the enterprise operating conditions in step S4 are as follows:
[0135] Step S401: For the non-Hermitian random matrix X = {x i,j} ∈ C N×T, whose elements are independent and identically distributed random variables, satisfying the mean μ(x) = 0 and variance σ 2 (x) = 1. Perform singular value decomposition on the matrix X to obtain the singular value equivalent matrix X u ∈C N×N . For L different matrices X i (i = 1, 2,..., L), the product of the corresponding L singular value equivalent matrices is
[0136] When N, T → ∞ and c = N / T ∈ (0, 1], the empirical spectral distribution of the time series is almost surely distributed according to the single-ring law, and the probability density function is:
[0137]
[0138] According to the single-ring theorem, the eigenvalues of the high-dimensional non-Hermitian matrix will be distributed between the outer ring radius of 1 and the inner ring radius of (1 - c) L / 2 , which indicates that the system is in a stable state.
[0139] Step S402: To make reasonable use of high-dimensional power consumption big data analysis, use the large-dimensional random matrix theory to identify abnormal working conditions. If the newly identified working condition is within the single ring, the reasonable working condition is updated in real time and added to the working condition library; otherwise, it is determined as an abnormal working condition for reporting or alarm processing. The evaluation result can reflect the typical working condition abnormalities of industrial large users, facilitate the analysis of the power consumption methods of industrial large users, and discover weak links for targeted energy efficiency improvement.
[0140] Markov means that the change of a certain random process depends only on the current state and has nothing to do with the past state. This characteristic attribute exactly conforms to the change of the enterprise production line over time. In addition, constructing a driving working condition based on the Markov method has stronger theoretical characteristics than traditional methods. In terms of segment selection, based on the state transition probability matrix that reflects the change law of data samples, it can intuitively represent the change law of the load curve within a specified time period, and can further improve the construction accuracy of typical working conditions. At the same time, the Markov chain can randomly generate representative working conditions of a specified duration, avoiding the adverse effects brought by sample clustering errors. Therefore, an auxiliary evaluation method for enterprise typical working conditions based on the Markov chain is proposed.
Claims
1. A method for auxiliary assessment of typical working conditions of enterprises based on Markov chain, characterized in that it includes the following steps: Step S1: At the entrance of the power enterprise, use non-invasive detection means to collect the active power with a time period of 15 minutes, and the daily load curve data of the enterprise for a total time length of 3 months. Preprocess these data, and at the same time conduct a periodic analysis of the load curve; Step S2: Divide the processed load curve data of the enterprise within a specified time unit into change points, use the dynamic time warping algorithm for clustering, and create state probability transition matrices between and within each class; Step S3: Use the state probability transition matrices between and within classes to synthesize Markov chain sub-fragments, so as to obtain a cluster of typical working condition curves of the enterprise; Step S4: Combine the single-ring theorem in random matrix theory to complete the auxiliary assessment of the operating conditions of the enterprise; Step S4 specifically includes the following steps: Step S401: For an \(N\times T\) -dimensional non - Hermitian random matrix \(X=\{x i,j \}\in\mathbb{C} N×T , whose elements are independent and identically - distributed random variables, satisfying the mean \(\mu(x) = 0\) and the variance \(\sigma 2 (x)=1\), perform singular - value decomposition on the matrix \(X\) to obtain the \(N\) - order singular - value equivalent matrix \(X u \in\mathbb{C} N×N \); for multiple different matrices \(X i , i = 1,2,\cdots,L\), the product of their corresponding \(L\) singular - value equivalent matrices is When N, T → ∞ and c = N / T ∈ (0, 1], the empirical spectral distribution of the time series is almost surely distributed according to the single-ring law, and the probability density function is: According to the single-ring theorem, the eigenvalues of a high-dimensional non-Hermitian matrix will be distributed between the outer-ring radius of 1 and the inner-ring radius of (1 - c). L / 2 This indicates that the system is in a stable state. Step S402: Use the large-dimensional random matrix theory to identify abnormal working conditions; if the newly identified working condition is within the single ring, the reasonable working condition is updated in real time and added to the working condition library; otherwise, it is determined as an abnormal working condition and reported or alarmed; the evaluation result reflects the abnormal situation of the typical working conditions of industrial large users.
2. A method for auxiliary assessment of typical working conditions of enterprises based on Markov chain according to claim 1, characterized in that Step S1 specifically includes the following steps: Step S101: Select two known point data x on each side of the missing value x q , x 1 , x 2 , x 3 , x 4 , and use Newton interpolation method to complete it to replace the missing value x q , and the Newton divided-difference interpolation polynomial is expressed as: P n f(x) = f[x 1 + a 2 (x - x 1 ) + a 2 (x - x 1 )(x - x 2 ) + a 3 (x - x 1 (x - x 2 (x - x 3 ) where the coefficient is: a k = f[x 1 , x 2 , …, x k , k = 0, 1, …; Step S102: The processing method for outlier rejection is to define a custom time window ω, starting from the first data of the load curve X m and taking the subsequent ω data in total. Then, the window is continuously shifted forward, and the standard deviation σ and mean μ of the data within the window are calculated each time. If the residual error v i of a certain measured value x i , 1 ≤ i ≤ n, satisfies the following formula: v i = |x i - μ| > 3σ Data that appears outside the range of μ ± 3σ is regarded as abnormal data. Take the average of the 6 data before and after the abnormal data as the replacement value to replace the abnormal data. Then continuously move the window forward, and replace the abnormal data each time until all the data within this feature is processed; Step S103: The method for judging whether the enterprise production has periodicity is as follows: Use the STL algorithm based on the decomposition model to conduct a time series periodicity test. If the time series has periodicity, calculate the period T of this sequence at the same time.
3. A method for auxiliary assessment of typical working conditions of enterprises based on Markov chain according to claim 1, characterized in that Step S2 specifically includes the following steps: Step S201: Since the dimension of this data set is relatively high and it takes a long time to perform analysis using a clustering algorithm, it is necessary to perform dimensionality reduction on this data set; assume there are M samples {X 1 , X 2 , …, X M}, and each sample has N-dimensional features is the average value of the feature dimension. Set multiple features to distinguish different segments after segmentation. The covariance coefficient between all different features is calculated as: The covariance matrix formed is: Next, calculate the eigenvalues and eigenvectors of the covariance matrix, sort the calculated eigenvalues, and retain the eigenvectors corresponding to the top N' largest eigenvalues so that the cumulative contribution rate exceeds 90%. Transform the original features into the new space constructed by the above N' eigenvectors. At this time, the data features are denoted as X ∈ R n×m ; Step S202: Use the dynamic time warping algorithm to calculate and cluster the load curve of the enterprise within a specified time period to determine the state transition space dimension between different classes of the Markov chain between classes. The specific method is as follows: When the lengths of two time series are inconsistent, the ordinary Euclidean method becomes invalid when calculating the distance; at this time, it is necessary to consider the points on the load curve within a specified time period of an enterprise, which can match multiple points on the load curve within a specified time period of another enterprise, and the load curve within a specified time period of another enterprise can also match multiple points on the load curve sequence within this specified time period. Then, the correct matching result of the elements on the load curve within the specified time period of the enterprise is to walk from the upper left corner to the lower right corner of the distance matrix, accumulate the elements passed by, and the minimum sum of the accumulation is the distance obtained by these two sequences using the dynamic programming algorithm. This is used as the similarity reference value; define the minimum cumulative distortion function as: Among them, the matching point pairs (i, j) and (i', j') respectively represent the point coordinates at different times on two curves to be compared; d(x i , y i ) represents the cumulative matching distance of the optimal path among all possible paths in front of the matching point pair (i, j); W n represents the weight of the matching path. If the search is optimized horizontally or vertically, its weight is 1. If the search is optimized diagonally, its weight is 2; The data X ∈ R after dimensionality reduction at this time n×m , X = {X 1 , X 2 ,..., X k} is used as the input for working condition clustering, and clustering is performed based on dynamic time warping (DTW). The number of clusters k is determined by the "elbow method". Furthermore, k centers in the space are clustered, and the objects closest to the centers are classified. By using an iterative method, the values of the cluster centers are updated successively, and the squared error is minimized as follows: Among them, is the mean vector of cluster X i ; finally, the clustering result M = [M 1 , M 2 , …, M n , where n is the number of clusters and also the state transition space dimension of the inter-class Markov chain; Step S203: Perform change point division within the class, segment and divide the load curve data within the specified time period of the enterprise after clustering processing, and determine the value of the in-class working condition state transition matrix. The specific method is as follows: Divide the load curve on an appropriate time scale by considering every possible partition of the time series and splitting it at the optimal location; for change point detection, use the sliding T-test. For a time series \(x\) with a sample size of \(n\), arbitrarily set a certain moment as the reference point, and the two subsequences \(\Delta\) 1 and \(\Delta\) 2 have sample sizes of \(n\) 1 and \(n\) 2 respectively. The average values of the two subsequences are and The variances are and Define the statistic as: where s is: Then test these two subsections Δ 1 and Δ 2 , if it exceeds the significance level, it is considered that a mutation has occurred, and check whether their approximation errors are lower than the specified threshold θ th ; if not, recursively continue to divide the subsequence until the approximation errors of all segments are lower than the threshold θ th ; recursively divide the active power time series of the power enterprise until a certain stop condition is met; For different M i ∈M, the active power values of all operating condition sub - segments in M i are statistically analyzed. Since the enterprise's electricity consumption data satisfies randomness and lack of after - effect, the operation mode of the enterprise's electricity consumption is regarded as a stochastic process of a Markov chain, which is defined as: where, P{X n+1 = j|X n = i} is the probability that the working condition X is in the i-th state at the n-th moment, and after a period of time, the working condition X is in the j-th state at the next moment n + 1; First, convert the active power-time change into a state-time change, calculate the transition probability between two adjacent states, establish a state probability transition matrix for the in-class working conditions, and establish a state transition probability matrix after statistics as: where P ij is the probability of transition from state i to j, m represents the number of states, F ij is the total number of transitions from state i to j within the class, F s is the number of all values in the segment, thereby determining the within-class transition probability value of the Markov chain; Step S204: According to the working conditions obtained in the previous step, for different M i ∈M, the active power values of all segments of M i are statistically counted; First, convert the active power-time change into a state-time change, calculate the transition probability between two adjacent states, and establish a state transition probability matrix after statistics as: where P' ij is the probability of transitioning from state i to j, m represents the number of states, F' ij is the total number of transitions from state i to j between classes, F' s is the number of all values in the segment, and the inter-class transition probability values of the Markov chain are determined.
4. A method for auxiliary evaluation of typical working conditions of enterprises based on Markov chain according to claim 1, characterized in that Step S3 specifically includes the following steps: Step S301: The process of constructing the working conditions based on the Markov chain is to first set the initial state, generate a random number λ, compare λ with the corresponding elements of the state transition matrix to determine the next state, and repeat the calculation in this way until the cumulative time reaches the target length; Step S302: Selection of the initial state; The first working condition sub-segment of different clustering centers, that is, the starting process of the enterprise's starting equipment, is used as the initial segment of the candidate working conditions; Step S303: Use the Gaussian distribution to construct the working conditions to form a curve band, and its probability density function is: where μ is the mean of each cluster center and σ' is the variance of each cluster center; when constructing the working conditions, the state value at the next moment j is determined by the state transition probability matrix. At each state jump, without repetition, according to the within-class state transition matrix P 1 and the random number λ to select the next within-class state; if the values before and after the time-varying point during splicing are different or there is a mutation, the average value of the front and back change points is taken as the connection point for splicing, and this point is treated as an intermediate transition link; at the same time, set thresholds for the upper and lower limits, and if the threshold is exceeded, re-select and splice the sub-fragments; Step S304: Due to the uncertainty of the length and time of the enterprise's working condition sub-segments, one of the following three situations occurs at the end of the working condition segment: When ΔT = Δt = 0, the working condition sub-segment is just spliced to complete a typical load working condition for one cycle, and no additional processing is required; When ΔT > Δt > 0, the length of the working condition segment ΔT exceeds the length Δt to be spliced, then the excess sequence length ΔT - Δt of the time series is used as the start segment of the load curve within the next specified time unit, and at the same time, the next inter-class state is selected according to the inter-class state transition matrix P' and the random number λ; When ΔT < Δt and Δt > 0, the length of the working condition segment ΔT is shorter than the length Δt to be spliced, and continue to splice the working condition segment; where, ΔT is the length of the time series of the working condition required for splicing at the cut-off point, and Δt is the length of the time series required for the last segment of the working condition to be spliced at the cut-off point; The working condition segments constructed in the above manner use the length of the typical working condition to be constructed as a constraint condition, select and combine segments from the segment library, and finally achieve that the generated working condition has the same time as the target working condition.
Citation Information
Patent Citations
Hybrid electric vehicle driving speed prediction method
CN112182962A
Electric vehicle driving condition construction and evaluation method
CN113222385A