A method and system for multi-frequency data identification in distributed photovoltaic distribution networks

By combining variational mode decomposition and multi-core support vector machine models, the challenges of data acquisition volatility and network attack identification in distributed photovoltaic distribution networks are addressed, achieving efficient and accurate anomaly data identification and attack classification, thereby enhancing the security of the power system.

CN115932466BActive Publication Date: 2025-10-31SHANGHE COUNTY POWER SUPPLY CO STATE GRID SHANDONG ELECTRIC POWER CO +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211422121.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2025-10-31
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

Data acquisition in distributed photovoltaic power distribution networks is subject to fluctuations and delays, making them vulnerable to network attacks. Traditional methods suffer from high computational costs and limited identification range when identifying abnormal data and network attacks, and also suffer from mode aliasing problems.

Method used

The signal is decomposed using variational mode decomposition (VMD), and data features are obtained by using the amplitude and frequency of intrinsic mode functions. Combined with the multi-core support vector machine (MSVM) model, a mapping relationship is established to identify abnormal data and set labels. The data labels are predicted using the decision boundary.

Benefits of technology

It effectively avoids modal aliasing, broadens the scope of abnormal data identification, improves the efficiency and accuracy of network attack classification, reduces computational load, and enhances the security of power systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115932466B_ABST
    Figure CN115932466B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for multi-frequency data identification in distributed photovoltaic (PV) distribution networks. It acquires multi-frequency detection signals from the distributed PV distribution network; decomposes the nonlinear signal using variational mode decomposition; and obtains data features using the amplitude and frequency of intrinsic mode functions. Based on these data features, a trained multi-core support vector machine model is used to define the detection signal. This invention overcomes the limitations of traditional methods in distributed PV distribution network data processing, reducing computational load, avoiding mode aliasing, broadening the identification range of abnormal PV data, and improving the efficiency and accuracy of network attack classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of distributed photovoltaic power distribution network data processing technology, and relates to a multi-frequency data identification method and system for distributed photovoltaic power distribution networks. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] With the continuous maturation and widespread construction of distributed photovoltaic (PV) power generation equipment technology, distributed PV is playing an increasingly important role in power distribution networks. To cope with the increasingly complex power systems, power networks are highly integrating computer and communication technologies, undergoing technological innovation at both the information and physical layers. Unlike other systems, in PV distribution networks, factors such as sunlight intensity and ambient temperature cause data to fluctuate significantly; the large data acquisition range leads to delays and acquisition interruptions in some data; and issues such as incorrect computer storage and excessively large sampling intervals can also affect the validity of the acquired data.

[0004] Because the data acquisition modules of distributed photovoltaic (PV) distribution networks are highly sensitive and susceptible to interference from various factors, they are easily targeted by cyberattacks. Some cyberattacks can inject false data while avoiding anomaly detection, preventing the attacked data from showing outliers and seriously jeopardizing the safe and reliable operation of the power system. To detect such anomaly data and mitigate cyberattack risks, it is first necessary to identify and process the anomaly data, thereby identifying the risks of cyberattacks. Traditional data processing methods generally use Empirical Mode Decomposition (EMD) or Ensemble Empirical Mode Decomposition (EEMD) for signal decomposition, which may lead to problems such as mode aliasing and high computational complexity. Traditional identification methods using Support Vector Machines (SVM) have limitations in identifying fluctuating and intermittent PV data, with a relatively small identification range. Summary of the Invention

[0005] To address the aforementioned problems, this invention proposes a multi-frequency data identification method and system for distributed photovoltaic distribution networks. This invention overcomes the limitations of traditional methods in distributed photovoltaic distribution network data processing, reduces computational load, avoids mode aliasing, broadens the identification range of abnormal photovoltaic data, and improves the efficiency and accuracy of network attack classification.

[0006] According to some embodiments, the present invention adopts the following technical solution:

[0007] A method for multi-frequency data identification in a distributed photovoltaic distribution network includes the following steps:

[0008] Acquire multi-frequency detection signals from distributed photovoltaic power distribution networks;

[0009] Variational mode decomposition is used to decompose nonlinear signals, and data characteristics are obtained by utilizing the amplitude and frequency of intrinsic mode functions;

[0010] Based on data features, a detection signal is defined using a trained multi-core support vector machine model.

[0011] As an alternative implementation, the training process of the multi-core support vector machine model includes: using variational mode decomposition to decompose historical abnormal signals, classifying abnormal data and setting labels according to the types of false data injection attacks that distributed photovoltaic distribution networks are susceptible to;

[0012] Based on the label type, the data characteristics are solved using the amplitude and frequency of the intrinsic mode functions;

[0013] Based on the extracted data features and labels, a mapping relationship is established using a multi-core support vector machine model, where different data features are mapped to different regions, thus separating the input vectors from each other.

[0014] As a further defined implementation, the categories of the abnormal data include: ramp attack, data exchange attack, scale attack, data loss attack, and spurious oscillation attack, with each type assigned a separate label.

[0015] As an alternative implementation, the process of decomposing a signal using variational mode decomposition includes decomposing the original frequency data into eigenmode functions with different finite bandwidths, each of which has its center frequency.

[0016] As a further step, the center frequency is adjusted to the predicted center frequency using a frequency shift coefficient; a quadratic penalty term and Lagrange multipliers are introduced, and the center frequency is continuously updated using the alternating direction method of the multiplier;

[0017] After obtaining the intrinsic mode functions of the signal, the amplitude and frequency of each intrinsic mode function are calculated according to the Hilbert transform.

[0018] As an alternative implementation, the data features include four types: the dominant frequency of each intrinsic mode function, the amplitude spectrum of the frequency data after removing the trend term, kurtosis, and envelope entropy.

[0019] As an alternative implementation, the multi-kernel support vector machine model uses kernel functions to map different feature data to different hyperplanes;

[0020] The kernel function includes at least one of linear kernel functions, polynomial kernel functions, sigmoid kernel functions, and radial basis function kernel functions.

[0021] As an alternative implementation, when the multi-core support vector machine model maps data features and labels, it establishes an objective function to learn the decision boundary. The objective function is solved by introducing Lagrange multipliers and the decision boundary is used to predict the label of the original multi-frequency data.

[0022] A multi-frequency data identification system for a distributed photovoltaic distribution network includes:

[0023] The data acquisition and monitoring module is configured to acquire multi-frequency detection signals from the distributed photovoltaic power distribution network;

[0024] The signal decomposition module is configured to decompose nonlinear signals using variational mode decomposition and to obtain data features using the amplitude and frequency of intrinsic mode functions.

[0025] The data definition module is configured to define the detection signal based on data features and using a trained multi-core support vector machine model.

[0026] A terminal device includes a processor and a computer-readable storage medium, the processor being configured to implement instructions; the computer-readable storage medium being configured to store a plurality of instructions adapted to be loaded by the processor and executed in accordance with the steps of the method described therein.

[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0028] This invention takes into account the types of fake data injection attacks that distributed photovoltaic distribution networks are susceptible to. It classifies abnormal data into ramp attacks, data exchange attacks, scale attacks, data loss attacks, and fake oscillation attacks, and sets each of them as a category label, effectively ensuring the accurate decomposition of multi-frequency data in distributed photovoltaic distribution networks.

[0029] This invention utilizes Variational Mode Decomposition (VMD) to decompose data into signals. It can adaptively determine the number of mode decompositions for a given sequence based on actual conditions, adaptively match the optimal center frequency and finite bandwidth of each mode, and achieve effective separation of intrinsic mode components, frequency domain division of the signal, and thus obtain the effective decomposition components of a given signal, ultimately obtaining the optimal solution to the variational problem. This invention can avoid mode aliasing while broadening the identification range of abnormal photovoltaic data and improving the efficiency and accuracy of network attack classification.

[0030] This invention addresses the characteristics of multi-frequency data from distributed photovoltaic (PV) distribution networks by selecting four types of features that effectively represent signal characteristics: the dominant frequency of intrinsic mode functions, the amplitude spectrum of frequency data after removing trend terms, kurtosis, and envelope entropy. These features facilitate rapid filtering and reduce computational load. The invention utilizes a multi-kernel support vector machine (MSVM) method to establish mapping relationships. Considering the unequal information among different features, the multi-kernel approach allows each feature to be mapped by a single kernel function. The proportions of different features are then adjusted, and an objective function is used to learn the decision boundary. This decision boundary is then used to predict the label of the original multi-frequency data, effectively improving decision-making efficiency and accuracy. Attached Figure Description

[0031] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0032] Figure 1 This is a flowchart illustrating one embodiment of the present invention. Detailed Implementation

[0033] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0034] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0035] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0036] A multi-frequency data identification method for distributed photovoltaic distribution networks is proposed, which overcomes the limitations of traditional methods in distributed photovoltaic distribution network data processing. It reduces the amount of computation, avoids mode mixing, and broadens the identification range of abnormal photovoltaic data, thereby improving the efficiency and accuracy of network attack classification.

[0037] Step S1: Data Preprocessing: Sensors in the distribution network collect power network information such as active power, voltage deviation, and harmonic content, and upload it to the data acquisition and monitoring control system via power line carrier communication. Abnormal data is then identified and classified to detect network attacks. To obtain the amplitude and frequency of the intrinsic mode functions (IMFs), variational mode decomposition (VMD) is used to decompose the data signal. Considering the types of spoofed data injection attacks that distributed photovoltaic distribution networks are susceptible to, abnormal data are classified into: (slope attacks, data exchange attacks, scale attacks, data loss attacks, spoofing attacks, etc.), and each type is assigned a label. Based on the label type, the IMFs of different types of data are converted into analytical functions to obtain the amplitude and frequency of each IMF.

[0038] This invention first preprocesses the data to obtain the amplitude and frequency of the intrinsic mode functions (EMFs) in order to solve for relatively significant features. To avoid mode aliasing and excessive computation, this invention employs Variational Mode Decomposition (VMD) to decompose the nonlinear signal. VMD is an adaptive, fully non-recursive method for mode variation and signal processing. This technique has the advantage of being able to determine the number of mode decompositions. Its adaptability is manifested in determining the number of mode decompositions for a given sequence based on the actual situation. It can adaptively match the optimal center frequency and finite bandwidth for each mode, and can achieve effective separation of intrinsic mode components, frequency domain partitioning of the signal, and thus obtain the effective decomposition components of a given signal, ultimately obtaining the optimal solution to the variational problem. VMD decomposes the original frequency data x(n) into EMFs with different finite bandwidths: b t (n), where t is the number of intrinsic mode functions and each function has its center frequency ω. t The decomposition of the original signal can be represented in the following form:

[0039]

[0040] Where δ(t) and * denote the Dirac distribution and convolution operation, respectively; using the frequency shift coefficient exp(-jω) t n) The center frequency ω t The adjustment is to predict the center frequency. To solve the variational problem, a quadratic penalty term and Lagrange multipliers are introduced, using the alternating direction method of the multipliers, with the center frequency ω... t and b t (n) can be continuously updated. Therefore, the output of VMD can be expressed as:

[0041]

[0042] Where b tr (n) is the residual mode component, which usually represents the trend of frequency data.

[0043] After obtaining the eigenmode functions of the signal, physical parameters such as the amplitude and frequency of each eigenmode function are calculated using the Hilbert transform. By adding the imaginary part, the eigenmode functions are transformed into analytic functions, which are H... b (n) is as follows:

[0044]

[0045] Where C a Indicates Cauchy's principal value. It is b t The imaginary part of (n) can be used to obtain the amplitude, phase and frequency of each eigenmode function.

[0046] Step S2: Data Feature Extraction: Based on the label type, data features are derived using the amplitude and frequency of the intrinsic mode functions (IMFs). Anomalies and cyberattacks are hidden within large amounts of photovoltaic data and are short-lived, requiring rapid screening. Four types of features—the dominant frequency of each IMF, the amplitude spectrum of the frequency data after removing the trend term, kurtosis, and envelope entropy—can effectively reflect signal characteristics and facilitate rapid screening.

[0047] The dominant frequency of each intrinsic mode function:

[0048]

[0049] Remove trend item b tr The amplitude spectrum of the frequency data of (n):

[0050]

[0051] Where N is the length of the frequency data x(n).

[0052] The amplitude spectrum after removing the residuals allows us to observe the changes in signal amplitude and frequency over time.

[0053] Kuroshi:

[0054]

[0055] Where μ b It is the average value of each eigenmode function, σ b It is the standard deviation of the intrinsic mode function. Since different types of data share similar characteristics, this invention utilizes kurtosis as a feature to describe the shape of the intrinsic mode function, thereby identifying different types of data and solving the difficulties in data identification.

[0056] Envelope entropy:

[0057]

[0058] Where a(j) represents bt Hilbert transition envelope of (n), s t It is the length of each eigenmode function. The purpose of using envelope entropy is to measure the irregularity of a signal, such as its sparsity.

[0059] Step S3: Construct the mapping model: Based on the four obtained features F d With tag y x The mapping relationship is established using the Multi-core Support Vector Machine (MSVM) method. F d Mapping to different hyperplanes separates indistinguishable input vectors. An objective function is established to learn the decision boundary, and this boundary is used to predict the labels of the original multi-frequency data.

[0060] Given a feature dataset D = (F d y x ), d = 1, 2, 3, 4, y x It is F d Tags. MSVM separates F by learning a hyperplane. d In MSVM, the purpose of kernel functions is to transfer F... d Mapping to different hyperplanes, these indistinguishable input vectors can be separated from each other after the mapping. Some of the most commonly used kernel functions include linear kernel functions, polynomial kernel functions, sigmoid kernel functions, and radial basis function kernel functions. For example, the radial basis function kernel function can be defined as:

[0061]

[0062] Where x i x j ∈x(n), σ k This represents the width kernel parameter used to adjust the kernel shape.

[0063] Considering the unequal information from different features, a multi-core method is proposed in MSVM, which can be defined as follows:

[0064]

[0065] Each input feature of the data can be mapped by a single kernel function; therefore, the kernel functions and parameters for all four features can be obtained. Represents the corresponding kernel function for the four input features; ω d Indicates kernel k d The weights are used to adjust the proportions of different features. In this embodiment, the model is trained based on historical data, and the weights used by the model are recorded as the final proportions when the prediction accuracy of the training set is the highest. Basically, k m It is a linear combination of multiple kernel functions.

[0066] Then, an objective function is established to learn the decision boundary Θ:ω T k m (x i x j )+b=0:

[0067]

[0068]

[0069] Where ω and b are the weights and biases of the learning decision boundary, respectively, and ξ i Here, is the slack variable, and C represents the regularization parameter. The objective function can be solved by introducing Lagrange multipliers, and the decision boundary can be used to predict the label of the original multi-frequency data to address the characteristics of unclear outlier features and large data volatility in a large amount of photovoltaic data.

[0070] Step S4: Measured Data Analysis: Measure the measured voltage, harmonics, and other data. Use Variational Mode Decomposition (VMD) to decompose the nonlinear signal, and then use the amplitude and frequency of the intrinsic mode function to obtain four data features. Finally, use the MSVM model to define the measured data by mapping the four data features to the data label.

[0071] The present invention also provides the following product examples:

[0072] A multi-frequency data identification system for a distributed photovoltaic distribution network includes:

[0073] The data acquisition and monitoring module is configured to acquire multi-frequency detection signals from the distributed photovoltaic power distribution network;

[0074] The signal decomposition module is configured to decompose nonlinear signals using variational mode decomposition and to obtain data features using the amplitude and frequency of intrinsic mode functions.

[0075] The data definition module is configured to define the detection signal based on data features and using a trained multi-core support vector machine model.

[0076] A terminal device includes a processor and a computer-readable storage medium, the processor being configured to implement various instructions; the computer-readable storage medium being configured to store a plurality of instructions adapted to be loaded by the processor and executed in the methods of the above embodiments.

[0077] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A multi-frequency data identification method for distributed photovoltaic distribution networks, characterized in that, Includes the following steps: Acquire multi-frequency detection signals from distributed photovoltaic power distribution networks; Variational mode decomposition is used to decompose nonlinear signals, and data characteristics are obtained by utilizing the amplitude and frequency of intrinsic mode functions; Based on data features, the detection signal is defined using a trained multi-core support vector machine model; The data features include four types: the dominant frequency of each intrinsic mode function, the amplitude spectrum of the frequency data after removing the trend term, kurtosis, and envelope entropy. The training process of the multi-core support vector machine model includes: using variational mode decomposition to decompose historical abnormal signals, classifying abnormal data and setting labels according to the types of false data injection attacks that distributed photovoltaic distribution networks are susceptible to; Based on the label type, the data characteristics are solved using the amplitude and frequency of the intrinsic mode functions; Based on the extracted data features and labels, a mapping relationship is established using a multi-core support vector machine model, where different data features are mapped to different regions, thus separating the input vectors from each other.

2. The multi-frequency data identification method for a distributed photovoltaic distribution network as described in claim 1, characterized in that, The categories of the abnormal data include: ramp attack, data exchange attack, scale attack, data loss attack, and spurious oscillation attack, with each type assigned a separate label.

3. The multi-frequency data identification method for a distributed photovoltaic distribution network as described in claim 1, characterized in that, The process of decomposing a signal using variational mode decomposition involves decomposing the original frequency data into eigenmode functions with different finite bandwidths, each of which has its center frequency.

4. The multi-frequency data identification method for a distributed photovoltaic distribution network as described in claim 3, characterized in that, The center frequency is adjusted to the predicted center frequency using a frequency shift coefficient; a quadratic penalty term and Lagrange multipliers are introduced, and the center frequency is continuously updated using the alternating direction method of the multiplier; After obtaining the intrinsic mode functions of the signal, the amplitude and frequency of each intrinsic mode function are calculated according to the Hilbert transform.

5. The multi-frequency data identification method for a distributed photovoltaic distribution network as described in claim 1, characterized in that, The multi-kernel support vector machine model uses kernel functions to map different feature data to different hyperplanes; The kernel function includes at least one of linear kernel functions, polynomial kernel functions, sigmoid kernel functions, and radial basis function kernel functions.

6. The multi-frequency data identification method for a distributed photovoltaic distribution network as described in claim 1, characterized in that, When the multi-core support vector machine model maps data features and labels, it establishes an objective function to learn the decision boundary. The objective function is solved by introducing Lagrange multipliers, and the decision boundary is used to predict the label of the original multi-frequency data.

7. A multi-frequency data identification system for a distributed photovoltaic distribution network, characterized in that, include: The data acquisition and monitoring module is configured to acquire multi-frequency detection signals from the distributed photovoltaic power distribution network; The signal decomposition module is configured to decompose nonlinear signals using variational mode decomposition and to obtain data features using the amplitude and frequency of intrinsic mode functions. The data definition module is configured to define the detection signal based on data features and using a trained multi-core support vector machine model; The data features include four types: the dominant frequency of each intrinsic mode function, the amplitude spectrum of the frequency data after removing the trend term, kurtosis, and envelope entropy. The training process of the multi-core support vector machine model includes: using variational mode decomposition to decompose historical abnormal signals, classifying abnormal data and setting labels according to the types of false data injection attacks that distributed photovoltaic distribution networks are susceptible to; Based on the label type, the data characteristics are solved using the amplitude and frequency of the intrinsic mode functions; Based on the extracted data features and labels, a mapping relationship is established using a multi-core support vector machine model, where different data features are mapped to different regions, thus separating the input vectors from each other.

8. A terminal device, characterized in that, It includes a processor and a computer-readable storage medium, the processor being used to implement various instructions; the computer-readable storage medium being used to store a plurality of instructions adapted to be loaded by the processor and executed as steps in the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Method and system for photovoltaic prediction

    CN107563561A

  • LED lamp power supply driving fault diagnosis method and system based on light output time-frequency characteristics

    CN113887322A