Power distribution facility multi-source heterogeneous information sensing method and system for low-carbon park

By improving the YOLOv5 and FP-Growth algorithm combined with the SVM classifier method, the problem of insufficient data coverage and processing in the information perception of power distribution facilities in low-carbon parks is solved, efficient abnormal identification and monitoring is achieved, and management efficiency is improved.

CN120339928APending Publication Date: 2025-07-18FUXIN POWER SUPPLY COMPANY STATE GRID LIAONING ELECTRIC POWER +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311707498.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

There are problems in the information perception of power distribution facilities in low-carbon parks with limited data coverage, low sampling frequency and insufficient data processing capabilities, resulting in low monitoring efficiency and inability to effectively manage and identify abnormal data.

Method used

Improved YOLOv5 is used to collect information of power distribution facilities, combine FP-Growth algorithm for correlation characteristic analysis and abnormal data set storage, and use SVM classifier to identify multi-source heterogeneous abnormal information, and improve data fusion and recognition accuracy through feature extraction and dimensionality reduction technologies.

Benefits of technology

It improves the perceived coverage rate and data abnormality identification accuracy of low-carbon parks, enhances the remote monitoring capabilities of power distribution facilities, and achieves more accurate abnormality detection and management decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The invention discloses a low-carbon park-oriented power distribution facility multi-source heterogeneous information sensing method and system. The method comprises the following steps of: acquiring power distribution facility information based on improved YOLOv5; power distribution facility information association characteristic analysis based on an FP-Growth algorithm; and sensing and identifying the multi-source heterogeneous abnormal information of the power distribution facility based on feature extraction. According to the method, data from different sources are fused, the perception coverage rate and the accuracy of data anomaly identification in the low-carbon park are improved, powerful data support is provided for accurate facility management and optimization decision, the problem of power distribution facility information perception of the low-carbon park is effectively solved, and a more accurate anomaly detection model can be constructed; the difference between the normal mode and the abnormal mode is identified through feature extraction, and the power distribution facility information sensing capability in the low-carbon park is enhanced, so that sensing and identification of abnormal information are realized, and remote monitoring and information sensing of the power distribution facility are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of equipment technology, and in particular to a multi-source heterogeneous information perception method and system for distribution facilities in a low-carbon park. Background Art

[0002] With the rapid development of the global economy, in order to address global climate change and curb the impact of greenhouse gases mainly composed of CO2 on the earth, low-carbon economy has become the development goal of governments and enterprises in various countries. As an important part of the low-carbon economy, low-carbon parks have become one of the important means to promote sustainable development in many countries and regions. A low-carbon park refers to a park that fully considers factors such as energy conservation, environmental protection, and resource recycling in park planning, design, construction, operation, and management, realizes the coordinated development of the park's economy, society, and environment, and achieves the development goals of low-carbon, environmental protection, and sustainability. The construction of a low-carbon park needs to fully consider the efficient use of energy and environmental protection, and the information perception of distribution facilities is the primary condition to ensure the safe and stable operation of the power system in the low-carbon park.

[0003] The distribution side of a low-carbon park is a key link to ensure the stable operation of the power system in the low-carbon park, and its operation efficiency is directly related to the stability of the power supply in the low-carbon park. Distribution side information perception covers the real-time acquisition of relevant data and status information of distribution facilities through various technical means, including power load, equipment operation status, power loss, fault diagnosis, etc. The purpose is to comprehensively monitor and accurately manage the operation status of various distribution facilities in the low-carbon park. In the context of a low-carbon park, the distribution system will be equipped with more power equipment to meet the needs of green energy and intelligent management. This poses a challenge: how to efficiently and real-time monitor and analyze the large amount of data generated by these devices, so as to optimize energy management and reduce waste. Currently, the perception of distribution side information in low-carbon parks usually adopts a combination of automatic reading of smart meters and manual inspections by staff. Although this method improves the monitoring efficiency to a certain extent, there are still the following deficiencies:

[0004] (1) Limited data coverage: The current distribution information perception technology in parks mainly relies on traditional measurement devices such as electricity meters and sensors. The deployment range of these devices is limited, and it is impossible to achieve a comprehensive perception of the entire park's distribution network. Especially for new energy access forms such as distributed energy and new distribution facilities, the deployment methods of traditional devices can no longer meet the requirements.

[0005] (2) Too low data sampling frequency: Traditional distribution information perception devices usually obtain data at a relatively low sampling frequency, resulting in insufficient perception ability of the changes and abnormalities in the power system. Especially in the case of load fluctuations and short-term faults on the distribution side, the data sampling frequency of sensors cannot capture all the detailed information, affecting the accurate grasp of the distribution network status.

[0006] (3) Insufficient data processing capabilities: The power distribution information perception system on the distribution side needs to process a large amount of real-time data and perform data analysis and judgment. However, the current data processing capabilities are still limited. In the process of data processing, there are deficiencies in data transmission, storage, cleaning, analysis and other links, resulting in an urgent need to improve the accuracy and reliability of information perception. Summary of the Invention

[0007] To solve the problem that a large amount of various data of power distribution facilities in a low-carbon park cannot be well integrated, information interaction management cannot be effectively carried out, and the monitoring and perception coverage rate and the accuracy of abnormal data identification are low, the present invention proposes a multi-source heterogeneous information perception method and system for power distribution facilities in a low-carbon park.

[0008] The technical solution of the present invention is as follows:

[0009] A multi-source heterogeneous information perception method for power distribution facilities in a low-carbon park, characterized by including:

[0010] Collecting power distribution facility information based on improved YOLOv5 to establish a multi-source heterogeneous information dataset of power distribution facilities;

[0011] Analyzing the correlation characteristics of power distribution facility information based on the FP-Growth algorithm for the collected multi-source heterogeneous information dataset of power distribution facilities, and storing the abnormal dataset and the association rule library composed of association rules of other elements different from the abnormal data;

[0012] Processing the stored abnormal dataset and matching it with the association rule library. After dimensionality reduction, it is input into the SVM classifier, calculated using the Polynomial kernel function, and the result of the kernel function calculation is input into the f(p) function to calculate the abnormal information recognition value, realizing the recognition of multi-source heterogeneous abnormal information.

[0013] Furthermore, the power distribution facility information collection based on improved YOLOv5 and the establishment of a multi-source heterogeneous information dataset of power distribution facilities include the following steps:

[0014] S1.1 Based on the location information of power distribution facilities in a low-carbon park, deploy intelligent induction cameras, current, voltage, temperature, and humidity sensors;

[0015] S1.2 Use the videos generated by real-time monitoring to separate each frame into images to generate a power distribution facility monitoring dataset;

[0016] S1.3 Analyze the fog and brightness conditions in the images according to the power distribution facility monitoring dataset;

[0017] S1.4 Image dehazing uses the dark channel processing method; for any input image J, its dark channel is shown in Equation (1):

[0018]

[0019] Where, on the left side of the equation, J dark is the dark channel map, and on the right side of the equation, J c represents each channel of the color image, Ω(x) is a square region; centered on pixel x; C represents one of the three channels of R (red), G (green), and B (blue), min represents the minimum value of the region, and x and y are variables;

[0020] S1.5 Through Equation (2), the image dehazing formula is derived to complete the removal of fog in the image, as shown in Equation (2):

[0021]

[0022] Where, t0 is the preset minimum value of the medium transmission coefficient, and max represents the maximum value of the region;

[0023] Based on the image dehazing principle, the atmospheric scattering model is adopted, as shown in Equation (3):

[0024] I(x) = J(x)t(x) + A(1 - t(x)) (3)

[0025] Where, I(x) is the observed pixel value, J(x) is the original brightness of the scene, t(x) is the transmittance, and A is the global atmospheric light.

[0026] S1.6 According to the result of Step S1.5, use the LabelMe tool to perform label binarization on the completed power distribution facility monitoring dataset;

[0027] S1.7 Use the improved YOLOV5 object detection model to detect the image. The model uses the feature map of the P-bifpn (PathBidirectional Feature Pyramid Network) aggregation scale to fuse the low-level features with the high-level features;

[0028] S1.8 Use the fully connected layer to extract the non-linear features emerging in the multi-dimensional data space, convert the power distribution facility monitoring dataset through the fully connected layer, realize the fusion with the collected sensor data, and establish a multi-source heterogeneous information dataset for power distribution facilities.

[0029] Furthermore, perform the analysis of the information association characteristics of power distribution facilities based on the FP-Growth algorithm and store the abnormal dataset and the association rule library composed of association rules different from other elements of the abnormal data, including:

[0030] S2.1 For the collected multi-source heterogeneous information dataset of power distribution facilities, use the Pandas library to delete duplicate values in the monitoring data, fill in missing values in the monitoring data using KNN, and replace outliers;

[0031] S2.2 Use the boundary value information of the sample space to linearly map the features to a specific range. The calculation equation for Min-Max normalization is shown in formula (5):

[0032]

[0033] where x is the value before sample normalization, x new is the value after normalization, x min is the minimum value of the sample space, x max is the maximum value of the sample space;

[0034] S2.3 Use generative adversarial networks (GANs) to fuse multi-source heterogeneous information in the low-carbon park and generate a synthetic model that can understand and reproduce the characteristics of different data sources;

[0035] S2.4 Import the processed multi-source heterogeneous information dataset of power distribution facilities and set the minimum support and minimum confidence;

[0036] S2.5 Conduct the first scan of the fused dataset to search for elements greater than or equal to the minimum support, which are frequent items. The scanned frequent items are composed into an abnormal data frequent item set, and an FP-Tree header table is constructed by arranging the elements in descending order of support;

[0037] S2.6 Establish an FP-tree, conduct the second scan of the frequent item set, establish branches for the FP-Tree header table, and insert the frequent items of each branch in descending order in turn to construct an FP-tree with Null (zero device) as the root node;

[0038] S2.7 Traverse each item from the bottom of the FP-Tree in turn, find the corresponding conditional FP-tree according to each frequent 1-item set in the item header table, mine the frequent item set, calculate the confidence between the elements of the frequent item set, retain the strong association rules greater than the minimum confidence, and conduct data mining until all conditional FP-trees are mined, and store the abnormal dataset and the association rule library composed of association rules different from other elements.

[0039] Furthermore, process the stored abnormal dataset and implement the identification of multi-source heterogeneous abnormal information, including:

[0040] S3.1 For different power distribution facility features X std of the abnormal dataset, conduct standardization processing as shown in formula (7):

[0041] X std ={x 1 ,x 2 ,…x n}(7)

[0042] where x n is the nth sample point of the distribution facility characteristics, and each sample has k-dimensional characteristics, as shown in formula (8):

[0043] x i =(x i1 ,x i2,… x ik )(8);

[0044] Each of the k-dimensional characteristics described in S3.2 corresponds to a performance parameter of a certain distribution facility; the distribution facility performance parameters include: the power generation of the server cabinet, the usage frequency of the industrial router, and the acquisition rate of the Internet of Things intelligent terminal acquisition device; by calculating the covariance matrix R of the facility characteristics, the mutual influence and correlation between different facilities are obtained and matched with the association rule library. The calculation formula of the covariance matrix R is as shown in (9):

[0045]

[0046] where X std represents the standardized data matrix, where each column represents a feature, that is, a certain performance index of the distribution facility, and each row represents an observation, that is, the performance reading at a specific time. (X std ) T is the transpose matrix of X std ; n is the number of data points collected every day in a month;

[0047] S3.3 performs eigenvalue decomposition on the covariance matrix R to obtain eigenvalues λ i and eigenvectors v i ; reorder the results of the eigenvalue decomposition from largest to smallest and select the eigenvectors corresponding to the largest n eigenvalues. Project each data point onto the eigenvector to obtain the value of the principal component, as shown in formula (10):

[0048] PC i =X std ·v i (10)

[0049] where i = 1, 2,..., n, PC i represents the ith principal component, representing the key operating modes of different facilities or the most significant energy consumption characteristics; v i represents the eigenvector corresponding to the ith eigenvalue of the covariance matrix R, indicating the direction in which the facility data changes the most; n represents the number of principal components selected to be retained;

[0050] S3.4 Calculate the cumulative information contribution rate η of the first p (p ≤ n) principal components in the eigenvalues p , the eigenvalue λ i 's information contribution rate y i is as shown in formula (11):

[0051]

[0052] S3.5 Calculate the dimensionality reduction result, using the transformation matrix T p = (t1, t2, …, t p ), which is composed of the eigenvectors corresponding to the first p eigenvalues. The result after dimensionality reduction is as shown in formula (12):

[0053] p i = T p x1 (12)

[0054] S3.6 Based on the data after dimensionality reduction in S3.5, input it into the SVM classifier, adopt the Polynomial kernel function, and the calculation result formula is as shown in (14): k(x, x i ) = [γ * (x · x i ) + coef] d (14)

[0055] where d is the order of the polynomial and coef is the bias coefficient; realize the identification of multi-source heterogeneous abnormal information.

[0056] S3.7 Let l i be the abnormal information identification value. Using the abnormal information prediction model is to find the relationship between p i and l i , as shown in formula (15):

[0057] l i = f(p i ) (15)

[0058] S3.8 Input the result k(x, x i ) calculated by the kernel function into the f(p) function to calculate the abnormal information identification value, and realize the identification of multi-source heterogeneous abnormal information. The calculation formula of f(p) is as shown in (16):

[0059]

[0060] where, a i , is the Lagrange multiplier and b is the bias term.

[0061] Further, step S1.7 includes: the P-bifpn module is divided into P3, P4, P5, and P4-td. On the basis of the original network, a residual edge of P4 pointing to P4-out is added to fuse features of more scales; according to the contribution of the input features to the output features, a bottom-up channel consisting of P3-out, P4-0ut and P5-out is added, and the trained weights are introduced at the nodes where multiple features are fused, and multiplied with the input features of each node to meet the different contributions of different input features to the final output. The fast normalization formula is used to train these weights, as shown in formula (4):

[0062]

[0063] where w i , w j Represents different feature weights. After obtaining each w i The Relu function is then introduced to ensure that w i >0,I i Represents the i-th input feature; to ensure numerical stability, ε=0.0001.

[0064] Furthermore, step S2.1 includes: using the K-Nearest Neighbors (KNN) algorithm to calculate the distance between each sample in the cleaned data, finding the K neighbor samples that are most similar to the samples with missing values, and estimating the missing values; if these missing values are continuous variables, taking the average or weighted average of the characteristic values of these neighbor samples as the estimate; if they are categorical variables, determining the most likely category through a voting mechanism; for the processing of outliers, the characteristic values of these neighbor samples are also used to replace them.

[0065] Further, step S2.3 includes: the generative adversarial network includes a G network and a D network, the G network is a learning network of a generative model, which is trained with the data of the sample set and generates a prediction model, denoted as G(z); D is a discriminant network, whose input parameter is x, x represents the accurate value, and the output represents the probability that x is the accurate value. If it is 1, it means that the prediction accuracy is 100%, and so on, as shown in formula (6):

[0066]

[0067] Where x is the true value, z is the input training sample, and G(z) represents the prediction model generated by the G network; D(G(z)) is the probability that the D network judges whether the prediction model generated by G is accurate. G should hope that the model it generates is as close to the accurate value as possible. When G takes the maximum value of D(G(z)), V(D,G) becomes smaller.

[0068] Further, in step S3.4, np The calculation formula is as shown in (13):

[0069]

[0070] A perception system adopting a multi-source heterogeneous information perception method for power distribution facilities in a low-carbon park as described above includes:

[0071] A multi-source heterogeneous information collection module, which collects power distribution facility information based on improved YOLOv5 and establishes a multi-source heterogeneous information dataset of power distribution facilities;

[0072] An abnormal dataset and associated rule base storage module, which analyzes the association characteristics of power distribution facility information based on the FP-Growth algorithm for the collected multi-source heterogeneous information dataset of power distribution facilities, and stores the abnormal dataset and the associated rule base composed of associated rules of other elements different from the abnormal data;

[0073] A multi-source heterogeneous abnormal information recognition module, which processes the stored abnormal dataset and matches it with the associated rule base. After dimensionality reduction, it is input into an SVM classifier, calculated using a Polynomial kernel function, and the result of the kernel function calculation is input into the f(p) function to calculate the abnormal information recognition value, realizing the recognition of multi-source heterogeneous abnormal information.

[0074] The beneficial effects of the present invention are:

[0075] 1. In the data collection link, through the power distribution facility information collection method based on improved YOLOV5, it is possible to extract image feature information at different distances or resolutions, fuse data from different sources, improve the perception coverage rate and the accuracy of data anomaly recognition under a low-carbon park, and provide strong data support for accurate facility management and optimization decisions.

[0076] 2. Based on the analysis of the association characteristics of power distribution facility information by the FP-Growth algorithm, it realizes the excavation of the association between various elements in the power distribution facility information, provides insights into the facility status and performance, effectively solves the problem of power distribution facility information perception in a low-carbon park, and helps to construct a more accurate anomaly detection model.

[0077] 3. Based on the feature extraction of multi-source heterogeneous abnormal information perception and recognition of power distribution facilities, it identifies the differences between normal and abnormal patterns through feature extraction, enhances the power distribution facility information perception ability under a low-carbon park, thereby realizing the perception and recognition of abnormal information, and realizing the remote monitoring and information perception of power distribution facilities. Description of the Drawings

[0078] Figure 1 is a flowchart of a method for identifying multi-source heterogeneous abnormal information of power distribution facilities;

[0079] Figure 2 It is a framework diagram for collecting distribution facility information;

[0080] Figure 3 It is a process diagram for detecting abnormal facilities by the YOLOV5 model;

[0081] Figure 4 It is an architecture diagram of the P-bifpn module;

[0082] Figure 5 It is a block diagram of a multi-source heterogeneous information perception system for distribution facilities.

[0083] Figure 6 It is a diagram for collecting current and voltage waveform information in Experiment 1 of the present invention;

[0084] Figure 7 It is a comparison diagram of the perception coverage rate with the target number of 20 for different algorithms;

[0085] Figure 8 It is a comparison diagram of the perception coverage rate of different sensing radius algorithms;

[0086] Figure 9 It is Figure 7 A comparison diagram of the recognition accuracy of abnormal data of distribution facilities. Specific implementation manners

[0087] The present invention will be described in detail below in conjunction with the accompanying drawings and specific examples:

[0088] A method for multi-source heterogeneous information perception of distribution facilities for a low-carbon park, as Figure 1 shown, includes the following steps:

[0089] Collect distribution facility information based on the improved YOLOv5, and establish a multi-source heterogeneous information dataset of distribution facilities;

[0090] Perform an analysis on the associated characteristics of distribution facility information based on the FP-Growth algorithm for the collected multi-source heterogeneous information dataset of distribution facilities, and store the abnormal dataset and the association rule library composed of association rules that are different from other elements of the abnormal data;

[0091] Process the stored abnormal dataset, match it with the association rule library, after dimensionality reduction, input it into the SVM classifier, calculate using the Polynomial kernel function, and input the result of the kernel function calculation into the f(p) function to calculate the abnormal information recognition value, realizing the recognition of multi-source heterogeneous abnormal information.

[0092] Furthermore, the collection of distribution facility information based on the improved YOLOv5 and the establishment of a multi-source heterogeneous information dataset of distribution facilities include the following steps:

[0093] S1.1 Based on the location information of the power distribution facilities in the low-carbon park, intelligent induction cameras, current, voltage, temperature, and humidity sensors are arranged around the power distribution facilities; the images are separated frame by frame to generate a monitoring data set for the power distribution facilities. The framework of the power distribution facility information acquisition method is as Figure 2 shown.

[0094] S1.2 As Figure 2 shown, collect the videos and images of the server cabinets, industrial energy routers, miniature circuit breakers, outgoing cabinets, and Internet of Things intelligent terminal information acquisition devices in the park. At the same time, use the intelligent induction camera to monitor the surroundings of the power distribution facilities in real time and generate videos, which are separated frame by frame into images to generate a monitoring data set for the power distribution facilities;

[0095] S1.3 According to the generated monitoring data set of the power distribution facilities, analyze the fog and light-dark conditions that appear in the images;

[0096] S1.4 Dark channel processing method is used for image dehazing; for any input image J, its dark channel is as shown in formula (1):

[0097]

[0098] where, on the left side of the equation, J dark is the dark channel map, on the right side of the equation, J c represents each channel of the color image, Ω(x) is a square area; centered on pixel x; C represents one of the three channels of R (red), G (green), and B (blue), and min represents the minimum value of the area, x and y are variables;

[0099] S1.5 Through formula (2), the image dehazing formula is derived to complete the removal of fog in the image, as shown in formula (2):

[0100]

[0101] where, t0 is the lowest preset value of the medium transmission coefficient, and max represents the maximum value of the area;

[0102] Based on the image dehazing principle, the atmospheric scattering model is adopted, as shown in formula (3):

[0103] I(x) = J(x)t(x) + A(1 - t(x)) (3)

[0104] where, I(x) is the observed pixel value, J(x) is the original brightness of the scene, t(x) is the transmittance, and A is the global atmospheric light.

[0105] S1.6 According to the result of step S1.5, use the LabelMe tool to perform label binarization on the completed power distribution facility monitoring data set;

[0106] S1.7 Use the improved YOLOV5 object detection model to detect images. The model structure is as follows Figure 3 shown. The model uses P-bifpn (Path Bidirectional Feature Pyramid Network) to aggregate feature maps of different scales, fusing low-level features with high-level features to enhance the localization information of high-level features while reducing the transmission loss of low-level feature information. The architecture of the P-bifpn module is as follows Figure 4 shown;

[0107] The P-bifpn module is divided into P3, P4, P5, and P4-td. Based on the original network, a residual edge from P4 to P4-out is added to fuse features of more scales. According to the contribution degree of the input features to the output features, a bottom-up channel composed of P3-out, P4-out, and P5-out is added. Training weights are introduced at the nodes where multiple features are fused and multiplied by the input features of each node to meet the different contributions of different input features to the final output. The fast normalization formula is used to train these weights, as shown in formula (4):

[0108]

[0109] where w i , w j represent different feature weights. After obtaining each w i , the Relu function is introduced to ensure that w i > 0. I i represents the i-th input feature; to ensure numerical stability, ε = 0.0001.

[0110] S1.8 Use the fully connected layer to extract the non-linear features emerging in the multi-dimensional data space, convert the distribution facility monitoring data set through the fully connected layer, realize the fusion with the collected sensor data, and establish a multi-source heterogeneous information data set for distribution facilities.

[0111] The present invention improves the YOLOV5 detection model and integrates the P-bifpn module, effectively fusing the feature information of different semantic levels, and solving the problems of limited data coverage and inaccurate data judgment that occur in the process of collecting distribution facility information.

[0112] Furthermore, perform the analysis of the correlation characteristics of distribution facility information based on the FP-Growth algorithm and store the abnormal data set and the association rule library composed of association rules different from other elements of the abnormal data, including:

[0113] S2.1 For the collected multi-source heterogeneous information data set of distribution facilities, the Pandas library is used to delete duplicate values of monitoring data, and the K-Nearest Neighbors (KNN) algorithm is used to calculate the distance between each sample in the cleaned data, find the K neighbor samples that are most similar to the samples with missing values, and estimate the missing values; if these missing values are continuous variables, the average or weighted average of the characteristic values of these neighbor samples is used as the estimate; if it is a categorical variable, the most likely category is determined through a voting mechanism; for the processing of outliers, the characteristic values of these neighbor samples are also used to replace them, so as to complete the filling of missing values in monitoring data and the replacement of outliers.

[0114] S2.2 uses the boundary value information of the sample space to linearly map the features to a specific range. The calculation equation of Min-Max normalization is shown in formula (5):

[0115]

[0116] Among them, x is the value of the sample before normalization, x new is the normalized value, x min is the minimum value of the sample space, x max is the maximum value of the sample space;

[0117] S2.3 Use generative adversarial networks (GANs) to fuse multi-source heterogeneous information of low-carbon parks and generate synthetic models that can understand and reproduce the characteristics of different data sources;

[0118] The generative adversarial network includes the G network and the D network. The G network is a learning network for generating models. It is trained with the data of the sample set and generates a prediction model, denoted as G(z). D is a discriminant network. Its input parameter is x, which represents the accurate value. The output represents the probability that x is the accurate value. If it is 1, it means that the prediction accuracy is 100%, and so on, as shown in formula (6):

[0119]

[0120] Where x is the true value, z is the input training sample, and G(z) represents the prediction model generated by the G network; D(G(z)) is the probability that the D network judges whether the prediction model generated by G is accurate. G should hope that the model it generates is as close to the accurate value as possible. When G takes the maximum value of D(G(z)), V(D,G) becomes smaller.

[0121] S2.4 uses the FP-Growth algorithm to import the processed multi-source heterogeneous information data set of power distribution facilities and sets the minimum support and minimum confidence;

[0122] S2.5 Construct the FP tree, scan the transaction set twice. In the first scan, fuse the data set, search for elements greater than or equal to the minimum support, i.e., frequent items, and form a frequent item set of abnormal data from the scanned frequent items. Build the FP-Tree header table in descending order according to the element support;

[0123] S2.6 Build the FP-tree. Conduct the second scan on the frequent item set, establish branches for the FP-Tree header table, and insert the frequent items of each branch in descending order successively to construct the FP-tree with Null (zero device) as the root node;

[0124] S2.7 Traverse each item from the bottom of the FP-Tree in turn, find the corresponding conditional FP-tree according to each frequent 1-item set in the item header table, mine the frequent item set, calculate the confidence between the elements of the frequent item set, retain the strong association rules greater than the minimum confidence, and conduct data mining until all conditional FP-trees are mined. Store the abnormal data set and the association rule library composed of association rules different from other elements.

[0125] Furthermore, process the stored abnormal data set and implement multi-source heterogeneous abnormal information recognition, including:

[0126] S3.1 For different distribution facility characteristics X of the abnormal data set std , conduct standardization processing as shown in formula (7):

[0127] X std ={x 1 ,x 2 ,…x n}(7)

[0128] Among them, x n is the nth sample point of the distribution facility characteristic, and each sample has k-dimensional characteristics, as shown in formula (8):

[0129] x i =(x i1 ,x i2, …x ik )(8);

[0130] S3.2 Each feature in the k-dimensional features corresponds to a performance parameter of a certain distribution facility; the performance parameters of the distribution facility include: the power generation of the server cabinet, the usage frequency of the industrial router, and the acquisition rate of the Internet of Things intelligent terminal acquisition device; by calculating the covariance matrix R of the facility characteristics, obtain the mutual influence and correlation between different facilities, and match with the association rule library. The calculation formula of the covariance matrix R is as shown in (9):

[0131]

[0132] Among them, X std represents the standardized data matrix, where each column represents a feature, i.e., a certain performance index of the power distribution facilities, and each row represents an observation, i.e., the performance reading at a specific time. (X std ) T is the transpose matrix of X std ; n is the number of data points collected every day in a month;

[0133] S3.3 Perform eigenvalue decomposition on the covariance matrix R to obtain the eigenvalues λ i and the eigenvectors v i ; Reorder the results of the eigenvalue decomposition from largest to smallest and select the eigenvectors corresponding to the largest n eigenvalues. Project each data point onto the eigenvectors to obtain the values of the principal components, as shown in formula (10):

[0134] PC i = X std · v i (10)

[0135] where i = 1, 2,..., n, PC i represents the i-th principal component, which represents the key operating modes or the most significant energy consumption characteristics of different facilities; v i represents the eigenvector corresponding to the i-th eigenvalue of the covariance matrix R, indicating the direction in which the facility data changes the most; n represents the number of principal components selected for retention;

[0136] S3.4 Calculate the cumulative information contribution rate η p of the first p (p ≤ n) principal components among the eigenvalues, and the information contribution rate y i of the eigenvalue λ i is as shown in formula (11):

[0137]

[0138] η p The calculation formula of is as shown in (13):

[0139]

[0140] S3.5 Calculate the dimensionality reduction result. Use the transformation matrix T p =(t1, t2,..., t p ), which is composed of the eigenvectors corresponding to the first p eigenvalues. The dimensionality reduction result is as shown in formula (12):

[0141] p i = T p x1 (12)

[0142] S3.6 Based on the data after dimensionality reduction in S3.5, input it into the SVM classifier, adopt the Polynomial kernel function, and the calculation result formula is as shown in (14): k(x,x i ) = [γ * (x · x i ) + coef] d (14)

[0143] where d is the order of the polynomial and coef is the bias coefficient; realize the identification of multi-source heterogeneous abnormal information.

[0144] S3.7 Let l i be the abnormal information identification value. Using the abnormal information prediction model is to find the relationship between p i and l i , as shown in formula (15):

[0145] l i = f(p i ) (15)

[0146] S3.8 Input the result k(x,x i ) calculated by the kernel function into the f(p) function to calculate the abnormal information identification value, and realize the identification of multi-source heterogeneous abnormal information. The calculation formula of f(p) is as shown in (16):

[0147]

[0148] where, a i , are Lagrange multipliers and b is the bias term.

[0149] Adopt a perception system for the multi-source heterogeneous information perception method of distribution facilities for a low-carbon park as described above, as Figure 5 shown, including:

[0150] A multi-source heterogeneous information collection module, which collects distribution facility information based on the improved YOLOv5 and establishes a multi-source heterogeneous information dataset of distribution facilities;

[0151] An abnormal dataset and associated rule library storage module, which analyzes the association characteristics of distribution facility information based on the FP-Growth algorithm for the collected multi-source heterogeneous information dataset of distribution facilities, and stores the abnormal dataset and the associated rule library composed of associated rules different from other elements of the abnormal data;

[0152] A multi-source heterogeneous abnormal information identification module, which processes the stored abnormal dataset, matches it with the associated rule library, performs dimensionality reduction, inputs it into the SVM classifier, calculates using the Polynomial kernel function, and inputs the result calculated by the kernel function into the f(p) function to calculate the abnormal information identification value, and realizes the identification of multi-source heterogeneous abnormal information.

[0153] To verify the effectiveness of the method, performance test experiments were conducted on the distribution facility information acquisition method based on the improved YOLOv5 and the multi-source heterogeneous abnormal information sensing and recognition of distribution facilities based on feature extraction. Experiment 1: The collected current and voltage waveform information is as follows Figure 6 shown; 20 target nodes are randomly deployed in the wireless sensor network. The sensing radius and sensor nodes are designed as variables, and the number of sensor nodes with a sensing radius of 20m is set to 40, 80, and 160 respectively. The comparison methods include: the improved YOLOV5 model, the Greedy strategy algorithm that selects the sensor node with the largest number of covered targets each time; the maximum lifetime target coverage rate algorithm (MLTC) that selects the sensor node with the highest remaining energy; the ALAA algorithm calculates the sensing coverage rate respectively, and the experimental results are as follows Figures 7 - 8 shown. It can be seen from Figures 7 - 8 that when the number of sensors increases and the coverage situation becomes complex, the effects of the other three algorithms are not good. When the sensing radius and sensor nodes are changed, the improved YOLOV5 model still outperforms the other algorithms. By improving the YOLOV5 model, the sensing coverage rate in the low-carbon park is increased to 91%, which can effectively help the managers of the low-carbon park to more accurately and comprehensively identify the concerned areas and distribution facilities.

[0154] Experiment 2: The multi-source heterogeneous data of distribution facilities after the FP-Growth algorithm and PCA dimensionality reduction are used as the sample data set. To test the abnormal information recognition effect, 5 groups of data with different data volumes are prepared for the experiment, which are: 2000 data, 4000 data, 8000 data, 10000 data, and each group contains 15% abnormal sample data. The experimental results are as follows Figure 9 shown. It can be seen from Figure 9 that SVM-FP-PCA can significantly improve the accuracy of abnormal information recognition of distribution facilities in the low-carbon park in different data volume cases, reaching more than 91%, improving the efficiency and reliability of the low-carbon park power distribution system in processing abnormal information, and enabling the managers of the low-carbon park to better manage the distribution facilities in the park.

[0155] The above are only specific embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multi-source heterogeneous information perception method for power distribution facilities in a low-carbon park, characterized in that Including: Collecting distribution facility information based on improved YOLOv5 to establish a multi-source heterogeneous information dataset for distribution facilities; Analyzing the correlation characteristics of distribution facility information based on the FP-Growth algorithm for the collected multi-source heterogeneous information dataset of distribution facilities, and storing the abnormal dataset and the association rule library composed of association rules of other elements different from the abnormal data; Processing the stored abnormal dataset, matching it with the association rule library, after dimensionality reduction, inputting it into the SVM classifier, calculating using the Polynomial kernel function, and inputting the result of the kernel function calculation into the f(p) function to calculate the abnormal information recognition value, realizing the recognition of multi-source heterogeneous abnormal information.

2. The multi-source heterogeneous information perception method for power distribution facilities in a low-carbon park according to claim 1, characterized in that, Collecting distribution facility information based on improved YOLOv5 and establishing a multi-source heterogeneous information dataset for distribution facilities, including: S1.1 Based on the location information of distribution facilities in the low-carbon park, deploying intelligent induction cameras, current, voltage, temperature, and humidity sensors; S1.2 Using the generated videos from real-time monitoring, separating each frame into images to generate a distribution facility monitoring dataset; S1.3 Analyzing the fog and brightness conditions in the images according to the distribution facility monitoring dataset; S1.4 Using the dark channel processing method for image dehazing; for any input image J, its dark channel is shown in formula (1): Among them, J on the left side of the equation dark is the dark channel map. On the right side of the equation, J c represents each channel of the color image. Ω(x) is a square region centered on pixel x. C represents one of the three channels of R (red), G (green), and B (blue). min represents the minimum value of the region. x and y are variables; S1.5 Deriving the image dehazing formula through formula (2) to complete the removal of fog in the image, as shown in formula (2): Among them, t0 is the preset minimum value of the medium propagation coefficient, and max represents the regional maximum value; Based on the principle of image dehazing, using the atmospheric scattering model, as shown in formula (3): I(x) = J(x)t(x) + A(1 - t(x)) (3) Among them, I(x) is the observed pixel value, J(x) is the original brightness of the scene, t(x) is the transmittance, and A is the global atmospheric light; S1.6 According to the result of step S1.5, using the LabelMe tool to perform label binarization on the processed distribution facility monitoring dataset; S1.7 Using the improved YOLOV5 object detection model to detect the images, and the model uses the feature map of the P-bifpn (PathBidirectional Feature Pyramid Network) aggregation scale to fuse the low-level features with the high-level features; S1.8 Using the fully connected layer to extract the non-linear features emerging in the multi-dimensional data space, converting the distribution facility monitoring dataset through the fully connected layer to realize the fusion with the collected sensor data, and establishing a multi-source heterogeneous information dataset for distribution facilities.

3. A multi-source heterogeneous information perception method for power distribution facilities in a low-carbon park according to claim 1, characterized in that, Performing the correlation characteristic analysis of distribution facility information based on the FP-Growth algorithm and storing the abnormal dataset and the association rule library composed of association rules of other elements different from the abnormal data, including: S2.1 For the collected multi-source heterogeneous information dataset of distribution facilities, using the Pandas library to delete duplicate monitoring data values, filling missing monitoring data values with KNN, and replacing abnormal values; S2.2 Utilize the boundary value information of the sample space to linearly map the features into a specific range. The calculation equation of Min - Max normalization is shown in formula (5): Among them, x is the value before sample normalization, and x new is the value after normalization, and x min is the minimum value of the sample space, and x max is the maximum value of the sample space; S2.3 Use generative adversarial networks (GANs) to fuse multi - source heterogeneous information in the low - carbon park and generate a synthetic model that can understand and reproduce the characteristics of different data sources; S2.4 Import the processed multi - source heterogeneous information dataset of distribution facilities and set the minimum support and minimum confidence; S2.5 Conduct the first scan of the fused dataset to search for elements greater than or equal to the minimum support, which are frequent items. The scanned frequent items form an abnormal data frequent item set, and a FP - Tree header is constructed by sorting the elements in descending order according to their support degrees; S2.6 Build a FP - tree, conduct the second scan of the frequent item set, establish branches for the FP - Tree header, and insert the frequent items of each branch in descending order successively. A FP - tree is constructed with Null (zero device) as the root node; S2.7 Traverse each item from the bottom of the FP - Tree in turn, find the corresponding conditional FP - tree according to each frequent 1 - item set in the item header table, mine the frequent item sets, calculate the confidence between the elements of the frequent item sets, retain the strong association rules with confidence greater than the minimum confidence, and conduct data mining until all conditional FP - trees are mined. Store the abnormal dataset and the association rule library composed of association rules different from other elements.

4. A multi-source heterogeneous information perception method for power distribution facilities in a low-carbon park according to claim 1, characterized in that Process the stored abnormal dataset and realize the identification of multi - source heterogeneous abnormal information, including: S3.1 For different power distribution facility features X of the abnormal data set std , perform normalization processing as shown in formula (7): X std = {x 1 , x 2 , … x n} (7) where x n is the nth sample point of the distribution facility characteristics, and each sample has k-dimensional characteristics, as shown in formula (8): x i = (x i1 , x i2 , … x ik ) (8); In the k - dimensional features described in S3.2, each feature corresponds to a performance parameter of a certain distribution facility; the distribution facility performance parameters include: the power generation of the server cabinet, the usage frequency of the industrial router, and the acquisition rate of the Internet of Things intelligent terminal acquisition device. By calculating the covariance matrix R of the facility features, the mutual influence and correlation between different facilities are obtained and matched with the association rule library. The calculation formula of the covariance matrix R is shown in (9): Among them, X std represents the standardized data matrix, where each column represents a feature, i.e., a certain performance index of the power distribution facilities, and each row represents an observation, i.e., the performance reading at a specific time. (X std ) T is the transpose matrix of X std ; n is the number of data points collected every day in a month; S3.3 Perform eigenvalue decomposition on the covariance matrix R to obtain eigenvalues λ i and eigenvectors v i ; reorder the results of the eigenvalue decomposition from largest to smallest, select the eigenvectors corresponding to the largest n eigenvalues, project each data point onto the eigenvectors, and obtain the values of the principal components, as shown in formula (10): PC i = X std · v i (10) where \(i = 1, 2, \ldots, n\), PC i represents the \(i\)-th principal component, which represents the key operating modes of different facilities or the most significant energy consumption characteristics; \(v\) i represents the eigenvector corresponding to the \(i\)-th eigenvalue of the covariance matrix \(R\), indicating the direction in which the facility data changes the most; \(n\) represents the number of principal components to be retained; S3.4 Calculate the cumulative information contribution rate η of the first p (p ≤ n) principal components among the eigenvalues p , eigenvalue λ i 's information contribution rate y i is as shown in formula (11): S3.5 Calculate the dimensionality reduction result using the transformation matrix T p =(t1, t2, …, t p ), which is composed of the eigenvectors corresponding to the first p eigenvalues. The result after dimensionality reduction is as shown in formula (12): p i = T p x1 (12) S3.6 Based on the data after dimensionality reduction in S3.5, input it into the SVM classifier, adopt the Polynomial kernel function, and the calculation result formula is as shown in (14): k(x,x i ) = [γ * (x · x i ) + coef] d (14) where d is the order of the polynomial and coef is the bias coefficient; realize the identification of multi - source heterogeneous abnormal information. S3.7 Set l i as the abnormal information recognition value, and use the abnormal information prediction model to find p i and l i The relationship between them is shown in formula (15): l i = f(p i )(15) S3.8 Input the result k(x, x i ) of the kernel function calculation into the f(p) function to calculate the abnormal information recognition value, realizing multi-source heterogeneous abnormal information recognition. The calculation formula of f(p) is shown in (16): where a i , is a Lagrange multiplier and b is a bias term.

5. A multi-source heterogeneous information perception method for power distribution facilities in a low-carbon park according to claim 2, characterized in that, Step S1.7 includes: The P - bifpn module is divided into P3, P4, P5, and P4 - td. On the basis of the original network, add a residual edge from P4 to P4 - out to fuse features of more scales; according to the contribution degree of the input features to the output features, add a bottom - up channel composed of P3 - out, P4 - out, and P5 - out. Introduce trained weights at the nodes where multi - features are fused, multiply them with the input features of each node to meet the different contributions of different input features to the final output, and use the fast normalization formula to train these weights, as shown in formula (4): where w i , w j represents different feature weights. After obtaining each w i , the Relu function is introduced to ensure that w i > 0, I i represents the i-th input feature; to ensure numerical stability, ε = 0.0001.

6. A multi-source heterogeneous information perception method for power distribution facilities in a low-carbon park according to claim 3, characterized in that Step S2.1 includes: Using the K-Nearest Neighbors (KNN) algorithm to calculate the distances between samples in the cleaned data, finding the K neighbor samples most similar to the sample with missing values, and estimating the missing values; if these missing values are continuous variables, taking the average or weighted average of the feature values of these neighbor samples as the estimate; if they are categorical variables, determining the most likely category through a voting mechanism; for the handling of outliers, also using the feature values of these neighbor samples to replace.

7. A multi-source heterogeneous information perception method for power distribution facilities in a low-carbon park according to claim 3, characterized in that, Step S2.3 includes: The generative adversarial network includes a G network and a D network. The G network is a learning network of a generative model, which accepts the data of the sample set for training and generates a prediction model, denoted as G(z); D is a discriminative network, whose input parameter is x, where x represents the accurate value, and the output represents the probability that x is the accurate value. If it is 1, it means the prediction accuracy is 100%, and so on, as shown in formula (6): Among them, x is the true value, z is the input training sample, and G(z) represents the prediction model generated by the G network; D(G(z)) is the probability that the D network judges whether the prediction model generated by G is accurate. G should hope that the model it generates is closer to the accurate value. The maximum value of G is taken as D(G(z)), and V(D, G) becomes smaller.

8. A multi-source heterogeneous information perception method for distribution facilities in a low-carbon park according to claim 4, characterized in that In step S3.4, η p is calculated according to the formula shown in (13):

9. A sensing system adopting a multi-source heterogeneous information sensing method for power distribution facilities in a low-carbon park as described in claim 1, characterized in that, It includes: The multi-source heterogeneous information acquisition module, which is used to collect the distribution facility information based on the improved YOLOv5 and establish a multi-source heterogeneous information dataset of distribution facilities; The abnormal dataset and association rule library storage module, which analyzes the association characteristics of the distribution facility information of the collected multi-source heterogeneous information dataset of distribution facilities based on the FP-Growth algorithm, and stores the abnormal dataset and the association rule library composed of association rules that are different from other elements of the abnormal data; The multi-source heterogeneous abnormal information recognition module, which processes the stored abnormal dataset, matches it with the association rule library, and after dimensionality reduction, inputs it into the SVM classifier, calculates using the Polynomial kernel function, and inputs the result of the kernel function calculation into the f(p) function to calculate the abnormal information recognition value, realizing the recognition of multi-source heterogeneous abnormal information.