Energy big data knowledge graph construction method based on multi-source data
By analyzing the disturbance and fluctuation characteristics of the operating data in the power and energy big data, adaptively adjusting the number of components of the Gaussian hybrid model, solving the problem of overfitting or underfitting in data cleaning, and improving the construction accuracy of the knowledge graph.
Patent Information
- Application Number
- CN202510070639.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
In power and energy big data, the diversity and potential connection of multi-source data leads to the problems of overfitting or underfitting when cleaning data, which affects the construction accuracy of the knowledge graph.
By analyzing the differences and dispersion between the operating data of each node and the fitting curve, combining volatility and complexity, the number of components of the Gaussian hybrid model is adaptively adjusted to determine the number of adaptive components of each node, and using the knowledge graph tool to build a knowledge graph.
Improve the accuracy of data cleaning and the accuracy of knowledge graph construction, and can more accurately identify abnormal data in nodes and the impact of network attacks.
Smart Images

Figure CN119988643A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of knowledge graph construction, and specifically to a method for constructing an energy big data knowledge graph based on multi-source data. Background Art
[0002] Energy big data refers to the data collection in the energy field. This plan focuses on electric energy big data, which refers to the multi-source data at each node in the process of electric energy production, distribution, consumption, etc. By constructing a knowledge graph for electric energy big data, it is helpful to integrate multi-source power data, allocate capacity resources at each node, and make power deployment decisions, thereby improving the intelligence level of the electric energy industry.
[0003] When constructing a knowledge graph based on multi-source data of each node in the power energy big data, the accuracy and perfection of the knowledge graph is heavily dependent on the accuracy of the collected data. In the data collection process, corresponding sensors are usually deployed on each power node, and data is collected by sensors and uploaded to the power energy management center using the information network. Due to the factors of power node information network attacks, the collected data may have defects and outliers, and data cleaning is required before building the knowledge graph. Usually, the Gaussian mixture model (GMM) based on the Gaussian mixture model (Gaussian Mixture Model) is used to implement data cleaning of power energy big data. In the GMM Gaussian mixture model, multiple Gaussian distributions are used to fit data based on a single type of variable, so the number of components usually needs to be set. However, the data types of each node in the power energy big data are diverse, and there are potential connections between each node. In the GMM Gaussian mixture model, the complexity of the model is adjusted by the number of components to determine the sensitivity of the model. If the number of components is too large, the model will learn abnormal data in the data, resulting in model overfitting; if the number of components is too small, the model will not be able to capture effective information in the data, resulting in underfitting, thereby affecting the accuracy of data cleaning and reducing the accuracy of knowledge graph construction. Summary of the invention
[0004] In order to solve the above technical problems, the present application provides a method for constructing an energy big data knowledge graph based on multi-source data to solve the existing problems.
[0005] The present invention adopts the following technical solution to construct a knowledge graph of energy big data based on multi-source data:
[0006] In the transmission of electric energy, all operation data of each electric device at each collection time within a preset time period, as well as the level and attributes of each electric device, are obtained, and the electric devices are recorded as nodes;
[0007] Perform curve fitting on each item of operating data in each node, and determine the disturbance degree of each item of operating data in each node based on the difference between each item of operating data and the fitting value on the corresponding fitting curve and the discrete degree of each item of operating data;
[0008] Determine the volatility of each item of operating data in each node based on the difference between each item of operating data in each node and all other items of operating data, and determine the volatility complexity of each item of operating data in each node in combination with the disturbance degree;
[0009] Based on the distance from each node to all nodes of the same level and attributes, the reference node of each node is determined; based on the distance from each node to any of its reference nodes, combined with the correlation between each node and all its reference nodes on various operating data, the fluctuation difference of each operating data in each node is determined; based on the fluctuation difference and fluctuation complexity, the fluctuation sensitivity of each operating data in each node is determined;
[0010] Based on the fluctuation sensitivity, the preset initial number of components and the preset scaling value, the number of adaptive components for each operating data in each node is determined, and a knowledge graph is constructed in combination with the knowledge graph tool.
[0011] Preferably, the expression of the disturbance degree of each operation data in each node is: i,j =σ i,j ×Δd i,j Where A i,j represents the disturbance degree of the j-th operation data in node i; σ i,j Indicates the degree of dispersion of all data in the jth running data of node i; Δd i,j Represents the mean of the differences between all data in the j-th running data in node i and the fitted values on the corresponding fitting curve.
[0012] Preferably, the fluctuation degree of each item of operating data in each node is an average level of the difference between each item of operating data in each node and all other items of operating data.
[0013] Preferably, the fluctuation complexity of each item of operating data in each node is a result of integrating the fluctuation degree and the disturbance degree of each item of operating data in each node.
[0014] Preferably, the method for determining the reference node of each node is:
[0015] Arrange the distances from each node to all nodes of the same level and attributes in ascending order, and select the first preset number of nodes in the result as reference nodes for each node.
[0016] Preferably, the method for determining the fluctuation difference of each item of operating data in each node is:
[0017] Based on the distance between each node and any of its reference nodes, determine the distance weight of any of the reference nodes of each node;
[0018] The fluctuation difference C of the j-th operation data in node i i,j The expression is: In the formula, represents the correlation between node i and its reference node k on the jth running data; represents the distance weight of the reference node k of node i; M i represents the total number of reference nodes of node i; ε represents a constant preset to be greater than 0.
[0019] Preferably, the distance weight of any reference node of each node is determined by:
[0020] The cumulative sum of the distances between each node and all its reference nodes is calculated, and the ratio of the distance from each node to any of its reference nodes to the cumulative sum is used as the distance weight of any reference node of each node.
[0021] Preferably, the fluctuation sensitivity of each item of operating data in each node is a normalized value of the product of the fluctuation complexity and the fluctuation difference of each item of operating data in each node.
[0022] Preferably, the expression for the number of adaptive components of each item of operating data in each node is: Where N i,j represents the number of adaptive components of the j-th running data in node i; represents the fluctuation sensitivity of the j-th operation data in node i; n0 represents the preset initial number of components; Indicates the preset zoom value; Represents the ceiling function.
[0023] Preferably, the process of constructing the knowledge graph is:
[0024] The running data of all items in each node are used as the input of the knowledge graph tool, wherein the Gaussian mixture model is used for data cleaning in the knowledge graph tool, and the adaptive number of components of each running data item in each node is used as the component number parameter in the Gaussian mixture model to output the knowledge graph.
[0025] An embodiment of the present application provides a method for constructing an energy big data knowledge graph based on multi-source data, the method comprising the following steps:
[0026] This application has at least the following beneficial effects:
[0027] The present application determines the fluctuation complexity of each item of operating data in each node by analyzing the difference between each item of operating data in each node and the fitting value on the corresponding fitting curve, combined with the discrete degree of each item of operating data, and the difference between each item of operating data in each node and all other items of operating data. Its beneficial effect is to identify whether there is abnormal data in the operating data of the node; the present application determines the reference node of each node based on the distance from each node to all nodes of the same level and attributes; determines the fluctuation difference of each item of operating data in each node based on the distance from each node to any of its reference nodes, and the correlation between each node and all its reference nodes in each item of operating data degree, and its beneficial effect is that it can accurately determine whether a node has been attacked abnormally, so as to identify abnormal data in the operating data of the node that has been attacked abnormally; the present application determines the fluctuation sensitivity of each operating data in each node based on the fluctuation difference and fluctuation complexity, and its beneficial effect is that it can accurately locate the degree to which each operating data of the node has been tampered with when it is attacked by a network; the present application determines the number of adaptive components of each operating data in each node based on the fluctuation sensitivity, the preset initial number of components and the preset scaling value, and constructs a knowledge graph in combination with the knowledge graph tool, and its beneficial effect is that it improves the accuracy of data cleaning, thereby improving the accuracy of knowledge graph construction. The present application obtains the number of adaptive components by analyzing the fluctuation distribution of the operating data of the node when it is attacked by a network, as well as the characteristics of the diffusion of adjacent nodes, thereby improving the accuracy of data cleaning in the construction of the knowledge graph and the accuracy of the knowledge graph construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0029] Figure 1 A flowchart of a method for constructing an energy big data knowledge graph based on multi-source data provided in one embodiment of the present application;
[0030] Figure 2 A schematic diagram of the knowledge graph construction principle provided for one embodiment of the present application;
[0031] Figure 3 A schematic diagram of an adaptive component number extraction process provided by an embodiment of the present application;
[0032] Figure 4 A comparison chart of data cleaning effects before and after improvement of the Gaussian mixture model provided in one embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to further explain the technical means and effects adopted by this application to achieve the predetermined invention purpose, the following is a detailed description of the specific implementation method, structure, features and effects of a method for constructing an energy big data knowledge graph based on multi-source data proposed in this application in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.
[0034] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0035] The following is a detailed description of a specific scheme for constructing an energy big data knowledge graph based on multi-source data provided by the present application in conjunction with the accompanying drawings.
[0036] An embodiment of the present application provides a method for constructing an energy big data knowledge graph based on multi-source data. Specifically, the following method for constructing an energy big data knowledge graph based on multi-source data is provided. Figure 1 , the method comprises the following steps:
[0037] Step S1: During the transmission of electric energy, all operation data of each electric device at each collection time within a preset time period, as well as the level and attributes of each electric device, are obtained.
[0038] In this embodiment, the overall electric energy big data of a city is taken as the research object. In the electric energy transmission, multiple power equipment such as power stations, relay stations, substations, transformers, etc. are included. Through the log data of the power equipment, all operating data of each power equipment at each collection time within the preset time t are obtained, including transmission voltage, node power and transmission frequency, wherein the data collection interval is set to T. At the same time, in order to simplify the expression, the power equipment is recorded as a node, and all the nodes mentioned below represent power equipment.
[0039] It should be noted that the values of the preset time length t and the data collection interval T are both manually set. In this embodiment, the value of the preset time length t is 1 day, and the value of the data collection interval T is 1 minute. The implementer can also set them according to the specific situation. This embodiment does not impose any special restrictions.
[0040] Furthermore, in order to eliminate the influence of the data unit dimension, the operation data of each node are normalized.
[0041] It should be understood that there are many commonly used data normalization methods. In this embodiment, the z-score normalization method is used to normalize the various operating data. The implementer can also use the maximum and minimum value normalization method to normalize the data. Regarding the selection of data normalization methods, this embodiment does not make any special restrictions.
[0042] Among them, the z-score normalization method is a well-known technology in the field of data processing, and the process of normalizing the data is not repeated here.
[0043] Furthermore, knowledge extraction is used to obtain the level and attributes of each node from all the operating data of each node. In this embodiment, the level of each node represents the working voltage of the power equipment, and the attribute represents the category of the power equipment. For example, for a substation operating at 110KV, 110KV is the level of the corresponding node, and the substation is the attribute of the corresponding node.
[0044] Among them, knowledge extraction is a well-known technology in the process of constructing knowledge graphs, and the specific process of obtaining node levels and attributes will not be repeated here.
[0045] Preferably, the schematic diagram of the knowledge graph construction principle provided in this embodiment is as follows Figure 2 shown.
[0046] Step S2: curve fitting is performed on each item of operating data in each node, and the disturbance degree of each item of operating data in each node is determined based on the difference between each item of operating data and the fitting value on the corresponding fitting curve and the discrete degree of each item of operating data.
[0047] When the equipment is operating normally, its corresponding operating data can basically remain consistent. Although affected by environmental noise, the collected data will only be slightly disturbed. However, when the node is attacked by the network, there will be abnormal data, which will break the original data distribution characteristics and cause large fluctuations in the node's operating data.
[0048] Therefore, in order to eliminate the influence of abnormal data existing during network attacks on the accuracy of knowledge graph construction, based on the difference between each operating data and the fitting value on the corresponding fitting curve, combined with the discrete degree of each operating data, the disturbance degree of each operating data in each node is determined to determine whether the operating data has been attacked by the network. Specifically:
[0049] (1) Fitting various operation data in each node to obtain fitting curves of various operation data;
[0050] It should be noted that there are many commonly used fitting methods. In this embodiment, the least squares polynomial fitting method is used to fit various operating data. The implementer can also use other nonlinear fitting methods such as polynomial regression. This embodiment does not impose any special restrictions on the selection of the fitting method.
[0051] Among them, the least squares polynomial fitting method is a well-known technology in the field of data fitting, and its specific fitting process will not be described in detail.
[0052] (2) Based on the difference between each operating data and the fitting value on the corresponding fitting curve, combined with the discrete degree of each operating data, the disturbance degree of each operating data in each node is determined, specifically:
[0053] The disturbance degree A of the jth running data in node i i,j The expression is: A i,j =σ i,j ×Δd i,j ; In the formula, σ i,j Indicates the degree of dispersion of all data in the jth running data of node i; Δd i,j Represents the mean of the differences between all data in the j-th running data in node i and the fitted values on the corresponding fitting curve.
[0054] It should be noted that there are many methods for measuring the difference between data. In this embodiment, the absolute value of the difference between the operating data and the fitting value on the corresponding fitting curve is calculated to measure the difference between the actually collected operating data and its fitting value. The implementer can also use other methods of measuring data differences such as ratios. This embodiment does not impose any special restrictions on the selection of methods for measuring data differences.
[0055] In addition, it should be understood that there are many ways to measure the degree of discreteness of a set of data. In this embodiment, the degree of discreteness of each operating data item is measured by calculating the standard deviation of all data in each operating data item in the node. The implementer may also use other methods of measuring the degree of discreteness, such as variance or dispersion coefficient. This embodiment does not impose any special restrictions on the selection of the method for measuring the degree of discreteness.
[0056] Furthermore, according to the disturbance degree of each item of operating data in each node, it can be understood that if the node is attacked by a network, there will be more abnormal data in the corresponding node, and the operating data will fluctuate greatly, making the standard deviation of the corresponding item of operating data larger. At the same time, due to the presence of more abnormal values in the operating data, it is difficult to fit the noise points corresponding to the abnormal data to the curve during curve fitting, resulting in a large difference between the actual operating data and the fitting value on the fitting curve, which ultimately makes the disturbance degree of the corresponding item of operating data of the node larger; conversely, if the node is not attacked by a network, the operating data will not fluctuate greatly, making the standard deviation of the corresponding item of operating data smaller, and the difference between the actual operating data and the fitting value on the fitting curve is smaller, which ultimately makes the disturbance degree of the corresponding item of operating data of the node smaller.
[0057] Step S3: Based on the difference between each item of operating data in each node and all other items of operating data, determine the volatility of each item of operating data in each node, and determine the volatility complexity of each item of operating data in each node in combination with the disturbance degree.
[0058] When the node is running normally, the environmental changes are consistent, so the fluctuation deviations between different items of operation data are relatively close. When the node is attacked by the network, most of the node operation data in the log is randomly forged data due to the attack. Therefore, the abnormal fluctuation of data is not only reflected in a single data item, but also in other data of the node due to the network attack. There are more random forged data, and the fluctuation deviations of the random forged data of different items of operation data are quite different.
[0059] Therefore, in order to identify whether a node has been attacked and there are a large number of abnormal data in the operation data, the fluctuation degree of each operation data in each node is determined according to the difference between each operation data in each node and all other operation data, and the fluctuation complexity of each operation data in each node is determined in combination with the disturbance degree, and the fluctuation deviation between different operation data is analyzed to identify abnormal data in the node operation data, specifically:
[0060] (1) Analyze the average level of the difference between each item of operating data in each node and all other items of operating data, and record it as the volatility of each item of operating data in each node.
[0061] It should be noted that there are many methods for measuring the differences between data groups. In this embodiment, the DTW (Dynamic Time Warping) distance between all data between each operating data in each node and any other operating data is calculated to measure the differences between different operating data. Implementers can also use other methods for measuring the differences between data groups, such as Euclidean distance and Manhattan distance. This embodiment does not impose any special restrictions on the selection of methods for measuring the differences between data groups.
[0062] The calculation process of the DTW distance is a well-known technique, and the specific calculation steps are not described in detail.
[0063] In addition, it should be understood that there are many ways to measure the average level of a set of data. In this embodiment, the average level of the difference is measured by calculating the mean of the difference between each item of operating data in each node and all other items of operating data. Implementers can also use other methods to measure the average level of data, such as the geometric mean. Regarding the selection of the method for measuring the average level of data, this embodiment does not impose any special restrictions.
[0064] (2) Further, based on the disturbance degree and the volatility, the volatility complexity of each item of operating data in each node is determined, specifically: the volatility complexity of each item of operating data in each node is the result of fusing the mean of the differences of each item of operating data in each node with the disturbance degree.
[0065] It should be understood that fusion refers to the result of combining two or more indicators through positive fusion, that is, combining two or more indicators by adding or multiplying them together to obtain a comprehensive indicator, so as to more comprehensively and accurately evaluate a phenomenon or problem. This fusion method is not limited to simple arithmetic operations, but can also include more complex statistical models and analysis methods, which can be selected by the implementer according to the specific situation, and this embodiment does not impose any special restrictions.
[0066] Preferably, in the present embodiment, the fluctuation complexity of each item of operating data in each node is the product of the fluctuation degree and the disturbance degree of each item of operating data in each node; in actual application, as other implementation methods, the fluctuation complexity of each item of operating data in each node is an exponential function value with a natural constant as the base and the sum of the fluctuation degree and the disturbance degree of each item of operating data in each node as the independent variable.
[0067] Furthermore, according to the fluctuation complexity of each item of operating data in each node, it can be understood that if the node is attacked by a network attack, the attack content is often to randomly forge various items of operating data, which causes a large deviation between different items of operating data. At the same time, there are more abnormal data in the operating data, and the greater the disturbance of the operating data, the greater the fluctuation complexity of the corresponding item of operating data of the node. On the contrary, if the current node is a normal node and has not been attacked by a network, the operating data of each item basically remains consistent, the deviation between different items of operating data is small, and the operating data is relatively stable. The smaller the disturbance, the smaller the fluctuation complexity.
[0068] Step S4: Determine the reference node of each node based on the distance from each node to all nodes of the same level and attributes; determine the distance weight of any reference node of each node based on the distance from each node to any reference node, and determine the fluctuation difference of each operating data in each node based on the correlation between each node and all its reference nodes in each operating data; determine the fluctuation sensitivity of each operating data in each node based on the fluctuation difference and fluctuation complexity.
[0069] Since there is a certain correlation between the nodes in the power energy system, when a power device is attacked by the network, the operating data of the device is false data forged by the attacker, and in order to ensure the normal operation of the power system, the abnormal situation of the current power device will spread to the surrounding operating devices. Therefore, the abnormal situation of each node's operating data can be further evaluated through the operating status of the current device and the associated devices, specifically:
[0070] In the power system, each node is interrelated. When the system is operating normally, the closer the node is to the current node, the more similar its operating status is to that node. However, due to the influence of power transmission attenuation, the farther the node is from the current node, the greater the difference in its operating status will be. However, if a single node is attacked by a network attack, the operating data of the node is forged abnormal data, and the power system will dynamically allocate power resources based on the operating status of each device node. Therefore, the abnormal data of the attacked node will spread to the related nodes, and the abnormal data is disordered, thus causing confusion in the operating data of the related device nodes, further reducing the correlation between the two, making the closer the nodes, the greater the difference, and the farther the nodes are from the current node, the less interference they receive, and the closer their operating status is.
[0071] Therefore, in order to determine whether a node has been attacked abnormally and identify abnormal data in the operating data, the distance weight of any reference node of each node is determined based on the distance between each node and any of its reference nodes, and the fluctuation difference of each operating data in each node is determined in combination with the correlation between each node and all of its reference nodes in each operating data; based on the fluctuation difference and fluctuation complexity, the fluctuation sensitivity of each operating data in each node is determined, specifically:
[0072] (1) Arrange the distances of each node to all nodes of the same level and attributes in ascending order, and select the first preset number of nodes as reference nodes for each node;
[0073] It should be noted that the value of the preset number is artificially set. In this embodiment, the value of the preset number is 5. The implementer can also set it by himself according to the specific situation. This embodiment does not impose any special restrictions.
[0074] In addition, it should be understood that the so-called same level and same attributes mean the same working voltage and the same category. For example, substations operating at the same voltage are of the same level and the same attributes.
[0075] (2) Further, the cumulative sum of the distances between each node and all its reference nodes is calculated, and the ratio of the distance from each node to any of its reference nodes to the cumulative sum is used as the distance weight of any of the reference nodes of each node;
[0076] (3) Further, based on the distance weight of any reference node of each node and the correlation between each node and all its reference nodes in each operation data, the fluctuation difference of each operation data in each node is determined, specifically:
[0077] The fluctuation difference C of the j-th operation data in node i i,j The expression is: In the formula, represents the correlation between node i and its reference node k on the jth running data; represents the distance weight of the reference node k of node i; M i Represents the total number of reference nodes of node i; ε represents a preset constant greater than 0, which is used to prevent the denominator from being 0. The value of ε is set manually. In this embodiment, the value of ε is 0.01. Under the premise of ensuring that the denominator is not 0 and does not excessively affect the calculation results, the implementer can also set it according to the specific situation. This embodiment does not impose any special restrictions.
[0078] It should be noted that there are many methods for measuring the correlation between data groups. In this embodiment, the Pearson correlation coefficient of each node and its reference node on each item of operating data is calculated to measure the correlation between each node and its reference node on the same item of operating data. Implementers can also use other methods for measuring the correlation of data groups, such as the Spearman correlation coefficient or the Kendall rank correlation coefficient. Regarding the selection of methods for measuring the correlation of data groups, this embodiment does not impose any special restrictions.
[0079] Furthermore, according to the fluctuation difference of each operation data in each node, it can be understood that by evaluating the relevant situation of the nodes associated with the same level and attributes as the current node in the operation data, the fluctuation difference is obtained. If the current node is attacked, the node closest to the current node will be more affected, that is, the distance weight of the reference node of the current node will be larger. Since the node is attacked, there are more abnormal data in the operation data. Therefore, the difference in operation data between the current node and its reference node is larger, and the correlation is smaller, so the final fluctuation difference is larger; conversely, if the current node is not attacked, the node closest to the current node will be less affected, that is, the distance weight of the reference node of the current node will be smaller, and the difference in operation data between the current node and its reference node will be smaller, and the correlation is greater, so the final fluctuation difference is smaller.
[0080] (4) Further, based on the fluctuation difference and fluctuation complexity, the fluctuation sensitivity of each item of operating data in each node is determined, and the fluctuation sensitivity of each item of operating data in each node is the normalized value of the product of the fluctuation complexity and the fluctuation difference of each item of operating data in each node.
[0081] According to the fluctuation sensitivity of each operating data in each node, it can be understood that for nodes that are attacked by network, the greater the fluctuation difference and fluctuation complexity of their operating data, the greater the fluctuation sensitivity of the operating data, which means that the distribution of the operating data is relatively disordered at this time. Therefore, it is necessary to appropriately increase the number of components so that the model can better identify false data. On the contrary, for nodes of normal power equipment, the smaller the fluctuation difference and fluctuation complexity of their operating data, the smaller the fluctuation sensitivity, which means that the distribution of the operating data of the corresponding nodes is basically the same. At this time, the number of components can be reduced to increase the model calculation speed.
[0082] Step S5: Based on the fluctuation sensitivity of each operating data item in each node, the preset initial component number and the preset scaling value, determine the number of adaptive components for each operating data item in each node, and construct a knowledge graph in combination with the knowledge graph tool.
[0083] Based on the fluctuation sensitivity obtained in step S4, and in combination with the preset initial component number and the preset scaling value, the number of adaptive components for each operation data in each node is determined, and the operation data is cleaned to construct a knowledge graph, specifically:
[0084] Based on the fluctuation sensitivity of each operation data in each node, the preset initial component number and the preset scaling value, the adaptive component number of each operation data in each node is determined, specifically:
[0085] The number of adaptive components N for the jth running data in node i i,j The expression is: In the formula, represents the fluctuation sensitivity of the j-th operation data in node i; n0 represents the preset initial number of components; Indicates the preset zoom value; Represents the ceiling function.
[0086] It should be noted that the values of the preset initial number of components and the preset scaling value are both manually set. In this embodiment, the value of the preset initial number of components is 1, and the value of the preset scaling value is 4. The implementer can also set them according to the specific situation. This embodiment does not impose any special restrictions.
[0087] Preferably, the schematic diagram of the adaptive component number extraction process provided in this embodiment is as follows: Figure 3 shown.
[0088] Furthermore, all the running data of each node are used as the input of the knowledge graph power, wherein the GMM Gaussian mixture model is used for data cleaning in the knowledge graph tool, and the adaptive number of components of each running data item in each node is used as the component number parameter in the Gaussian mixture model to output the knowledge graph.
[0089] Among them, the GMM Gaussian mixture model and the process of building a knowledge graph using a knowledge graph tool are both well-known technologies, and the specific process of data cleaning and building a knowledge graph will not be repeated here.
[0090] So far, this embodiment obtains the fluctuation sensitivity of the power equipment nodes by analyzing the abnormal fluctuations of the operating data when the power equipment nodes are attacked by the network and the diffusion characteristics of the adjacent power equipment nodes, and determines the number of adaptive components so that the number of components can be better matched with the abnormal data, thereby improving the accuracy of data cleaning and the accuracy of knowledge graph construction during the knowledge graph construction process.
[0091] Preferably, the comparison diagram of data cleaning effect before and after the improvement of the Gaussian mixture model provided in this embodiment is as follows: Figure 4 As shown; the horizontal axis in the figure represents the amount of data; the vertical axis represents the F value. The method for obtaining the F value in this embodiment is: in this embodiment, the F1 score (F1-score) is used as an evaluation index to determine the data cleaning accuracy. The larger the output F value, the better the effect of cleaning the data, and vice versa. The dotted line represents the curve of the traditional Gaussian mixture model on the data cleaning effect; the solid line represents the curve of the data cleaning effect after improving the number of components in the Gaussian mixture model in the solution provided in this embodiment, wherein the calculation process of the F1 score is a well-known technology and will not be described in detail.
[0092] It should be noted that the above sequence of the embodiments of the present application is for description only and does not represent the advantages and disadvantages of the embodiments. The above is a description of a specific embodiment of this specification. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0093] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.
[0094] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Modifications to the technical solutions recorded in the aforementioned embodiments, or equivalent replacement of some of the technical features therein, do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for constructing an energy big data knowledge graph based on multi-source data, characterized in that: The method comprises the following steps: In the transmission of electric energy, all operation data of each electric device at each collection time within a preset time period, as well as the level and attributes of each electric device, are obtained, and the electric devices are recorded as nodes; Perform curve fitting on each item of operating data in each node, and determine the disturbance degree of each item of operating data in each node based on the difference between each item of operating data and the fitting value on the corresponding fitting curve and the discrete degree of each item of operating data; Determine the volatility of each item of operating data in each node based on the difference between each item of operating data in each node and all other items of operating data, and determine the volatility complexity of each item of operating data in each node in combination with the disturbance degree; Based on the distance from each node to all nodes of the same level and attributes, the reference node of each node is determined; based on the distance from each node to any of its reference nodes, combined with the correlation between each node and all its reference nodes on various operating data, the fluctuation difference of each operating data in each node is determined; based on the fluctuation difference and fluctuation complexity, the fluctuation sensitivity of each operating data in each node is determined; Based on the fluctuation sensitivity, the preset initial number of components and the preset scaling value, the number of adaptive components for each operating data in each node is determined, and a knowledge graph is constructed in combination with the knowledge graph tool.
2. The method for constructing an energy big data knowledge graph based on multi-source data according to claim 1, characterized in that: The expression of the disturbance degree of each operation data in each node is: i,j =σ i,j ×Δd i,j Where A i,j represents the disturbance degree of the j-th operation data in node i; σ i,j Indicates the degree of dispersion of all data in the jth running data of node i; Δd i,j Represents the mean of the differences between all data in the j-th running data in node i and the fitted values on the corresponding fitting curve.
3. The method for constructing an energy big data knowledge graph based on multi-source data according to claim 1, characterized in that: The fluctuation degree of each item of operation data in each node is the average level of the difference between each item of operation data in each node and all other items of operation data.
4. The method for constructing an energy big data knowledge graph based on multi-source data according to claim 1, characterized in that: The fluctuation complexity of each item of operating data in each node is the result of the fusion of the fluctuation degree and the disturbance degree of each item of operating data in each node.
5. The method for constructing an energy big data knowledge graph based on multi-source data according to claim 1, characterized in that: The method for determining the reference node of each node is: Arrange the distances from each node to all nodes of the same level and attributes in ascending order, and select the first preset number of nodes in the result as reference nodes for each node.
6. The method for constructing an energy big data knowledge graph based on multi-source data according to claim 1, characterized in that: The method for determining the fluctuation difference of each operation data in each node is as follows: Based on the distance between each node and any of its reference nodes, determine the distance weight of any of the reference nodes of each node; The fluctuation difference C of the j-th operation data in node i i,j The expression is: In the formula, represents the correlation between node i and its reference node k on the jth running data; represents the distance weight of node i to the reference node k; M i represents the total number of reference nodes of node i; ε represents a constant preset to be greater than 0.
7. The method for constructing an energy big data knowledge graph based on multi-source data according to claim 6, characterized in that: The method for determining the distance weight of any reference node of each node is as follows: The cumulative sum of the distances between each node and all its reference nodes is calculated, and the ratio of the distance from each node to any of its reference nodes to the cumulative sum is used as the distance weight of any reference node of each node.
8. The method for constructing an energy big data knowledge graph based on multi-source data according to claim 1, characterized in that: The fluctuation sensitivity of each item of operating data in each node is a normalized value of the product of the fluctuation complexity and the fluctuation difference of each item of operating data in each node.
9. The method for constructing an energy big data knowledge graph based on multi-source data according to claim 1, characterized in that: The expression of the number of adaptive components of each operation data in each node is: Where N i,j represents the number of adaptive components of the j-th running data in node i; represents the fluctuation sensitivity of the j-th operation data in node i; n0 represents the preset initial number of components; Indicates the preset zoom value; Represents the ceiling function.
10. The method for constructing an energy big data knowledge graph based on multi-source data according to claim 1, characterized in that: The process of constructing the knowledge graph is as follows: The running data of all items in each node are used as the input of the knowledge graph tool, wherein the Gaussian mixture model is used for data cleaning in the knowledge graph tool, and the adaptive number of components of each running data item in each node is used as the component number parameter in the Gaussian mixture model to output the knowledge graph.
Citation Information
Patent Citations
Intelligent production deployment method and system based on industrial knowledge graph
CN119005651A
Power grid equipment standard knowledge graph generation method and system
CN119150973A
Techniques for reconstructing supply chain networks using pair-wise correlation analysis
US20050096958A1