Power grid data processing system based on multi-source data fusion
By designing a power grid data processing system based on multi-source data fusion, the problem of multi-source data dispersion and integration in the power grid is solved, efficient data fusion and intelligent decision-making are achieved, and the power grid management and service level is improved.
Patent Information
- Application Number
- CN202510914621.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In modern power grids, multi-source data are scattered and in different formats, making them difficult to integrate effectively. Traditional data processing systems lack dynamic update capabilities and are unable to adapt to the new requirements of power grid development, which limits the process of intelligent upgrading.
A power grid data processing system based on multi-source data fusion is designed, which includes a data collection module, a data transmission module, a knowledge graph construction module and a data fusion module. Through real-time data collection, transmission, cleaning, entity recognition and relationship extraction, a knowledge graph is constructed to achieve efficient fusion and integration of multi-source data.
It achieves efficient fusion and integration of multi-source data, improves the value of data utilization, assists scientific decision-making, supports power grid management and optimization services, and improves the intelligence level of the power grid.
Smart Images

Figure CN120805042A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of power grid data processing based on multi-source data fusion, and particularly relates to a power grid data processing system based on multi-source data fusion. BACKGROUND
[0002] Modern power grids cover a large number of different types of data sources, such as power grid asset data in a data asset management platform, geographic information data in a South Grid Smart View platform, dispatching and operation data in a power grid management platform, and power consumption and feedback data in a customer service platform. These data are scattered in various independent platforms, have different formats and standards, and are difficult to effectively integrate, resulting in difficulty in fully mining the value of the data and in providing comprehensive support for the unified management and decision-making of the power grid. With the rapid development of power grid technology, new equipment and new technologies are emerging, and the structure and operation mode of the power grid are continuously changing. The traditional data processing system lacks dynamic updating and adaptive capabilities, and is unable to timely incorporate new knowledge and changes into the data processing process, making it difficult to flexibly adjust to meet the new requirements of the development of the power grid, thereby limiting the intelligent upgrading process of the power grid. In order to solve the above-mentioned defects, the present application provides a technical solution. SUMMARY
[0003] In order to solve the technical problems raised in the background art, the present application provides a power grid data processing system based on multi-source data fusion.
[0004] The purpose of the present application can be achieved by the following technical solution: the present application is a power grid data processing system based on multi-source data fusion, comprising a data collection module, a data transmission module, a knowledge graph construction module, a data fusion module, an application layer, and a data center. The application layer comprises a data asset management platform, a South Grid Smart View platform, a power grid management platform, and a customer service platform. The data collection module is connected to the bottom layer architecture of the data center and stores various basic power grid data in real time from multiple data sources in the application layer. The specific process is as follows: The ports of the data collection module and the asset management platform port, the South Grid Smart View platform port, the power grid management platform port, and the customer service platform port in the application layer are connected. The data collection module generates collection instructions and sends them to the platform ports in the application layer in sequence. The asset management platform port receives the collection instructions, takes the last received collection instruction as the initial time point, and connects to obtain the collection time interval T1 from the initial time point to the current time point. If there is no collection instruction, the installation time of the system is taken as the initial time point, and the power grid asset data in T1 is obtained. The power grid asset data includes equipment account, maintenance record, and life cycle information. The geographic information data in the T1 of the South Grid Smart Panorama platform is also acquired, and the geographic information data includes the geographic position of the power grid equipment, the topological relationship and the data of the power grid operation state; by analogy, the power grid management platform port receives the collection instruction, and the dispatching data and the operation and maintenance data in the T1 are acquired, the dispatching data includes the power grid dispatching instruction, the load prediction and the fault information, and the operation and maintenance data includes the equipment inspection record and the fault processing record; finally, the customer power consumption data and the customer feedback data in the T1 of the customer service platform are acquired, the power consumption data includes the power consumption, the power consumption behavior and the payment record, and the customer feedback data includes the complaint, the suggestion and the satisfaction survey information; when the data collection of the customer service platform is completed, the application layer generates a collection completion signal and sends it to the data transmission module and the data center, and the data center stores the collected power grid data according to the time partition of each type of data in the T1.
[0005] The data transmission module acquires the transmission rate and the packet loss rate in real time according to the received collection completion signal through the established data transmission monitoring mechanism, and if any index exceeds the normal range, a warning measure is taken, and the specific process is as follows: A monitoring unit and a warning adjustment unit are arranged in the data transmission module; The monitoring unit starts from the transmission time point, sets the calculation interval to 3 seconds, acquires the bit quantity zs in the transmission process every 3 seconds through the network interface card, and calculates the data transmission rate Rt every 3 seconds according to the formula The preset transmission rate range of the data center is extracted, and if the real-time data transmission rate is less than the minimum value of the transmission rate range, a low buffer signal is generated and sent to the warning adjustment unit; the total amount of data sc sent by each platform port and the total amount of data rw received by the receiving side are counted through the network switch, and the packet loss rate sw of this transmission is calculated by using the formula The preset packet loss rate range of the data center is extracted, and if the real-time packet loss rate is greater than the maximum value of the preset packet loss rate range, an overrun signal is generated and sent to the warning adjustment unit; When the warning adjustment unit receives the low buffer signal or the overrun signal, the TCP connection is re-established, and the specific process is as follows: first, each platform port sends a synchronization sequence number SYN packet to the knowledge graph, and it is necessary to point out that the packet carries the initial sequence number of each platform port, which is used to synchronize the sequence numbers of both sides; after the knowledge graph receives the SYN packet, a SYN+ACK synchronization confirmation packet is returned, which contains the initial sequence number of the knowledge graph and the confirmation information of the SYN packet of each platform port, and until each platform port sends an ACK packet to complete the establishment of the connection; the power grid data is retransmitted according to the TCP protocol and the retransmission mechanism.
[0006] The knowledge graph construction module extracts the text according to the entity recognition and relationship extraction model based on intelligent learning, and creates corresponding nodes and edges through the acquired data to build a knowledge graph library, and the specific process is as follows: The knowledge graph construction module is internally provided with an identification unit, an extraction unit and a construction unit; The identification unit cleans up the obtained power grid data through natural language processing technology, removes noise, stop words and special characters, and converts the power grid text data; the entities in the processed power grid text data are labeled, and the labeled entities include various device entities, customer information and power supply companies, the various device entities include transformers, circuit breakers and voltage regulators, the customer information includes residential and industrial user classification, electricity behavior and payment records, and the power supply formula includes power supply and equipment maintenance; the labeled power grid text data is input into the pre-trained NER model, and the NER model can identify the entity set of the input text content, and then the identified first entity text set is sent to the extraction unit as input; The extraction unit converts the first entity text set into a word vector sequence through a pre-trained word embedding model, specifically: the first entity text set is set as T={ , ,...., }, n represents the total number of texts, and it is assumed that represents any word in the text, each word is mapped to a fixed-length vector by a word embedding model, where , R represents a real set, and d is the dimension of the word vector. The word vectors are arranged in order to form a word vector sequence V= input into the CNN relationship extraction model, for example, in the text "transformer installed in substation" describing the relationship of the power grid equipment, "transformer", "installed in", and "substation" three words will be converted into corresponding word vectors respectively, and the word embedding model is Word2Vec; The relationship extraction model based on CNN includes convolution layer, pooling layer and full connection layer, specifically: the convolution kernel is set as , where h is the window size of the convolution kernel, that is, the number of currently processed words, and the convolution operation is performed on the word vector sequence V, and the formula is used to calculate to obtain the i-th element of the feature map C, where is the j-th row vector of the convolution kernel K, and b represents the bias term; thus, each feature map is obtained; the maximum pooling layer is selected and the pooling window size s is set, the maximum pooling operation is performed on the feature map C, and the formula is used to calculate to obtain the k-th element of the pooled feature vector P; Then the full connection layer integrates the feature vector output by the pooling layer, sets the weight matrix of the full connection layer as , and the bias vector as , where p is the dimension of the pooled feature vector, and m is the number of relationship categories, and the formula is used to calculate to obtain the output vector y of the full connection layer, where is an activation function softmax, each element of y represents the probability value of the text belonging to the jth relationship category; for example, for the relationship categories such as "transformer-installed in-substation", "user-electricity in-some regional power grid", etc., the full connection layer outputs the probability value of each relationship category, which is arranged in descending order, and the category corresponding to the maximum probability value is determined as the predicted relationship; The construction unit creates corresponding nodes in the knowledge graph according to the entities identified by the identification unit, takes each specific device instance as an independent node, and assigns attribute information to each node. For customer information nodes, residential users and industrial users are created as nodes, and the node attributes include user number, name, contact information, electricity address, and account opening time information. Similarly, the electricity consumption behavior data and power supply company nodes are also associated as attributes to the corresponding nodes. Then, according to the predicted relationship obtained, the corresponding entity nodes are connected by edges to construct the knowledge graph.
[0007] The data fusion module matches and associates multi-source data based on the knowledge graph, updates and feedback to adapt to the development of the power grid. The specific process is as follows: The data fusion module is provided with a matching unit and an updating unit; The matching unit collects the equipment maintenance records in the power grid management platform and the real-time operation data of the Nantong Smart Platform, and marks them as and The knowledge graph is marked as KG, where the second entity set is and the relationship set is , c represents the total number of entities, m represents the total number of relationships, and the unique identifier of the equipment is set as ID. In , the data item of the equipment maintenance record can be represented as , where represents the equipment number, represents other attributes of the equipment maintenance record, including maintenance time and maintenance content, etc. In , the data item of the real-time operation data is represented as , where represents the equipment number, represents other attributes of the real-time operation data of the equipment, including voltage, current, and temperature, etc. The matching model is: When , the and are associated to obtain the fusion data item , and the same steps are followed to obtain the data fusion of any two or more data sources in different platforms according to the matching unit.
[0008] The updating unit adds new knowledge into the knowledge graph if it is monitored in each platform, specifically: if the new knowledge involves a new device entity or a change in entity attribute, update the entity set E in the knowledge graph, set the new device entity as Then , EP represents the updated graph, if the attribute of the entity changes, update its attribute information; if the new knowledge involves a new entity relationship, update the relationship set in the knowledge graph as , set the new relationship as .
[0009] Compared with the prior art, the beneficial effects of the present application are: 1. Efficient fusion and integration of multi-source data: The data collection module can obtain different types of basic power grid data from multiple application layer platforms in real time, such as device account of asset management platform, geographic information of Nengzhikan platform, etc., and store them centrally; the data fusion module matches and associates data from different platforms with the help of knowledge graph, such as device maintenance records of power grid management platform and real-time operation data of Nengzhikan platform, breaking down the barriers of scattered data, making all kinds of data form an organic whole, providing comprehensive data support for power grid management, and improving the utilization value of data; 2. Deeply mining data value, assisting scientific decision-making: The knowledge graph construction module uses intelligent learning technology to identify entities and extract relationships from power grid data, and the constructed knowledge graph presents the complex relationships between devices, customers and power supply company entities, which helps to deeply understand the internal laws of power grid operation, and also provides valuable reference for power grid planning, device maintenance, customer service optimization, etc. For example, by analyzing device relationships and operation data, device failure can be predicted in advance, and maintenance plan can be reasonably arranged; according to customer electricity behavior and feedback data, personalized service strategy can be formulated to improve customer satisfaction. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description, the following drawings are not deliberately drawn according to the actual size, etc. Proportion, the emphasis is on showing the main idea of the present application.
[0011] Figure 1 The figure is a schematic diagram of the system structure of the present application. DETAILED DESCRIPTION
[0012] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings, obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor also belong to the scope of protection of the present application.
[0013] Referring to Figure 1 The application is a power grid data processing system based on multi-source data fusion, comprising a data collection module, a data transmission module, a knowledge graph construction module, a data fusion module, an application layer and a data center, the application layer comprising a data asset management platform, a South Grid intelligent platform, a power grid management platform and a customer service platform; The data collection module is connected to the bottom layer architecture of the data center and stores various basic power grid data in real time from multiple data sources in the application layer, and the specific process is as follows: The port of the data collection module is connected to the port of the asset management platform, the port of the South Grid intelligent platform, the port of the power grid management platform and the port of the customer service platform in the application layer, the data collection module generates collection instructions and sends them to the platform ports in sequence, the asset management platform port receives the collection instructions, takes the last time when the collection instruction was received as the initial time point, and obtains the power grid asset data within the collection time interval T1 from the current time point, if there is no collection instruction, takes the time when the system is installed as the initial time point, the power grid asset data includes equipment account, maintenance record and life cycle information; Similarly, the South Grid intelligent platform obtains geographic information data within T1, the geographic information data includes the geographic location, topological relationship and power grid operation state data of the power grid equipment; by analogy, the power grid management platform port receives the collection instruction and obtains the dispatching data and operation and maintenance data within T1, the dispatching data includes power grid dispatching instructions, load prediction and fault information, the operation and maintenance data includes equipment inspection records and fault handling records; finally, the customer service platform obtains customer power consumption data and customer feedback data within T1, the power consumption data includes power consumption, power consumption behavior and payment records, the customer feedback data includes complaints, suggestions and satisfaction survey information; when the customer service platform data collection is completed, the application layer generates a collection completion signal and sends it to the data transmission module and the data center, the data center stores the collected power grid data according to T1.
[0014] The data transmission module collects the transmission rate and packet loss rate in real time through the established data transmission monitoring mechanism according to the received collection completion signal, if any index exceeds the normal range, a warning measure is taken, and the specific process is as follows: The data transmission module is provided with a monitoring unit and a warning adjustment unit; The monitoring unit starts from the transmission time point, sets the calculation interval to 3 seconds, and obtains the number of bits zs in the transmission process every 3 seconds through the network interface card, and calculates the transmission rate according to the formula The data transmission rate Rt every 3 seconds is calculated, the preset transmission rate range of the data center is extracted, and if the real-time data transmission rate is less than the minimum value of the transmission rate range, a low buffer signal is generated and sent to the early warning adjustment unit; the total amount of data sent by each platform port sc and the total amount of data received by the receiving party rw are counted through the network switch, and the packet loss rate sw of this transmission is calculated by the formula The preset packet loss rate range of the data center is extracted, and if the real-time packet loss rate is greater than the maximum value of the preset packet loss rate range, an out-of-limit signal is generated and sent to the early warning adjustment unit; When the early warning adjustment unit receives the low buffer signal or the out-of-limit signal, the TCP connection is re-established, specifically: first, each platform port sends a synchronization sequence number SYN package to the knowledge graph, it needs to be explained that the package carries the initial sequence number of each platform port, which is used to synchronize the sequence numbers of both parties; after receiving the SYN package, the knowledge graph replies with a SYN+ACK synchronization confirmation package, which contains the initial sequence number of the knowledge graph and the confirmation information of the SYN package of each platform port, until each platform port sends an ACK package to complete the establishment of the connection; the power grid data is retransmitted according to the TCP protocol and the retransmission mechanism.
[0015] The knowledge graph construction module extracts text based on an intelligent learning-based entity recognition and relationship extraction model, and creates corresponding nodes and edges based on the obtained data to build a knowledge graph library, and the specific process is as follows: The knowledge graph construction module is provided with a recognition unit, an extraction unit and a construction unit; The recognition unit cleans, removes noise, stop words and special characters, and converts the obtained power grid data into power grid text data through natural language processing technology; the entities in the processed power grid text data are labeled, and the labeled entities include various device entities, customer information and power supply companies, various device entities include transformers, circuit breakers and voltage regulators, customer information includes residential and industrial user classification, power consumption behavior and payment records, and power supply formula includes power supply and equipment maintenance; the labeled power grid text data is input into a pre-trained NER model, which can identify the entity set of the input text content, and then the identified first entity text set is sent to the extraction unit as input; The extraction unit converts the first entity text set into a word vector sequence through a pre-trained word embedding model, specifically: the first entity text set is set as T={ , ,...., }, n represents the total number of texts, and it is assumed that represents any word in the text, each word is mapped to a fixed-length vector by a word embedding model, where R, d is the dimension of the word vector, the word vectors are arranged in order to form a word vector sequence input into the CNN relation extraction model n2 represents the number of word vectors, for example, in the text describing the relationship of power grid equipment "transformer is installed in substation", the three words "transformer", "is installed in" and "substation" are respectively converted into corresponding word vectors, and the word embedding model is Word2Vec; The CNN-based relation extraction model includes convolution layer, pooling layer and full connection layer, specifically: let the convolution kernel be where h is the window size of the convolution kernel, i.e. the number of currently processed words, the convolution operation is performed on the word vector sequence V, and the formula is used to calculate to obtain the i-th element of the feature map C, where is the j-th row vector of the convolution kernel K, and b represents the bias term; thus, each feature map is obtained; the maximum pooling layer is selected and the pooling window size s is set to be s, the maximum pooling operation is performed on the feature map C, and the formula is used to calculate to obtain the k-th element of the pooled feature vector P; Then the full connection layer integrates the feature vector output by the pooling layer, the weight matrix of the full connection layer is set to and the bias vector is where p is the dimension of the pooled feature vector, and m is the number of relation categories, the formula is used to calculate to obtain the output vector y of the full connection layer, where is the activation function softmax, and each element of y represents the probability value of the text belonging to the j-th relation category; for example, for the relation categories "transformer-is installed in-substation" and "user-electricity in-some regional power grid", the full connection layer outputs the probability value of each relation category, which is arranged in descending order, and the category corresponding to the maximum probability value is determined as the predicted relation; The construction unit creates corresponding nodes in the knowledge graph according to the entities identified by the recognition unit, takes each specific device instance as an independent node, and assigns attribute information to each node. For customer information nodes, residential users and industrial users are created as nodes, and the node attributes include user number, name, contact information, electricity address and account opening time information. Similarly, the electricity consumption behavior data and power supply company nodes are also associated as attributes to the corresponding nodes. Then, according to the predicted relationship obtained, the corresponding entity nodes are connected by edges, for example, when the model determines that the "transformer-is installed in-substation" relationship is established, a directed edge is created between the transformer node and the corresponding substation node, and the label of the edge is clearly "installed in", so as to construct the knowledge graph.
[0016] The data fusion module matches and correlates multi-source data according to the knowledge graph, updates and feeds back to adapt to the development of the power grid, and the specific process is as follows: The data fusion module is provided with a matching unit and an updating unit; The matching unit acquires any two data sources in different platforms, such as the equipment maintenance record in the power grid management platform and the real-time operation data of the Nangnet intelligent platform, marked as and , and the knowledge graph is marked as KG, wherein the second entity set is , the relationship set is , c represents the total number of entities, m represents the total number of relationships, the unique identifier of the equipment is set as ID, and the data item of the equipment maintenance record in can be expressed as , wherein represents the equipment number, represents other attributes of the equipment maintenance record, and the other attributes include maintenance time and maintenance content, etc., and the data item of the real-time operation in is expressed as , wherein represents the equipment number, represents other attributes of the equipment real-time operation data, including voltage, current and temperature, etc. The matching model is: When , the and are associated to obtain the fusion data item It should be noted that any two or more data sources in any platform can be substituted and matched according to this step; The updating unit adds new knowledge generated in each platform to the knowledge graph if it is monitored, specifically: if the new knowledge involves new equipment entities or changes in entity attributes, update the entity set E in the knowledge graph, set the new equipment entity as , then , EP represents the updated graph, and if the attributes of the entity change, update the attribute information thereof; if the new knowledge involves new entity relationships, update the relationship set in the knowledge graph as , set the new relationship as .
[0017] The foregoing is illustrative of the present application, and is not to be construed as limiting thereof. While a number of exemplary embodiments of the application have been described, those skilled in the art will readily appreciate that many modifications are possible in the exemplary embodiments without materially departing from the novel teachings and advantages of the application. Accordingly, all such modifications are intended to be included within the scope of the present application as defined in the claims. It is to be understood that the above description is intended to be illustrative, and not restrictive. Many other embodiments will be apparent to those of skill in the art upon reviewing the above description. The scope of the application should, therefore, be determined not with reference to the above description, but should instead be determined with reference to the appended claims, along with their full scope of equivalents.
Claims
1. , including a knowledge graph construction module and a data fusion module; its characteristics are: The knowledge graph construction module is provided with an identification unit, an extraction unit, and a construction unit. The text is extracted according to the entity recognition and relationship extraction model based on intelligent learning, and then the corresponding nodes and edges are created based on the acquired data to construct the knowledge graph library. The data fusion module is provided with a matching unit and an updating unit, which matches and associates multiple source data through the knowledge graph and performs fusion processing on multiple data. Specifically, the matching unit collects the equipment maintenance records in the power grid management platform and the real-time operation data of the Southern Power Grid Zhikan platform, and marks them as and , mark the knowledge graph as KG, where the second entity set is The relationship set is , c represents the total number of entities, m represents the total number of relationships, and the unique identifier of the device is set as ID. The data items of the equipment maintenance record are expressed as ,in Indicates the device number, Indicates other attributes of the equipment maintenance record, including maintenance time and maintenance content. The data items running in real time are represented as ,in Indicates the device number, Other attributes representing the real-time operating data of the device, including voltage, current, and temperature; Match the data, the matching model is: ,when When and Perform association to obtain fused data items ; And so on, obtain any two or more data sources in different platforms and perform data fusion according to the steps of the matching unit.
2. A power grid data processing system based on multi-source data fusion according to claim 1, characterized in that: The updating unit updates and feeds back the graph, and the specific process is as follows: If new knowledge is generated in each platform, it will be added to the knowledge graph. Specifically, if the new knowledge involves a new device entity or a change in entity attributes, the entity set E in the knowledge graph will be updated and the new device entity will be set as ,but , EP represents the updated graph. If the attributes of the entity change, its attribute information is updated; if the new knowledge involves new entity relationships, the relationship set in the updated knowledge graph is , set the new relationship to be, then .
3. The power grid data processing system based on multi-source data fusion according to claim 1, characterized in that: The system also includes a data transmission module, which is equipped with a monitoring unit and an early warning adjustment unit. Upon receiving a collection completion signal, the system collects the transmission rate and packet loss rate in real time through an established data transmission monitoring mechanism. If any indicator exceeds the normal range, early warning measures are taken. The specific process is as follows: The monitoring unit starts at the transmission time point, sets the calculation interval to 3 seconds, and obtains the number of bits zs during the transmission process every 3 seconds through the network interface card. According to the formula Calculate the data transmission rate Rt every 3 seconds, extract the preset transmission rate range of the data center, and generate a low-speed signal if the real-time data transmission rate is less than the minimum value of the transmission rate range and send it to the early warning adjustment unit; The total amount of data sc sent by each platform port and the total amount of data rw received by the receiver are counted through the network switch, and the formula is used to calculate The packet loss rate sw of the transmission is obtained, and the preset packet loss rate range of the data center is extracted. If the real-time packet loss rate is greater than the maximum value of the preset packet loss rate range, an over-limit signal is generated and sent to the early warning adjustment unit.
4. A power grid data processing system based on multi-source data fusion according to claim 3, characterized in that: When the early warning adjustment unit receives a low-speed signal or an over-limit signal, the TCP connection is re-established as follows: First, each platform port sends a synchronization sequence number SYN packet to the knowledge graph. The SYN packet carries the initial sequence number of each platform port and is used to synchronize the sequence numbers of both parties. After receiving the SYN packet, the knowledge graph replies with a SYN+ACK synchronization confirmation packet, which contains the initial sequence number of the knowledge graph and confirmation information for the SYN packet of each platform port, until each platform port sends an ACK packet again to complete the connection establishment. The power grid data is retransmitted according to the TCP protocol and retransmission mechanism.
5. The power grid data processing system based on multi-source data fusion according to claim 1, characterized in that: The knowledge graph construction module extracts text based on deep learning-based entity recognition, and then creates corresponding nodes and edges from the acquired data to build a knowledge graph. The specific process is as follows: The recognition unit uses natural language processing technology to clean the acquired power grid data, remove noise, stop words, and special characters, and convert it into power grid text data. It then labels entities in the processed power grid text data. The labeled entities include various equipment entities, customer information, and power supply companies. Equipment entities include transformers, circuit breakers, and transformers. Customer information includes residential and industrial user classifications, electricity usage behavior, and payment records. Power supply formulas include power supply and equipment maintenance. The annotated power grid text data is input into the pre-trained NER model. The NER model can identify the entity set of the input text content, and then send the identified first entity text set as input to the extraction unit.
6. A power grid data processing system based on multi-source data fusion according to claim 5, characterized in that: The extraction unit converts the first entity text set into a word vector sequence through a pre-trained word embedding model, and then outputs a predicted probability value according to the relationship extraction model, specifically: Set the first entity text set to T={ , , ...., }, n represents the total number of texts, set Represents any word in the text, and maps each word into a vector of fixed length through the word embedding model ,in , R represents a real number set, d is the word vector dimension, and each word vector is arranged in order to form a word vector sequence for the input CNN relation extraction model , n2 represents the number of word vectors, and the word embedding model is Word2Vec; The CNN-based relationship extraction model includes convolutional layers, pooling layers, and fully connected layers. Specifically, let the convolution kernel be , where h is the window size of the convolution kernel, that is, the number of words currently processed, and the convolution operation is performed on the word vector sequence V, and the formula is calculated Get the i-th element of the feature map C, where is the j-th row vector of the convolution kernel K, and b is represented as the bias term; each feature map is obtained from this; the maximum pooling layer is selected and the pooling window size is set to s, and the maximum pooling operation is performed on the feature map C, and the calculation is performed according to the formula Get the kth element of the pooled feature vector P; Then the fully connected layer integrates the feature vectors output by the pooling layer and sets the weight matrix of the fully connected layer to , the bias vector is , where p is the dimension of the feature vector after pooling, and m is the number of relationship categories, calculated according to the formula Get the output vector y of the fully connected layer, where is the activation function softmax, each element of y It represents the probability value of the text belonging to the jth relationship category. All probability values are arranged from large to small, and the category corresponding to the maximum probability value is determined as the predicted relationship.
7. The power grid data processing system based on multi-source data fusion according to claim 5, characterized in that: The construction unit creates corresponding nodes in the knowledge graph based on the entities identified by the identification unit, and treats each specific device instance as an independent node, specifically: Assign attribute information to each node. For customer information nodes, create nodes for residential users and industrial users respectively. The node attributes include user number, name, contact information, electricity address and account opening time information. Similarly, the electricity consumption behavior data and power supply company nodes are also associated with the corresponding nodes as attributes. Then, according to the obtained prediction relationship, establish edges with the corresponding entity nodes to construct a knowledge graph.
8. The power grid data processing system based on multi-source data fusion according to claim 1, characterized in that: It also includes a data collection module, an application layer, and a data center. The application layer includes a data asset management platform, a Southern Power Grid Zhikan platform, a power grid management platform, and a customer service platform. The data collection module, based on docking with the underlying architecture of the data center, obtains various basic power grid data from multiple data sources in the application layer in real time for storage. The specific process is as follows: The port of the data collection module establishes connections with the asset management platform port, the Southern Power Grid Zhikan platform port, the power grid management platform port, and the customer service platform port in the application layer. The data collection module generates a collection instruction and sends it to each platform port in sequence. The asset management platform port receives the collection instruction and uses the last time the collection instruction was received as the initial time point. The connection is made to the current time point to obtain the collection time interval T1. If there is no collection instruction, the initial time point is the time when the system is installed, and the power grid asset data within T1 is obtained. The power grid asset data includes equipment ledgers, maintenance records, and life cycle information. Also obtain geographic information data from the Southern Power Grid Zhikan Platform T1, which includes the geographic location, topological relationships, and grid operation status of power grid equipment; By analogy, the power grid management platform port receives the collection instruction and obtains the dispatching data and operation and maintenance data in T1. The dispatching data includes power grid dispatching instructions, load forecasts and fault information, and the operation and maintenance data includes equipment inspection records and fault handling records; finally, it obtains the customer electricity consumption data and customer feedback data in the customer service platform T1. The electricity consumption data includes electricity consumption, electricity consumption behavior and payment records, and the customer feedback data includes complaints, suggestions and satisfaction survey information; when the customer service platform data collection is completed, the application layer generates a collection completion signal and sends it to the data transmission module and the data center. The data center stores the collected power grid data in time partitions according to T1.