Knowledge graph construction method and system based on distributed energy heterogeneous data

By preprocessing and extracting the distributed photovoltaic access data, the knowledge graph in the field of distributed photovoltaics is constructed, and the problems of multi-source heterogeneous data integration and knowledge management are solved, and efficient knowledge organization and dynamic updates are achieved.

CN120012885APending Publication Date: 2025-05-16ELECTRIC POWER RES INST OF GUANGXI POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411814850.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

During distributed photovoltaic access, multi-source heterogeneous data is difficult to efficiently integrate and process, and lack of systematic knowledge classification and management methods, which makes it difficult to effectively organize and utilize knowledge, and difficult to automatically identify and extract entity relationships. The existing technology lacks effective mechanisms for dynamic updates and maintenance of knowledge graphs.

Method used

By collecting the data after distributed photovoltaic access for preprocessing, the relationship between core knowledge entities and entities in the photovoltaic field is defined, the entity recognition is used using keyword matching and support vector machine algorithms, the entity relationship is extracted using relation extraction templates and neural network classification models, and the knowledge graph is constructed in the graph database, and it is regularly evaluated and updated.

Benefits of technology

It improves the data integration efficiency during distributed photovoltaic access, improves the organization of knowledge management, improves the accuracy of entity recognition and relationship extraction, and ensures dynamic updates and long-term stability of the knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012885A_ABST
    Figure CN120012885A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph construction method and system based on distributed energy heterogeneous data, and relates to the technical field of knowledge graph construction, and the method comprises the steps: collecting distributed energy data after distributed photovoltaic access; preprocessing the collected distributed energy data; defining a core knowledge entity of the photovoltaic field according to the knowledge classification of the distributed photovoltaic access field; defining a relationship between the core knowledge entities; extracting the core knowledge entity and the relationship; and based on the extracted core knowledge entities and relationships, constructing a knowledge graph in the distributed photovoltaic field. According to the knowledge graph construction method based on the distributed energy heterogeneous data, the multi-source heterogeneous data is preprocessed, the data integration efficiency is improved, photovoltaic field core knowledge entities and relationships are systematically defined, the entity recognition accuracy is improved in combination with an SVM algorithm, the knowledge graph is constructed by using the graph database, efficient query and maintenance are achieved, and the knowledge graph construction efficiency is improved. And the real-time performance and the consistency of the data are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graph construction, and specifically to a knowledge graph construction method and system based on distributed energy heterogeneous data. Background Art

[0002] In the distributed energy access, the distributed photovoltaic access has large differences in the implementation of standards in terms of electrical connection, the standard requirements are not implemented in some equipment communication links, and there is a lack of security authentication in equipment access. On this basis, it is necessary to set up a photovoltaic field knowledge graph for distributed energy (photovoltaic) access. Therefore, how to construct a knowledge graph based on multi-source heterogeneous data captured by distributed energy has become a technical difficulty. The present invention is based on multi-source heterogeneous data and is constructed on the basis of the knowledge structure based on the field of distributed photovoltaic access to realize the knowledge graph construction process of distributed photovoltaic energy. Summary of the invention

[0003] In view of the above-mentioned problems, the present invention is proposed.

[0004] Therefore, the technical problems solved by the present invention are: first, multi-source heterogeneous data are difficult to integrate and process efficiently, and different types of data (including structured, semi-structured and unstructured data) differ greatly in data format and standard execution, resulting in the inability to unify analysis and management; second, the lack of systematic knowledge classification and management methods makes it difficult to effectively organize and utilize knowledge in the field of distributed photovoltaics, and the scalability and operability of knowledge are poor; in addition, in the process of distributed photovoltaic access, the complex relationships between entities are difficult to automatically identify and extract, and the existing methods are inefficient and inaccurate when processing entity relationships in heterogeneous data; finally, the existing technology lacks an effective mechanism to dynamically update and maintain the data in the knowledge graph, making it difficult to ensure the continuous accuracy and consistency of knowledge.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: a knowledge graph construction method based on distributed energy heterogeneous data, comprising:

[0006] Collect distributed energy data after distributed photovoltaic access;

[0007] Pre-process the collected distributed energy data;

[0008] Based on the knowledge classification of distributed photovoltaic access, define the core knowledge entities in the photovoltaic field;

[0009] defining relationships between the core knowledge entities;

[0010] Extracting the core knowledge entities and relationships;

[0011] Based on the extracted core knowledge entities and relationships, a knowledge graph in the field of distributed photovoltaics is constructed.

[0012] As a preferred solution of the knowledge graph construction method based on distributed energy heterogeneous data described in the present invention, the collection of distributed energy data after distributed photovoltaic access includes collecting real-time operation data, meteorological data, maintenance data and energy transaction data of photovoltaic power stations through photovoltaic inverters and meteorological monitoring stations, and recording the data at predetermined time intervals.

[0013] As a preferred solution of the method for constructing a knowledge graph based on distributed energy heterogeneous data described in the present invention, the preprocessing includes cleaning the collected data to remove incomplete or erroneous data records;

[0014] Convert and merge structured, semi-structured and unstructured data from different data sources;

[0015] Label the cleaned and fused data.

[0016] As a preferred solution of the knowledge graph construction method based on distributed energy heterogeneous data described in the present invention, wherein: the core knowledge entity in the photovoltaic field is defined according to the knowledge classification of the distributed photovoltaic access field, and the core knowledge entity includes basic entities, equipment entities and data entities, wherein the basic entity includes scope, normative reference documents and technical location, the equipment entity includes reactive capacity, voltage regulation, start and stop, operation applicability and safety and protection, and the data entity includes power quality, general technical requirements, power metering, communication and signal and system detection;

[0017] Defining relationships between the core knowledge entities, including connection relationships between basic entities and device entities and data entities, measurement relationships between device entities and data entities, and generation relationships between device entities;

[0018] For each core knowledge entity, its related attributes are defined, including voltage range and frequency range attributes of operation applicability, and basic requirement attributes of power quality.

[0019] As a preferred solution of the knowledge graph construction method based on distributed energy heterogeneous data described in the present invention, wherein: the definition of the relationship between the core knowledge entities includes identifying the basic entity through a keyword matching rule, and the keyword matching rule is used to detect specific keywords in the data. If a predefined keyword appears in the data, the corresponding content is identified as a basic entity;

[0020] Using a support vector machine algorithm to identify device entities and data entities in text data, the process includes labeling the data and training an SVM model based on the labeled data, the model predicting the types of photovoltaic device entities and data entities in the text based on vocabulary, part of speech and context information in the text;

[0021] Extracting the relationship between entities from the preprocessed data using a relationship extraction template, wherein the relationship extraction identifies the relationship between power quality and voltage deviation data by analyzing specific timestamp information in the data;

[0022] The entity pairs in the labeled data are trained through a relationship classification model of a neural network. The model classifies the relationship between entity pairs according to the attributes, context information and features of the entities, and determines the specific relationship type between the entities.

[0023] As a preferred solution of the method for constructing a knowledge graph based on distributed energy heterogeneous data described in the present invention, wherein: the construction of a knowledge graph in the field of distributed photovoltaics includes constructing a knowledge graph in a graph database, storing entities as nodes based on the results of entity recognition and relationship extraction, and storing the relationships between entities as edges;

[0024] Set corresponding attributes for each node, connect the nodes together according to the relationship between them, and form a complete knowledge graph structure.

[0025] As a preferred solution of the method for constructing a knowledge graph based on distributed energy heterogeneous data described in the present invention, the knowledge graph is evaluated regularly, and the evaluation includes comparing the attributes and connection relationships of entity nodes, checking and correcting nodes or relationships that do not meet preset conditions, so as to ensure that the data in the knowledge graph remains consistent.

[0026] Another object of the present invention is to provide a knowledge graph construction system based on distributed energy heterogeneous data, which can solve the problems of inconsistent data integration, unsystematic knowledge classification, and low accuracy of relationship extraction in existing distributed photovoltaic access data processing by constructing a photovoltaic field knowledge management system based on the knowledge graph.

[0027] To solve the above technical problems, the present invention provides the following technical solutions: a knowledge graph construction system based on distributed energy heterogeneous data, comprising: a data acquisition module, a data processing module, an entity definition module, a relationship definition module, a relationship extraction module and a graph construction module; the data acquisition module is used to collect distributed energy data after distributed photovoltaic access; the data processing module is used to pre-process the collected distributed energy data; the entity definition module is used to define the core knowledge entities in the photovoltaic field according to the knowledge classification of the distributed photovoltaic access field; the relationship definition module is used to define the relationship between the core knowledge entities; the relationship extraction module is used to extract the core knowledge entities and relationships; the graph construction module is used to construct a knowledge graph in the distributed photovoltaic field based on the extracted core knowledge entities and relationships.

[0028] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method for constructing a knowledge graph based on heterogeneous data of distributed energy are implemented as described above.

[0029] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for constructing a knowledge graph based on heterogeneous data of distributed energy resources as described above.

[0030] Beneficial effects of the present invention: The knowledge graph construction method based on distributed energy heterogeneous data provided by the present invention effectively improves the data integration efficiency of different data sources in the distributed photovoltaic access process by introducing a preprocessing method for multi-source heterogeneous data, and reduces the processing difficulties caused by inconsistent data formats. By systematically defining the core knowledge entities in the photovoltaic field, the orderliness of knowledge management in the distributed photovoltaic access process is improved, so that the knowledge in complex data can be clearly organized and quickly accessed. The present invention combines keyword matching and support vector machine (SVM) algorithms in the entity recognition process, so that the accuracy of entity recognition can be improved when processing complex text data, especially when facing unstructured data. In terms of relationship extraction, the use of relationship extraction templates and neural network classification models has better solved the problems of low efficiency and insufficient accuracy in extracting complex relationships between entities. Through the knowledge graph construction method based on the graph database, the present invention ensures the visual storage and efficient query of entity and relationship data. In addition, the knowledge graph is regularly evaluated and updated to reduce the problem of outdated or inconsistent knowledge and ensure the long-term stability of the system and the real-time nature of the data. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0032] Figure 1 An overall flow chart of a method for constructing a knowledge graph based on distributed energy heterogeneous data provided by an embodiment of the present invention.

[0033] Figure 2 A schematic diagram of knowledge modeling of a method for constructing a knowledge graph based on distributed energy heterogeneous data provided by an embodiment of the present invention.

[0034] Figure 3 An overall structural diagram of a knowledge graph construction system based on distributed energy heterogeneous data provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0035] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0036] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0037] Example 1

[0038] Reference Figure 1-Figure 2 , which is an embodiment of the present invention, provides a knowledge graph construction method based on distributed energy heterogeneous data, including:

[0039] Step 1: Collect distributed energy-related data after distributed photovoltaic access, namely real-time operation data of photovoltaic power stations, meteorological data, maintenance and operation data of photovoltaic power stations, and energy trading data, etc. These data can be directly read from photovoltaic inverters, meteorological monitoring stations and other equipment, and are usually recorded at certain time intervals.

[0040] Step 2: Since the data obtained is heterogeneous, data preprocessing is required, which includes data cleaning, data fusion and data labeling. Labeling is to label photovoltaic equipment and data to provide a data basis for constructing entities and relationships in the knowledge graph.

[0041] Step 3: 1. Determine the type of knowledge entity, and define the core entities in the photovoltaic field according to the knowledge classification of distributed photovoltaic access, including basic entities such as scope, normative application documents, technical location, equipment entities such as reactive capacity and voltage regulation, start and stop, operation suitability, safety and protection, and data entities such as power quality, general technical requirements, power metering, communication and signal, and system detection.

[0042] 2. Define relationship types, such as establishing connection relationships between basic entities and equipment entities and data entities, defining measurement relationships between equipment entities and data entities, and establishing generation relationships between equipment entities.

[0043] 3. Define attributes. Define relevant attributes for each entity. For example, the attributes of operational suitability can be reported as voltage range and frequency range, and the attributes of power quality can be reported as basic requirements (as shown in the diagram above).

[0044] Step 4: 1. Entity recognition: Use keyword matching rules to identify basic entities. If the data contains keywords such as "range" and "file", it can be identified as the corresponding basic entity;

[0045] 2. Use the support vector machine algorithm to identify device entities and data entities in text data. First, annotate the data to mark the entity type and data category, and use the annotated data to train the SVM model. For example, use the SVM algorithm to train a photovoltaic device entity recognition model. The model can predict whether the text contains a photovoltaic device entity and its type based on the features in the text (such as vocabulary, part of speech, context, etc.);

[0046] 3. Relationship extraction: define some relationship extraction templates to extract the relationship between entities based on the structure and characteristics of the data. For example, for power quality and voltage deviation data, if the timestamp of the power quality data is the same as the acquisition time of the voltage deviation data, it can be considered that there is a "produce" relationship between them. In addition, the labeled data can be trained according to the relationship classification model of the neural network to learn the relationship pattern between entities. First, label the entity pairs in the data and the relationship type between them, and then use the characteristics of the entity pairs (entity attributes, contextual information, etc.) as input and the relationship type as output to train the relationship classification model. For example, use a multi-layer perceptron to build a relationship classification model, input the relevant attributes of power quality and voltage deviation, and their contextual information in the text, and output the relationship between them (produce or no connection, etc.).

[0047] Step 5: Use the graph database to build a knowledge graph. Based on the results of entity recognition and relationship extraction, use entities as nodes and relationships as edges to add data to the graph database. For example, for the power quality node, add attributes to it as a basic requirement, and then connect it to the power metering node based on the connection relationship.

[0048] Step 6: Regularly evaluate the quality of the knowledge graph to check the accuracy, completeness, and consistency of entities and relationships.

[0049] It should be noted that the knowledge graph consists of four parts: concepts, entities, attributes and relationships. It is divided into two levels: concept layer and data layer. The concept layer is mainly used to organize the structural hierarchy of knowledge, while the data layer is the detailed specific knowledge.

[0050] 1. Knowledge structure model: A high-quality knowledge structure model can improve the efficiency of knowledge graph construction, avoid many unnecessary repetitive tasks, and effectively reduce the cost of domain knowledge graph construction and maintenance.

[0051] (1) Since the present invention is aimed at a specific energy field with a relatively small scope, the knowledge modeling adopts a top-down design method. Taking distributed photovoltaic access in distributed energy as an example, the domain concept modeling is carried out based on the knowledge classification of distributed photovoltaic access, such as Figure 2 shown.

[0052] (2) Relationship is one of the main components of knowledge graph. It is also a bridge to describe the connection between concepts, between concepts and entities, and between entities. It is indispensable for the formation of semantic network. For the above-mentioned distributed photovoltaic access domain concept modeling, the classification relationship and non-classification relationship between domain knowledge are established. The classification relationship is a hierarchical relationship. For example, the relationship between scope and operational applicability is a non-classification relationship, and the relationship between operational adaptability and voltage range is a classification relationship. By analogy, the entities, relationships and attributes between all knowledge domain concepts are established, and then the domain knowledge structure diagram is constructed based on the entity concepts, relationship concepts and attributes.

[0053] 2. Entity extraction: According to the scope and degree of certainty of distributed photovoltaics, there are heterogeneous data in the process of acquiring data, namely structured data, unstructured data and semi-structured data. The relevant entity concepts and heterogeneous data can be divided into three categories. The first type of entity concept is a concept with few instances and certainty, such as scope and normative reference documents. Such entities are directly determined according to the existing classification system or professional experience, or they can be directly converted and determined by structured data. The second type of entity concept is a concept with many instances but certainty, such as operational applicability. There are many concepts of such entities, so there is no need to list them one by one. A small part of them can be extracted through the existing classification system or using rules, and then the data set can be trained using machine learning methods to identify new entities. The third type of entity concept has a wide range and uncertainty, such as power quality. This entity has strong personalized characteristics and changes over time. It is necessary to combine the specific corpus and entity description characteristics, and adopt different methods for entity recognition according to actual conditions. If it is structured data, the direct conversion method can be used to directly extract relevant information, and the rule extraction method and machine learning method can be used to extract information from unstructured text; if the unstructured text has obvious characteristics, relevant information can be extracted by manually constructing rules; if the rules in the unstructured text are inconvenient to construct, the machine learning text features can be used to build a model to achieve entity discrimination and extraction.

[0054] 3. Relationship extraction: The purpose of relationship extraction is to identify the target relationship of entities in the text, that is, to obtain entity relationship facts from unstructured text and add them to the knowledge graph, which is one of the important links in building the knowledge graph. Its main tasks are relationship classification and open relationship extraction: relationship classification refers to relying on knowledge, experience and related theories to pre-given relationships in related fields, classifying and matching entities, and obtaining entity relationships; open relationship extraction is to extract relationships from structured texts and then map them to pre-set relationships. Relationship extraction methods are mainly divided into rule-based extraction and machine learning-based extraction. Rule-based extraction requires a large number of rules to be manually formulated, which consumes a lot of manpower and time, and is difficult to fully describe. Machine learning-based extraction includes supervised relationship extraction and remote supervised relationship extraction. Supervised relationship extraction refers to labeling data in large-scale data and further training the model. This is currently the best method. This type of relationship extraction is essentially a multi-classification task, but it takes a lot of time and cost to complete data labeling; remote supervised relationship extraction can effectively solve the shortcomings of insufficient supervised data, but it will produce a large number of labeling errors. Each method has its own advantages and disadvantages, so this paper adopts a method of extracting in multiple ways to improve the comprehensiveness and accuracy of relationship extraction.

[0055] Furthermore, relation extraction mainly improves relation extraction based on description rules:

[0056] (1) Define relationship types: They are divided into entity relationships between the same entity types and entity relationships between different entity types. Entity relationships between the same entity types mainly include class relationships, inclusion relationships, and equivalence relationships; entity relationships between different entity types mainly include subordinate relationships, inclusion relationships, reference relationships, adoption relationships, and participation relationships.

[0057] (2) Constructing a relationship type dictionary: Based on the entities around the located feature words, entities that meet the relationship type can be screened out, and the feature words can be summarized by constructing a type dictionary.

[0058] (3) Data cleaning and relationship extraction: Use feature words to filter data, use feature words as trigger words, and filter out sentences that do not contain trigger words. Use the regular expression "text = re.sub(r"\s+","",text)" to merge redundant spaces, and use Python to remove stop words according to the stop word list. Use the command line based on the trigger word to judge the extracted relationship. If it does not meet the preset relationship, it will be corrected. The specific process is as follows: First, input the data set that needs to be extracted, and use ";,?,!,." as delimiters to split the sentences in the input data set to obtain sentences. Find the trigger word of the sentence, retain the entity pair with the closest distance, remove other content, and judge the relationship type according to the relationship type feature word to which the trigger word belongs. Judge the extracted relationship type. If the relationship type is correct, the relationship extraction task of the sentence ends, and go to the next sentence, and repeat the above steps; if the relationship type is wrong, re-extract the relationship until it is correct, then the relationship extraction of the sentence ends, and go to the next sentence. By traversing the database, the first column represents the serial number, the second column represents the head entity, the third column represents the file location of the head entity, the fourth column represents the tail entity, the fifth column represents the file location of the tail entity, and the sixth column is the relationship between the head entity and the tail entity.

[0059] 4. Knowledge fusion: The main tasks of knowledge fusion are entity alignment and entity disambiguation. Entity alignment is the main task in the knowledge fusion stage. Its purpose is to determine whether the entities in different data sources represent the same object in the real world. If so, it is necessary to establish an alignment relationship between these entities to ensure that the understanding of the same thing remains consistent, and to complement and unify the information contained in the entities to avoid repeated appearances of entities in the knowledge graph or incomplete entity information. Entity disambiguation is to eliminate the phenomenon of polysemy based on contextual information and solve problems such as unclear pointing in heterogeneous data.

[0060] 5. Build a graph: Based on the above acquired data, perform entity extraction, relationship extraction and entity fusion to obtain a large amount of knowledge information in the form of triples. The triple information is stored, and the storage methods include storage methods based on table structure, storage methods based on RDF structure and storage methods based on graph structure. Neo4j is used here to store the knowledge graph in the field of distributed photovoltaics. Neo4j is a directed graph composed of nodes and edges, and nodes are composed of nodes, relationships and attributes. It is represented by circles, and relationships are represented by edges. Relationships connect nodes. Each node and relationship can be set with one or more attributes. Attributes are key-value pairs, expressed as "name: value". Based on the underlying use of Neo4j to store nodes and relationships in a graph manner, the connection between different nodes can be efficiently queried.

[0061] The present invention analyzes the knowledge classification in the field of distributed photovoltaic energy, and establishes a knowledge structure model in the field of distributed photovoltaics based on the analysis of distributed photovoltaic knowledge needs and knowledge sources.

[0062] According to the different data sources and the degree of structuring of heterogeneous data, a classification system is set up in the entity extraction process to make the efficiency and accuracy of entity extraction higher.

[0063] Improvements are made to the description rule-based relationship extraction, mainly by setting detailed rules for relationship extraction to make the relationship extraction more accurate.

[0064] Example 2

[0065] Reference Figure 3 , which is an embodiment of the present invention, provides a knowledge graph construction system based on distributed energy heterogeneous data, including:

[0066] Data collection module 100, data processing module 200, entity definition module 300, relationship definition module 400, relationship extraction module 500 and graph construction module 600;

[0067] The data acquisition module 100 is used to collect distributed energy data after distributed photovoltaic access;

[0068] The data processing module 200 is used to pre-process the collected distributed energy data;

[0069] The entity definition module 300 is used to define core knowledge entities in the photovoltaic field according to the knowledge classification in the field of distributed photovoltaic access;

[0070] The relationship definition module 400 is used to define the relationship between the core knowledge entities;

[0071] The relationship extraction module 500 is used to extract the core knowledge entities and relationships;

[0072] The graph construction module 600 is used to construct a knowledge graph in the field of distributed photovoltaics based on the extracted core knowledge entities and relationships.

[0073] Example 3

[0074] An embodiment of the present invention is different from the first two embodiments in that:

[0075] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.

[0076] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0077] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0078] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0079] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A knowledge graph construction method based on distributed energy heterogeneous data, characterized in that: include: Collect distributed energy data after distributed photovoltaic access; Pre-process the collected distributed energy data; Based on the knowledge classification of distributed photovoltaic access, define the core knowledge entities in the photovoltaic field; defining relationships between the core knowledge entities; Extracting the core knowledge entities and relationships; Based on the extracted core knowledge entities and relationships, a knowledge graph in the field of distributed photovoltaics is constructed.

2. The method for constructing a knowledge graph based on distributed energy heterogeneous data according to claim 1, characterized in that: The collection of distributed energy data after distributed photovoltaic access includes collecting real-time operation data, meteorological data, maintenance data and energy transaction data of the photovoltaic power station through photovoltaic inverters and meteorological monitoring stations, and recording the data at predetermined time intervals.

3. The method for constructing a knowledge graph based on distributed energy heterogeneous data according to claim 2, characterized in that: The preprocessing includes cleaning the collected data to remove incomplete or erroneous data records; Convert and merge structured, semi-structured and unstructured data from different data sources; Label the cleaned and fused data.

4. The method for constructing a knowledge graph based on distributed energy heterogeneous data according to claim 3, characterized in that: The definition of core knowledge entities in the photovoltaic field includes defining core knowledge entities in the photovoltaic field according to the knowledge classification of the distributed photovoltaic access field, and the core knowledge entities include basic entities, equipment entities and data entities, wherein the basic entities include scope, normative reference documents and technical positions, the equipment entities include reactive capacity, voltage regulation, start and stop, operation applicability and safety and protection, and the data entities include power quality, general technical requirements, power metering, communication and signal and system detection; Defining relationships between the core knowledge entities, including connection relationships between basic entities and device entities and data entities, measurement relationships between device entities and data entities, and generation relationships between device entities; For each core knowledge entity, its related attributes are defined, including voltage range and frequency range attributes of operation applicability, and basic requirement attributes of power quality.

5. The method for constructing a knowledge graph based on distributed energy heterogeneous data according to claim 4, characterized in that: Defining the relationship between the core knowledge entities includes identifying basic entities through keyword matching rules, wherein the keyword matching rules are used to detect specific keywords in the data, and if predefined keywords appear in the data, the corresponding content is identified as a basic entity; Using a support vector machine algorithm to identify device entities and data entities in text data, the process includes labeling the data and training an SVM model based on the labeled data, the model predicting the types of photovoltaic device entities and data entities in the text based on vocabulary, part of speech and context information in the text; Extracting the relationship between entities from the preprocessed data using a relationship extraction template, wherein the relationship extraction identifies the relationship between power quality and voltage deviation data by analyzing specific timestamp information in the data; The entity pairs in the labeled data are trained through a relationship classification model of a neural network. The model classifies the relationship between entity pairs according to the attributes, context information and features of the entities, and determines the specific relationship type between the entities.

6. The method for constructing a knowledge graph based on distributed energy heterogeneous data according to claim 5, characterized in that: The construction of the knowledge graph in the field of distributed photovoltaics includes constructing a knowledge graph in a graph database, storing entities as nodes based on the results of entity recognition and relationship extraction, and storing the relationships between entities as edges; Set corresponding attributes for each node, connect the nodes together according to the relationship between them, and form a complete knowledge graph structure.

7. The method for constructing a knowledge graph based on distributed energy heterogeneous data according to claim 6, characterized in that: The knowledge graph is evaluated regularly, and the evaluation includes comparing the attributes and connection relationships of entity nodes, checking and correcting nodes or relationships that do not meet preset conditions, and ensuring that the data in the knowledge graph remains consistent.

8. A system using the method for constructing a knowledge graph based on distributed energy heterogeneous data as described in any one of claims 1 to 7, characterized in that: include: A data collection module (100), a data processing module (200), an entity definition module (300), a relationship definition module (400), a relationship extraction module (500) and a graph construction module (600); The data collection module (100) is used to collect distributed energy data after distributed photovoltaic access; The data processing module (200) is used to pre-process the collected distributed energy data; The entity definition module (300) is used to define core knowledge entities in the photovoltaic field according to the knowledge classification in the field of distributed photovoltaic access; The relationship definition module (400) is used to define the relationship between the core knowledge entities; The relationship extraction module (500) is used to extract the core knowledge entities and relationships; The graph construction module (600) is used to construct a knowledge graph in the field of distributed photovoltaics based on the extracted core knowledge entities and relationships.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for constructing a knowledge graph based on distributed energy heterogeneous data as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for constructing a knowledge graph based on distributed energy heterogeneous data as described in any one of claims 1 to 7 are implemented.