Construction method of multi-source heterogeneous data knowledge graph for power marketing

By acquiring, preprocessing and integrating power marketing data, building a power marketing knowledge graph is solved, and the problem of low efficiency in power marketing data management and utilization is achieved, efficient data integration and quality improvement of knowledge graphs are achieved.

CN120106204APending Publication Date: 2025-06-06WEIHAI POWER SUPPLY COMPANY OF STATE GRID SHANDONG ELECTRIC POWER COMPANY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510229480.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

It is difficult to carry out unified centralized management of power marketing work in the existing technology, and the data related to power marketing is numerous and chaotic, making it difficult to use efficiently.

Method used

By obtaining power marketing data, preprocessing and structure, determining entities, attributes and relationships, integrating data using deep fusion methods to form a consistent knowledge representation, and storing the data in the Neo4j graph database to build a power marketing knowledge graph.

Benefits of technology

It effectively integrates various data related to power marketing, improves data utilization and accuracy, eliminates duplication and contradictions in the data, and improves the quality and reliability of the knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106204A_ABST
    Figure CN120106204A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power marketing-oriented multi-source heterogeneous data knowledge graph construction method, and belongs to the technical field of power grids, and the method comprises the following steps: obtaining electric power marketing data; preprocessing the power marketing data to form structured power marketing data; extracting the structured power marketing data, determining entities, entity attributes and entity relationships, and constructing a dominant network in the field of power marketing through entity recognition and relationship extraction; integrating the obtained entities, attributes and relationships by adopting a deep fusion method, eliminating repetition and contradiction, and forming consistent knowledge representation; storing the fused power marketing data in a Neo4j graph database to form a power marketing knowledge graph, and updating the knowledge graph regularly; according to the method, various data related to power marketing are effectively integrated, the utilization rate and accuracy of the data are improved, repetition and contradiction in the data are eliminated, consistent knowledge representation is formed, and the quality and reliability of the knowledge graph are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power grid technology, and in particular relates to a method for constructing a multi-source heterogeneous data knowledge graph for power marketing. Background Art

[0002] Power marketing is the purpose of power companies to meet people's power consumption needs in a changing market environment. Through a series of market-related business activities, they provide power products and corresponding services that meet consumer needs, thereby achieving corporate goals.

[0003] However, in the existing technology, it is difficult to manage power marketing in a unified and centralized manner, and the data sources related to power marketing are numerous and messy, making it difficult to use them efficiently. For example, similar concepts such as electricity charges and electricity prices cannot be accurately judged as the same concept, resulting in the need to improve the efficiency and quality of knowledge graph construction. Summary of the invention

[0004] In view of the deficiencies in the prior art, the present invention provides a subject matter to solve the above-mentioned problems.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: A method for constructing a multi-source heterogeneous data knowledge graph for power marketing, comprising the following steps: Obtaining power marketing data; wherein, power marketing data includes power customer data, power grid company data, and power plant data; Preprocessing the power marketing data to form structured power marketing data; the preprocessing of the power marketing data includes data cleaning, data classification, data structure transformation and data integration; Extract structured power marketing data, identify entities, entity attributes and entity relationships, and construct an explicit network in the power marketing field through entity identification and relationship extraction; entities include power customers, power grid companies, and power plants; entity relationships include power supply and demand relationships, power consumption relationships, and power user affiliation relationships; entity attributes include power consumption information, power supply information, and power generation information; Adopt deep fusion methods to integrate the acquired entities, attributes and relationships, eliminate duplications and contradictions, and form a consistent knowledge representation; The integrated power marketing data is stored in the Neo4j graph database to form a power marketing knowledge graph, and the knowledge graph is updated regularly; The deep fusion method is used to integrate the acquired entities, attributes and relationships, eliminate duplications and contradictions, and form a consistent knowledge representation, which specifically includes the following steps: Calculate the likelihood between entities through the features between different entities; Determine the similarity between entities based on the likelihood values ​​between them; If the similarity between entities is high, the likelihood value between entity attributes is calculated through the features between entity attributes; According to the likelihood values ​​between entity attributes, the similarity between entity attributes is judged; If the similarity between entity attributes is high, the entities and entity attributes will be merged.

[0006] On the basis of the above technical solution, the present invention also provides the following optional technical solution: Further technical solution: The acquisition of power marketing data specifically includes: Collect data from data sources such as electricity fee management platform, historical electricity consumption data, electricity price policy data, etc. through API interface or data import tool.

[0007] Further technical solution: The power marketing data is preprocessed to form structured power marketing data, specifically the following steps: Perform data cleaning on power marketing data; data cleaning includes outlier identification, data error correction and data missing value processing; Summarize the knowledge content of power marketing and classify the data of power marketing; the summarization of power marketing knowledge refers to the relationship between important concepts, key attributes and entities in the field of power marketing; Convert non-numeric data in power data into numeric data; Perform data scale conversion on numerical data; data scale conversion includes standardization and normalization.

[0008] Further technical solution: The likelihood values ​​between entities are calculated by using the features between different entities, specifically including the following steps: Get the number of times each entity appears in the power marketing data, marked as count(E1) and count(E2) respectively; According to the number of times each entity appears in the power marketing data, the co-occurrence matrix count(E1, E2) is constructed; Standardize the co-occurrence matrix to obtain the likelihood value score (E1, E2) between entities; The method of normalizing the co-occurrence matrix is ​​specifically as follows: By formula , generate the likelihood value score (E1, E2) between entities; In the formula, count(E1,E2) represents the number of times entities E1 and E2 appear in the power marketing data.

[0009] Further technical solution: The similarity between entities is judged according to the likelihood value between the entities, and the judgment method is specifically as follows: Compare the likelihood value score (E1, E2) between entities with the entity likelihood value setting value; If the likelihood value score (E1, E2) between entities is greater than the entity likelihood value setting value; it means that the larger the likelihood value, the higher the similarity between entities, and the similarity between the entities is judged to be high.

[0010] A further technical solution: if the similarity between entities is high, the likelihood values ​​between the entity attributes are calculated by using the features between the entity attributes, which specifically includes the following steps: Get the value of each attribute of entity E1 and entity E2, and mark each attribute in entity E1 as A 1 , A 2 , ..., A n ; Label each attribute in entity E2 as B 1 , B 2 , ..., B n ; Get the value of the ith attribute of entity E1 and entity E2, marked as A respectively i1 and A i2 ; By formula , generate the likelihood value Cosine Similarity (E1, E2) between entity attributes; In the formula, A i1 It represents the value of the i-th attribute in entity E1. i2 It represents the value of the i-th attribute in entity E2.

[0011] Further technical solution: The similarity between entity attributes is determined based on the likelihood values ​​between entity attributes, and the determination method is specifically as follows: Compare the likelihood value Cosine Similarity (E1, E2) between entity attributes with the entity attribute likelihood value setting value; If the likelihood value Cosine Similarity (E1, E2) between entity attributes is less than or equal to the entity attribute likelihood value setting value, it means that the smaller the likelihood value Cosine Similarity (E1, E2) between entity attributes, the lower the similarity between entity attributes, and it is determined that the similarity between entity attributes is low; If the likelihood value Cosine Similarity (E1, E2) between entity attributes is greater than the entity attribute likelihood value setting value, it means that the larger the likelihood value Cosine Similarity (E1, E2) between entity attributes is, the higher the similarity between entity attributes is. Then it is determined that the similarity between entity attributes is high, and entity E1 and entity E2 need to be fused.

[0012] Further technical solution: The integrated power marketing data is stored in the Neo4j graph database to form a power marketing knowledge graph, and the knowledge graph is updated regularly, specifically including the following steps: The integrated power marketing data is stored in the Neo4j graph database in the form of "entity-attribute-attribute" to form a power marketing knowledge graph; Introduce new power marketing data and determine whether the new knowledge data is consistent with the existing knowledge. If consistent, first determine the correctness of the new knowledge and then update the relevant entities and attributes.

[0013] The present invention provides a method for constructing a multi-source heterogeneous data knowledge graph for power marketing, which has the following beneficial effects compared with the prior art: The present invention effectively integrates various types of data related to power marketing by processing, classifying and transforming multi-source heterogeneous data of power marketing, thereby improving data utilization and accuracy. It also eliminates duplications and contradictions in the data by performing similarity judgment on the integrated data, forms a consistent knowledge representation, and improves the quality and reliability of the knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 A flowchart of a method for constructing a multi-source heterogeneous data knowledge graph for power marketing provided in an embodiment of the present invention.

[0015] Figure 2 A type diagram of multi-source heterogeneous data provided by the present invention.

[0016] Figure 3 This is a flowchart of step S40 provided by the present invention.

[0017] Figure 4 This is a structural diagram of the building platform device provided by the present invention. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0019] The specific implementation of the present invention is described in detail below in conjunction with specific embodiments.

[0020] like Figure 1 and Figure 2 As shown, a method for constructing a multi-source heterogeneous data knowledge graph for power marketing provided by an embodiment of the present invention includes the following steps: Step S10: Acquire power marketing data; wherein the power marketing data includes power customer data, power grid company data, and power plant data; Step S20: pre-processing the power marketing data to form structured power marketing data; wherein the pre-processing of the power marketing data includes data cleaning, data classification, data structure conversion and data integration; Step S30: extracting structured power marketing data, determining entities, entity attributes and entity relationships, and constructing an explicit network in the power marketing field through entity identification and relationship extraction; wherein entities include power customers, power grid companies, and power plants; entity relationships include power supply and demand relationships, power consumption relationships, and power user affiliation relationships; entity attributes include power consumption information, power supply information, and power generation information; Step S40: using a deep fusion method to integrate the acquired entities, attributes, and relationships, eliminate duplications and contradictions, and form a consistent knowledge representation; Step S50: Store the integrated power marketing data in the Neo4j graph database to form a power marketing knowledge graph, and update the knowledge graph regularly.

[0021] As a preferred embodiment of the present invention, the step S10 specifically includes: Collect data from data sources such as electricity fee management platform, historical electricity consumption data, electricity price policy data, etc. through API interface or data import tool.

[0022] As a preferred embodiment of the present invention, the step S20 comprises the following steps: Step S21: performing data cleaning on the power marketing data; wherein data cleaning includes outlier identification, data error correction and data missing value processing; Step S22: Summarize the power marketing knowledge content and classify the power marketing data; wherein the summarization of power marketing knowledge refers to the relationship between important concepts, key attributes and entities in the power marketing field; Power marketing data classification refers to the classification of power marketing data into numerical data and non-numerical data; for example, power marketing data such as user electricity bills and power plant power generation can be classified as numerical data, while user complaint data (content) and user demand content (voice data, picture data, etc.) can be classified as non-numerical data; Step S23: converting non-numeric data in the power data into numeric data; The conversion methods include but are not limited to using the VALUE function to convert text into numbers, using the DATEVALUE function to convert text into dates, using the TEXT function to convert numbers or dates into text, or using data conversion tools, etc.; for example, the user's address can be converted through longitude and latitude or by creating a coordinate system and output as numerical data (longitude and latitude coordinates, coordinate points); Step S24: performing data scale conversion on the numerical data; wherein the data scale conversion includes standardization and normalization.

[0023] like Figure 3 As shown, as a preferred embodiment of the present invention, the step S40 specifically includes the following steps: Step S41: Calculate the likelihood values ​​between entities through the features between different entities; Step S42: judging the similarity between entities according to the likelihood values ​​between the entities; Step S43: If the similarity between entities is high, the likelihood values ​​between the entity attributes are calculated based on the features between the entity attributes; Step S45: determining the similarity between entity attributes according to the likelihood values ​​between entity attributes; Step S46: If the similarity between the entity attributes is high, the entities and entity attributes are merged.

[0024] As a preferred embodiment of the present invention, the step S41 specifically includes: Step S41.1: Get the number of times each entity appears in the power marketing data, marked as count(E1) and count(E2) respectively; Step S41.2: construct a co-occurrence matrix count(E1, E2) based on the number of times each entity appears in the power marketing data; Step S41.3: Standardize the co-occurrence matrix to obtain the likelihood value score (E1, E2) between entities; The method of normalizing the co-occurrence matrix is ​​specifically as follows: By formula , generate the likelihood value score (E1, E2) between entities; In the formula, count(E1,E2) represents the number of times entities E1 and E2 appear in the power marketing data.

[0025] As a preferred embodiment of the present invention, the determination method of step S42 is: Compare the likelihood value score (E1, E2) between entities with the entity likelihood value setting value; If the likelihood value score (E1, E2) between entities is less than or equal to the entity likelihood value setting value; it means that the smaller the likelihood value, the lower the similarity between entities, and the similarity between the entities is judged to be low; If the likelihood value score (E1, E2) between entities is greater than the entity likelihood value setting value; it means that the larger the likelihood value, the higher the similarity between entities, and the similarity between the entities is judged to be high; It should be explained that the method for obtaining the entity likelihood value setting value is consistent with the method for obtaining the likelihood value score (E1, E2), and its value is set by relevant personnel in this field.

[0026] As a preferred embodiment of the present invention, the step S43 specifically includes: Step S43.1: Get the value of each attribute of entity E1 and entity E2, and mark each attribute in entity E1 as A 1 , A 2 , ..., A n ; Label each attribute in entity E2 as B 1 , B 2 , ..., B n ; Step S43.2: Get the values ​​of the ith attribute of entity E1 and entity E2, marked as A and i1 and A i2 ; By formula , generate the likelihood value Cosine Similarity (E1, E2) between entity attributes; In the formula, A i1 It represents the value of the i-th attribute in entity E1. i2 It represents the value of the i-th attribute in entity E2.

[0027] As a preferred embodiment of the present invention, the determination method of step S44 is specifically as follows: Compare the likelihood value Cosine Similarity (E1, E2) between entity attributes with the entity attribute likelihood value set value; If the likelihood value Cosine Similarity (E1, E2) between entity attributes is less than or equal to the entity attribute likelihood value setting value, it means that the smaller the likelihood value Cosine Similarity (E1, E2) between entity attributes is, the lower the similarity between entity attributes is, and it is determined that the similarity between entity attributes is low; If the likelihood value Cosine Similarity (E1, E2) between entity attributes is greater than the entity attribute likelihood value setting value, it means that the larger the likelihood value Cosine Similarity (E1, E2) between entity attributes is, the higher the similarity between entity attributes is. Then it is determined that the similarity between entity attributes is high, and it is necessary to merge entity E1 and entity E2. It should be explained that the method for obtaining the setting value of the entity attribute likelihood value is consistent with the method for obtaining the likelihood value Cosine Similarity (E1, E2) between entity attributes, and its value is set by relevant personnel in this field.

[0028] As a preferred embodiment of the present invention, step S50 specifically includes the following steps: Step S51: storing the integrated power marketing data in a Neo4j graph database in an “entity-attribute-attribute” manner to form a power marketing knowledge graph; Step S52: Introduce new power marketing data and determine whether the new knowledge data is consistent with the existing knowledge. If consistent, first determine the correctness of the new knowledge and then update the relevant entities and attributes.

[0029] like Figure 4 As shown, as a preferred embodiment of the present invention, the present invention also provides a platform device for constructing a multi-source heterogeneous data knowledge graph for power marketing, and the platform device includes a memory, a processor, and a computer program stored in the memory and executable on the processor.

[0030] In addition, the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the above-mentioned method for constructing a multi-source heterogeneous data knowledge graph for power marketing.

[0031] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for constructing a knowledge graph of multi-source heterogeneous data for power marketing, characterized in that: The following steps are involved: Obtain power marketing data; Preprocessing the power marketing data to form structured power marketing data; the preprocessing of the power marketing data includes data cleaning, data classification, data structure transformation and data integration; Extract structured power marketing data, identify entities, entity attributes and entity relationships, and construct an explicit network in the power marketing field through entity identification and relationship extraction; entities include power customers, power grid companies, and power plants; entity relationships include power supply and demand relationships, power consumption relationships, and power user affiliation relationships; entity attributes include power consumption information, power supply information, and power generation information; Adopt deep fusion methods to integrate the acquired entities, attributes and relationships, eliminate duplications and contradictions, and form a consistent knowledge representation; The integrated power marketing data is stored in the Neo4j graph database to form a power marketing knowledge graph, and the knowledge graph is updated regularly; The deep fusion method is used to integrate the acquired entities, attributes and relationships, eliminate duplications and contradictions, and form a consistent knowledge representation, which specifically includes the following steps: Calculate the likelihood between entities through the features between different entities; Determine the similarity between entities based on the likelihood values ​​between them; If the similarity between entities is high, the likelihood value between entity attributes is calculated through the features between entity attributes; According to the likelihood values ​​between entity attributes, the similarity between entity attributes is judged; If the similarity between entity attributes is high, the entities and entity attributes will be merged.

2. The method for constructing a multi-source heterogeneous data knowledge graph for power marketing according to claim 1 is characterized in that: The obtaining of power marketing data specifically includes: Collect data from data sources such as electricity fee management platform, historical electricity consumption data, electricity price policy data, etc. through API interface or data import tool.

3. The method for constructing a multi-source heterogeneous data knowledge graph for power marketing according to claim 1 is characterized in that: The preprocessing of the power marketing data to form structured power marketing data comprises the following specific steps: Perform data cleaning on power marketing data; data cleaning includes outlier identification, data error correction and data missing value processing; Summarize the knowledge content of power marketing and classify the data of power marketing; the summarization of power marketing knowledge refers to the relationship between important concepts, key attributes and entities in the field of power marketing; Convert non-numeric data in power data into numeric data; Perform data scale conversion on numerical data; data scale conversion includes standardization and normalization.

4. The method for constructing a multi-source heterogeneous data knowledge graph for power marketing according to claim 1 is characterized in that: The method of calculating the likelihood values ​​between entities by using the features between different entities specifically includes the following steps: Get the number of times each entity appears in the power marketing data, marked as count(E1) and count(E2) respectively; According to the number of times each entity appears in the power marketing data, the co-occurrence matrix count(E1, E2) is constructed; Standardize the co-occurrence matrix to obtain the likelihood value score (E1, E2) between entities; The method of normalizing the co-occurrence matrix is ​​specifically as follows: By formula , generate the likelihood value score (E1, E2) between entities; In the formula, count(E1,E2) represents the number of times entities E1 and E2 appear in the power marketing data.

5. The method for constructing a multi-source heterogeneous data knowledge graph for power marketing according to claim 4 is characterized in that: The similarity between entities is determined based on the likelihood values ​​between the entities, and the determination method is specifically as follows: Compare the likelihood value score (E1, E2) between entities with the entity likelihood value setting value; If the likelihood value score (E1, E2) between entities is greater than the entity likelihood value setting value; it means that the larger the likelihood value, the higher the similarity between entities, and the similarity between the entities is judged to be high.

6. The method for constructing a multi-source heterogeneous data knowledge graph for power marketing according to claim 1 is characterized in that: If the similarity between entities is high, the likelihood values ​​between the entity attributes are calculated by using the features between the entity attributes, which specifically includes the following steps: Get the value of each attribute of entity E1 and entity E2, and label each attribute in entity E1 as A1, A2, ..., A n ; Label each attribute in entity E2 as B1, B2, ..., B n ; Get the value of the ith attribute of entity E1 and entity E2, marked as A respectively i1 and A i2 ; By formula , generate the likelihood value Cosine Similarity (E1, E2) between entity attributes; In the formula, A i1 It represents the value of the i-th attribute in entity E1. i2 It represents the value of the i-th attribute in entity E2.

7. The method for constructing a multi-source heterogeneous data knowledge graph for power marketing according to claim 6, characterized in that: The similarity between entity attributes is determined based on the likelihood values ​​between entity attributes, and the determination method is specifically as follows: Compare the likelihood value Cosine Similarity (E1, E2) between entity attributes with the entity attribute likelihood value set value; If the likelihood value Cosine Similarity (E1, E2) between entity attributes is less than or equal to the entity attribute likelihood value setting value, it means that the smaller the likelihood value Cosine Similarity (E1, E2) between entity attributes is, the lower the similarity between entity attributes is, and it is determined that the similarity between entity attributes is low; If the likelihood value Cosine Similarity (E1, E2) between entity attributes is greater than the entity attribute likelihood value setting value, it means that the larger the likelihood value Cosine Similarity (E1, E2) between entity attributes is, the higher the similarity between entity attributes is. Then it is determined that the similarity between entity attributes is high, and entity E1 and entity E2 need to be fused.

8. The method for constructing a multi-source heterogeneous data knowledge graph for power marketing according to claim 1 is characterized in that: The integrated power marketing data is stored in the Neo4j graph database to form a power marketing knowledge graph, and the knowledge graph is updated regularly, specifically including the following steps: The integrated power marketing data is stored in the Neo4j graph database in the form of "entity-attribute-attribute" to form a power marketing knowledge graph; Introduce new power marketing data and determine whether the new knowledge data is consistent with the existing knowledge. If consistent, first determine the correctness of the new knowledge and then update the relevant entities and attributes.