Hydropower station knowledge graph construction method and system
By standardizing the format and semantic normalizing the multi-source heterogeneous data in the hydropower station field, combined with data classification and extraction models, the problems of incomplete data processing and insufficient field adaptability in the existing technology are solved, high-quality knowledge graph construction is achieved, and data accuracy and completeness are improved.
Patent Information
- Application Number
- CN202510056350.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
The existing knowledge graph construction method is difficult to effectively process multi-source heterogeneous data in the field of hydropower stations, resulting in lost or incomplete information processing, and lack of adaptation to domain-specific term specifications and semantic correlation characteristics, resulting in high missed detection and false detection rates during naming entity recognition, relationship extraction and attribute matching.
A method for building a knowledge graph for hydropower stations is proposed. By collecting multi-source heterogeneous data for format standardization and semantic normalization, the basic data is generated; then the basic data is classified, and the corresponding extraction model is matched for extraction of entities, relationships and attributes; finally, the extraction results are consistently checked to generate a high-quality knowledge graph.
It significantly improves the unity and standardization of data, improves the accuracy and comprehensiveness of the extraction of entities, relationships and attributes, ensures the semantic consistency and logical rationality of the knowledge graph, and the generated knowledge graph has high accuracy and completeness, and is suitable for application scenarios such as intelligent question-and-answer and operation monitoring.
Smart Images

Figure CN119990274A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge graphs, and specifically to a method for constructing a knowledge graph of a hydropower station and a system for constructing a knowledge graph of a hydropower station. Background Art
[0002] With the rapid development of modern information technology, knowledge graphs, as a structured knowledge management and expression tool, are widely used in scenarios such as intelligent question answering and decision support. However, in the field of hydropower station technology, due to data complexity and domain expertise, existing general knowledge graph construction solutions face significant challenges in practical applications.
[0003] The data sources in the field of hydropower stations are diverse, covering engineering design documents, operation monitoring logs, environmental assessment reports, etc., in the form of various heterogeneous data such as text, tables and images. However, the existing knowledge graph construction methods mainly target a single type of data (such as natural language text) and lack the ability to jointly process heterogeneous data, especially in the analysis of tabular data, image text information extraction and cross-modal information integration, where information is often lost or incompletely processed. This limitation directly affects the extraction effect of entities, relationships and attributes, and cannot fully reflect the complex technical system of hydropower stations.
[0004] Furthermore, the expression form and semantic rules of data in the hydropower station field are highly specialized. For example, the expressions of terms such as equipment name, operating status, and technical parameters are diverse and complex, and the same entity or relationship may be expressed in multiple ways in different documents. Most existing methods are based on general language models and simple rule matching technologies, which are unable to adapt to the field-specific terminology specifications and semantic association characteristics, resulting in high missed detection rates and false detection rates in named entity recognition, relationship extraction, and attribute matching. In addition, the lack of a rationality verification mechanism for domain logic also leads to the possibility of semantic conflicts or logically inconsistent relationships and attributes in the knowledge graph.
[0005] Therefore, in order to solve the problems of heterogeneous data processing and insufficient domain adaptability, there is an urgent need for a knowledge graph construction method for the field of hydropower station technology, which can achieve comprehensive extraction of multi-source heterogeneous data and generate high-quality knowledge graphs based on domain rules and logical constraints, providing a reliable data foundation for subsequent intelligent applications. Summary of the invention
[0006] The purpose of the embodiments of the present invention is to provide a method and system for constructing a knowledge graph of a hydropower station, so as to at least solve the problems of heterogeneous data processing and insufficient domain adaptability in the construction of a knowledge graph of a hydropower station.
[0007] In order to achieve the above-mentioned objectives, the first aspect of the present invention provides a method for constructing a knowledge graph of a hydropower station, the method comprising: collecting multi-source heterogeneous data related to the hydropower station, and performing format standardization and semantic normalization processing on the data to obtain basic data; performing data classification on the basic data to obtain multiple data sets; matching the corresponding extraction models based on each data set to respectively perform entity extraction, relationship extraction and attribute extraction to obtain an extraction result set; performing consistency check on the extraction result set to filter out duplicate extraction results to obtain a graph construction data set; performing knowledge graph construction of a ternary array based on the graph construction data set to obtain and output a knowledge graph of a hydropower station.
[0008] Optionally, the rules for format standardization are as follows: use a unified character encoding format for text data, and replace content containing special symbols or non-standard characters; use a word segmentation tool to perform refined word segmentation on the text, and perform part-of-speech tagging based on a predefined dictionary and language model; perform structural analysis on table data, extract column headers and establish a one-to-one mapping relationship with cell contents, and process merged cells and nested header structures to reorganize data; for image data, extract text information through optical character recognition technology, and generate a text label mapping matrix based on the spatial position information of elements in the image.
[0009] Optionally, the rules for the semantic normalization processing are: standardizing the mapping of synonymous terms in the text based on the domain term list and the word vector embedding model; semantically disambiguating polysemous words based on the context relevance model, and determining the optimal interpretation by analyzing the position and dependency of the words in the sentence; combining the rule base and the corpus to perform normalization processing on data expressions involving specific fields, and unify the dimensions and unit formats of numerical data; completing the normalization mapping of non-standard expression terms through similarity calculation of the context.
[0010] Optionally, data classification is performed on the basic data to obtain multiple data sets, including: classification training of text data based on a support vector machine model, and construction of high-dimensional feature vectors through semantic features, key phrases and / or contextual distribution information of the text; classification of column headers and cell contents in tabular data based on regular expressions and keyword matching algorithms, and mapping column fields of different categories to specific data sets; for mixed data, by constructing a semantic clustering model based on word embeddings, entities are divided into different semantic clusters according to their semantic similarities, and classification boundaries are optimized by dynamically adjusting cluster centers.
[0011] Optionally, the entity extraction rules include: preliminary identification of entities in text and table data based on predefined regular expression templates, and matching standardized entity expressions; parsing non-standard expressions in the text using a dependency syntax analysis model, and extracting subject-predicate-object structures and modifiers by constructing a syntax tree to supplement missing entity information; combining contextual semantic similarity calculations to adjust entities with blurred boundaries or overlapping expressions, and using a domain knowledge base to verify and complete uncertain entities; for potential entities in table and image data, extracting and annotating entities through cross-modal mapping technology combined with image spatial labels and text features.
[0012] Optionally, the relationship extraction rules include: constructing dependency paths for text data based on syntactic dependency analysis, extracting potential entity pairs and their contextual relationships; using a predefined domain rule library to filter candidate relationship pairs, and retaining only entity pairs that meet the rule constraints as the initial candidate set; for the candidate relationship set, using an optimized deep learning model, in which the input layer enhances the entity pair position encoding by introducing special marker characters, the middle layer combines the self-attention mechanism to capture the contextual semantic information between entities, and the output layer completes the relationship type classification; combined with the prior knowledge base, the semantic consistency of the initially extracted relationship types is verified, and low-confidence relationships are re-evaluated through a cross-validation mechanism.
[0013] Optionally, the attribute extraction rules include: constructing a syntactic dependency tree for the basic data, and generating an embedded representation of attribute triples based on a graph neural network as candidate attribute triples; extracting contextual information of entities, attribute names and attribute values through window segmentation technology, and calculating semantic relevance based on word order similarity and dependency weights; for candidate attribute triples, selecting the optimal matching combination in combination with a sorting algorithm, and for entries without clear matches, using extended rules to infer possible attribute names and update the candidate set; for numerical attributes, performing unit normalization and dimensional conversion, and applying standard time format template alignment to time attributes; and finally formatting the updated attribute set for output.
[0014] Optionally, a consistency check is performed on the extraction result set, including: calculating the semantic similarity of entities in the extraction results based on a hash vectorization method, grouping and merging entities whose semantic similarity is greater than a preset similarity threshold; detecting semantic consistency of relationships in the extraction results through logical constraint rules; wherein the semantic consistency includes verifying the compatibility of device type and operating status and the logical continuity of time series data; detecting attribute values based on preset range constraints, and removing abnormal entries that exceed the range; generating an inconsistency report for entities, relationships or attributes that fail the check, and pushing it to the user end.
[0015] The second aspect of the present invention provides a hydropower station knowledge graph construction system, the system comprising: a collection unit, used to collect multi-source heterogeneous data related to the hydropower station, and perform format standardization processing and semantic normalization processing on this data to obtain basic data; a classification unit, used to perform data classification on the basic data to obtain multiple data sets; an extraction unit, based on matching the corresponding extraction model of each data set, to perform entity extraction, relationship extraction and attribute extraction respectively, to obtain an extraction result set; a processing unit, used to perform consistency check on the extraction result set, screen out duplicate extraction results, and obtain a graph construction data set; a construction unit, used to perform knowledge graph construction of a ternary array based on the graph construction data set, to obtain and output a knowledge graph of a hydropower station.
[0016] On the other hand, the present invention provides a computer-readable storage medium, which stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned method for constructing a knowledge graph of a hydropower station.
[0017] Through the above technical scheme, the scheme of the present invention aims at the complexity of multi-source heterogeneous data in the field of hydropower stations, and significantly improves the uniformity and standardization of data through format standardization and semantic normalization processing, laying a high-quality foundation for subsequent processing. In the data classification stage, a targeted strategy is adopted to divide the basic data into different categories, so that the data can match the most suitable extraction model, thereby improving the extraction accuracy and comprehensiveness of entities, relationships and attributes. In addition, through consistency verification, duplicate and conflicting information is effectively screened out to ensure the semantic consistency and logical rationality of the extraction results, further improving the reliability of the data. In the process of knowledge graph construction, the structured organization method based on triples realizes the efficient storage and association of domain knowledge, and the generated hydropower station knowledge graph has high accuracy and completeness. This construction method not only solves the problem of heterogeneous data integration, but also significantly improves the practicality and scalability of the graph in application scenarios such as intelligent question and answer and operation monitoring, providing reliable data support and technical guarantee for intelligent applications in the field of hydropower stations.
[0018] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings:
[0020] Figure 1 It is a flowchart of the steps of a method for constructing a knowledge graph of a hydropower station provided by an embodiment of the present invention;
[0021] Figure 2It is a system structure diagram of a hydropower station knowledge graph construction system provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0022] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.
[0023] Figure 1 This is a flow chart of a method for constructing a knowledge graph of a hydropower station provided by an embodiment of the present invention. Figure 1 As shown, an embodiment of the present invention provides a method for constructing a knowledge graph of a hydropower station, the method comprising:
[0024] Step S10: Collect multi-source heterogeneous data related to the hydropower station, and perform format standardization and semantic normalization processing on it to obtain basic data.
[0025] Specifically, the rules for format standardization are as follows: use a unified character encoding format for text data, and replace content containing special symbols or non-standard characters; use word segmentation tools to perform refined word segmentation operations on text, and perform part-of-speech tagging based on predefined dictionaries and language models; perform structural analysis on table data, extract column headers and establish a one-to-one mapping relationship with cell contents, and process merged cells and nested header structures to reorganize data; for image data, extract text information through optical character recognition technology, and generate a text label mapping matrix based on the spatial position information of elements in the image.
[0026] Furthermore, the rules of the semantic normalization processing are: standardize the mapping of synonymous terms in the text based on the domain term list and the word vector embedding model; semantically disambiguate polysemous words based on the context relevance model, and determine the optimal interpretation by analyzing the position and dependency of the words in the sentence; combine the rule base and corpus to perform normalization processing on data expressions involving specific fields, and unify the dimensions and unit formats of numerical data; complete the normalization mapping of non-standard expression terms through similarity calculation of context.
[0027] In the embodiment of the present invention, as an important clean energy infrastructure, a large amount of multi-source heterogeneous data is accumulated in the operation management and design evaluation process of the hydropower station. These data include engineering design documents, operation monitoring logs, environmental assessment reports, drawings and table data, etc., which are diverse and complex. In order to achieve high-quality knowledge graph construction, these data need to be format standardized and semantically normalized first, so as to generate basic data for further processing.
[0028] Further, the rules for format standardization processing include:
[0029] 1) Text data processing: The first step in format standardization is to unify the character encoding of text data. Data in the hydropower station field comes from various sources, and different sources may use different character encoding formats, such as UTF-8, GBK, etc. The unified character encoding format can avoid information loss or parsing errors caused by inconsistent encoding. Subsequently, special symbols and non-standard characters in the text data are replaced. For example, the unit symbols "㎡" and "square meters" may be encountered during the processing, which need to be uniformly replaced with standard expressions. The word segmentation tool is further used to perform word segmentation operations on the text, and the accurate segmentation and recognition of domain terms are ensured by combining predefined professional dictionaries and language models. In addition, part-of-speech tagging can further clarify the function of words in sentences, such as distinguishing the different semantics of "running" as a noun and a verb.
[0030] 2) Tabular data processing: Tabular data is widely used in hydropower station technical documents, recording key parameters and operating data. Format standardization includes structural analysis of tabular data, extracting the mapping relationship between column headers and cell contents, and ensuring data structuring. For complex situations where headers merge cells or nested headers, rule-based splitting and reorganization technology is used to convert non-standard tables into standardized two-dimensional tables.
[0031] 3) Image data processing: The design drawings and schematic diagrams of hydropower stations contain a lot of valuable information. Through optical character recognition (OCR) technology, text information can be extracted from images, and the spatial position information of image elements can be combined to generate a text label mapping matrix. This processing method enables the content in the image to be combined with text data to achieve cross-modal information integration.
[0032] Furthermore, the rules for semantic normalization processing include:
[0033] 1) Standardized mapping of synonymous terms: There are many different expressions of terms in the hydropower station field. For example, "installed capacity" and "equipment capacity" may have the same meaning. Through the domain terminology table and word vector embedding model, these terms can be mapped to standard terms to ensure semantic consistency. This step helps to eliminate ambiguity caused by differences in data sources.
[0034] 2) Semantic disambiguation of polysemous words: Some terms may have different meanings in different contexts. For example, "water level" can refer to both real-time monitoring data and design standards. Using the context relevance model, the optimal interpretation can be determined by analyzing the position and dependencies of words in a sentence. This technology is particularly suitable for processing unstructured texts such as operation logs.
[0035] 3) Standardization of data expression: Hydropower station data involves a large number of numerical parameters and unit information, such as "500 kilowatts" and "0.5 megawatts", which need to be standardized. In addition, the physical quantities involved in the data are standardized by combining the rule base and corpus. For example, "dam height 120m" will be standardized as "dam height: 120 meters".
[0036] 4) Normalization of non-standard terms: Non-standard expressions may have large variability due to recorders or format constraints. For example, "the design water level is 80 meters" may be described as "the design elevation is 80 meters". Through contextual similarity calculation, these variants can be identified and normalized to standard terms.
[0037] Based on the solution of the present invention, unified coding and format processing eliminates redundancy and conflict in data and improves data integrity and consistency. It successfully converts heterogeneous data such as text, tables, and images into structured, processable basic data to achieve cross-modal information integration. Based on professional dictionaries and context analysis, it ensures accurate parsing and unified expression of terms, providing high-quality input for subsequent knowledge extraction. Through automated format standardization and normalization rules, manual intervention and error rates are significantly reduced.
[0038] Step S20: performing data classification on the basic data to obtain multiple data sets.
[0039] Specifically, text data is classified and trained based on the support vector machine model, and a high-dimensional feature vector is constructed through the semantic features, key phrases and / or contextual distribution information of the text; column headers and cell contents in tabular data are classified based on regular expressions and keyword matching algorithms, and column fields of different categories are mapped to specific data sets; for mixed data, a semantic clustering model based on word embedding is constructed to divide entities into different semantic clusters according to their semantic similarities, and the classification boundaries are optimized by dynamically adjusting the cluster centers.
[0040] In the embodiment of the present invention, the data mainly includes engineering design documents, operation logs and environmental assessment reports, etc. Its classification relies on the support vector machine (SVM) model, and the specific steps are as follows:
[0041] 1) Feature extraction: Extract the semantic features, key phrases, and contextual distribution information of the text through preprocessing. For example, for the text describing the status of hydropower station equipment, extract phrases such as "generator failure" and "pump operation" as key features.
[0042] 2) Feature vector construction: Use term frequency-inverse document frequency (TF-IDF) or word embedding technology to convert the extracted features into high-dimensional vectors to form input data that can be processed by the SVM classifier.
[0043] 3) Model training: Based on the annotated domain corpus, the SVM model is trained so that it can classify text data into categories such as "engineering design", "operation monitoring", and "environmental assessment" according to semantic features.
[0044] 4) Classification output: Map the classification results to the corresponding data set to ensure that the classified text data can efficiently match the subsequent extraction model.
[0045] Furthermore, the tabular data contains key operating parameters and indicators, but its structure is complex, and the semantics of column headers and cell contents are diverse. Regular expressions and keyword matching algorithms are used to classify tabular data. The specific process is as follows:
[0046] 1) Column title analysis: Perform regular matching on the column titles of the table. For example, fields such as "water level" and "flow" are identified through rule matching and preliminarily classified as operation monitoring data.
[0047] 2) Keyword matching: Combine keywords in column headers and cell contents to further confirm the semantic category of the field. For example, columns containing "temperature" and "pressure" are assigned to the engineering design dataset.
[0048] 3) Mapping relationship construction: Based on the parsing and matching results, the semantic categories of column titles and contents are mapped to specific data sets to ensure the consistency and logic of data after classification.
[0049] Furthermore, mixed data usually contains text and tables or other formats of data, and its classification requires stronger semantic analysis capabilities. The semantic clustering model based on word embedding is used to achieve classification:
[0050] 1) Embedding representation: Use word embedding technology (such as Word2Vec or BERT) to convert text and table content in mixed data into semantic vector representation.
[0051] 2) Semantic similarity calculation: Calculate semantic similarity based on the embedding vector and preliminarily divide it into different semantic clusters.
[0052] 3) Dynamically adjust cluster centers: Dynamically adjust cluster centers through algorithms such as K-means, optimize classification boundaries, and ensure the semantic consistency of each data cluster.
[0053] Based on the solution of the present invention, combined with machine learning and rule matching methods, the classification results of text, table and mixed data are ensured to be accurate, and the interference of cross-category data is reduced. The automated classification method significantly reduces the need for manual intervention and improves the efficiency of large-scale data processing. Special classification strategies are adopted for different data types, laying a high-quality foundation for subsequent entity, relationship and attribute extraction. The classified data set can be seamlessly connected to the subsequent knowledge graph construction steps to support the application requirements of various intelligent scenarios of hydropower stations, such as operation monitoring and fault prediction.
[0054] Step S30: Matching the corresponding extraction model based on each data set to perform entity extraction, relationship extraction and attribute extraction respectively to obtain an extraction result set.
[0055] Specifically, the entity extraction rules include: preliminary identification of entities in text and table data based on predefined regular expression templates, matching standardized entity expressions; parsing non-standard expressions in text using a dependency syntax analysis model, extracting subject-predicate-object structures and modifiers by constructing a syntax tree to supplement missing entity information; combining contextual semantic similarity calculations to adjust entities with blurred boundaries or overlapping expressions, and using domain knowledge bases to verify and complete uncertain entities; for potential entities in table and image data, extracting and annotating entities through cross-modal mapping technology combined with spatial labels and text features of images.
[0056] In an embodiment of the present invention, in the field of hydropower stations, entity extraction is an important basis for constructing a knowledge graph, which involves extracting key entities such as equipment name, operating status, geographic information, etc. from text, table and image data. Due to the heterogeneous data sources and diverse expressions, entity extraction needs to combine rule matching, syntactic analysis and semantic computing technology to ensure the comprehensiveness and accuracy of the extraction. In text and tabular data, standardized entity expressions are usually relatively fixed, such as "dam height: 120 meters" or "installed capacity: 500MW". These standard format data can be quickly matched through predefined regular expression templates. For example, for equipment parameters, templates are used to identify key entities and their attribute values. This method is efficient and accurate, and is suitable for processing structured or semi-structured tabular data.
[0057] Furthermore, for entities with non-standard expressions in natural language texts, they are parsed through the dependency syntactic analysis model. First, a syntactic tree is constructed to extract the subject-predicate-object structure. For example, in the sentence "The capacity of the generator set is 500 megawatts", "generator set" is the subject and "capacity" is the attribute modified by the predicate. The parsing of the syntactic tree can locate the subject and modifiers, thereby supplementing the entities that regular expressions cannot recognize. In addition, for multi-layer modification structures (such as "left bank main dam elevation"), the dependency relationship can be parsed layer by layer and the complete entity information can be extracted.
[0058] Furthermore, in some cases, the boundaries of entities are fuzzy or there are overlapping expressions, such as "Generator No. 1" and "Generator". By calculating the contextual semantic similarity, it is possible to determine whether there is a nested relationship between entities and optimize the boundaries. For example, a word embedding model (such as Word2Vec or BERT) is used to generate a semantic vector of the entity context, and the semantic similarity is compared to decide whether to adjust the entity boundary. At the same time, for entities with uncertain expressions, they can be verified and completed in conjunction with the domain knowledge base. For example, if there is a definition of "generator" in the knowledge base, the extraction results are aligned with it. The design drawings and schematics of hydropower stations contain a large number of potential entities, such as equipment names, parameter values, and spatial positions. Through cross-modal mapping technology, the text information extracted by OCR is aligned with the spatial labels in the image. For example, in the schematic diagram, the text labeled "Pump A" may be adjacent to the equipment location label. Using this spatial relationship, the text and position features are combined to accurately extract entities in the image. At the same time, for entities in the table data, entity annotations containing semantic information are generated through the matching relationship between the header and the cell content.
[0059] Furthermore, the relationship extraction rules include: constructing dependency paths for text data based on syntactic dependency analysis, extracting potential entity pairs and their contextual relationships; using a predefined domain rule library to filter candidate relationship pairs, and only retaining entity pairs that meet the rule constraints as the initial candidate set; for the candidate relationship set, using an optimized deep learning model, in which the input layer enhances the entity pair position encoding by introducing special marker characters, the middle layer combines the self-attention mechanism to capture the contextual semantic information between entities, and the output layer completes the relationship type classification; combined with the prior knowledge base, the semantic consistency of the initially extracted relationship types is verified, and the low-confidence relationships are re-evaluated through a cross-validation mechanism.
[0060] In the embodiment of the present invention, in the process of constructing the knowledge graph of the hydropower station, relationship extraction is one of the key links. It aims to identify the relationship between entities from various data such as text and tables, and accurately classify them into specific semantic types, such as "equipment operation status", "geographic location association", etc. Due to the diverse expressions and complex semantics of hydropower station data, relationship extraction needs to combine syntactic analysis, domain rules and deep learning technology to ensure the accuracy and reliability of the results.
[0061] Specifically, a dependency path is constructed for text data through a syntactic dependency analysis model to extract potential entity pairs and their contextual relationships from sentences. Dependency analysis can parse the grammatical structure of sentences, such as subject-verb-object relationships, modification relationships, etc. For sentences describing equipment status, such as "Pump A is running under high load", the model can parse the dependency path between "Pump A" and "running under high load", thereby extracting entity pairs and their context. In addition, the analysis results can reflect the modification relationship between entities, such as the association between the attributive "high load" and the verb "run", providing important contextual information for subsequent relationship classification.
[0062] Furthermore, in order to reduce the interference of invalid relationships, the candidate relationships are filtered using a predefined domain rule library. The rule library contains common relationship patterns in the specific field of hydropower stations, such as "equipment-operating status", "equipment-geographic location", etc. By matching the entity type and context features of the candidate relationships, only the relationship pairs that meet the rule constraints are retained as the initial candidate set. For example, the rule library can filter out sentence relationships that describe equipment operation and filter out irrelevant content that describes background information. This process improves the accuracy of the relationship candidate set and significantly reduces the computational burden of the subsequent classification model.
[0063] Furthermore, when classifying the candidate relationship set, an optimized deep learning model is used to complete the multi-category classification task. The specific process includes:
[0064] 1) Input layer processing: Special marker characters are introduced for candidate entity pairs, such as [E1] and [E2], to mark the position of the entity in the sentence. This position encoding method helps the model capture the dependency relationship between entities more accurately.
[0065] 2) Intermediate layer optimization: Combined with the self-attention mechanism to capture the contextual semantic information between entities. This mechanism can effectively focus on key semantic fragments, such as verbs or modifiers that describe relationships, and ignore redundant information.
[0066] 3) Output layer classification: The model output layer classifies the relationship into specific types, such as "controls", "is located in", or "leads to", based on the extracted semantic features. The classification results are directly applied to the triple construction of the knowledge graph.
[0067] Furthermore, to ensure the semantic consistency of the classification results, the initially extracted relationship types are verified in combination with the hydropower station prior knowledge base. For example, for the extracted relationship "Pump A-Running-High Load", check whether there is a similar relationship pattern in the knowledge base. For low-confidence relationships, their semantic rationality is re-evaluated through a cross-validation mechanism, including re-analyzing the contextual semantic features and the correlation with other entity relationships to ensure that the final output relationship is accurate and logical.
[0068] Furthermore, the attribute extraction rules include: constructing a syntactic dependency tree for the basic data, and generating an embedded representation of attribute triples based on a graph neural network as candidate attribute triples; extracting contextual information of entities, attribute names, and attribute values through window segmentation technology, and calculating semantic relevance based on word order similarity and dependency weights; for candidate attribute triples, selecting the optimal matching combination in combination with a sorting algorithm, and for entries without clear matches, using extended rules to infer possible attribute names and update the candidate set; for numerical attributes, performing unit normalization and dimensional conversion, and applying standard time format template alignment to time attributes; and finally performing formatted output on the updated attribute set.
[0069] In the embodiment of the present invention, attribute extraction plays a key role in the construction of the knowledge graph of the hydropower station, aiming to extract detailed description information of entities from multi-source data, such as technical parameters, operating status and environmental impact factors of equipment. The accurate extraction of this information can significantly improve the comprehensiveness and practicality of the knowledge graph. The following are the technical details and effect description of the attribute extraction process.
[0070] Specifically, the first step of attribute extraction is to construct a syntactic dependency tree for the basic data. The syntactic dependency tree parses the grammatical structure of the text and clarifies the dependency relationships between entities, attribute names, and attribute values in the sentence. For example, in the sentence "The operating efficiency of pump A is 85%", the dependency tree can clearly identify "pump A" as an entity, "operating efficiency" as an attribute name, and "85%" as an attribute value. These dependencies lay the foundation for the generation of candidate attribute triples. On this basis, a graph neural network (GNN) is used to embed the syntactic dependency tree, encode the semantic representation of entities, attribute names, and attribute values into vectorized features, and generate candidate attribute triples. This method can capture complex contextual semantics and dependencies and improve the accuracy of candidate triples.
[0071] Furthermore, in order to optimize the candidate attribute triples, the window segmentation technology is used to extract the contextual information of entities, attribute names and attribute values, and the semantic relevance is calculated based on word order similarity and dependency weights. For example, in the sentence "The rated power of the generator is 500 kilowatts", the window segmentation technology can place "generator", "rated power" and "500 kilowatts" in an analysis window to ensure semantic focus when calculating relevance. By combining word order similarity and dependency weights, the semantic consistency of the triples is further quantified. For example, "generator-rated power-500 kilowatts" has a higher relevance score than "generator-working environment-500 kilowatts", thereby ensuring more accurate attribute extraction results.
[0072] Furthermore, for entries without clear matches in the candidate triples, the extension rules are used to infer possible attribute names and update the candidate set. For example, if the sentence lacks an explicit attribute name, but the context can infer that the entry describes "design flow", the extension rule will dynamically supplement the attribute name and update the triple. For numerical attributes, unit normalization and dimension conversion are performed. For example, "500 kilowatts" is unified as "500kW" to avoid repeated extraction problems caused by unit differences. For time attributes, such as "January 10, 2023", standard time format templates (such as ISO 8601) are used for alignment to ensure the consistency and parsability of time information. After all update operations are completed, the final attribute set is formatted and output. The output format is usually in the form of structured key-value pairs, such as JSON or RDF, to ensure that the attribute information can be seamlessly integrated into the knowledge graph.
[0073] Step S40: Perform consistency check on the extraction result set, filter out duplicate extraction results, and obtain a graph construction data set.
[0074] Specifically, a consistency check is performed on the extraction result set, including: calculating the semantic similarity of entities in the extraction results based on a hash vectorization method, grouping and merging entities whose semantic similarity is greater than a preset similarity threshold; detecting semantic consistency of relationships in the extraction results through logical constraint rules; wherein the semantic consistency includes verifying the compatibility of device types and operating states and the logical continuity of time series data; detecting attribute values based on preset range constraints, and removing abnormal entries that exceed the range; generating an inconsistency report for entities, relationships or attributes that fail the check, and pushing it to the user end.
[0075] In an embodiment of the present invention, the goal of consistency verification is to generate a high-quality graph to construct a data set by eliminating redundancy, verifying logical consistency, and removing outliers. In the extraction results, multiple entities may appear repeatedly due to different forms of expression. For example, "Pump A" and "Pump No. 1" may point to the same object. Through the hash vectorization method, the semantic information of each entity is converted into a vector representation of a fixed length, and the semantic similarity between entities is calculated using cosine similarity. If the similarity is higher than a preset threshold (such as 0.85), these entities are grouped together and merged into a unified entity. In the hydropower station scenario, this method can effectively reduce the redundancy of repeated entities such as equipment names and operating status. For example, "Generator Unit 2" and "Unit 2" can be merged into "Generator Unit 2" after similarity calculation to ensure that each entity in the data set is unique.
[0076] Furthermore, for the relations in the extraction results, the semantic consistency is checked through logical constraint rules. These rules are defined based on the operation logic and technical specifications of the hydropower station field, including:
[0077] 1) Compatibility of equipment type and operating status: Verify that the equipment operating status matches its type. For example, a "generator set" should not have an operating status of "opening the gate to release floodwater".
[0078] 2) Logical continuity of time series data: Check whether the time-related relationship is logical. For example, the event of "water level rise" should occur after "reservoir refilling", not before.
[0079] These logical constraint rules can significantly reduce semantic conflicts in relationship extraction and improve the accuracy of graph construction.
[0080] Furthermore, anomaly detection is performed on attribute values based on preset range constraint rules. For example:
[0081] 1) Numerical attributes: For example, “installed capacity” should be within a reasonable range (e.g. 50MW to 500MW). Entries outside the range will be removed.
[0082] 2) Time attributes: For example, the date format must comply with the standard (such as ISO 8601), otherwise it will be marked as an exception.
[0083] This process is particularly critical in the processing of operational monitoring data, which can effectively eliminate noise data and ensure the accuracy of attribute information.
[0084] Furthermore, for entities, relationships or attributes that fail the verification, a detailed inconsistency report is generated, including a description of the problem and recommended actions, and pushed to the user for manual confirmation. For example, if a relationship fails the logical verification, the report will indicate "the relationship type does not match the entity" and provide corresponding context information for correction.
[0085] Step S50: Execute knowledge graph construction of the ternary array based on the graph construction data set, and obtain and output the knowledge graph of the hydropower station.
[0086] Specifically, the core of constructing the hydropower station knowledge graph is to integrate the extracted entities, relationships and attribute data into a structured triple form (subject-predicate-object), and generate a knowledge graph that can be queried and reasoned through reasonable organization and storage. This process needs to take into account both domain characteristics and technical specifications to ensure that the generated knowledge graph has high accuracy and practicality.
[0087] Specifically, the graph is used to construct the entity, relationship and attribute information in the data set, and triples that conform to the knowledge graph structure are generated. For example, for the operating status of "Pump A", a triple ("Pump A", "Operational Status", "High Load") is generated. In this process, the entity and attribute values need to be semantically mapped and normalized. For example, "Generator Unit 1" is unified as "Generator Unit 1" to ensure that all entity expressions are consistent. The attribute values need to be unified in units and formats, such as converting "500 kilowatts" to "500kW". According to the knowledge rules in the field of hydropower stations, common relationship types (such as "located in", "controlled", and "dependent on") are defined, and the semantic associations between entities are mapped into triples. For example, for the data describing "Reservoir A is located in a certain basin", ("Reservoir A", "located in", "certain basin") are generated. This process relies on the results of the relationship extraction stage, and is further verified and supplemented by the domain knowledge base.
[0088] Furthermore, the knowledge graph uses a graph database (such as Neo4j or GraphDB) to store triple data. Each entity is represented as a node in the graph, and each relationship is represented as an edge connecting the nodes. For example, "Generator 1" is a node, and "Operation Status" is an edge connected to the "Normal Operation" node. The advantages of graph databases are efficient query and scalability, and are particularly suitable for the storage and association of complex data in hydropower stations. Additional information such as timestamp, data source, and confidence is added to each triple. For example, record that a piece of data comes from the "2023 Operation Log" and attach a confidence score (such as 90%). This information is particularly important in the dynamic update and reasoning process.
[0089] Furthermore, the knowledge graph supports multiple output formats, such as JSON, RDF, and CSV, to meet the needs of different application scenarios. For example, the JSON format is easy to connect to the front-end system, while RDF is suitable for semantic query and reasoning. During the output process, data integrity and consistency must also be ensured. For example, verify the semantic logic of all entities, relationships, and attributes to avoid isolated nodes or incomplete relationship chains. After building the graph, the structure and content of the graph are displayed through visualization tools such as Gephi or D3.js. For example, dynamically display the operating status of a hydropower station and the dependencies between its equipment to intuitively present data associations.
[0090] Figure 2 1 is a system structure diagram of a hydropower station knowledge graph construction system provided by an embodiment of the present invention. Figure 2As shown, an embodiment of the present invention provides a hydropower station knowledge graph construction system, and the system includes: a collection unit, which is used to collect multi-source heterogeneous data related to the hydropower station, and perform format standardization processing and semantic normalization processing on this data to obtain basic data; a classification unit, which is used to perform data classification on the basic data to obtain multiple data sets; an extraction unit, which matches the corresponding extraction model based on each data set to perform entity extraction, relationship extraction and attribute extraction respectively to obtain an extraction result set; a processing unit, which is used to perform consistency check on the extraction result set, screen out duplicate extraction results, and obtain a graph construction data set; a construction unit, which is used to perform knowledge graph construction of a ternary array based on the graph construction data set, and obtain and output a knowledge graph of a hydropower station.
[0091] An embodiment of the present invention also provides a computer-readable storage medium, on which instructions are stored, which, when executed on a computer, enable the computer to execute the above-mentioned method for constructing a knowledge graph of a hydropower station.
[0092] Those skilled in the art will understand that all or part of the steps in the method for implementing the above-mentioned embodiments can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including several instructions for making a single-chip microcomputer, a chip or a processor (processor) perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0093] The optional embodiments of the present invention are described in detail above in conjunction with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the technical concept of the embodiments of the present invention, the technical scheme of the embodiments of the present invention can be subjected to a variety of simple modifications, and these simple modifications all belong to the protection scope of the embodiments of the present invention. It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the embodiments of the present invention will not further describe various possible combinations.
[0094] In addition, various embodiments of the present invention may be arbitrarily combined, and as long as they do not violate the concept of the embodiments of the present invention, they should also be regarded as the contents disclosed in the embodiments of the present invention.
Claims
1. A method for constructing a knowledge graph of a hydropower station, characterized in that: The method comprises: Collect multi-source heterogeneous data related to hydropower stations, and perform format standardization and semantic normalization processing on them to obtain basic data; Performing data classification on the basic data to obtain multiple data sets; Based on matching the corresponding extraction models of each data set, entity extraction, relationship extraction and attribute extraction are performed respectively to obtain the extraction result set; Perform consistency check on the extraction result set, filter out duplicate extraction results, and obtain the graph construction data set; Based on the graph construction dataset, the knowledge graph of the ternary array is constructed to obtain and output the knowledge graph of the hydropower station.
2. The method according to claim 1, characterized in that The rules for format standardization are: Use unified character encoding format for text data and replace content containing special symbols or non-standard characters; Use word segmentation tools to perform refined word segmentation on text and perform part-of-speech tagging based on predefined dictionaries and language models; Perform structural analysis on table data, extract column headers and establish a one-to-one mapping relationship with cell contents, and handle merged cells and nested header structures to reorganize data; For image data, optical character recognition technology is used to extract text information, and a text label mapping matrix is generated by combining the spatial position information of elements in the image.
3. The method according to claim 1, characterized in that The rules for the semantic normalization process are: Standardize the mapping of synonymous terms in the text based on the domain glossary and word vector embedding model; Disambiguate the meaning of polysemous words based on the context relevance model, and determine the optimal interpretation by analyzing the position and dependency relationship of words in the sentence; Combine the rule base and corpus to perform normalization processing on data expressions related to specific fields and unify the dimensions and unit formats of numerical data; For non-standard expressions, the normalized mapping is completed by calculating the similarity of the context.
4. The method according to claim 1, characterized in that: Perform data classification on the basic data to obtain multiple data sets, including: Classify and train text data based on a support vector machine model, and construct a high-dimensional feature vector through the semantic features, key phrases, and / or context distribution information of the text; Classify column headers and cell contents in table data based on regular expressions and keyword matching algorithms, and map column fields of different categories to specific data sets; For mixed data, a semantic clustering model based on word embedding is constructed to divide entities into different semantic clusters according to their semantic similarity, and the classification boundaries are optimized by dynamically adjusting the cluster centers.
5. The method according to claim 1, characterized in that The entity extraction rules include: Perform preliminary recognition of entities in text and table data based on predefined regular expression templates, matching standardized entity expressions; The dependency syntax analysis model is used to parse non-standard expressions in the text, and the subject-verb-object structure and modifiers are extracted by constructing a syntax tree to supplement the missing entity information; Combined with contextual semantic similarity calculation, entities with blurred boundaries or overlapping expressions are adjusted, and domain knowledge base is used to verify and complete uncertain entities; For potential entities in table and image data, cross-modal mapping technology is used to combine the spatial labels and text features of the image to extract and annotate the entities.
6. The method according to claim 1, characterized in that The relationship extraction rules include: Based on syntactic dependency analysis, a dependency path is constructed for text data to extract potential entity pairs and their contextual relationships; Use the predefined domain rule base to filter candidate relationship pairs and only retain entity pairs that meet the rule constraints as the initial candidate set; For the candidate relationship set, an optimized deep learning model is used, in which the input layer enhances the entity position encoding by introducing special marker characters, the middle layer combines the self-attention mechanism to capture the contextual semantic information between entities, and the output layer completes the relationship type classification; Combined with the prior knowledge base, the semantic consistency of the initially extracted relationship types is verified, and the low-confidence relationships are re-evaluated through a cross-validation mechanism.
7. The method according to claim 1, characterized in that The attribute extraction rules include: Construct a syntactic dependency tree for the basic data, and generate an embedding representation of attribute triples based on a graph neural network as candidate attribute triples; The context information of entities, attribute names and attribute values is extracted through window segmentation technology, and the semantic relevance is calculated based on word order similarity and dependency weights; For candidate attribute triples, the best matching combination is selected in combination with the sorting algorithm. For entries without clear matches, the possible attribute names are inferred using the expansion rules and the candidate set is updated. For numerical attributes, unit normalization and dimension conversion are performed, and standard time format template alignment is applied to time attributes; finally, formatted output is performed on the updated attribute set.
8. The method according to claim 1, characterized in that Perform consistency checks on the extracted result set, including: The semantic similarity of entities in the extraction results is calculated based on the hash vectorization method, and entities with semantic similarity greater than the preset similarity threshold are grouped and merged; For the relations in the extraction results, the semantic consistency is detected through logical constraint rules; among them, The semantic consistency includes verifying the compatibility of equipment types and operating states and the logical continuity of time series data; Detect attribute values based on preset range constraints and remove abnormal entries that are out of range; For entities, relationships, or attributes that fail verification, an inconsistency report is generated and pushed to the user end.
9. A hydropower station knowledge graph construction system, characterized in that: The system comprises: The acquisition unit is used to collect multi-source heterogeneous data related to the hydropower station, and perform format standardization and semantic normalization processing on it to obtain basic data; A classification unit, used for performing data classification on the basic data to obtain multiple data sets; The extraction unit matches the corresponding extraction model based on each data set to perform entity extraction, relationship extraction and attribute extraction respectively to obtain an extraction result set; A processing unit, used to perform consistency check on the extraction result set, filter out duplicate extraction results, and obtain a graph construction data set; A construction unit is used to execute knowledge graph construction of a ternary array based on a graph construction data set, and obtain and output a knowledge graph of a hydropower station.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the method for constructing a knowledge graph of a hydropower station as described in any one of claims 1 to 8.
Citation Information
Cited By
Artificial intelligence-based table generation method, system and equipment and storage medium
CN120409440A
Network multi-source financial text big data processing method
CN120578767A
Automatic historical text classification method based on natural language processing
CN120973942A
Historical material text automatic classification method based on natural language processing
CN120973942B
Effect matrix construction method based on topic clustering and two rounds of questions and answers
CN121256033A