Entity extraction method and system for hydropower station data

By identifying attribute features and classifying data of hydropower station data, calling targeted entity extraction models, and merging entities through hash vectorization and confidence scoring mechanisms, the problem of insufficient adaptability of entity extraction in the hydropower station field in the existing technology is solved, and high-precision and consistency entity extraction results are achieved.

CN119962536APending Publication Date: 2025-05-09GUODIAN DADU RIVER POWER ENG
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510051104.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing entity extraction methods are insufficient in processing professional data in the field of hydropower stations, resulting in insufficient accuracy of extraction results.

Method used

By collecting and preprocessing hydropower station data, identifying attribute features and performing data classification, calling entity extraction models for different data sets, and combining hash vectorization and confidence scoring mechanisms for entity merging.

Benefits of technology

It improves the accuracy and comprehensiveness of entity extraction, solves the problem of repeated entities caused by diversified expressions, and ensures the integrity and consistency of the knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962536A_ABST
    Figure CN119962536A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an entity extraction method and system for hydropower station data, and belongs to the technical field of knowledge graph construction. The entity extraction method for hydropower station data comprises the following steps: collecting original hydropower station data, and preprocessing the original hydropower station data to obtain to-be-extracted data; performing attribute feature recognition on the to-be-extracted data, and performing data classification based on an attribute feature recognition result to obtain a plurality of data sets; calling a corresponding entity extraction model based on each data set, and outputting an entity extraction result based on the corresponding entity extraction model; and based on the entity extraction result of each data set, performing same entity merging, and outputting the entity extraction result of the original hydropower station data. According to the scheme, the characteristics of complexity and diversity of hydropower station data are effectively handled, accurate entity identification and unified management are realized, and a high-quality data basis is provided for subsequent knowledge graph construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graph construction, and specifically to a method for extracting entities from hydropower station data and a system for extracting entities from hydropower station data. Background Art

[0002] With the rapid development of modern information technology, hydropower stations, as important clean energy infrastructure, have accumulated a large amount of data in their operation management and design evaluation. These data cover a variety of types, including engineering design reports, operation monitoring records, environmental assessment documents, etc., in the form of unstructured natural language texts, as well as semi-structured or structured data such as tables and drawings. These data contain rich engineering information, real-time monitoring parameters and environmental characteristics, and are an important basis for knowledge management and intelligent applications of hydropower stations. However, to extract meaningful knowledge from these complex and diverse data and build an efficient knowledge graph, accurate entity extraction technology is essential.

[0003] Existing entity extraction methods have made some progress in general fields, usually using rule-based or deep learning-based models. These methods can handle common named entities (such as names of people, places, and organizations) well, but their adaptability is insufficient when facing professional data in the field of hydropower stations. On the one hand, hydropower station data presents the characteristics of diversified domain terminology and complex data forms. Engineering design data contains equipment names, technical parameters, etc., operation monitoring data involves dynamic time series information, and environmental assessment data contains complex content such as geographic entities and ecological indicators. On the other hand, the entity characteristics of different types of data vary significantly, and general entity extraction models do not work well when processing such data. In addition, the column headers and cell contents in tabular data usually lack unified specifications. The existing solutions have simple processing methods for tabular data, making it difficult to accurately classify and match them to appropriate entity extraction models.

[0004] In addition, there are multiple ways to express entities in hydropower station data. For example, the same equipment or location may be described using different terms or names. This diverse expression leads to a lack of uniformity in the existing methods when dealing with repeated entities, making it difficult to achieve accurate merging, thus affecting the integrity and consistency of the knowledge graph. At the same time, the problem of insufficient classification and model matching for data in specific fields has not been effectively solved. Existing methods usually use a single model to process different types of data, lack a differentiation strategy based on data characteristics, and cannot meet the actual needs of hydropower station data.

[0005] Therefore, an improved entity extraction scheme is urgently needed to address the special data characteristics in the hydropower station field. Summary of the invention

[0006] The purpose of the embodiments of the present invention is to provide a method and system for entity extraction of hydropower station data, so as to at least solve the problem of insufficient extraction result accuracy existing in existing entity extraction schemes for hydropower station data.

[0007] In order to achieve the above-mentioned purpose, the first aspect of the present invention provides an entity extraction method for hydropower station data, the method comprising: collecting original hydropower station data, and performing preprocessing on the original hydropower station data to obtain data to be extracted; performing attribute feature recognition on the data to be extracted, and performing data classification based on the attribute feature recognition results to obtain multiple data sets; calling a corresponding entity extraction model based on each data set to output an entity extraction result based on the corresponding entity extraction model; based on the entity extraction results of each data set, merging the same entities and outputting the entity extraction result of the original hydropower station data.

[0008] Optionally, preprocessing is performed on the original hydropower station data, including: performing word segmentation, stop word removal and dimension unification on the original text; performing format parsing on the table data, extracting column headers and establishing a mapping relationship with the cell content; for chart data, identifying text information in the chart through OCR technology, and extracting entity features in combination with position relationship annotations.

[0009] Optionally, attribute feature recognition is performed on the extracted data, and data classification is performed based on the attribute feature recognition results to obtain multiple data sets, including: preliminary structured analysis of the text and tabular data in the extracted data to divide them into engineering design data, operation monitoring data and environmental assessment data; wherein, during the classification process, a support vector machine model is used to perform semantic feature classification on the text content to extract key descriptive features; column headers and cell contents in the tabular data are assigned to corresponding categories using specific rules; wherein the specific rules are any one or more of keyword matching, regular expression matching and semantic similarity analysis.

[0010] Optionally, the extraction rules of the entity extraction model of the engineering design data are as follows: a rule-based entity recognition model matches standard format data in the engineering design data through a predefined regular expression template; based on a domain-specific terminology dictionary, entity extraction is performed on project names, equipment names, and design parameters in the engineering design data; for engineering design data with non-standard expressions, sentence structure parsing is performed based on the dependency syntactic analysis method to supplement the identification of missed entity information.

[0011] Optionally, the extraction rules of the entity extraction model of the operation monitoring data are as follows: based on a sequence labeling model combining a bidirectional long short-term memory network and a conditional random field, entity labeling is performed on the time series data in the operation monitoring data; wherein, during model training, the entity extraction model of the operation monitoring data uses time labels as specific features; and long texts are segmented based on a windowed sliding technique to output extracted entities with time series attributes based on the segmentation results.

[0012] Optionally, the extraction rules of the entity extraction model of the environmental assessment data are: constructing a nested entity recognition model, extracting nested entities in the central environmental assessment data based on a sparse representation method; through multi-level decoding, labeling and associating geographic entities and related environmental factors respectively as extracted entities; wherein, during the decoding process, entity extraction is performed in combination with a self-attention mechanism.

[0013] Optionally, the entity extraction results based on each data set are used to merge identical entities, including: performing similarity matching on entities extracted from different data sets based on an entity similarity calculation method based on hash vector representation; identifying entity names whose semantic relevance is greater than a preset relevance threshold by calculating cosine similarity, and grouping them into the same entity; during the merging process, introducing a confidence scoring mechanism to give priority to dictionary matching results or entities whose occurrence frequency is greater than a preset frequency as the final merging results.

[0014] Optionally, after completing the entity merger, the method also includes: entity linking the merged entity with existing entities in the domain knowledge base, and automatically generating preliminary attributes for unmatched entities; using a knowledge graph-based reasoning engine to check the semantic consistency of the merged entities, and correcting entities with attribute conflicts or boundary errors; for unmatched entities, pushing them to the user end, and recovering user confirmation results based on the user end, and performing corresponding entity corrections based on the confirmation results.

[0015] A second aspect of the present invention provides an entity extraction system for hydropower station data, the system comprising: a collection unit, used to collect original hydropower station data, and perform preprocessing on the original hydropower station data to obtain data to be extracted; a classification unit, used to perform attribute feature recognition on the data to be extracted, and perform data classification based on the attribute feature recognition results to obtain multiple data sets; an entity extraction unit, used to call a corresponding entity extraction model based on each data set to output an entity extraction result based on the corresponding entity extraction model; an output unit, used to merge the same entities based on the entity extraction results of each data set, and output the entity extraction result of the original hydropower station data.

[0016] On the other hand, the present invention provides a computer-readable storage medium having instructions stored thereon, which, when executed on a computer, enables the computer to execute the above-mentioned entity extraction method of hydropower station data.

[0017] Through the above technical scheme, the scheme of the present invention standardizes the input data format by collecting and preprocessing the original hydropower station data, laying the foundation for subsequent processing. Through attribute feature recognition, the data is accurately classified, the data is divided into multiple categories, and the corresponding entity extraction model is called according to the data features of different categories, which improves the pertinence and accuracy of entity recognition. The classified data set can be matched with a more suitable model processing, thereby improving the accuracy and comprehensiveness of the entity extraction results. Finally, by merging the same entities in the extraction results of each data set, the problem of repeated entities caused by diversified expressions is solved, ensuring the consistency and integrity of the extraction results. The overall method effectively responds to the complex and diverse characteristics of hydropower station data, realizes accurate entity recognition and unified management, and provides a high-quality data foundation for the subsequent construction of knowledge graphs.

[0018] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings:

[0020] Figure 1 It is a flowchart of the steps of a method for extracting entities from hydropower station data provided by one embodiment of the present invention;

[0021] Figure 2 It is a system structure diagram of a hydropower station data entity extraction system provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0022] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.

[0023] Figure 1 FIG. 1 is a flow chart of a method for extracting entities from hydropower station data provided by an embodiment of the present invention. Figure 1 As shown, an embodiment of the present invention provides a method for extracting entities from hydropower station data, the method comprising:

[0024] Step S10: collecting original hydropower station data, and performing preprocessing on the original hydropower station data to obtain data to be extracted.

[0025] Specifically, preprocessing is performed on the original hydropower station data, including: performing word segmentation, stop word removal, and dimension unification processing on the original text in sequence; performing format parsing on the tabular data, extracting column headers, and establishing a mapping relationship with the cell content; for chart data, identifying the text information in the chart through OCR technology, and annotating and extracting entity features in combination with the positional relationship.

[0026] In the embodiment of the present invention, standardization processing is performed on the original text data. The processing of the text includes three main steps: word segmentation processing, stop word removal processing, and dimension unification processing. Word segmentation processing is for the special needs of Chinese texts, which divides continuous Chinese character sequences into independent word units. Through word segmentation, key terms in the text can be effectively identified, such as "installed capacity", "basin area", etc., enhancing the semantic expression ability of the text. Stop word removal processing reduces interference information and improves the recognition efficiency of subsequent extraction models by removing high-frequency words without practical meaning such as "of", "and". Dimension unification processing further standardizes the numerical expression forms in the text, for example, unifying "kilowatt" and "megawatt" into standard units, thereby avoiding data errors or duplicate recognition problems caused by inconsistent units. These text processing steps not only enhance the consistency of the corpus but also provide a clear context for entity recognition in specific fields.

[0027] Furthermore, format parsing and structuring processing are performed on the tabular data. The tables in the hydropower station data often record key parameters and indicators, such as water level, flow rate, power generation, etc., and have highly structured characteristics. In the preprocessing step, by analyzing the column headers and cell content of the table, a mapping relationship between the two is established. Specifically, the column headers represent the attributes of the data (such as "water level", "time"), while the cell content records the specific parameter values. Through this mapping relationship, not only can clear feature positioning be provided for subsequent entity extraction, but also classification errors caused by non-standard table structures or ambiguous headers can be avoided. For complex structures in tabular data, such as merged cells or cross-list headers, effective information is extracted and a standardized table is reconstructed through rule parsing and pattern matching.

[0028] Furthermore, for the chart data, OCR (optical character recognition) technology is used to extract the text information therein. Hydropower station data often contain design drawings, trend charts and schematic diagrams. These charts carry a lot of important information, but they usually exist in the form of images and cannot be processed directly. Through OCR technology, the text content in the chart can be identified, and the key information can be structured and extracted in combination with the specific positional relationship marked in the figure. For example, in the design drawings of the hydropower station, OCR can identify descriptions such as "dam height: 120 meters" or "basin area: 500 square kilometers" and extract them into two entities of "dam height" and "basin area" and their corresponding values. Combined with the annotation of the positional relationship, the contextual association of the entity can be further judged, such as aligning the annotation in the figure with the table or text content to ensure the completeness and accuracy of the extraction results.

[0029] Based on the scheme of the present invention, the information expression in the original text is made more standardized through word segmentation, stop word removal and dimension unification, providing high-quality input for subsequent semantic analysis and entity extraction. By establishing a mapping relationship between column headers and cell contents, the attributes and semantics of the data are clarified, so that the important information in the table can be accurately classified and matched to a suitable entity extraction model. OCR technology is used to extract text information in charts, and structured output is achieved in combination with position annotations, so that important information implicit in the charts can be effectively extracted, making up for the lack of image data processing capabilities of traditional methods. In the preprocessing step, by cleaning up redundant information and standardizing the data format, the interference of invalid data in the original data on the subsequent extraction process is effectively reduced, and the overall extraction accuracy and efficiency are improved. The preprocessed data has a unified structural form and feature expression, which is convenient for the subsequent merging of entity recognition results from multiple data sources and the construction of knowledge graphs.

[0030] Step S20: Attribute feature recognition is performed on the data to be extracted, and data classification is performed based on the attribute feature recognition results to obtain multiple data sets.

[0031] Specifically, a preliminary structured analysis is performed on the text and table data in the extracted materials to divide them into engineering design data, operation monitoring data and environmental assessment data; wherein, in the classification process, a support vector machine model is used to perform semantic feature classification on the text content and extract key descriptive features; the column headers and cell contents in the table data are assigned to corresponding categories using specific rules; wherein the specific rules are any one or more of keyword-based matching, regular expression-based matching and semantic similarity analysis.

[0032] In an embodiment of the present invention, a preliminary structured analysis is performed on the text and table data in the extracted data. Text data includes various forms such as engineering design reports, operation logs, and environmental assessment documents. Tabular data usually records key operating parameters, design indicators, and evaluation data. Through structured analysis, the explicit semantic features of text and tables can be extracted to provide accurate input for classification. For example, for text data, the core words in the sentence are extracted during the analysis process, such as "installed capacity", "power generation", "pollution index", etc.; for tabular data, the analysis process extracts the column headers and the corresponding cell contents, and converts the structural information in the table into the form of attribute-value pairs, such as "water level: 80 meters", "watershed area: 500 square kilometers".

[0033] In the classification process, the support vector machine (SVM) model is used to classify the semantic features of the text content. Support vector machine is a commonly used supervised learning method, which is suitable for data classification tasks with small samples and high-dimensional features. In this process, the text is first preprocessed, including word segmentation, stop word removal and TF-IDF vectorization, to convert the text into a feature vector that can be processed by the SVM model. Then, combined with the classification requirements in the field of hydropower stations, the SVM model is trained to perform semantic classification on the text data. For example, text containing "dam height" and "unit capacity" is classified as engineering design data, text containing "real-time water level" and "power generation" is classified as operation monitoring data, and text containing "pollution" and "ecological assessment" is classified as environmental assessment data. Through this classification step, the text data can be accurately assigned to the corresponding category, providing high-quality input for the subsequent entity extraction model.

[0034] For tabular data, a specific rule-based method is used in the classification process, combining three strategies: keyword matching, regular expression matching, and semantic similarity analysis to ensure efficient parsing and classification of complex table structures. First, the keyword matching method is used to preliminarily screen the column headers and cell contents of the table through a predefined domain vocabulary. For example, table columns containing keywords such as "watershed" and "water level" are directly classified as environmental assessment data. Secondly, regular expression matching is used to process content with specific formats in the table, such as "^\d+m$" matches water level information, and "^\d+m3 / s$" matches flow data. These columns are assigned to the operation monitoring data category. Finally, for some complex or ambiguous table contents, a semantic similarity analysis method is used. Semantic vectors of column headers and domain keywords are generated through BERT or Word2Vec, the similarity is calculated, and the table columns are assigned to the category closest to their semantics. For example, the vector of the "rainfall" column header has the highest similarity with the central vector of the "environmental assessment" category, so it is classified as environmental assessment data.

[0035] The entire classification process combines multiple rules with machine learning methods to fully explore the semantic information of the data to be extracted and accurately divide it into three categories: engineering design data, operation monitoring data, and environmental assessment data. This separation method provides the possibility of matching each type of data with a more suitable entity extraction model, significantly improving the accuracy of entity recognition.

[0036] Based on the scheme of the present invention, by combining the SVM model and multiple rules, the classification accuracy of text and table data has been significantly improved. Different types of data are accurately assigned to engineering design, operation monitoring and environmental assessment categories, laying a high-quality foundation for the subsequent entity extraction model. The introduction of specific rules, especially keyword matching and regular expression matching, solves the classification difficulties caused by the diversity of table column headers and cell contents, and ensures the efficient processing of table data. By processing complex and ambiguous table column headers through semantic similarity analysis, the classification method can adapt to the diverse characteristics of data in the hydropower station field, further improving the robustness and coverage of classification. The classified data sets call targeted entity extraction models according to the category characteristics, avoiding the adaptability problems that occur when a single model processes diversified data, thereby greatly improving the accuracy and efficiency of entity recognition. After the complex data to be extracted is clearly divided into three data sets, the subsequent knowledge graph construction process is more organized and targeted, and the final generated graph data is more complete and consistent.

[0037] Step S30: calling a corresponding entity extraction model based on each data set to output an entity extraction result based on the corresponding entity extraction model.

[0038] Specifically, the extraction rules of the entity extraction model of the engineering design data are as follows: a rule-based entity recognition model matches standard format data in the engineering design data through a predefined regular expression template; based on a domain-specific dictionary, entity extraction is performed on the project name, equipment name, and design parameters in the engineering design data; for engineering design data with non-standard expressions, sentence structure parsing is performed based on the dependency syntactic analysis method to supplement the identification of missed entity information.

[0039] In an embodiment of the present invention, engineering design data, as an important part of hydropower station data, usually contains a large amount of key information describing the project name, equipment name and design parameters. These data provide the core foundation for the construction of the knowledge graph. However, the expression of engineering design data is diverse, including both formatted structured data and non-standardized natural language descriptions. In view of these characteristics, the entity extraction model of engineering design data adopts rule-based entity recognition technology, combined with regular expression templates, domain-specific term dictionaries, and dependency syntax analysis methods, to form an accurate and efficient entity extraction process.

[0040] The rule-based entity recognition model matches the standard format content in the engineering design data through a predefined regular expression template. These contents are usually expressed in a fixed grammatical structure or format. For example, structured descriptions such as "dam height: 120 meters" and "unit capacity: 1000MW" can be accurately matched through regular expressions. For example: for data describing the dam height, the regular expression ^\s*dam height\s*[::]? \s*\d+\s*meters\s*$ can match content similar to "dam height: 120 meters". For data describing installed capacity, the regular expression ^\s*unit capacity\s*[::]? \s*\d+\s*(KW|MW)\s*$ can match the expression of "unit capacity: 1000MW" or "unit capacity: 500KW". This method has high efficiency and accuracy in recognizing standard format data, and can quickly extract common engineering design parameters.

[0041] Furthermore, the model combines the domain-specific terminology dictionary to identify professional terms such as project names and equipment names. There are a large number of domain terms and proper nouns in hydropower station projects, such as "left bank generator set" and "spillway gate". By establishing a domain-specific terminology dictionary and mapping these names to standardized entity categories, the extraction model's ability to recognize professional terms can be significantly improved. For example, when "left bank unit put into operation" appears in the sentence, the dictionary can directly identify "left bank unit" as an equipment entity and mark the category as "unit".

[0042] Furthermore, for engineering design data with non-standard expressions, the dependency syntactic analysis method is used to perform sentence structure analysis to supplement the identification of missing entity information. In engineering design documents, some data are expressed in complex or irregular natural languages. For example, in "the dam site is located above the main river channel, and the engineering design height is 120 meters", "dam site" and "design height" are entities, but their expressions do not have a fixed format. Dependency syntactic analysis determines the key entities and their attributes by parsing the subject-predicate-object, modifier and other relationships in the sentence. For example, through parsing, "dam site" can be extracted as the subject, "design height" can be extracted as the modifier, and the value "120 meters" can be used as its corresponding attribute. This method can make up for the shortcomings of rule matching and achieve more flexible processing of non-standard data.

[0043] Through the combination of the above three technologies, the entity extraction model of engineering design data can comprehensively cover formatted data, domain terminology and non-standard descriptions, ensuring high accuracy and high recall rate of entity recognition.

[0044] Preferably, the extraction rules of the entity extraction model of the operation monitoring data are as follows: based on a sequence labeling model combining a bidirectional long short-term memory network and a conditional random field, entity labeling is performed on the time series data in the operation monitoring data; wherein, during model training, the entity extraction model of the operation monitoring data uses time labels as specific features; and long texts are segmented based on a windowed sliding technique to output extracted entities with time series attributes based on the segmentation results.

[0045] In an embodiment of the present invention, operation monitoring data is an important data category in the field of hydropower stations, and usually contains dynamic time series information, such as water level, flow, power generation, and temperature. These data record the parameters that change in real time during the operation of the hydropower station, and are crucial to monitoring and managing the operation of the hydropower station. However, operation monitoring data usually exists in the form of text or logs, with complex content and strong time attributes, and traditional entity extraction methods are difficult to adapt to this feature. To solve this problem, the entity extraction model of operation monitoring data adopts a sequence annotation model that combines a bidirectional long short-term memory network (BiLSTM) with a conditional random field (CRF), and combines the time label characteristics and window sliding technology to specifically improve the accuracy and time series correlation of entity extraction.

[0046] The entity extraction model for operation monitoring data adopts the architecture of combining BiLSTM and CRF. As a deep learning model, BiLSTM can capture the dependencies between the context in sequence data and is particularly suitable for processing time series data. Through the bidirectional network, the model can simultaneously obtain the contextual information before and after the current entity to ensure accurate judgment of the entity boundary. For example, for descriptions such as "the water level reaches 80 meters in December 2023" in the log, BiLSTM can simultaneously use the time feature of "December 2023" and the contextual relationship of "the water level reaches 80 meters" to effectively extract the time entity and the numerical entity. Subsequently, the CRF layer further uses the dependencies between the predicted labels to optimize the sequence labeling results. For example, through the CRF layer, it can be ensured that the time entity is reasonably followed by a dynamic numerical entity (such as water level or flow) to avoid unreasonable labeling results.

[0047] Furthermore, the model uses time tags as specific features during the training phase to enhance the perception of time series attributes. In operational monitoring data, time information such as "December 2023" or "8 a.m." is often a key annotation target, and is also the contextual condition for other dynamic data. Therefore, during model training, time tags are explicitly added to the input vector as a specific feature, giving the model a stronger sense of time. For example, through the explicit representation of time features, the model can accurately associate "the water level reaches 80 meters" with "December 2023", thereby generating a complete entity with time series attributes.

[0048] Furthermore, in order to cope with the long text characteristics of the operation monitoring log, the model uses window sliding technology to segment the text. Operation monitoring data often contains large sections of continuous time series records. Directly processing the entire text may cause the model input to be too long, thus affecting the recognition effect. Through the window sliding technology, the text is segmented into segments of fixed length, for example, each segment contains 50 characters. The processing results of each segment are merged through overlapping areas to ensure entity integrity. For example, the long sentence "The water level reaches 80 meters at 8 am on December 1, 2023, and the flow is 200 cubic meters / second" can be divided into two windows: "The water level reaches" and "reaches 80 meters at 8 am on December 1, 2023, and the flow is 200 cubic meters / second". This method not only avoids model input overload, but also ensures the continuity of entities across windows. Through the above technical means, the entity extraction model of operation monitoring data can not only accurately extract dynamic operation data, but also effectively associate time series characteristics, providing high-quality basic data for subsequent analysis and application.

[0049] Preferably, the extraction rules of the entity extraction model of the environmental assessment data are: constructing a nested entity recognition model, extracting nested entities in the central environmental assessment data based on a sparse representation method; through multi-level decoding, labeling and associating geographic entities and related environmental factors respectively as extracted entities; wherein, during the decoding process, entity extraction is performed in combination with a self-attention mechanism.

[0050] In an embodiment of the present invention, environmental assessment data is an important component of hydropower station data, mainly including key information such as geographical entities (such as river basins, ecological protection areas) and environmental factors (such as pollution indicators, animal and plant populations). These data are usually presented in the form of text or reports, with highly nested and complex semantic relationships. For example, in "the ecological protection area in the Yangtze River Basin contains multiple polluted areas", "Yangtze River Basin" is a geographical entity, "ecological protection area" is its sub-entity, and "polluted area" is an environmental factor. In view of this feature, the entity extraction model of environmental assessment data adopts nested entity recognition technology, which accurately extracts nested entities and annotates their association relationships through sparse representation methods, self-attention mechanisms and multi-level decoding.

[0051] The model uses sparse representation methods to centralize nested entities. Sparse representation methods reduce feature redundancy and convert text content into sparse feature vectors, which helps the model capture the core semantics between nested entities more efficiently. For example, in "The ecological protection area in the Yangtze River Basin contains multiple polluted areas", sparse representation can mark "Yangtze River Basin" and "Ecological Protection Area" as parent-child relationships, and further associate the environmental factor "polluted area". Through this method, complex nested semantics are simplified into a hierarchical representation, which effectively solves the dimensionality problem of nested entity recognition.

[0052] Furthermore, the model uses a multi-level decoding strategy to label and associate nested entities and their relationships. Multi-level decoding decomposes the extraction task into multiple stages, and each stage decodes entities at different levels. For example, the first stage of decoding identifies geographical entities such as "Yangtze River Basin"; the second stage of decoding identifies its sub-entities such as "ecological protection area"; the third stage of decoding associates environmental factors such as "pollution indicators". This layer-by-layer decoding strategy can not only clearly express the hierarchical structure of entities, but also ensure that the semantic relationships between different entities are accurately labeled. For example, through the decoding process, the model can associate the "Yangtze River Basin" with the "ecological protection area", and further associate the "pollution area" as an environmental factor with the "ecological protection area".

[0053] During the decoding process, the model introduces a self-attention mechanism to enhance the ability to capture long-distance semantic relationships. The self-attention mechanism can dynamically adjust the model's attention to different text fragments, and is especially suitable for situations where the distance between nested entities is far. For example, in "pollutants exceed the standard in the Yangtze River Basin Ecological Protection Area", there is a distant semantic relationship between "pollutants" and "Yangtze River Basin", but through the self-attention mechanism, the model can effectively capture this association and achieve accurate labeling. In addition, the self-attention mechanism can also identify the mutual influence between nested entities, such as the definition of "protected area" is constrained by the geographical scope of the "Yangtze River Basin". Through the above techniques, the model can efficiently handle complex nested semantic structures in environmental assessment data and ensure the hierarchical and semantic integrity of entity extraction results.

[0054] Step S30: Based on the entity extraction results of each data set, the same entities are merged and the entity extraction results of the original hydropower station data are output.

[0055] In an embodiment of the present invention, a method for calculating entity similarity based on hash vector representation is used to perform similarity matching on entities extracted from different data sets; by calculating cosine similarity, entity names whose semantic relevance is greater than a preset relevance threshold are identified and grouped as the same entity; during the merging process, a confidence scoring mechanism is introduced to give priority to dictionary matching results or entities whose occurrence frequency is greater than a preset frequency as the final merging results.

[0056] In an embodiment of the present invention, in the hydropower station data, due to the diversity of data sources and different expressions, different names or descriptions representing the same entity may appear in different data sets. For example, "Three Gorges Dam" and "Three Gorges Project Dam" may refer to the same entity; "left bank unit" and "left bank generator unit" are also different expressions of the same equipment. In order to ensure the uniqueness and consistency of entities in the knowledge graph, it is necessary to merge the entity results extracted from each data set for the same entity. In an embodiment of the present invention, an entity similarity calculation method based on hash vectorization representation is combined with cosine similarity and confidence scoring mechanism to achieve efficient and accurate entity merging.

[0057] The extracted entities are vectorized and encoded through hash vectorization representation. This method maps the entity name into a sparse vector of fixed length, which not only reduces the computational complexity of high-dimensional features, but also retains the semantic information of the entity. For example, after hash vectorization, the vectors generated by "Three Gorges Dam" and "Three Gorges Project Dam" have a certain similarity, which provides a basis for subsequent similarity calculations.

[0058] Furthermore, the semantic relevance of the identified entities is calculated through cosine similarity. In vector space, cosine similarity reflects the directional similarity of two vectors and is suitable for evaluating the semantic relevance of text data. For any two entity vectors, when their cosine similarity is greater than the preset relevance threshold, the model considers that the two entities may be the same entity. For example, when the similarity between "watershed area" and "river basin area" exceeds the set threshold (such as 0.8), the system marks them as semantically related and processes them further.

[0059] In order to ensure the accuracy of the merged results, a confidence scoring mechanism is introduced in the entity merging process. First, for entities that pass the dictionary matching results, a higher confidence score is assigned, because dictionary matching usually has a higher accuracy. Secondly, for entities that fail to directly match the dictionary, by counting their frequency of appearance in the data set, entities with higher frequencies are given a higher confidence. For example, if "left bank unit" appears frequently in multiple data sets, and "left bank power generation equipment" appears less frequently, "left bank unit" is preferred as the merged result. In addition, the reliability of confidence calculation can be further improved through contextual information, such as combining the attributes or associations of the entities to perform secondary verification of the results.

[0060] Through the above method, the same entity merging not only solves the entity redundancy problem caused by diversified expressions, but also ensures the semantic integrity of the entity and the unity of the knowledge graph. For example, the final output merged result classifies "left bank unit" and "left bank generator set" as the same entity and marks it as "equipment type: generator set", ensuring the consistency of the entity in different scenarios.

[0061] Preferably, after completing the entity merger, the method also includes: entity linking the merged entity with the existing entities in the domain knowledge base, and automatically generating preliminary attributes for the unmatched entities; using a reasoning engine based on a knowledge graph to check the semantic consistency of the merged entities, and correcting entities with attribute conflicts or boundary errors; for unmatched entities, pushing them to the user end, and recovering the user confirmation results based on the user end, and performing corresponding entity corrections based on the confirmation results.

[0062] In an embodiment of the present invention, after completing the entity merger, in order to further improve the semantic integrity and consistency of the data, the method improves the attributes and boundary information of the merged entities through entity linking, knowledge graph reasoning and user confirmation mechanism, thereby ensuring the accuracy and practicality of the knowledge graph.

[0063] The merged entity is entity-linked with the existing entities in the domain knowledge base. The domain knowledge base contains the standardized entities and their attributes and relationships commonly used in hydropower station projects, such as dam site information, equipment names and operating parameters. Entity linking matches the entities in the knowledge base by comparing the names, attributes and semantic features of the merged entities. During the matching process, the similarity calculation method based on embedding representation is adopted to calculate the cosine similarity between the merged entity and the knowledge base entity through the semantic embedding vector generated by BERT. When the similarity is higher than the preset threshold, the two are considered to match. For example, the merged entity "Three Gorges Project of the Yangtze River" matches "Three Gorges of the Yangtze River" in the knowledge base and inherits the relevant standard attributes (such as "geographic type: watershed engineering"). For entities that fail to match the knowledge base, the system marks them as newly added entities and automatically generates preliminary attributes (such as source, preliminary category, etc.) for further improvement.

[0064] Furthermore, the semantic consistency of the merged entities is checked through a reasoning engine based on the knowledge graph. The reasoning engine uses the entity-attribute-relationship structure in the knowledge graph to automatically check the semantic consistency of the merged entities. For example, if the "dam height" in the associated attribute of an entity is "200 meters", and its geographical location is marked as "plain area", the reasoning engine can identify this contradiction and mark it as an attribute conflict. In addition, the reasoning engine can also detect entity boundary errors, such as "left bank unit" and "right bank unit" are incorrectly labeled as the same entity in some contexts. For the detected semantic conflicts and boundary problems, the system automatically corrects them, such as by reconstructing the attribute range of the entity or adjusting its association relationship to ensure the semantic consistency of the merged entity.

[0065] Furthermore, for entities that still cannot be matched or have great uncertainty, the system pushes them to the user end and collects user confirmation results to complete the entity information. Users review and supplement the pushed entities through the interactive interface, such as confirming whether "Longtan Reservoir belongs to the Yangtze River Basin Project" or supplementing the technical parameters of new equipment. After the user's confirmation results are collected, the system uses them as new training data to further optimize the entity recognition and merging model. In addition, based on the feedback results confirmed by the user, the system can dynamically adjust the knowledge base, add new entities and their attributes, and enhance the adaptability of the knowledge base. Through the above steps, the results after the entity merger have been further verified, supplemented and optimized, providing a high-quality semantic foundation for the construction of the knowledge graph.

[0066] Figure 2 FIG. 1 is a system structure diagram of a hydropower station data entity extraction system provided by an embodiment of the present invention. Figure 2 As shown, an embodiment of the present invention provides an entity extraction system for hydropower station data, the system comprising: a collection unit, used to collect original hydropower station data, and perform preprocessing on the original hydropower station data to obtain data to be extracted; a classification unit, used to perform attribute feature recognition on the data to be extracted, and perform data classification based on the attribute feature recognition results to obtain multiple data sets; an entity extraction unit, used to call a corresponding entity extraction model based on each data set to output an entity extraction result based on the corresponding entity extraction model; an output unit, used to merge the same entities based on the entity extraction results of each data set, and output the entity extraction result of the original hydropower station data.

[0067] The embodiment of the present invention further provides a computer-readable storage medium, on which instructions are stored, which, when executed on a computer, enable the computer to execute the above-mentioned entity extraction method of hydropower station data.

[0068] Those skilled in the art will understand that all or part of the steps in the method for implementing the above-mentioned embodiments can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including several instructions for making a single-chip microcomputer, a chip or a processor (processor) perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0069] The optional embodiments of the present invention are described in detail above in conjunction with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the technical concept of the embodiments of the present invention, the technical scheme of the embodiments of the present invention can be subjected to a variety of simple modifications, and these simple modifications all belong to the protection scope of the embodiments of the present invention. It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the embodiments of the present invention will not further describe various possible combinations.

[0070] In addition, various embodiments of the present invention may be arbitrarily combined, and as long as they do not violate the concept of the embodiments of the present invention, they should also be regarded as the contents disclosed in the embodiments of the present invention.

Claims

1. A method for extracting entities from hydropower station data, characterized in that: The method comprises: Collecting original hydropower station data, and performing preprocessing on the original hydropower station data to obtain data to be extracted; Attribute feature recognition is performed on the extracted data, and data classification is performed based on the attribute feature recognition results to obtain multiple data sets; Calling a corresponding entity extraction model based on each data set to output an entity extraction result based on the corresponding entity extraction model; Based on the entity extraction results of each data set, the same entities are merged and the entity extraction results of the original hydropower station data are output.

2. The entity extraction method of hydropower station data according to claim 1 is characterized in that: Preprocessing the original hydropower station data includes: The original text is processed by word segmentation, stop word removal and dimension unification; Perform format parsing on table data, extract column headers and establish a mapping relationship with cell content; For chart data, OCR technology is used to identify text information in the image, and entity features are extracted in combination with positional relationship annotations.

3. The entity extraction method of hydropower station data according to claim 1 is characterized in that: Attribute feature recognition is performed on the extracted data, and data classification is performed based on the attribute feature recognition results to obtain multiple data sets, including: The text and table data in the extracted materials are preliminarily structured and analyzed, and divided into engineering design data, operation monitoring data and environmental assessment data; among them, In the classification process, the support vector machine model is used to classify the semantic features of the text content and extract key descriptive features; The column headers and cell contents in the table data are assigned to corresponding categories using specific rules; The specific rules are any one or more of keyword matching, regular expression matching and semantic similarity analysis.

4. The entity extraction method of hydropower station data according to claim 3 is characterized in that: The extraction rules of the entity extraction model of the engineering design data are: The rule-based entity recognition model matches standard format data in engineering design data through predefined regular expression templates; Based on the domain-specific term dictionary, entity extraction is performed on the project names, equipment names, and design parameters in the engineering design data; For engineering design data with non-standard expressions, sentence structure parsing is performed based on the dependency parsing method to supplement the identification of missing entity information.

5. The entity extraction method of hydropower station data according to claim 3 is characterized in that: The extraction rules of the entity extraction model of the operation monitoring data are: Based on the sequence labeling model combining bidirectional long short-term memory network and conditional random field, the time series data in the operation monitoring data are labeled; The entity extraction model of the operation monitoring data uses the time label as a specific feature during model training; Segment long texts based on windowed sliding technology to output extracted entities with temporal attributes based on the segmentation results.

6. The entity extraction method of hydropower station data according to claim 3 is characterized in that: The extraction rules of the entity extraction model of the environmental assessment data are: Construct a nested entity recognition model to extract nested entities from central environmental assessment data based on sparse representation method; Through multi-level decoding, geographic entities and related environmental factors are labeled and associated as extracted entities; During the decoding process, entity extraction is performed in conjunction with the self-attention mechanism.

7. The entity extraction method of hydropower station data according to claim 1 is characterized in that: The entity extraction results based on each data set are used to merge the same entities, including: An entity similarity calculation method based on hash vector representation is used to perform similarity matching on entities extracted from different data sets; By calculating the cosine similarity, we identify entity names whose semantic relevance is greater than a preset relevance threshold and group them into the same entity. During the merging process, a confidence scoring mechanism is introduced to give priority to dictionary matching results or entities with a frequency greater than a preset frequency as the final merging results.

8. The entity extraction method of hydropower station data according to claim 1 is characterized in that: After completing the entity merger, the method further includes: Perform entity linking between the merged entities and existing entities in the domain knowledge base, and automatically generate preliminary attributes for unmatched entities; Use the reasoning engine based on the knowledge graph to check the semantic consistency of the merged entities and correct entities with attribute conflicts or boundary errors; For entities that cannot be matched, they are pushed to the user end, and based on the user end, the user confirmation result is recovered, and the corresponding entity correction is performed based on the confirmation result.

9. A hydropower station data entity extraction system, characterized in that: The system comprises: A collection unit, used for collecting original hydropower station data, and performing preprocessing on the original hydropower station data to obtain data to be extracted; A classification unit is used to perform attribute feature recognition on the data to be extracted, and classify the data based on the attribute feature recognition results to obtain multiple data sets; An entity extraction unit, used to call a corresponding entity extraction model based on each data set to output an entity extraction result based on the corresponding entity extraction model; The output unit is used to merge the same entities based on the entity extraction results of each data set and output the entity extraction results of the original hydropower station data.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the entity extraction method for hydropower station data as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Knowledge graph generation method and device based on SAR satellite and storage medium

    CN120235231A

  • Special equipment data identification method and device, equipment and medium

    CN120853182A