Method and system for fusing and sharing multi-source heterogeneous data of high-altitude hydropower engineering
Through multi-source data acquisition, preprocessing and knowledge graph modeling, combined with large language models, the data integration and sharing problems of the hydropower engineering data platform throughout the life cycle is solved, efficient integration and management of multimodal data is achieved, and data quality and query efficiency are improved.
Patent Information
- Application Number
- CN202510418921.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-04
AI Technical Summary
The existing hydropower engineering data platforms have shortcomings in data collection, fusion processing and sharing management, especially in the engineering design, construction and other stages, there is little attention to data, lack of targeted design of multimodal data, and the data island problem is serious, making it difficult to achieve efficient integration and sharing throughout the life cycle.
Multi-source data acquisition, preprocessing, knowledge graph modeling and large language modeling technology are used to build a data sharing interactive question-and-answer system to realize the element-level, feature-level and decision-making fusion of structured, BIM, GIS, video and text data, combined with Fourier spectrum analysis and moving average method to clean data, and use search enhancement generation technology to support natural language query.
It improves the organization and interpretability of data, optimizes data quality, and improves data query efficiency. It is suitable for multi-source data management throughout the life cycle of high-altitude hydropower projects, and supports engineering design, construction, operation and operation and maintenance.
Smart Images

Figure CN120256583A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of data fusion. More specifically, the present invention relates to a method and system for multi-source heterogeneous data fusion and sharing in high-altitude hydropower projects. Background Art
[0002] With the continuous development of high-altitude hydropower projects, the technology of multi-source heterogeneous data fusion and sharing has been increasingly widely applied in engineering design, construction, and operation and maintenance. However, there are still many deficiencies in the existing technologies in aspects such as data acquisition, fusion processing, and sharing management, as follows:
[0003] In terms of data scenarios, most existing hydropower project data platforms focus on the data in the project operation stage and pay less attention to the data throughout the entire life cycle such as design and construction stages;
[0004] In terms of data modalities, most existing platform systems are designed around basin monitoring data and lack targeted design for multi-modal data such as text, images, GIS, and BIM;
[0005] In terms of data organization, existing platforms only perform simple integration of data, lack modeling of the semantic relationships between data, and there is a problem of "data islands", which is not conducive to data retrieval, sharing, and scenario applications. Summary of the Invention
[0006] In order to at least solve the technical problems described in the above background art section, the present invention proposes a method and system for multi-source heterogeneous data fusion and sharing in high-altitude hydropower projects, which can be applicable to the management of multi-source data throughout the entire life cycle of high-altitude hydropower projects and can provide reliable data support for engineering design, construction, operation, and maintenance. In view of this, the present invention provides solutions in the following aspects.
[0007] The first aspect of the present invention provides a method for multi-source heterogeneous data fusion and sharing in high-altitude hydropower projects, including: collecting and storing multi-source data, where the multi-source data includes structured data, BIM data, GIS data, video and picture data, and text data; performing data pre-inspection on the multi-source data to ensure that the data format and quality meet relevant requirements; and performing preprocessing on structured monitoring data by using the methods of moving average and Fourier spectrum analysis; constructing a knowledge graph to model the semantic relationships between multi-source heterogeneous data; divided into two stages of defining the schema layer and filling the data layer: in the stage of defining the schema layer, using a top-down method, systematically sorting out the concept hierarchy and semantic relationships of hydropower projects and setting corresponding attributes; in the stage of filling the data layer, using a bottom-up method, designing corresponding knowledge extraction schemes according to data of different formats to complete the fusion and storage of knowledge; combining the large language model LLM and the retrieval augmented generation RAG technology to construct a data sharing interactive question and answer system based on the knowledge graph.
[0008] In one embodiment, collecting and storing multi-source data includes: for structured data with a relatively low update frequency and semi-structured data including spatial data, BIM data, video, and file data, adopting a data regular pulling strategy to achieve data collection and update; for structured data with a relatively high update frequency, adopting a real-time synchronization strategy to achieve data collection and update; for information data such as engineering design and construction based on BIM technology, adopting a digital handover method to complete data collection and update.
[0009] In one embodiment, the method of using moving average and Fourier spectrum analysis to preprocess structured monitoring data includes: for the case where there is obvious periodic noise in the data, using the method based on Fourier spectrum analysis to clean the data:
[0010]
[0011] where X[k] represents the frequency-domain signal, k represents the frequency index, and x[n] represents the discrete time-domain signal;
[0012] For data with accidental errors or systematic error data, use the moving average method to clean the data:
[0013] F MA =(f(x1)+f(x2)+f(x3)+…+f(x N )) / N#(2)
[0014] where F MA represents the value of the variable x i (i = 1, 2, …, N) after moving average calculation, N is the size of the moving average window, and f(x i )(u = 1, 2, …, N) represents the average calculation function.
[0015] In one embodiment, the construction of the knowledge graph includes: defining the schema layer of the knowledge graph, systematically sorting out the concept hierarchy and semantic relationships in the field of high-altitude hydropower engineering, and establishing a complete knowledge system; filling the data layer, converting multi-format data into entities and relationships in the knowledge graph through different knowledge extraction schemes, and performing semantic fusion: for BIM and GIS type data, manually screen the useful information and import it into the knowledge graph; for structured monitoring data, store it in a time series database, and store the API of the database in the attributes of the monitoring object nodes; for unstructured text data, adopt a method based on prompt engineering to extract triples, that is, design prompt words according to requirements, and call the large model API or website to obtain triple information in the text; finally, perform text disambiguation, and save the extracted triples in JSON format; store and manage the knowledge graph, write a Python script, and import the JSON format triples into the Neo4j graph database. Implement complex knowledge graph query requirements through the Cypher query language.
[0016] In one embodiment, the text disambiguation includes: using the FastTex word embedding model to convert the text description of the entity into a vector representation, and using cosine similarity to calculate the semantic similarity of the text:
[0017]
[0018] where e1 and e2 respectively represent the vector representations of the entities, and the larger the value of cos(e1, e2), the higher the semantic similarity of the entities corresponding to e1 and e2; during the knowledge fusion process, merge the entities with similarity higher than the threshold.
[0019] In one embodiment, constructing a data sharing and interactive Q&A system based on the knowledge graph includes:
[0020] For the text q input by the user, use a pre-trained model to perform text embedding on it:
[0021]
[0022] where, z q represents the text vector, and PLM represents the pre-trained model for text embedding;
[0023] Calculate the cosine similarity between z q and the standard question vector in the case library:
[0024]
[0025] where i represents the scenario category;
[0026] Select the scenario with the highest cosine similarity as the classification of the input text, which will determine the rules for subsequent knowledge graph node retrieval, the type of database to be called, and the template of the prompt words;
[0027] Use a word segmentation tool to split the text q input by the user, extract the nouns related to watershed monitoring, and form a keyword set
[0028] The knowledge graph is defined as
[0029] Among them, ε represents the set of entities in the knowledge graph, represents the set of relationships in the knowledge graph;
[0030] Set the threshold of cosine similarity to τ, and then use cosine similarity to match each noun in the list K with the entity names in ε. Nodes with a similarity exceeding τ are selected into the target node set ε t in:
[0031]
[0032] According to different application scenarios i, adopt the corresponding reasoning strategy R i Obtain the relevant information of the target node e in the knowledge graph, including node attributes, triples, etc., to form a subgraph
[0033]
[0034] In the subgraph , the general expression of the node attribute is (entity, attribute name, attribute value), and its template converted to a string is "{entity}'s {attribute name} is {attribute value}"; the general expression of the triple is (entity1, relationship name, entity2), and its template converted to a string is "{entity1}{relationship name}{entity2}";
[0035] After converting all the node attributes and triple information in the graph into strings, the concatenated string is defined as D g ; Define the general specification text obtained from the document database under scenario i as T i , and concatenate the user question, the real-time monitoring data of the target node, the knowledge graph information, and the general scheduling rules into a complete prompt word Prompt(q), and its definition is as follows:
[0036]
[0037] Input this prompt word into the large model to obtain the answer to the user input text q.
[0038] The second aspect of the present invention provides a multi-source heterogeneous data fusion and sharing system for high-altitude hydropower projects, which runs any of the above-mentioned multi-source heterogeneous data fusion and sharing methods for high-altitude hydropower projects.
[0039] The present invention has the following advantages: 1. Multi-source data fusion: The present invention adopts a knowledge graph modeling method to perform element-level, feature-level, and decision-level fusion on structured data, BIM, GIS, video, and text data, improving the organization and interpretability of data. 2. Data quality optimization: Combining Fourier spectrum analysis and the moving average method can effectively clean periodic noise and outliers in monitoring data, improving the accuracy and reliability of data. 3. Efficient data query: Based on the Retrieval-Augmented Generation (RAG) technology and combined with large language models, it supports users to query and retrieve data in natural language form, improving data utilization. 4. Wide range of application scenarios: This method is applicable to the management of multi-source data throughout the life cycle of high-altitude hydropower projects and can provide reliable data support for engineering design, construction, operation, and maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0041] Figure 1 shows a multi-source heterogeneous data fusion and sharing method according to an embodiment of the present invention;
[0042] Figure 2 shows a multi-source data acquisition method according to an embodiment of the present invention;
[0043] Figure 3 shows a data processing flow according to an embodiment of the present invention;
[0044] Figure 4 shows a structured data cleaning flow according to an embodiment of the present invention;
[0045] Figure 5 shows a knowledge graph construction flow according to an embodiment of the present invention;
[0046] Figure 6 shows a schematic of the Retrieval-Augmented Generation method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] It should be understood that the terms "first", "second", "third", "fourth", etc. in the claims, the description and the drawings of the present invention are used to distinguish different objects rather than to describe a specific order. The terms "including" and "comprising" used in the description and claims of the present invention indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0049] It should also be understood that the terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the description and claims of the present invention, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms. It should be further understood that the term "and / or" used in the description and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0050] As used in this specification and the claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0051] The present invention is directed to the scenario of data governance throughout the life cycle of high-altitude hydropower projects. Corresponding data collection, cleaning, and storage solutions are proposed according to different data types. The semantic relationships between multi-source heterogeneous data are modeled based on a knowledge graph, and then the integrated storage and efficient sharing of data are realized.
[0052] The present invention proposes a method and system for multi-source heterogeneous data fusion and sharing in high-altitude hydropower projects, aiming to solve the problems of efficient integration, management, and sharing of various data types throughout the entire life cycle of hydropower projects. This method covers five categories of data, namely structured data, BIM data, GIS data, video (picture) data, and text data. For different data sources, efficient data collection, processing, storage, and query solutions are proposed.
[0053] The method of the present invention first designs a targeted data collection mechanism, including three methods: regular pulling, real-time synchronization, and digital handover, to ensure the integrity and real-time nature of the data. Secondly, Fourier spectrum analysis and the moving average method are used to preprocess the structured monitoring data to eliminate periodic noise and accidental errors and improve data quality. At the data fusion level, a knowledge graph is constructed to model the semantic relationships between multi-source heterogeneous data, and combined with large language model (LLM) and retrieval-augmented generation (RAG) technologies, it supports users to query data through natural language, improving data accessibility and sharing efficiency.
[0054] The following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings.
[0055] In the first aspect of the present invention, a method for multi-source heterogeneous data fusion and sharing in high-altitude hydropower projects is provided. The following will Figure 1 illustrate the method for multi-source heterogeneous data fusion and sharing in hydropower projects of the present invention. According to the Figure 1 , the method for multi-source heterogeneous data fusion and sharing of the present invention includes steps S100 - S400:
[0056] Step S100: Collect and store multi-source data, where the multi-source data includes structured data, BIM data, GIS data, video and picture data, and text data;
[0057] Step S200: Conduct data pre-inspection on the multi-source data to ensure that the data format and quality meet relevant requirements; and use the methods of moving average and Fourier spectrum analysis to preprocess the structured monitoring data;
[0058] Step S300: Construct a knowledge graph to model the semantic relationships between multi-source heterogeneous data; it is divided into two stages: defining the schema layer and populating the data layer: In the stage of defining the schema layer, a top-down method is adopted to systematically sort out the concept hierarchy and semantic relationships of hydropower projects and set corresponding attributes; in the stage of populating the data layer, a bottom-up method is adopted to design corresponding knowledge extraction schemes according to different formats of data, and complete the fusion and storage of knowledge;
[0059] Step S400: Combine the large language model LLM and the retrieval-augmented generation RAG technology to construct a data sharing interactive Q&A system based on the knowledge graph.
[0060] In a preferred embodiment of the present invention, the multi-source data acquisition method in the above step S100 is specifically as follows:
[0061] For the monitoring and management data from sources such as the basin comprehensive monitoring platform and the business management information system, a message service mechanism is adopted to complete data acquisition and on-demand update synchronization, and the basic database is maintained; for the engineering design data from sources such as the engineering digital system, digital handover means are adopted to complete data storage.
[0062] The message service adopted by the present invention includes two parts: regular data pulling and real-time synchronization. The specific steps are as follows:
[0063] 1. In terms of regular data pulling: (1) Data source classification and pulling requirement confirmation. First, confirm the data source classification, which mainly includes structured data (such as databases), semi-structured data (such as JSON, XML), and unstructured data (such as spatial data, BIM data, videos, files). Subsequently, confirm the types of interfaces open to the data source system, such as API, JDBC, FTP, etc., and clarify the data pulling frequency, data volume, and data format. Finally, confirm the metadata scope, including business metadata (such as data source, data business attributes), technical metadata (such as data structure, storage path), operation / management metadata (such as data access rights, update time), etc. (2) Structured data pulling. First, configure the data access interface, using interface protocols such as RESTful API or SOAP. If the business system cannot provide an API interface, the read permission of the database table can be opened, and data can be obtained through SQL queries. The present invention uses DataX as the ETL tool and Airflow to schedule tasks regularly to achieve periodic data pulling, and sets the pulling frequency at the daily / hourly / minute level according to the data volume. (3) Semi-structured and unstructured data pulling. According to the file API interface provided by the business system, semi-structured data is pulled through HTTP / HTTPS. For spatial data, GIS data is obtained in formats such as GeoJSON and Shapefile, and is written into the database or graph data storage through ETL conversion; for BIM data, the IFC format model data is pulled, and at the same time, the structural information and component information are parsed and extracted; for video data, the video data file is pulled, and at the same time, relevant metadata (such as file format, frame rate, recording time, etc.) is obtained. The Kafka message queue is used to achieve metadata synchronization and transmission.
[0064] 2. In terms of real-time data synchronization, the present invention mainly adopts the following three implementation methods: (1) Customized data interface at the receiving end. Customized data receiving interface based on the Restful API open to the data platform. After acquiring monitoring data, the downstream business system (such as the integrated monitoring platform for the watershed) automatically pushes the data to the data platform interface. After receiving the data, the data platform parses the data structure and stores the data in the specified database. (2) Real-time message queue synchronization. After the business end data source produces data, it pushes it to the message queue (the present invention uses Kafka to implement it). The data platform receiving end writes a consumer program to consume messages from the queue and parse structured, semi-structured or unstructured data. For structured data, the message body is directly parsed and written into the data warehouse; for unstructured data, metadata and file location information are extracted, and file pull is initiated after parsing. (3) Real-time database synchronization. The business system opens database access rights to the data platform for data pull. Use tools such as Canal, Maxwell, and FlinkCDC to monitor the database Binlog log in real time. Generate SQL statements for data change operation logs and execute them on the platform end to achieve data synchronization. For structured data, it is directly synchronized to the data asset platform database; for unstructured data, the business system needs to store metadata in a relational database and carry file location information.
[0065] Digital handover is a form of engineering data delivery based on BIM (Building Information Model). Through lightweight, structured and standardized processing, it associates and manages multi-source data generated during planning, design and construction, and provides basic data for engineering life cycle management and big data applications. The specific implementation steps are as follows:
[0066] (1) Digital delivery rule making
[0067] First, define the scope of delivery, covering data from planning and design, material procurement, construction, equipment manufacturing, installation and commissioning. The core object is to build a power plant-level digital twin based on the BIM model. Subsequently, establish the coding and classification rules for deliverables. Establish coding rules for engineering objects, attribute information, documents, models and other data to ensure a unified data structure. Clarify the classification standards for various types of data based on industry specifications such as the "BIM Model Delivery Standard".
[0068] (2) Digital delivery plan formulation
[0069] Focusing on digital delivery, we formulate data integration strategies, convert data from different systems into a unified data format through standard interfaces or ETL tools (Extract, Transform, Load), and associate them with the BIM model. Data types include geographic information data, attribute data, documents, 3D models, and other types of data.
[0070] (3) Information integration and quality verification
[0071] Mainly conduct verification on the accuracy of model data. According to the geometric accuracy levels G1 - G4 in the "Specification for Design and Delivery of Hydropower Engineering Information Model", verify the model accuracy: G1 represents symbolic identification requirements, G2 represents spatial occupancy and rough identification requirements, G3 represents construction and installation processes and procurement requirements, and G4 represents high-precision rendering and manufacturing processing requirements. Generate a "Quality Review Report" based on the review results to ensure that the data meets the delivery standards.
[0072] (4) Digital handover and acceptance
[0073] The scope of data handover includes BIM models, attribute data, engineering drawings, survey data, design reports, digital delivery item change tables, etc. The data format specifications are strictly delivered according to mainstream formats such as IFC, RVT, DWG, DXF, SHP, KML, DOC, XLS, XML, etc. Conduct multi-dimensional verifications including geometric accuracy, attribute information, data integrity, etc., and form a "Digital Acceptance Report". Archive all delivery data and acceptance reports as data assets in the production and operation stage.
[0074] The multi-source data collection plan is as attached Figure 2 shown. According to the different data types and update requirements, design three types of data collection and update plans: regular pull, real-time synchronization, and digital handover.
[0075] Regular data pull: For structured data with a low update frequency (such as data stored in mainstream relational databases like MySQL), through the API interfaces opened by the business system or database read permissions, complete one-time or regular data pull operations. For semi-structured data (such as spatial data, BIM data, video, and file data), the data collection needs to include the file body and its metadata information, and the data interface needs to comply with the technical standards formulated by relevant units.
[0076] Data real-time synchronization: For structured data with a high update frequency, adopt a real-time synchronization strategy to achieve data collection and update. Data real-time synchronization is achieved through a customized receiving-end data interface. Specifically, the business-side data source pushes data to the message queue opened by the platform, and the receiving end of the data asset platform parses the queue messages and writes them into the data warehouse. At the same time, the business system opens the access permission to the backup database to the platform, and by real-time monitoring the database change log, parses and executes the SQL statements corresponding to operations such as addition, deletion, and modification, so as to achieve efficient real-time synchronization of data.
[0077] Digital handover: The information summary of engineering design, construction, etc. based on BIM technology is completed through the digital handover system. The multi-source data generated during the engineering planning, design, and construction processes has little subsequent update requirement. After being lightweighted, structured, and standardized, it is directly handed over to the data platform. This process also establishes associations between the BIM model and relevant engineering data (including structured and unstructured data), thereby forming a basic data system to support the full life cycle management of the project and big data applications.
[0078] In a preferred embodiment of the present invention, the preprocessing and storage of multi-source data in the above step S200 are specifically as follows:
[0079] The preprocessing process of multi-source data is as shown in the appendix Figure 3 After data collection, data pre-inspection is carried out to ensure that the data format and quality meet the relevant requirements. First, the collected data needs to meet the relevant regulations of data storage, exchange, and asset catalog. Secondly, it should follow the data object coding system and standards of the data platform management unit to ensure unified coding. In addition, the data format and precision need to meet the specific requirements of the data platform management unit. Finally, for the spatial data such as drawings, BIM 3D designs, surveying and mapping geographic information (GIS), and equipment and facilities collected by each unit, it should be ensured that they follow the unified spatial reference system of the data platform management unit.
[0080] After completing the data pre-inspection, a processing plan needs to be designed according to the data format. For spatial data, information desensitization needs to be completed according to relevant requirements, and then the data format, geometric features such as scale and precision, and attribute items are unified. Specifically, it includes: First, identify sensitive information, including spatial data such as geographic coordinates, specific engineering locations, and key equipment locations. The coordinate mapping method is adopted, and the real coordinates are mapped to pseudo-coordinates using an encrypted mapping function to achieve data desensitization. Subsequently, the data is converted into standard spatial data formats such as SHP, KML, GEOJSON, DXF, and GML. The scale standard is determined according to the project stage, and the data precision is adjusted to the G1-G4 level that meets the engineering specifications. The attribute field names and formats are standardized according to the "BIM / GIS Data Delivery Standard".
[0081] For BIM data, multi-project integration at different spatial locations and multi-disciplinary integration covering different fields such as geology, hydraulic engineering, and mechanical and electrical engineering are carried out. Specifically, it includes: for BIM data at different locations, using a unified coordinate reference (such as WGS84 or CGCS2000) to unify the spatial positions of BIM models of different projects; using BIM model management tools such as Revit for model splicing and coordinate alignment. Convert data of different disciplines such as geology, hydraulic engineering, and mechanical and electrical engineering into IFC format, unify the coordinate system, create a main model using BIM software, import each discipline model into the main model in sequence, perform model merging under the same coordinate system, carry out hierarchical model management and collision detection to achieve multi-disciplinary integration.
[0082] For data such as videos and documents, multi-format conversion is supported. Structured data are mostly on-site monitoring data obtained through sensors. Due to the influence of uncertain factors at the engineering site, such as power outages, network outages, and sensor failures, data collection is discontinuous or there are mutations. At the same time, due to reasons such as deviation in the buried position of sensors and instrument failures, there are overall numerical deviations or regular fluctuations in the data.
[0083] Therefore, the present invention uses the methods of moving average and Fourier spectrum analysis to preprocess structured monitoring data. The above-mentioned structured monitoring data is a special type of structured data, automatically collected using sensors and remote monitoring devices, usually time-series data, where each data point contains a timestamp and a monitoring value, and is usually stored in a time-series database in the form of timestamp + monitoring value.
[0084] A schematic diagram of the related method is as shown in the appendix Figure 4 shown. For cases where there is obvious periodic noise in the data, such as vibration monitoring of structures and mechanical and electrical equipment, a method based on Fourier spectrum analysis can be used to clean the data. Fourier analysis decomposes the time-series signal into a superposition of sine waves of different frequencies to identify and filter out noise of specific frequencies. In the actual monitoring process, the discrete Fourier transform is often used, as shown in Equation (1), where X[k] represents the frequency-domain signal, k represents the frequency index, and x[n] represents the discrete time-domain signal. During the data cleaning process, by analyzing the spectrogram, the specific frequency components corresponding to the noise can be identified, and the frequency-domain signals within the noise frequency band index range are set to 0, which can effectively reduce the interference of periodic noise and retain the main features of the signal.
[0085]
[0086] In most cases, the accidental errors brought by sudden situations such as on-site power failure and unstable network to the original monitoring data, as well as the systematic errors caused by instrument embedding deviation and monitoring system failure, seriously affect the quality of data acquisition. Accidental errors have the property of mutual compensation and usually conform to the normal distribution; systematic errors have strong regularity. For such monitoring data of cascade reservoirs in a basin, the moving average method is used for data cleaning. On the one hand, the hidden high-frequency accidental errors are eliminated, and on the other hand, the sensors with abnormal values are dynamically identified and calibrated so that they do not participate in the average calculation, thereby reducing the regular systematic errors existing in the monitoring data. The moving average formula is shown in Equation (2), where F MA represents the variable x i (i = 1, 2, …, N) after moving average calculation, N is the size of the moving average window, and f(x i )(i = 1, 2, …, N) represents the average calculation function.
[0087] F MA =(f(x1)+f(x2)+f(x3)+…+f(x N )) / N#(2)
[0088] In a preferred embodiment of the present invention, the construction of the knowledge graph in the above step S300 is specifically as follows:
[0089] The construction process of the knowledge graph is as shown in the appendix Figure 5 and is divided into two stages: defining the schema layer and populating the data layer. In the stage of defining the schema layer, a top-down method is adopted to systematically sort out the concept hierarchy and semantic relationships of hydropower projects and set corresponding attributes. Specifically, it includes:
[0090] (1) Defining the schema layer of the knowledge graph
[0091] Systematically sort out the concept hierarchy and semantic relationships in the field of high-altitude hydropower projects and establish a complete knowledge system. First, identify the key fields involved in the whole life cycle of hydropower projects, such as geology, hydraulic engineering, mechanical and electrical engineering, etc. Then divide the concept levels from large to small. For example, natural objects include rivers and lakes, slopes, etc., engineering objects include dams, units, powerhouses, etc., and social objects include responsible persons, contractors, etc. Establish the semantic relationships between entities. For example, the dam "is located in" the river, the unit "is installed in" the powerhouse, and the responsible person "is responsible for" the equipment operation, etc. Define relevant attribute fields for different types of entities. For example, the attributes of the river include name, length, basin area, etc.
[0092] (2) Populating the data layer
[0093] Through different knowledge extraction schemes, multi-format data is transformed into entities and relationships in the knowledge graph and semantic fusion is carried out. For BIM and GIS type data, useful information is manually screened (such as engineering parameters in BIM design drawings, river lengths in GIS data, etc.) and imported into the knowledge graph. For structured monitoring data, it is stored in a time-series database, and the API of this database is stored in the attributes of the monitoring object nodes. For unstructured text data such as engineering reports and standard specifications, a method based on prompt engineering is used to extract triples, that is, prompt words are designed according to requirements, and the large model API or website is called to obtain triple information in the text segment. Finally, text disambiguation is performed. The extracted triples are saved in JSON format.
[0094] (3) Knowledge Graph Storage and Management
[0095] Write a Python script to import the JSON format triples into the Neo4j graph database. Complex knowledge graph query requirements are realized through the Cypher query language.
[0096] To comprehensively cover the data related to the full life cycle of high-altitude hydropower projects, the schema layer of the knowledge graph designed by the present invention is as Figure 5 shown. The entity types include natural objects such as rivers and lakes, slopes, etc., engineering objects such as dams, units, etc., and social objects such as responsible persons. The relevant object information is stored using the entity attribute list.
[0097] In the stage of filling the data layer, a bottom-up method is adopted. According to different formats of data, corresponding knowledge extraction schemes are designed to complete the fusion and storage of knowledge. The structured monitoring data is extremely large in volume and high in update frequency, and it is not suitable to be stored in a graph database. The relevant database interface information is stored as the attributes of the knowledge graph nodes. For massive text data, the traditional supervised learning-based method has a high cost of data annotation, and a method based on prompt engineering is adopted to carry out knowledge extraction.
[0098] For GIS, BIM, and structured monitoring data, due to their high degree of structuring, no further fusion is required after element-level fusion. For text data, due to its wide range of sources, there may be problems of entity reference ambiguity. During the construction of the knowledge graph, further semantic disambiguation and fusion of text data are required. The invention uses the FastTex word embedding model to transform the text description of entities into vector representations, and uses cosine similarity to calculate the semantic similarity of texts, as shown in Equation (3), where e1 and e2 respectively represent the vector representations of entities, and the larger the value of cos(e1, e2), the higher the semantic similarity of the entities corresponding to e1 and e2. During the knowledge fusion process, entities with a similarity higher than the threshold are merged.
[0099]
[0100] The present invention selects the open-source graph database software Neo4j for the storage and management of the knowledge graph. This software supports the efficient graph query language Cypher, which can achieve fast query and retrieval of the knowledge graph and support downstream applications.
[0101] In a preferred embodiment of the present invention, the retrieval enhancement generation system in the above step S400 is specifically as follows:
[0102] The present invention constructs a data sharing interactive Q&A system based on the knowledge graph, which reduces the threshold for data sharing entities to obtain data, thereby improving the efficiency of engineering data sharing. This interactive Q&A system is based on the retrieval enhancement generation technology, retrieves data related to the user input from the knowledge graph and the external database, and utilizes the powerful text generation ability of the large language model to optimize the interactive experience. The schematic diagram of the method is shown in the appendix Figure 6 as shown.
[0103] For the text q input by the user, use the pre-trained model to perform text embedding on it, as shown in Equation (4), where z q represents the text vector, and PLM represents the pre-trained model for text embedding.
[0104]
[0105] Calculate the cosine similarity between z q and the standard problem vector in the case base, as shown in Equation (5). Where i represents the scenario category, and the present invention considers application scenarios such as information retrieval and reservoir operation. Select the scenario with the highest cosine similarity as the classification of the input text, and this classification will determine the rules for subsequent knowledge graph node retrieval, the type of database to be called, and the template of the prompt word.
[0106]
[0107] Use the word segmentation tool to split the text q input by the user, extract the nouns related to watershed monitoring therein, and form a keyword set The knowledge graph of the present invention is defined as where ε represents the entity set in the knowledge graph, represents the relationship set in the knowledge graph. Set the threshold of the cosine similarity to τ, and then use the cosine similarity to match each noun in the list K with the entity names in ε. The nodes with a similarity exceeding τ are all selected into the target node set ε t as defined in Equation (6).
[0108]
[0109] According to different application scenarios i, adopt the corresponding inference strategy Ri Obtain relevant information about the target node e in the knowledge graph, including node attributes, triples, etc., to form a subgraph As shown in Equation (7). In the subgraph , the general expression of node attributes is (entity, attribute name, attribute value), and the template for converting it into a string is "{entity}'s {attribute name} is {attribute value}"; the general expression of triples is (entity1, relationship name, entity2), and the template for converting it into a string is "{entity1}{relationship name}{entity2}".
[0110]
[0111] After converting all node attribute and triple information in the graph into strings, the string obtained by concatenating all of them is defined as D g . Define the general specification text obtained from the document database in scenario i as T i . Concatenate the user question, real-time monitoring data of the target node, knowledge graph information, and general scheduling rules into a complete prompt Prompt(q), and its definition is as shown in Equation (8). Input this prompt into the large model to obtain an answer to the user input text q
[0112]
[0113] The second aspect of the present invention discloses a multi-source heterogeneous data fusion and sharing system for high-altitude hydropower projects, using the above-disclosed multi-source heterogeneous data fusion and sharing method for high-altitude hydropower projects
[0114] Although this specification has shown and described multiple embodiments of the present invention, it is obvious to those skilled in the art that such embodiments are provided only by way of example. Those skilled in the art can think of many changes, alterations, and alternative ways without departing from the spirit and concept of the present invention. It should be understood that various alternative solutions of the embodiments of the present invention described herein can be adopted during the practice of the present invention. The appended claims are intended to define the protection scope of the present invention and thus cover the module compositions, equivalents, or alternative solutions within the scope of these claims
Claims
1. A multi-source heterogeneous data fusion and sharing method for high-altitude hydropower projects, characterized in that, Including: Collect and store multi-source data, where the multi-source data includes structured data, BIM data, GIS data, video and image data, and text data; Perform data pre-check on the multi-source data to ensure that the data format and quality meet the relevant requirements; And carry out preprocessing on the structured monitoring data by using the methods of moving average and Fourier spectrum analysis; Construct a knowledge graph to model the semantic relationships between multi-source heterogeneous data; It is divided into two stages: defining the schema layer and filling the data layer: In the stage of defining the schema layer, adopt a top-down method to systematically sort out the concept hierarchy and semantic relationships of hydropower projects, and set corresponding attributes; In the stage of filling the data layer, adopt a bottom-up method to design corresponding knowledge extraction schemes according to different formats of data, and complete the fusion and storage of knowledge; Combine the large language model LLM and the retrieval-augmented generation RAG technology to construct a data sharing and interactive Q&A system based on the knowledge graph.
2. The multi-source heterogeneous data fusion and sharing method for high-altitude hydropower projects according to claim 1, wherein The collecting and storing of multi-source data includes: For structured data with a low update frequency and semi-structured data including spatial data, BIM data, video, and file data, adopt a data regular pull strategy to achieve data collection and update; For structured data with a high update frequency, adopt a real-time synchronization strategy to achieve data collection and update; For engineering design, construction and other information data based on BIM technology, adopt a digital handover method to complete data collection and update.
3. The multi-source heterogeneous data fusion and sharing method for high-altitude hydropower projects according to claim 1, characterized in that The preprocessing of the structured monitoring data by using the methods of moving average and Fourier spectrum analysis includes: For the situation where there is obvious periodic noise in the data, adopt the method based on Fourier spectrum analysis to clean the data: Among them, X[k] represents the frequency-domain signal, k represents the frequency index, and x[n] represents the discrete time-domain signal; For the data with accidental errors or systematic error data, adopt the moving average method to clean the data: F MA = (f(x1) + f(x2) + f(x3) + … + f(x N )) / N #(2) Among them, F MA represents the value of the variable x i (i = 1, 2, …, N) after moving average calculation, where N is the size of the moving average window, and f(x i )(u = 1, 2, …, N) represents the average calculation function.
4. The multi-source heterogeneous data fusion and sharing method for high-altitude hydropower projects according to claim 1, characterized in that The construction of the knowledge graph includes: Define the schema layer of the knowledge graph, systematically sort out the concept hierarchy and semantic relationships in the field of high-altitude hydropower projects, and establish a complete knowledge system; Fill the data layer, through different knowledge extraction schemes, convert multi-format data into entities and relationships in the knowledge graph, and perform semantic fusion: For BIM and GIS type data, manually screen the useful information and import it into the knowledge graph; For structured monitoring data, store it in a time-series database, and store the API of this database in the attributes of the monitoring object nodes; For unstructured text data, adopt the method based on prompt engineering to extract triples, that is, design prompt words according to the requirements, and call the large model API or website to obtain the triple information in the text segment; Finally, perform text disambiguation and save the extracted triples in JSON format; Knowledge graph storage and management, write Python scripts to import the JSON format triples into the Neo4j graph database; Implement complex knowledge graph query requirements through the Cypher query language.
5. The multi-source heterogeneous data fusion and sharing method for high-altitude hydropower projects according to claim 4, wherein The text disambiguation includes: Use the FastTex word embedding model to convert the text description of the entity into a vector representation, and use cosine similarity to calculate the semantic similarity of the text: Among them, e1 and e2 respectively represent the vector representations of entities. The larger the value of cos(e1, e2), the higher the semantic similarity between the entities corresponding to e1 and e2; During the knowledge fusion process, entities with a similarity higher than the threshold are merged.
6. The multi-source heterogeneous data fusion and sharing method for high-altitude hydropower projects according to claim 1, wherein Constructing a data sharing interactive Q&A system based on the knowledge graph includes: For the text q input by the user, use a pre-trained model to perform text embedding on it: Among them, represents a text vector, and PLM represents a pre-trained model for text embedding; Calculation with the standard problem vector in the case base for cosine similarity: Among them, i represents the scenario category; Select the scenario with the highest cosine similarity as the classification of the input text. This classification will determine the rules for subsequent knowledge graph node retrieval, the type of database to be called, and the template of the prompt word; Use a word segmentation tool to split the text q input by the user, extract the nouns related to watershed monitoring, and form a keyword set A knowledge graph is defined as Among them, ε represents the set of entities in the knowledge graph, represents the set of relationships in the knowledge graph; Set the threshold of cosine similarity to τ, and then use cosine similarity to match each noun in list K with the entity names in ε. All nodes with a similarity exceeding τ are selected into the target node set In: According to different application scenarios i, adopt the corresponding inference strategy R i Obtain information about the target node e in the knowledge graph, including node attributes, triples, etc., to form a subgraph In the subgraph the general expression of node attributes is (entity, attribute name, attribute value), and the template for converting it into a string is "{entity}'s {attribute name} is {attribute value}"; the general expression of triples is (entity 1, relationship name, entity 2), and the template for converting it into a string is "{entity 1}{relationship name}{entity 2}"; After converting all node attributes and triple information in Figure into strings and concatenating them all, the resulting string is defined as D g ; Define the general specification text obtained from the document database in scenario i as T i . Concatenate the user question, real-time monitoring data of the target node, knowledge graph information, and general scheduling rules into a complete prompt Prompt(q), and its definition is as follows: Input the prompt word into the large model to obtain an answer to the user input text q.
7. A multi-source heterogeneous data fusion and sharing system for high-altitude hydropower projects, characterized in that Use the multi-source heterogeneous data fusion and sharing method for high-altitude hydropower projects described in any one of claims 1-6.
Citation Information
Patent Citations
Natural language artificial intelligence system and method for water and electricity operation and maintenance knowledge
CN115563968A
Traffic prediction method and traffic prediction device
CN116911421A
DCS intelligent decision-making method and system fusing large language model and knowledge graph
CN118820778A
Method for constructing knowledge graph based on large language model and vector library
CN119129722A
Method and system for generating questions and answers based on knowledge graph
KR102697127B1
Cited By
Full-life-cycle equipment test intelligent calibration system
CN121093961A