River network modeling and question answering method based on time sequence knowledge graph
By constructing a multi-dimensional semantic ontology system based on time-series knowledge graphs, the problem of insufficient fusion of multi-source data and coordination of static and dynamic information in river network systems is solved. This enables dynamic semantic modeling and intelligent question answering in river network systems, improving data consistency and question answering accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HOHAI UNIV
- Filing Date
- 2025-11-14
- Publication Date
- 2026-08-04
AI Technical Summary
Traditional GIS systems struggle to support semantic management of river network systems, especially in expressing time-series hydrological events and dynamic states, and multi-source heterogeneous data is difficult to integrate and fused.
We employ a temporal knowledge graph-based approach, constructing a multi-dimensional semantic ontology system and combining it with a large language model and graph structure reasoning mechanism to achieve multi-source data collection, preprocessing, knowledge extraction, and intelligent question answering.
Dynamic semantic modeling and knowledge reasoning of the river network system were realized, which improved data consistency and interpretability, and enhanced the accuracy of question-and-answer results and the practical value of the system.
Smart Images

Figure CN121350276B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a river network modeling and question-answering method based on a time-series knowledge graph that integrates multi-source heterogeneous data, spatial topology modeling, semantic reasoning, and intelligent question-answering functions, and belongs to the field of knowledge graph construction technology. Background Technology
[0002] River networks, as a crucial structure for surface water distribution, fulfill multiple functions including water resource allocation, flood control and disaster reduction, and ecological governance. In plain areas, river networks are complex and densely interwoven, requiring high-quality data fusion and intelligent system support for their management and analysis. Traditional GIS systems focus primarily on the geometric structure of river networks, struggling to support semantic management of objects such as rivers, reservoirs, and sluices, and particularly lacking effective mechanisms for representing time-series hydrological events and dynamic states.
[0003] Meanwhile, the heterogeneity of data sources such as remote sensing, water conservancy databases, monitoring stations, and text records makes data integration difficult. Inconsistent naming, different spatial granularities, and varying update frequencies among data make traditional data center architectures unable to meet the needs of deep reasoning at the knowledge level.
[0004] Knowledge graphs, as structured semantic networks, have been widely used in intelligent recommendation and intelligent question answering scenarios in recent years, providing a new knowledge representation paradigm for information systems in the water conservancy industry. However, when dealing with river network systems, they still lack the ability to fuse and represent three types of data: geographic, hydrological, and temporal, and cannot achieve advanced functions such as entity evolution path modeling, multimodal knowledge fusion, and semantic reasoning. Summary of the Invention
[0005] Purpose of the invention: This invention provides a river network modeling and question answering method based on temporal knowledge graphs. It aims to address the difficulties in integrating multi-source data of river networks and the lack of coordination between static and dynamic information. By introducing temporal knowledge graph construction technology, large language models and graph structure reasoning mechanisms, it realizes a fully automated support system from data collection, ontology modeling, knowledge extraction, knowledge fusion to intelligent question answering.
[0006] Technical solution: A river network modeling and question answering method based on temporal knowledge graph, comprising the following steps: Step 1): Collect and preprocess multi-source river network data. The data includes two categories: structured data and unstructured data. Structured data includes geographic vector data from geographic information databases and publicly available station monitoring data from the platform; unstructured data includes text data from publicly available hydrological reports. The geographic vector data and unstructured text data are then processed separately to obtain geospatial topological relationship data and a cleaned and segmented text corpus.
[0007] Step 2): Construct a multi-dimensional semantic ontology system that integrates static and dynamic elements, covering geographic entities, hydrological and hydrodynamic data entities, temporal semantic entities in the river network system, and their topological and temporal relationships.
[0008] Step 3): Extract knowledge from the multi-source data collected and preprocessed in Step 1). Specifically, use a field-relationship mapping template language to generate temporal knowledge quadruples consistent with the domain ontology from structured data such as geospatial topology data and site monitoring data. ,in For the head entity, For the relationship, For tail entities, Using time anchors, the unstructured text data is extracted using the OneRel joint extraction model to improve extraction accuracy and generalization ability. After extraction, entity names, relational predicates, and time expressions are standardized, including unified naming rules, time formats, and semantic tags. Based on this, a geographic topological sub-map is generated from geospatial topological relationship data, a monitoring time-series sub-map is generated from station monitoring data, and a hydrological event sub-map is generated from the unstructured text extraction results. These three sub-maps serve as inputs for subsequent fusion.
[0009] Step 4): Design the SelfKG-T cross-modal entity alignment algorithm. After candidate entity pairs are generated, a temporal path attention module is added during the graph encoding stage to perform weighted aggregation of the timestamped relationship paths connecting candidate pairs, obtaining a time-consistent path representation, and jointly calculating the alignment score with the entity representations at both ends. The entities in the geographic topology sub-graph, monitoring temporal sub-graph, and hydrological event sub-graph obtained in Step 3) are aligned and merged in terms of naming, attributes, and time dimensions to form a unified entity and relationship representation, ultimately obtaining a complete river network knowledge graph. This provides a unified knowledge foundation for subsequent semantic retrieval and intelligent question answering, enabling dynamic semantic modeling and knowledge reasoning support for the river network system.
[0010] Step 5): Based on the RAG architecture and large language model, the entities of the river network knowledge graph are transformed into retrieval corpus to achieve question-answering recall and semantic generation.
[0011] Step 1) of river network multi-source data acquisition and preprocessing includes the following steps: 1-1) Import geographic vector data using the GeoJSON spatial data format, including geographic units such as river segments, estuaries, cross sections, and reservoirs. Use the nine-intersection model to extract the spatial topological relationships between these geographic units, forming a Boolean topological matrix. Let the set of spatial entities be... Its spatial relationship matrix Describe the interaction between units, such as adjacency, inclusion, and overlap, to provide data support for the subsequent establishment of spatial topological semantics in the ontology.
[0012] 1-2) Store the monitoring data from publicly available sites on the platform in JSON or XML format, and standardize the observed values such as water level, flow velocity, and precipitation into a unified structure using a data standardization template. ,in As a unique identifier for the site, For observation time, To standardize numerical values under a unified dimension, The corresponding units are (e.g., m, m / s, mm). To adapt to time series modeling, a sliding window is used to resample the monitoring data of the publicly available stations on the platform. The observations within the window are aggregated according to the index using the appropriate aggregation method (e.g., average value for water level / flow velocity, summation for precipitation) to generate a standard time series with equal intervals.
[0013] 1-3) Unstructured text data mainly comes from hydrological logs and hydrological reports. The raw text undergoes data cleaning to remove special characters and redundant content, and is then segmented according to themes. Subsequently, based on key entities defined by semantic paragraph annotations, and through a unified encoding format and segmentation identifiers, a set of independent semantic fragments is ultimately formed. .
[0014] Step 2) of constructing a multidimensional semantic ontology system includes the following process: 2-1) The cleaned unstructured text data set—a set of semantic fragments Semantic encoding is performed using a pre-trained SBERT model to obtain sentence vector representations. ,in It is a set The total number of semantic segments was determined. Then, the K-Means clustering algorithm was used to cluster all semantic vectors, resulting in a set of semantic centers. Each cluster center represents a potential event type vocabulary and entity / relation candidate set to guide the initial design of the water resources ontology.
[0015] 2-2) Structured geographic vector data and publicly available site monitoring data from the platform are transformed into core concepts and attributes in the ontology using a mapping template language. Site monitoring data are defined as observation record nodes in the map and associated with corresponding geographic entities such as rivers and stations to obtain structured results.
[0016] 2-3) Combining the semantic clustering centers from 2-1) and the structured results from 2-2), and referring to water conservancy standards such as the "General Rules for Classification and Coding of Water Conservancy Objects," define the geographical and hydrodynamic ontology structure: Geographic ontology layer: includes geographic entities such as “rivers”, “lakes”, “sluice gates”, and “stations”; and their geographic topological relationships such as “flowing into”, “belonging to”, and “located in”.
[0017] Hydrodynamic entity layer: Defines dynamic observation entities such as "monitoring records" and "hydrological events", configures attributes such as "water level", "flow velocity", "flow rate" and "precipitation", and establishes associations with relevant geographic entities through semantic relationships such as "has observed value" and "exceeds threshold".
[0018] 2-4) Introducing "semantic temporal entities," operational time periods such as "flood season," "dispatch period," and "historical flood period" are modeled as independent entities. These entities are associated with monitoring records and hydrological event entities through relationships such as "occurred in" and "belong to," together forming a temporal semantic ontology. The temporal semantic ontology, together with the geographic ontology and the hydrodynamic ontology, constitutes a multi-dimensional semantic ontology system, forming a complete river network temporal knowledge graph ontology.
[0019] 2-5) Finally, the constructed multidimensional river network time-series knowledge graph ontology will be output as an OWL format file, which supports the representation of all entities and relations with Chinese tags. The semantic inferencer HermiT will be used to perform consistency checks, hierarchical inheritance verification and inference rule testing to ensure that the ontology structure is complete, inferable and scalable.
[0020] The knowledge extraction in step 3) includes the following steps: 3-1) The structured data collected and standardized in step 1), namely geographic vector data and site monitoring data, are processed by field and relation mapping to generate knowledge quadruples that conform to the ontology structure. Let the set of quadruplets be . Define field-relationship mapping function As shown in equation (1): in, , These represent the head and tail entities, respectively. For semantic relations, For context time.
[0021] The set of field template rules is defined as shown in equation (2): in, For the first A field template function is used to combine field groups. Map each field record using template functions. Obtain temporal knowledge quadruples consistent with the ontology structure. .
[0022] 3-2) For the set of semantic fragments obtained in step 1), The OneRel model is used to perform semantic encoding and joint prediction for each sentence.
[0023] The encoding function of the pre-trained language model is defined as shown in equation (3): in, This represents a BERT-based semantic encoding function used to encode the input sentence. Convert to a hidden vector representation; For the first A text fragment or sentence; This is the corresponding word-level hidden feature matrix. Indicates sentence length. Indicates the dimension of the encoded vector.
[0024] Let the set of relations be For each relation Conditional semantic vectors are generated, and the relation-aware entity boundary prediction distribution is constructed as shown in equation (4): in, and These represent the start and end boundary indices of the head and tail entities in the sentence, with values ranging from {1, ..., L}. For the first One relationship; For the Pointer Network module: given a token, represent Relationship Conditions Output four probability distributions of length L, corresponding to the "starting point" and "ending point" position distributions of the head and tail entities, respectively. The logarithm of each of the four probabilities is summed to define the joint score of the candidate pairs.
[0025] All candidate pairs with a joint score not lower than a preset threshold are combined into a result set, which is shown in Equation (5): in, and These represent the head entity and the tail entity at the 1st and 2nd respectively. The and the first a sentence Start and end boundary indexes in; The first one defined in the ontology / relation schema The relation type; the time anchoring corresponding to this sentence. The time recognition module described in equation (6) Extract from sentences and map to a normalized time domain: Finally, the structured and unstructured extraction results are integrated to form a unified knowledge set as shown in equation (7): It serves as the basic quadruple input in the graph construction process.
[0026] The entity alignment algorithm in step 4) includes the following steps: 4-1) Generate a set of candidate entity pairs and construct pseudo-labels. Use the set of temporal knowledge quadruples obtained in step 3) as an example. For direct input: First, from Obtain each entity as the initial node set, and then combine the four-tuples... As the initial edge set, The records are represented by edge timestamps and time indices; simultaneously, the source markers (geographic topology, monitoring time series, hydrological events) of each quadruple are retained, forming a "time-based" base graph with three subgraphs. Let the multi-source entity set be... ,in For geographical entities, For hydrological and hydrodynamic entities. A set of candidate entity pairs is constructed using heuristic rules such as adjacent context, temporal overlap, and spatial co-occurrence, as shown in Equation (8): Based on the idea of self-supervision, pseudo-labels are generated by cross-semantic view comparison, and the positive and negative sample sets are constructed as shown in Equation (9): in This represents the candidate entity pairs constructed in each subgraph (i.e., geographic topology, monitoring time series, and hydrological events), and their pseudo-labels. The self-supervised method for cross-semantic view comparison automatically identifies two entities as pseudo-aligned positive samples when they meet consistency requirements in name / alias, semantic context, and temporal anchor. Conversely, it is used as a negative sample, i.e. The positive and negative sample sets do not require manual annotation and are automatically constructed using Bootstrap Sampling and entity clustering similarity methods.
[0027] 4-2) Perform multimodal feature fusion and temporal path attention modeling. For each pair of candidate entities... Extracting multimodal similarity features, including: The string similarity feature uses normalized edit distance, as shown in Equation (10). This indicates the minimum number of edit steps required to change the entity name. The larger of the two string lengths is used; the ratio of the two lengths yields a similarity measure in the [0,1] interval. The semantic similarity feature uses the cosine similarity of equation (11). , in For encoding functions based on pre-trained language models: Path characteristics: For the set of quadruples output in step 3). The multi-source knowledge graph is constructed from... arrive The set of directed paths As shown in equation (12), the path set is obtained by performing a limited number of hops traversal / enumeration on the graph: For each path Construct path embedding vectors Furthermore, a time-series decay function is introduced to apply decay over time, as shown in equation (13): Among them Time span in the path, is the attenuation factor. The resulting path attention distribution is shown in equation (14): in, To query the entity feature vector, For learnable weight matrix, For the first Embedded representation of a path It is a path index variable, and the denominator represents the index of all paths. The algorithm performs an exponentially weighted summation on each path in the candidate path set. This formula is used to calculate the normalized weights of different time paths under the path attention mechanism, reflecting the relative importance of each path to the entity alignment score.
[0028] The final path similarity is shown in equation (15): In the formula, To be from the entity arrive The set of paths Let be the path feature mapping function. The comprehensive path similarity is obtained through weighted summation. The final joint similarity feature vector is shown in equation (16): Joint similarity feature vectors The input is fed into the alignment discrimination function to determine entity pairs, and the output is the matching probability. .
[0029] 4-3) Finally, the merged edge construction is aligned with the entity graph output. For the judgment value... The entity pairs generate the aligned edge set as shown in equation (17): The resulting entity alignment edge set The merged graph is incorporated into the graph structure to form a fused graph as shown in equation (18): in For a collection of entities, This is the set of relationships in the original graph. For the newly generated set of entity alignment relations, The temporal dimension edge labels or index structure of the graph are represented by the quadruples of temporal anchors generated in step 3). After writing the attributes of each edge in step 4-1), a unified and abstracted knowledge graph is generated to record the occurrence time or time interval of each relation edge, thereby ensuring the temporal consistency of cross-source entities. By merging it with the original relation set, a unified knowledge graph with unified entities and interconnected paths is finally output. .
[0030] Step 5) involves transforming graph entities into retrieval corpora based on the RAG architecture and large language model to achieve question-answering retrieval and semantic generation, including the following steps: 5-1) Textify the graph and generate a retrieval corpus. For the data obtained in step 4) All of them Quadruple That is, converting static relationships between entities or temporal relationships with time stamps into a collection of documents in natural language form. , This represents a natural language corpus (document) transformed from a knowledge graph triple or quadruple. Each document corresponds to a factual expression, as shown in equation (19): Construct template function The quadruple is transcribed into the retrieval corpus as shown in equation (20): All sentences constitute the document index set, which is used during the retrieval phase.
[0031] 5-2) Constructing a RAG retrieval and generation architecture and performing knowledge retrieval. The Retrieval-Augmented Generation (RAG) architecture is introduced, combining a vector retrieval unit with a generative language model. The RAG consists of two parts: a retrieval unit and a generator. The retrieval unit constructs a dense semantic index based on vectorization. Let the corpus of documents transformed from the graph be... The query vector is The corpus vector is The retrieval system recalls previous results using maximum cosine similarity. The document is shown in equation (21): The generator introduces a language model , will use the natural language query input by the user The results are concatenated with the recall results to form a prompt word sequence, generating the final answer as shown in equation (22): 5-3) Finally, construct and semantically fuse the prompt words. When constructing the prompt words, the user's natural language query will be incorporated. As the main question, it is rewritten and linked with entities to generate a structured search intent. If the query contains a graph entity name... Then extract its relevance from the graph. Skip-adjacency graph Generate multiple auxiliary statements This constitutes a complete sequence of prompt words, as shown in equation (23): Ultimately, the complete input consists of the user's question and multiple structured knowledge fragments, which serve as the contextual input for the large language model, resulting in a semantically consistent and context-complete answer.
[0032] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the river network modeling and question-answering method based on time-series knowledge graphs as described above.
[0033] A computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the steps of the river network modeling and question-answering method based on time-series knowledge graphs as described above.
[0034] Beneficial effects: (1) At the data fusion level, existing river network semantic modeling methods generally lack the ability to uniformly express and semantically align multi-source heterogeneous data. This invention introduces a semantic modeling framework based on temporal knowledge graphs to achieve multimodal fusion and temporal unified modeling of geospatial data, hydrological monitoring data, and textual hydrological data, which significantly improves the consistency and interpretability of the data.
[0035] (2) At the temporal reasoning level, existing knowledge graphs are mostly static structures and cannot effectively express the dynamic changes in river networks. This invention, by constructing a temporal path attention model, can capture the evolution of river network structure and hydrological events in the time dimension, significantly improving the ability of dynamic semantic modeling and trend reasoning.
[0036] (3) At the entity alignment level, existing graph fusion algorithms often rely on manual rules or single semantic similarity, resulting in insufficient accuracy. The SelfKG-T alignment algorithm proposed in this invention fuses structural, semantic and temporal multimodal features to achieve high-precision entity matching under unsupervised conditions, thereby improving the accuracy of cross-source data fusion.
[0037] (4) At the question-and-answer application level, existing water conservancy question-and-answer systems mostly rely on keyword retrieval or template matching, which is difficult to handle complex semantics. This invention combines the RAG framework with a large language model to transform the temporal knowledge graph into a searchable corpus, realize semantic understanding and knowledge retrieval of natural language questions, and improve the accuracy and interpretability of question-and-answer results.
[0038] (5) In terms of engineering applications, the present invention adopts a modular design, which has good scalability and versatility. It can be deployed and applied in multiple scenarios such as flood control scheduling and water resource optimization management, which significantly enhances the practical value of the system. Attached Figure Description
[0039] Figure 1 This is a flowchart of the construction method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of an actual application system according to an embodiment of the present invention. Detailed Implementation
[0040] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.
[0041] A method for river network modeling and question answering based on temporal knowledge graphs, including the following steps: Step 1: Multi-source river network data acquisition and preprocessing, collecting geographic vector data, hydrological and hydrodynamic monitoring data, and unstructured text data, unifying spatial and temporal formats, and constructing a multi-source dataset with consistent structure; Step 2: Construct a multidimensional semantic ontology model, including a geographic ontology, a dynamic hydrodynamic ontology, and a temporal semantic ontology, and complete the ontology hierarchy definition and semantic relationship modeling; Step 3: Heterogeneous data knowledge extraction. Construct field mapping rules for structured data, and use a deep joint extraction model to extract entity relationships and time anchors for unstructured text, generating knowledge quadruples. Step 4: Multi-source knowledge fusion and entity alignment, construct a candidate entity pair set, introduce the SelfKG-T algorithm, use the path attention mechanism to calculate similarity and perform entity fusion, and construct a unified temporal graph structure; Step 5: Implement a graph-driven intelligent question-answering system, transforming the graph into natural language corpus, and realizing semantic retrieval and question-answer generation based on the RAG architecture and large language model.
[0042] Multi-source river network data acquisition and preprocessing includes the following steps: The first step is to import vector geographic units such as river segments, cross sections, reservoirs, and estuaries in GeoJSON format from the geographic database. Spatial relationships are extracted using a nine-intersection model, and a Boolean spatial topological matrix is constructed to represent spatial semantics such as "adjacent," "intersecting," and "containing." The second step involves acquiring hydrological and hydrodynamic data in JSON or XML format from hydrological databases and publicly available monitoring websites, including observed indicators such as water level, flow velocity, and precipitation. This data is then time-standardized, and a sliding window mechanism is introduced to uniformly resample data from different sampling frequencies. The third step involves obtaining unstructured text data from documents such as hydrological logs and dispatch briefings. First, the text is cleaned to remove invalid symbols and formatting errors. Then, regular expressions and keyword methods are used to extract semantic paragraphs, which are labeled with hydrological objects, indicator names, and time expressions, forming a set of semantic fragments. .
[0043] The construction of the multi-dimensional semantic ontology structure upon which the temporal knowledge graph depends includes the following steps: The first step is to analyze the set of semantic fragments. Semantic encoding is performed using the SBERT model to obtain a vector set. ; The second step is to use the K-Means clustering algorithm to perform semantic clustering on the vector set to obtain a set of semantic centers. Each cluster represents a high-frequency event type or semantic concept; The third step involves using Mapping Template Language to map structured geographic data and monitoring data into concept nodes and attribute nodes, thereby constructing a static geographic ontology and a dynamic hydrodynamic ontology. The fourth step is to introduce semantic time entities, model business time periods such as "flood season" and "scheduling period" as independent entities, and establish semantic connection relationships such as "occurred in" and "corresponding period" with other entity nodes; The fifth step is to export the completed ontology model as an OWL format file, which supports Chinese tags, and use the semantic inference engine HermiT to verify the consistency of the ontology structure and test the inference rules.
[0044] Extracting structurally consistent knowledge quadruples for graph construction from structured and unstructured data includes the following steps: The first step is to define a set of field mapping rules for structured data. Indicates the first A set of field mapping rules for structured data. Each rule maps a group of fields to a quadruple. ,in and For entities, For the relationship, For time information; The second step involves using the OneRel model for joint extraction of unstructured text. For each text sentence... Using an encoder The semantic representation is obtained, and candidate entity boundaries and relation pairs are generated through relation condition vectors. The extraction form is shown in Equation (22): in, This indicates that the OneRel model is used in sentences. and relation types The relation feature scoring function calculated above is used to measure the probability score of whether candidate entity pairs in a sentence satisfy the relation.
[0045] The third step involves using regular expressions and lexical rules to extract and standardize time expressions in sentences, forming time anchors. Combined with the extraction results, a set of time-series quadruples is generated as shown in equation (23): The fourth step is to merge the structured and unstructured extraction results into a unified knowledge set, as shown in equation (24): The process of multi-source knowledge fusion and entity alignment includes the following steps: The first step is to construct a set of candidate entity pairs. Entity candidate pairs are matched based on heuristic rules such as string similarity, semantic vector distance, and spatial adjacency. The second step involves introducing the SelfKG-T alignment algorithm and employing a path attention mechanism in the graph to model the temporal relationships between entities. A path set is defined. To be from the entity arrive The set of directed paths, with path embeddings represented as shown in equation (25): in, This represents the time span between path nodes. This is the time decay factor; The third step is to integrate semantic features, path features, and string features, input them into the alignment discrimination model, and output the alignment score. The fourth step is to separate entity pairs whose scores are above a threshold. Included in the fusion edge set And generate the final fusion map structure. .
[0046] The process of transforming temporal knowledge graphs into searchable corpora and implementing semantic answers to user questions based on a large language model includes the following steps: The first step is to analyze the quadruplets in the graph. Applying a natural language template function, as shown in equation (26): Get document collection ; The second step is to construct a dense semantic vector index. (User input issue) Then, recalling previous users through semantic matching. Related documents ; The third step is to combine the issue and the recall content into a sequence of prompt words. Input into a large language model and output the final answer. .
[0047] This invention constructs a river network temporal knowledge graph with temporal reasoning and intelligent question answering capabilities based on multi-source river network data fusion and semantic modeling technology. The method generally consists of a data acquisition and preprocessing module, an ontology modeling module, a knowledge extraction module, an entity alignment and fusion module, and a question answering generation module. The data acquisition module supports the access of static geographic information, hydrological and hydrodynamic observation data, and unstructured text data, unifying spatial, temporal, and semantic representations. The ontology modeling module constructs a three-layer ontology structure including geographic entities, hydrodynamic features, and temporal semantics, supporting semantic reasoning and inheritance relationship verification. The knowledge extraction module combines template mapping and a deep joint extraction model to automatically generate knowledge quadruples. The entity alignment and fusion module, based on the SelfKG-T algorithm, introduces a temporal path attention mechanism to fuse and match multi-source heterogeneous entities. The question answering module adopts a RAG retrieval-enhanced generation architecture, utilizing a large language model to complete graph-based semantic retrieval and answer generation, supporting intelligent responses to users' natural language questions.
[0048] This embodiment is applicable to smart water conservancy business scenarios such as river network evolution modeling, scheduling auxiliary decision-making, and hydrological information Q&A. It helps to improve the information expression capability of the river network system, enhance the temporal knowledge organization effect, and build a more intelligent and semantically interactive application platform.
[0049] Obviously, those skilled in the art will understand that the steps of the present invention described above can be implemented by general-purpose computing devices. These modules can be integrated on a single server or distributed across a network platform composed of heterogeneous systems. Each module can be implemented on a general-purpose processor based on an executable program, or it can be integrated and deployed in a microservices manner. Therefore, the present invention is not limited to any specific combination of hardware and software, nor is it limited to a fixed execution order of the steps. All equivalent modifications or extensions made based on the core ideas of the present invention should fall within the protection scope of the claims of this patent.
Claims
1. A river network modeling and question-answering method based on temporal knowledge graphs, characterized in that, Includes the following steps: Step 1): Collect and preprocess multi-source river network data, which includes two types of data: structured data and unstructured data. Structured data includes geographic vector data from geographic information databases and publicly available station monitoring data from the platform. Unstructured data includes text data from publicly available hydrological reports. Process the geographic vector data and unstructured text data respectively to obtain geospatial topological relationship data and a cleaned and segmented text corpus. Step 2): Construct a multi-dimensional semantic ontology system that integrates static and dynamic elements, covering geographic entities, hydrological and hydrodynamic data entities, temporal semantic entities in the river network system, and their topological and temporal relationships; Step 3): Extract knowledge from the multi-source data collected and preprocessed in Step 1); Geographic topology sub-maps are generated based on geospatial topology data, monitoring time series sub-maps are generated based on station monitoring data, and hydrological event sub-maps are generated based on unstructured text extraction results. Step 4): Design a cross-modal entity alignment algorithm, SelfKG-T; after candidate entity pairs are generated, a temporal path attention module is added in the graph encoding stage to perform weighted aggregation of the timestamped relationship paths connecting candidate pairs, obtain time-consistent path representations, and jointly calculate the alignment score with the entity representations at both ends; align and merge the entities in the geographic topology sub-graph, monitoring temporal sub-graph, and hydrological event sub-graph obtained in Step 3) in terms of naming, attributes, and time dimensions to form a unified entity and relationship representation, and finally obtain a complete river network knowledge graph; Step 5): Based on the RAG architecture and large language model, the entities of the river network knowledge graph are transformed into retrieval corpus to achieve question-answering retrieval and semantic generation.
2. The river network modeling and question answering method based on time-series knowledge graphs according to claim 1, characterized in that, Knowledge extraction is performed on the multi-source data collected and preprocessed in step 1); time-series knowledge quadruples consistent with the domain ontology are generated from geospatial topological relationship data and site monitoring data using the field-relationship mapping template language. ,in For the head entity, For the relationship, For tail entities, The data is anchored to time. Unstructured text data is extracted using the OneRel joint extraction model to improve extraction accuracy and generalization ability. After extraction, entity names, relational predicates and time expressions are standardized, including unified naming rules, time format and semantic tags.
3. The river network modeling and question answering method based on time-series knowledge graphs according to claim 1, characterized in that, Step 1) of river network multi-source data acquisition and preprocessing includes the following steps: 1-1) Import geographic vector data using the GeoJSON spatial data format, including geographic units: river segment, river mouth, cross section, and reservoir. Use the nine-intersection model to extract the spatial topological relationships between the above geographic units and form a Boolean topological matrix. 1-2) Store the monitoring data from publicly available sites on the platform in JSON or XML format, and standardize the observations into a uniform structure using a data standardization template. ,in As a unique identifier for the site, For observation time, To standardize numerical values under a unified dimension, For corresponding units; to adapt to time series modeling, a sliding window is used to resample the monitoring data of the publicly available sites on the platform, and the observations within the window are aggregated according to the indicators to generate standard time series with equal intervals; 1-3) Unstructured text data is cleaned to remove special characters and redundant content, and then segmented according to themes; Subsequently, based on the key entities defined by semantic paragraph annotations, a set of independent semantic fragments is ultimately formed through a unified encoding format and segmentation identifiers. .
4. The river network modeling and question-answering method based on time-series knowledge graphs according to claim 1, characterized in that, Step 2) involves constructing a multidimensional semantic ontology system, which includes the following processes: 2-1) The cleaned unstructured text data set—a set of semantic fragments Semantic encoding is performed using a pre-trained SBERT model to obtain sentence vector representations. , ,in It is a set The total number of semantic segments was determined; then, the K-Means clustering algorithm was used to cluster all semantic vectors to obtain the set of cluster centers. Each cluster center represents a potential vocabulary of event types and a candidate set of entities / relationships. 2-2) The structured geographic vector data and the publicly available site monitoring data of the platform are transformed into concepts and attributes in the ontology using a mapping template language; the site monitoring data are defined as observation record nodes in the map and associated with geographic entity objects to obtain structured results; 2-3) Combining the cluster centers from 2-1) with the structured results from 2-2), and referencing standards in the water resources field, define the geographical and hydrodynamic ontological structure: Geographic ontology layer: includes geographic entities and their geographic topological relationships; Hydrodynamic ontology layer: Defines dynamic observation entities, configures attributes, and establishes associations with related geographic entities through semantic relationships; 2-4) Introduce business time periods as independent entities; establish associations between these independent entities and monitoring records and hydrological event entities through the relationships of "occurred in" and "belong to the period", and jointly form a temporal semantic ontology; the temporal semantic ontology, together with the geographic ontology and the hydrodynamic ontology, constitutes a multi-dimensional semantic ontology system, forming a complete river network temporal knowledge graph ontology; 2-5) Finally, the constructed multidimensional river network time-series knowledge graph ontology will be output as an OWL format file, which supports the representation of all entities and relations with Chinese tags, and uses the semantic inferencer HermiT to perform consistency checks, hierarchical inheritance verification and inference rule testing.
5. The river network modeling and question answering method based on time-series knowledge graphs according to claim 1, characterized in that, Step 3) of knowledge extraction includes the following steps: 3-1) Perform field and relation mapping processing on geographic vector data and site monitoring data to generate knowledge quadruples that conform to the ontology structure; Define a set of field template rules; 3-2) For the set of semantic segments, the OneRel model is used to perform semantic encoding and joint prediction on each sentence; all candidate pairs with joint scores not lower than the preset threshold are combined into the output as the result set.
6. The river network modeling and question-answering method based on time-series knowledge graphs according to claim 1, characterized in that, Step 4) of the entity alignment algorithm includes the following steps: 4-1) Generate a set of candidate entity pairs and construct pseudo-labels, using a set of temporal knowledge quadruples. For direct input: First, from Obtain each entity as the initial node set, and then combine the four-tuples... As the initial edge set, Record the edge's time label and time index. For the head entity, For the relationship, For tail entities, As the time anchor point; while retaining the source marker of each quadruple, forming a time-based graph with three subgraphs; let the multi-source entity set be... ,in For geographical entities, For hydrological and hydrodynamic entities; construct a set of candidate entity pairs using heuristic rules based on adjacent contexts, temporal overlap, and spatial co-domain. Based on the concept of self-supervision, pseudo-labels are generated by cross-semantic view comparison to construct a set of positive and negative samples; 4-2) Perform multimodal feature fusion and temporal path attention modeling; for each pair of candidate entities... Extracting multimodal similarity features, including: The string similarity feature uses normalized edit distance, as shown in Equation (10). This indicates the minimum number of edit steps required to change the entity name. The larger of the two string lengths is used; the ratio of the two lengths yields a similarity measure in the [0,1] interval. The semantic similarity feature uses the cosine similarity of equation (11). ,in For encoding functions based on pre-trained language models: Path characteristics: For the set of quadruples The multi-source knowledge graph is constructed from... arrive The set of directed paths As shown in equation (12), the path set is obtained by performing a limited number of hops traversal / enumeration on the graph: For each path Construct path embedding vectors Furthermore, a time-series decay function is introduced to apply decay over time, as shown in equation (13): in The time span in the path, The attenuation factor is used; thus, the path attention distribution is obtained as shown in equation (14): in, To query the entity feature vector, For learnable weight matrix, For the first Embedded representation of a path It is a path index variable, and the denominator represents the index of all paths. Perform exponential weighted summation; Equation (14) is used to calculate the normalized weights of different time paths under the path attention mechanism, reflecting the relative importance of each path to the entity alignment score; The final path similarity is shown in equation (15): In the formula, To be from the entity arrive The set of paths The path feature mapping function is used to obtain the comprehensive path similarity through weighted summation, and the final joint similarity feature vector is shown in equation (16): Joint similarity feature vectors The input is fed into the alignment discrimination function to determine entity pairs, and the output is the matching probability. ; 4-3) Finally, the merged edge construction is aligned with the entity graph output; for the judgment value The entity pairs generate the aligned edge set as shown in equation (17): The resulting entity alignment edge set The merged graph is incorporated into the graph structure to form a fused graph as shown in equation (18): in For a collection of entities, This is the set of relationships in the original graph. For the newly generated set of entity alignment relations, This represents the time dimension edge labels or index structure of the knowledge graph; by merging the entity-aligned edge set with the original relation set, a unified knowledge graph with unified entities and interconnected paths is finally output. .
7. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the river network modeling and question-answering method based on time-series knowledge graphs as described in any one of claims 1-6.
8. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that: When the computer program / instruction is executed by the processor, it implements the steps of the river network modeling and question answering method based on time-series knowledge graphs as described in any one of claims 1-6.