Reservoir hydropower station carbon accounting key information extraction method and system based on large language model

By constructing a knowledge graph and multimodal document parsing in the field of carbon accounting for reservoirs and hydropower stations, and combining enhanced retrieval technology with a large language model, the problem of data acquisition and integration difficulties in carbon accounting for reservoirs and hydropower stations has been solved. This has enabled efficient and accurate extraction of carbon accounting parameters, which is applicable to water conservancy and environmental science research in multiple fields.

CN121743367APending Publication Date: 2026-03-27CHONGQING INST OF GREEN & INTELLIGENT TECH CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies for carbon accounting in reservoirs and hydropower stations suffer from problems such as difficulty in data acquisition, heterogeneous information sources and low utilization efficiency, and difficulty in data integration, resulting in low efficiency and high cost of carbon accounting work.

Method used

By employing a large language model-based approach, a knowledge graph for carbon accounting in reservoirs and hydropower stations is constructed. This is followed by multimodal document parsing and vectorized knowledge base construction. Combined with enhanced retrieval techniques, this enables efficient and accurate extraction of core parameters for carbon accounting.

Benefits of technology

It improves the completeness and accuracy of data collection, reduces deployment costs, and can automate the processing of complex environmental impact assessment reports. It is suitable for hydropower carbon accounting research in small and medium-sized water conservancy design institutes, research institutions, and universities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121743367A_ABST
    Figure CN121743367A_ABST
Patent Text Reader

Abstract

The invention relates to a reservoir hydropower station carbon accounting key information extraction method and system based on a large language model, and belongs to the technical field of artificial intelligence and environmental science. The method mainly comprises the following steps that firstly, domain entities, attributes and relations are extracted through a large language model, and a reservoir hydropower station carbon accounting domain knowledge graph is constructed; secondly, core concept nodes are recognized based on the knowledge graph, retrieval keywords are generated, and target literatures are accurately obtained from the multi-source heterogeneous data; thirdly, performing multi-modal analysis on the literature, extracting unstructured text and table data, and converting the unstructured text and table data into a structured document; secondly, performing semantic partitioning and vectorization processing on the document content, and constructing a vector knowledge base; and finally, responding to user query, realizing enhanced retrieval through multi-path recall and reordering, and accurately obtaining carbon accounting information meeting multiple limitation conditions. According to the method provided by the invention, efficient and accurate extraction of the carbon accounting core parameters can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and environmental science technology, and relates to a method and system for extracting key information for carbon accounting of reservoirs and hydropower stations based on a large language model. Background Technology

[0002] Carbon accounting refers to the process of quantifying and assessing greenhouse gas emissions and their absorption using scientific methods and models. It is a key technological tool for addressing global climate change and achieving the goals of "carbon peaking and carbon neutrality." Hydropower stations, as important sources of clean energy and key water conservancy projects, have become a focus of attention in environmental science and energy fields due to their net greenhouse gas emissions during operation—the carbon flux effect—and their full life-cycle carbon footprint covering construction, operation, and decommissioning. Conducting scientific and systematic carbon accounting for hydropower can not only objectively assess the environmental benefits of hydropower but also provide data support for strategies to improve energy structure and mitigate climate change.

[0003] Currently, carbon accounting data for reservoirs and hydropower stations mainly rely on field measurement methods, model simulation methods, and life cycle assessment methods. However, these methods face significant bottlenecks in practice: (1) Data acquisition is difficult and costly. Long-term field monitoring (such as static box method and eddy covariance method) requires huge investment and high manpower costs, and can only cover a limited time section and geographical location. For the large number of existing and under-construction reservoirs with different geographical locations and reservoir characteristics, it is almost impossible to form a comprehensive and continuous measured database, resulting in the problem of high sparsity and insufficient representativeness of existing data. (2) Information sources are heterogeneous and utilization efficiency is low. A large number of key parameters required for carbon accounting are widely distributed in a large number of heterogeneous documents such as environmental impact assessment reports, feasibility study reports, hydrological situation monitoring data, academic journal papers, design specifications, and operation logs. This information is mostly in the form of unstructured natural language or semi-structured tables, which relies on manual review and manual extraction. The process is cumbersome, inefficient, and prone to errors, which seriously restricts the large-scale development of carbon accounting work. (3) Data integration is difficult. The data from the aforementioned heterogeneous sources exhibit numerous inconsistencies in parameter naming (e.g., "reservoir area" can be expressed as "submerged area under normal water level" or "water surface area corresponding to total reservoir capacity"), units of measurement, and statistical caliber. The workload for manual data cleaning, alignment, and standardization is enormous, representing a core challenge in current data integration. Therefore, to address these difficulties, a knowledge base for reservoir carbon flux and hydropower carbon footprint needs to be designed for data integration. Furthermore, key information extraction techniques should be used to extract crucial information for carbon accounting, providing data support for relevant researchers.

[0004] To address the challenges of information extraction, various information extraction technologies have been explored by academia and industry. These key information extraction methods can be categorized into rule-based, statistical, and deep learning-based methods. However, when these existing technologies are applied to the carbon accounting scenario of reservoirs and hydropower stations, there is a significant technological gap. (1) Rule-based methods require experts to design several matching rules to complete the matching of key information. However, the matching rules are relatively fixed, and the parameters and terminology involved in reservoir carbon accounting are professional and complex. Their expression in the text is flexible and varied, often accompanied by complex numerical values, units, and limiting conditions. Designing and maintaining a complete set of extraction rules for these situations is extremely costly, resulting in poor rule completeness and weak transferability. (2) Although statistical and deep learning-based methods reduce the design of manual rules, these methods rely on large-scale, high-quality domain-annotated corpora for model training. They also depend on the amount of data, information content, and key information distribution of the training data, resulting in poor robustness and generalization ability of key information extraction. In the highly specialized interdisciplinary field of reservoir and hydropower carbon accounting, the cost of constructing such annotated datasets is extremely high, and expert resources are scarce. This results in insufficient training samples for the model, making it difficult to meet application requirements in terms of extraction accuracy and robustness for non-standard tables, complex sentences, and out-of-vocabulary words in reports.

[0005] In recent years, large language models (LLMs), represented by GPT and LLaMA, have demonstrated powerful contextual understanding, logical reasoning, and zero-shot / few-shot learning capabilities, bringing new technological opportunities to address the aforementioned challenges. However, there is currently a lack of systematic research on deeply integrating the capabilities of large language models with expertise in reservoir hydropower carbon accounting, and no dedicated model or technical solution has emerged that can automatically and accurately construct a carbon accounting knowledge base and efficiently extract key parameters from it. Therefore, developing a deeply customized large language model application solution specifically for this scenario has become an urgent technical challenge to be solved in this field. Summary of the Invention

[0006] In view of this, the purpose of this invention is to provide a method and system for extracting key information for carbon accounting of reservoirs and hydropower stations based on a large language model. This method and system address the shortcomings of existing technologies in terms of low efficiency in acquiring carbon accounting information from reservoirs and hydropower stations, strong data heterogeneity, and difficulty in data integration. Through a systematic process that includes domain knowledge graph construction, multimodal document parsing, vectorized knowledge base construction, and retrieval-enhanced reasoning, the core parameters of carbon accounting can be extracted efficiently and accurately.

[0007] To achieve the above objectives, the present invention provides the following technical solution: A method for extracting key information for carbon accounting of reservoirs and hydropower stations based on a large language model, the method specifically includes the following steps: S1. Construct a knowledge graph for carbon accounting of reservoirs and hydropower stations: Use a large language model to extract entities, attributes and relationships in the field of carbon accounting of reservoirs and hydropower stations from external data sources, and construct a domain knowledge graph that includes core entities such as reservoirs and greenhouse gases, as well as key attributes such as reservoir age, installed capacity and emission factors. S2. Obtain target literature in the field of carbon accounting for reservoirs and hydropower stations: Based on the knowledge graph of the field, core concept nodes including methane emissions, eutrophication of water bodies, and carbon source effect of reservoirs are identified by analyzing node centrality. Based on the core concept nodes, search keyword combinations are automatically generated to accurately obtain target literature from heterogeneous data sources including environmental impact assessment reports, academic papers, and technical specifications. S3. Analyze the carbon accounting content of reservoirs and hydropower stations: Perform multimodal content analysis on the acquired target documents, identify unstructured text in the documents, including carbon flux calculation formulas and tables with emission data, and convert the extracted carbon accounting content into a structured document format. S4. Construct a vector knowledge base for carbon accounting of reservoirs and hydropower stations: Semantically segment the content of the structured document, whereby the segmentation aims to aggregate contextual information describing specific facts of carbon accounting within the same text block. These specific facts include the name of a specific reservoir or hydropower station, the type of greenhouse gas, and environmental constraints of the emission pathway. Subsequently, the segmented text content is vectorized and stored in a vector database to construct a vector knowledge base. S5. Perform enhanced retrieval: In response to user queries regarding reservoirs, hydropower stations, greenhouse gases and emissions, use the query and the knowledge base to perform multi-path recall and reordering to retrieve the candidate text set that most completely responds to all the limiting conditions in the query. S6. Extract key information for carbon accounting: Input the query and candidate text set into the large language model. The large model will extract the numerical value, unit, and limiting conditions of the query from the candidate text set and output them in a structured form.

[0008] Furthermore, in step S1, the external data source is obtained by calling a network search tool using a large language model, and the extracted entities are merged and deduplicated; the relationship types in the knowledge graph include causal and correlation logic unique to the field of carbon accounting.

[0009] Furthermore, the merging and deduplication process specifically involves: using a large language model to determine whether semantically similar entities belong to the same entity, and merging multiple statements that are determined to be the same entity.

[0010] Furthermore, in step S2, the core concept nodes are entities that have a significant impact on the field, identified from the knowledge graph based on centrality calculations, such as methane, submerged vegetation, and eutrophication of water bodies; the search keyword generation strategy is to semantically combine these core nodes with their neighboring relational nodes to generate highly relevant professional search term combinations such as "reservoir methane diffusion flux" or "submerged vegetation carbon emissions".

[0011] Furthermore, in step S3, the multimodal content parsing includes using a deep learning model to identify specific content elements in the document. The content elements include at least a data table containing emission factors or water chemistry parameters, and a mathematical formula representing a greenhouse gas flux calculation model. The identified content is then converted into Markdown or TeX format.

[0012] Furthermore, in step S4, the semantic block strategy aims to aggregate contextual information describing a single carbon accounting fact within the same text block, the defining conditions of which include the name of a specific reservoir or hydropower station, the type of greenhouse gas, and the emission path.

[0013] Furthermore, in step S5, the enhanced retrieval includes: performing intent analysis and rewriting on the user query to generate an optimized query instruction; executing a multi-path recall strategy, including vector retrieval based on semantic similarity and precise matching based on entity keywords; and finally, using a re-ranking model, prioritizing the selection of text blocks that simultaneously satisfy semantic and entity constraints and have the most complete information as the candidate set.

[0014] Furthermore, the multi-path recall strategy also includes: using the domain knowledge graph constructed in step S1 to perform reasoning retrieval through the relationship paths between entities, so as to integrate information scattered in different sources and use it to answer complex queries that require multi-step reasoning.

[0015] Furthermore, in step S6, the large language model is guided to extract information through a preset prompt word template, which includes a structured data model customized for reservoir carbon accounting; the large language model determines information fields based on the query, and correctly associates and fills the information scattered in the candidate text set into the data model through contextual reasoning.

[0016] This invention also provides a system for extracting key information for carbon accounting of reservoirs and hydropower stations based on a large language model.

[0017] The beneficial effects of this invention are as follows: This invention boasts high data integrity. By introducing a customized knowledge graph in the field of reservoir and hydropower carbon accounting, the system can understand the professional concepts and complex relationships within this field. This allows for the collection of comprehensive literature on reservoir and hydropower stations, fundamentally improving the completeness of information collection and effectively solving the challenge of accurately extracting data from heterogeneous and unstructured reports. Secondly, the data processing workflow of this invention is systematic and traceable. It provides an end-to-end automated solution from literature acquisition and intelligent parsing to knowledge base construction. All extracted data points, such as a specific methane flux value, can be accurately traced back to their original position in environmental impact assessment reports or academic papers, fully meeting the stringent requirements for data source credibility in scientific research review and official carbon audits. Furthermore, this invention has low application threshold and cost. It eliminates the need for costly retraining or fine-tuning of the model itself. Through ingenious system design and Prompt engineering, high-precision extraction can be achieved, significantly reducing deployment costs. This makes it affordable for small and medium-sized water conservancy design institutes, research institutions, and even university teams to conduct large-scale, systematic hydropower carbon accounting research. Finally, by employing retrieval enhancement and chain-thinking strategies, this invention not only overcomes the limitations of context length and illusion problems when large models process long and professional environmental impact assessment reports, but also successfully constrains and guides its powerful general reasoning capabilities to solve the scientific problem of reservoir carbon accounting, which is characterized by sparse data and complex issues. This provides a typical example for the in-depth application of general artificial intelligence technology in the interdisciplinary field of water conservancy and environmental science.

[0018] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0019] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 This is a flowchart of the key information extraction process for carbon accounting of reservoirs and hydropower stations based on a large language model in this invention. Figure 2 This is a flowchart of the knowledge graph-based literature collection process in this invention; Figure 3 This is a flowchart of the multimodal document structure conversion method in this invention; Figure 4 This is a flowchart of the process for retrieving relevant information and extracting key information for carbon accounting of reservoirs and hydropower stations in this invention; Figure 5This is a flowchart of the key information extraction process using the Three Gorges Reservoir as an example in this invention. Detailed Implementation

[0020] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0021] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0022] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0023] Figure 1 The flowchart shown in the figure illustrates the key information extraction process for carbon accounting of reservoirs and hydropower stations based on a large language model in this invention. The method provided by this invention specifically includes the following: Professional literature collection: Before conducting professional literature collection, it is first necessary to query network knowledge related to reservoir hydropower stations, including news reports, professional literature, technical reports, etc. These contents will help define the key entities in the field of reservoir hydropower stations and their interrelationships. This process mainly includes four steps: the large model calls the Web search tool, entity relationship extraction, entity merging and deduplication, and knowledge graph analysis. Through these steps, effective literature retrieval keywords can be generated. Specifically, in this embodiment, based on the globally recognized greenhouse gas inventory guidelines (such as the supplementary methodology for reservoirs in the IPCC 2019), and an authoritative carbon flux calculation model, a basic domain ontology is manually defined. This ontology includes core entities (such as reservoirs, hydropower stations, greenhouse gases, etc.), attributes (such as reservoir age, geographical location, power generation, etc.), and relationships (such as located in, generated by, affected by, etc.). Then, the large language model is used to call external search tools to consult relevant web pages and literature. Let the collection of web pages and literature obtained by the large model through the Web search tool be , where represents web pages or literature in the relevant field. Next, the large model is used again to extract entities and relationships based on the query materials, obtaining a set of entity sets , a set of attribute sets , and a set of relationship sets . These entities, attributes, and relationships reveal the key elements and interconnections in the carbon accounting of reservoir hydropower stations. There is a problem of duplicate entities in the extracted entities, so it is necessary to merge the same entities. For similar entities , the large language model is used to judge whether they belong to the same entity and merge them. For example, "Three Gorges Hydropower Station" and "ThreeGorges Hydropower Station" are the same entity. Through the large model judgment mechanism, the merged entity is obtained, where . This process can eliminate redundant entities and improve the quality and accuracy of the knowledge graph. The merged entity, attribute, and relationship triples will constitute a large knowledge graph in the field of carbon accounting of reservoir hydropower stations , where is the entity node set, is the connection edge between entities, is the relationship between entities. This knowledge graph shows the research boundaries and core contents of this field, providing a structured knowledge framework for further analysis and retrieval. Finally, by analyzing the connectivity in the knowledge graph, the most core research keywords can be extracted. These keywords help to further search and obtain the most crucial knowledge. For example, methane bubble emissions, methane degassing emissions, reservoir age, etc. will show high centrality and can be used as query keywords. For this reason, the present invention calculates the centrality of each entity based on the knowledge graph. The calculation formula for degree centrality is:

[0024] in This represents the weight of the edge; here, it only takes the values ​​0 and 1, representing the two cases where the edge does not exist and exists, respectively. Represents a node centrality, Represents a node Adjacent nodes. In this way, the present invention can determine the importance of a node in the graph based on the number of connected edges, and further extract research keywords based on these core nodes. These keywords will further optimize literature retrieval and provide data support for the establishment of a knowledge base. Figure 2 This is a flowchart of the knowledge graph-based literature collection process in this invention.

[0025] For document content extraction and format conversion, this invention uses a multimodal document structure conversion method to convert heterogeneous environmental impact assessment reports, academic journal articles, government technical guidelines and other PDF documents into a unified, model-friendly Markdown format. First, the input PDF is type-identified and divided into scanned version and text version. For scanned version, the OCR engine is called to perform text recognition and extraction and then directly perform structure reconstruction. For text version, it mainly includes two key steps: deep layout parsing and structure reconstruction. The core task is to accurately identify and parse the key information carriers in the carbon accounting report, such as carbon flux calculation formula, data tables containing emission factors, and geographical location maps indicating inundation areas. (1) In the layout parsing step, this invention uses a deep learning-based target detection algorithm to parse the document page. It not only identifies elements such as text and images, but more importantly, it accurately locates the bounding boxes of complex elements such as tables and mathematical formulas, and marks each element with a type label to form a document structure metadata containing precise location information. (2) Structure reconstruction mainly includes two steps: content encoding and content synthesis. For the parsed table and formula areas, specialized recognition models are invoked respectively. For the formulas, the models are trained to correctly identify greenhouse gas chemical formulas. , ) and complex physical units (such as The model losslessly converts data to TeX code, fundamentally solving the problem of conventional OCR models mistaking subscripts and superscripts for ordinary characters, leading to the loss of formula semantics. For tables, the model accurately reconstructs their row and column structures, especially emission factor data tables containing complex headers and units. In the content synthesis step, based on the layout detection results, text, images, tables, formulas, etc., are reorganized into a Markdown file according to the original document's structure. The final output Markdown file is a high-fidelity digital copy that preserves the semantics and logical hierarchy of the original document, providing a solid data foundation for deep understanding and zero-illusion question answering in downstream large-scale language models.

[0026] The construction of a knowledge base for reservoir carbon flux and hydropower carbon footprint mainly includes three steps: text data preprocessing, word vectorization technology, and the construction of a knowledge base for reservoir and hydropower station carbon accounting. For the collected Markdown format data, low-density information such as the cover, table of contents, and references are first removed, followed by the removal of image links, non-standard text, etc. Non-standard text includes garbled characters, misspellings, and semantically ambiguous text. The preprocessed text data is then segmented, and text fragments are embedded using vectorization technology to obtain embedding vectors. The text fragments and their embedding vectors are stored in a vector database for easy subsequent retrieval. During the construction of the knowledge base, it is categorized according to journal articles, reports, news, etc., to facilitate the maintenance and updating of each sub-knowledge base. Simultaneously with the construction of the reservoir and hydropower station carbon accounting knowledge base, LLM is used to extract entities and relationships from each text fragment, constructing a knowledge graph for querying.

[0027] The retrieval enhancement method, which integrates domain knowledge, includes context-aware rewriting of query commands, dual recall based on vectors and keywords, context-aware content reordering, construction of extracted prompts, and execution. The context-aware rewriting of query commands begins with the user's original query command. Using an LLM (Limited Language Management) analysis of the current and historical context, the intent of the instruction is identified, supplementing it with rich semantic information and carbon accounting-specific constraints, thereby generating an optimized query instruction. Optimized query command The semantic representation is more user-friendly and accurate for downstream retrieval and extraction tasks. In the dual recall step based on vectors and keywords, the optimized query command... A dual-path recall process is implemented, retrieving relevant information from the established reservoir and hydropower station knowledge base. This step includes two stages: broad semantic recall and deep keyword filtering. In the broad semantic recall stage, a text embedding model is used to interpret the query commands. Convert to vector representation Then, semantic similarity retrieval is performed in the vector database to recall an initial candidate text set. Then, in the deep keyword filtering stage, LLM is used in parallel from... The process identifies and extracts core keywords (such as greenhouse gases, emission factors, reservoir parameters, and hydropower station operating time) and constructs a keyword list. This keyword list will serve as a filter for the candidate set. Perform precise matching, retaining only text that explicitly contains key information, generating a more focused candidate set. ,in In the context-aware content re-ranking process, to ensure that the information ultimately provided to the large model is the most relevant and least noisy, the system employs a re-ranking model for the candidate set. The model performs processing. It evaluates not only individual text fragments but also query instructions. The relevance of the text also considers the logical coherence between text fragments, prioritizing those that best match the query instruction. The text fragments are from a reservoir hydropower station carbon accounting project. The final reordering model outputs an optimized list of the top N most relevant text fragments. ,in In the process of building and executing the extractive prompting engineering step, the system will use optimized query commands. and a list of reordered text fragments Combined, this constructs a structured, context-rich extraction cue. This message was submitted to the core LLM for processing. Because... The data already contains clear instructions and highly relevant background knowledge, from which LLM can accurately and reliably extract the key carbon accounting parameters required by the user, such as specific emission flux values, the models or emission factors on which the calculations are based, and the original sources of the data.

[0028] = +

[0029] The system also includes an optional enhancement module. For some complex carbon accounting problems, information may be scattered across different documents or data sources, requiring multi-step reasoning to obtain, such as querying the indirect carbon emissions of steel used in the construction of a hydropower station. To address this, the system integrates a knowledge graph-based recall path. During the recall phase, the system utilizes the query knowledge graph built during the knowledge base construction to perform reasoning-based retrieval through the relationship paths between entities, integrating information scattered across different sources to answer complex questions that cannot be solved by ordinary retrieval methods.

[0030] In this embodiment, carbon accounting data extraction from the Three Gorges Reservoir is taken as the research object. As the world's largest water conservancy project, the Three Gorges Reservoir has a long research period, diverse data sources, and complex data formats, making it an ideal case for testing the capabilities of this invention. The overall implementation process can be referred to... Figure 5 .

[0031] refer to Figure 5 The implementation process of this invention is as follows: First, a knowledge graph for the carbon flux of the Three Gorges Reservoir is systematically constructed. Using a large language model (LLM) and a web search interface, various articles and materials related to "carbon flux of the Three Gorges Reservoir," including news reports, research abstracts, and technical reports, are automatically retrieved. Based on a pre-defined domain ontology based on IPCC guidelines, the LLM extracts key entities and relationships from unstructured text, such as "carbon flux of the Three Gorges Reservoir," "methane," and "carbon dioxide," and constructs semantic relationships between them, such as "carbon dioxide is a component of the carbon flux of the Three Gorges Reservoir" and "static box method is used for measuring water-air interface diffusion flux." Furthermore, the system extracts different attributes for each entity, such as the reservoir's construction time and installed capacity, thus forming a structured, machine-understandable knowledge network. After the knowledge graph is constructed, the system performs a centrality analysis to identify core nodes within the domain, such as "carbon flux of the Three Gorges Reservoir," "reservoir carbon source effect," and "water-air interface diffusion." Based on these core nodes, the system further automatically generates a series of search keywords according to neighboring nodes, including "Three Gorges Reservoir carbon flux", "reservoir carbon sink effect", "nitrogen cycle and N2O flux in the drawdown zone", "spatiotemporal heterogeneity", "carbon emissions during operation", and "water-air interface diffusion", to ensure that all relevant Three Gorges Reservoir carbon flux data can be accessed comprehensively and accurately.

[0032] Subsequently, the system followed Figure 3 The process shown employs multimodal intelligent parsing, where this module extracts all content from the document and converts it into a structured Markdown file. This file not only preserves the integrity of the original information but also makes the large amount of professional data and complex formulas programmatically processable. Next, the system semantically segments the Markdown file, ensuring that each text block is a complete semantic unit. Each semantic block is then embedded into a vector and stored along with the original text blocks in a high-performance vector database, laying the foundation for subsequent fast and accurate retrieval.

[0033] Finally, this system follows Figure 4The intelligent question-answering process shown responds to user queries. When a user enters a natural language query at the system front end, such as "What is the methane diffusion flux during the initial operation of the Three Gorges Reservoir? What factors are involved?", the LLM first rewrites the query into a more specific one, such as "Query the literature regarding the specific measured values, units, and error ranges of methane surface diffusion flux during the initial operation of the Three Gorges Reservoir, as well as discussions of the key factors affecting the early methane emission flux after the Three Gorges Reservoir's impoundment, such as submerged vegetation decomposition, water temperature, or eutrophication." The system then performs a rapid search in the knowledge base and quickly locates relevant key text information. In the reordering stage, the system prioritizes the Top N text blocks that are highly semantically relevant to the user's intent and simultaneously contain precise numerical values, time constraints, and specific locations to ensure the relevance and accuracy of the information. Finally, the system fills the user's query intent and these high-scoring text blocks into a preset prompt word template and sends it to the LLM for final extraction. Through in-depth analysis and reasoning, the LLM provides the precise extraction results required by the user.

[0034] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for extracting key information for carbon accounting of reservoirs and hydropower stations based on a large language model, characterized in that, The method specifically includes the following steps: S1. Construct a knowledge graph for carbon accounting of reservoirs and hydropower stations: Use a large language model to extract entities, attributes and relationships in the field of carbon accounting of reservoirs and hydropower stations from external data sources, and construct a domain knowledge graph that includes core entities such as reservoirs and greenhouse gases, as well as key attributes such as reservoir age, installed capacity and emission factors. S2. Obtain target literature in the field of carbon accounting for reservoirs and hydropower stations: Based on the knowledge graph of the field, core concept nodes including methane emissions, eutrophication of water bodies, and carbon source effect of reservoirs are identified by analyzing node centrality. Based on the core concept nodes, search keyword combinations are automatically generated to accurately obtain target literature from heterogeneous data sources including environmental impact assessment reports, academic papers, and technical specifications. S3. Analyze the carbon accounting content of reservoirs and hydropower stations: Perform multimodal content analysis on the acquired target documents, identify unstructured text in the documents, including carbon flux calculation formulas and tables with emission data, and convert the extracted carbon accounting content into a structured document format. S4. Construct a vector knowledge base for carbon accounting of reservoirs and hydropower stations: Semantically segment the content of the structured document, whereby the segmentation aims to aggregate contextual information describing specific facts of carbon accounting within the same text block. These specific facts include the name of a specific reservoir or hydropower station, the type of greenhouse gas, and environmental constraints of the emission pathway. Subsequently, the segmented text content is vectorized and stored in a vector database to construct a vector knowledge base. S5. Perform enhanced retrieval: In response to user queries regarding reservoirs, hydropower stations, greenhouse gases and emissions, use the query and the knowledge base to perform multi-path recall and reordering to retrieve the candidate text set that most completely responds to all the limiting conditions in the query. S6. Extract key information for carbon accounting: Input the query and candidate text set into the large language model. The large model will extract the numerical value, unit, and limiting conditions of the query from the candidate text set and output them in a structured form.

2. The method for extracting key information for carbon accounting of reservoirs and hydropower stations based on a large language model according to claim 1, characterized in that, In step S1, the external data source is obtained by calling the network search tool using the large language model, and the extracted entities are merged and deduplicated; the relationship types in the knowledge graph include causal and correlation logic unique to the field of carbon accounting.

3. The method for extracting key information for carbon accounting of reservoirs and hydropower stations based on a large language model according to claim 2, characterized in that, The merging and deduplication process specifically involves: using a large language model to determine whether semantically similar entities belong to the same entity, and merging multiple statements that are determined to be the same entity.

4. The method for extracting key information for carbon accounting of reservoirs and hydropower stations based on a large language model according to claim 3, characterized in that, In step S2, the core concept nodes are entities that have a significant impact on the field and are identified from the knowledge graph based on centrality calculation; the search keyword generation strategy is to generate highly relevant professional search term combinations by semantically combining these core nodes with their neighboring relation nodes.

5. The method for extracting key information for carbon accounting of reservoirs and hydropower stations based on a large language model according to claim 4, characterized in that, In step S3, the multimodal content parsing includes using a deep learning model to identify specific content elements in the document. The content elements include at least a data table containing emission factors or water chemistry parameters, and a mathematical formula representing a greenhouse gas flux calculation model. The identified content is then converted into Markdown or TeX format.

6. The method for extracting key information for carbon accounting of reservoirs and hydropower stations based on a large language model according to claim 5, characterized in that, In step S4, the semantic segmentation strategy aims to aggregate contextual information describing a single carbon accounting fact within the same text block, the defining conditions of which include the name of a specific reservoir or hydropower station, the type of greenhouse gas, and the emission pathway.

7. The method for extracting key information for carbon accounting of reservoirs and hydropower stations based on a large language model according to claim 6, characterized in that, In step S5, the enhanced retrieval includes: performing intent analysis and rewriting on the user query to generate an optimized query instruction; executing a multi-path recall strategy, including vector retrieval based on semantic similarity and exact matching based on entity keywords; and finally, using a re-ranking model, prioritizing the selection of text blocks that simultaneously satisfy semantic and entity constraints and have the most complete information as the candidate set.

8. The method for extracting key information for carbon accounting of reservoirs and hydropower stations based on a large language model according to claim 7, characterized in that, The multi-path recall strategy also includes: using the domain knowledge graph constructed in step S1 to perform reasoning retrieval through the relationship paths between entities, so as to integrate information scattered in different sources and use it to answer complex queries that require multi-step reasoning.

9. The method for extracting key information for carbon accounting of reservoirs and hydropower stations based on a large language model according to claim 7, characterized in that, In step S6, the large language model is guided to extract information by a preset prompt word template, which includes a structured data model customized for reservoir carbon accounting. The large language model determines the information fields based on the query and correctly associates the information scattered in the candidate text set and fills it into the data model through contextual reasoning.

10. A system for extracting key information for carbon accounting of reservoirs and hydropower stations based on a large language model, characterized in that, The system employs the method described in any one of claims 1 to 9.

Citation Information

Cited By

  • Carbon inspection domain knowledge graph construction method and system based on prompt engineering

    CN122072837A