Methane emission facility geographic information extraction method and system based on large language model
Patent Information
- Application Number
- CN202611047351.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-07-15
AI Technical Summary
[0039] (1) This invention targets methane emission facilities in energy industries such as coal mines and oil and gas. It organizes facility type terms, industry terms, place names, location descriptions and enterprise names into a retrieval task set according to the combination retrieval rules. It also automatically collects, extracts and preprocesses text from multiple sources of public information such as government announcements, enterprise public information, news reports, bidding announcements, environmental impact assessment documents, industry website information and public map markings through a web page information targeted acquisition program. This can gather facility clues from scattered sources into a candidate corpus set, reducing the repetitive work of manual item-by-item retrieval and sorting.
Smart Images

Figure CN122614970B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of methane emission monitoring and geographic information technology, specifically relating to a method and system for extracting geographic information of methane emission facilities based on a large language model. Background Technology
[0002] Accurately grasping the geographical information of methane emission facilities in the energy sector is a crucial foundation for conducting methane monitoring, emission accounting, and emission reduction supervision. For the coal mining industry (such as ventilation shafts, gas extraction pumping stations, and gas power plants), and the oil and gas industry (such as gas gathering stations, compressor stations, processing plants, well sites, and combined stations), their spatial location, type attributes, and operational status information directly affect the accuracy and completeness of methane emission source identification, emission estimation, and supervision of key areas.
[0003] Official geological survey data and corporate environmental reports from some countries or regions are often difficult to obtain, or may be incomplete, outdated, or inconsistent in their presentation. Existing methods rely heavily on manual searches, expert judgment, or official data sources, making it difficult to systematically and comprehensively acquire information on methane emission facilities in the global energy sector, which can easily lead to a lack of basic data for emissions accounting.
[0004] Geographic information on methane emission facilities in the energy sector is typically scattered across multiple sources of publicly available information, including government announcements, corporate documents, news reports, bidding documents, and map annotations. This results in complex information sources, diverse data formats, and a significant amount of unstructured text. Furthermore, descriptions of facility names, addresses, spatial relationships, and operational information in these publicly available texts are often inconsistent, with discrepancies between different sources regarding names, incomplete address descriptions, and ambiguous location representations. This makes it difficult to directly structure and represent the relevant facility information in coordinate form, further increasing the challenges of facility identification, geographic analysis, and unified organization.
[0005] Furthermore, existing methods mostly rely on manual retrieval and organization, which is not only inefficient and has limited coverage, but also lacks effective verification of the authenticity, completeness, and timeliness of facility information. This can lead to biases, omissions, or duplications in the extraction results, making it difficult to form a unified, standardized, and reliable geographic information database for methane emission facilities in the energy industry. Consequently, this affects the quality of the basic data for subsequent methane emission accounting, key emission source identification, and regulatory applications. Summary of the Invention
[0006] To address the aforementioned shortcomings in existing technologies, this invention provides a method and system for extracting geographic information of methane emission facilities based on a large language model. Targeting methane emission facilities in the energy industry, the method involves targeted acquisition of publicly available online information, extraction and parsing of facility entities and location information from candidate corpora using a large language model, and auxiliary comparison using map services, energy infrastructure databases, and remote sensing observation results. This results in the location and attribute information of potential methane emission facilities, thereby improving the automation, coverage, and application efficiency of acquiring geographic information for methane emission facilities. It solves the problems of scattered information sources, low efficiency of manual retrieval, difficulty in directly structuring publicly available text, inconsistencies between facility names and location descriptions, and difficulty in quickly pinpointing the spatial location of facilities in existing energy industry methane emission facility geographic information acquisition processes.
[0007] To achieve the aforementioned objectives, the technical solution adopted by this invention is: a method for extracting geographic information of methane emission facilities based on a large language model, comprising the following steps:
[0008] S1. Determine the target facility type range based on methane emission facilities in the energy industry, construct a search term list and generate combined search rules to form a search task set, conduct targeted retrieval and crawling of publicly available information from multiple sources based on the search task set, preprocess the extracted text content, and generate a candidate corpus set.
[0009] S2. Input the candidate corpus into the large language model parsing module to perform semantic understanding, entity recognition, attribute extraction and location parsing, extract facility name, facility type, address information and spatial location description, and combine with context information to perform name merging, address completion and location reasoning to form a potential facility structured information record including candidate coordinates and / or candidate spatial range.
[0010] S3. Record the structured information of potential facilities and input it into the spatial location and auxiliary comparison module. Use network map services or open map services to perform geocoding, coordinate lookup or spatial range delineation, and combine hyperspectral satellite remote sensing methane plume observation results, international energy infrastructure database and network map information for auxiliary comparison, and output a set of potential methane emission facilities results.
[0011] Furthermore: S1 includes the following sub-steps:
[0012] S11. For methane emission facilities in the energy industry, the target facility types are determined according to the needs of methane emission monitoring and regulation. The target facility types include ventilation shafts, gas extraction pumping stations, gas power plants and related auxiliary facilities in the coal mining industry, as well as gas gathering stations, compressor stations, processing plants, well sites and combined stations in the oil and gas industry.
[0013] S12. Construct a search term list based on the target facility type, combine the search term list according to preset rules to generate combined search rules for targeted acquisition of publicly available information on the Internet, and organize the target facility type, search term list and combined search rules to form a search task set;
[0014] S13. Perform targeted retrieval and crawling of publicly available information from multiple sources of the network based on the retrieval task set, and generate publicly available web page information, including government announcements, publicly available corporate information, news reports, bidding documents, environmental impact assessment documents, industry website information and publicly available map labeling information;
[0015] S14. Extract text from publicly available information on web pages, preprocess the extracted text content, and generate a candidate corpus.
[0016] Furthermore: In S12, the search term list includes facility type terms, industry terminology terms, place names, location description terms, and company name terms;
[0017] Among them, facility type terms are used to characterize the standard names, common alternative names, abbreviations, acronyms, and corresponding Chinese and English expressions of different facilities; industry terminology terms are used to characterize professional terms related to facility construction, operation, expansion, renovation, shutdown, emissions, and related engineering activities; place names include the names of geographical units such as countries, states, provinces, cities, counties, townships, mining areas, oil fields, gas fields, and industrial parks; location descriptive terms are words that indicate spatial and distance relationships, including "located in," "located in," "situated on," "east side," "west side," "nearby," "adjacent to," and "approximately ... kilometers away"; and company name terms include the full name, abbreviation, and historical name of the company.
[0018] The combined search rules can take the form of “place name + facility type”, “company name + facility type”, “place name + industry term + facility type”, and “company name + location description + facility type”.
[0019] Furthermore, S2 includes the following sub-steps:
[0020] S21. Input the candidate corpus into the large language model parsing module, perform semantic analysis on the candidate corpus, and extract attribute information, including facility name, facility type, affiliated enterprise, affiliated project, administrative division, and operating status.
[0021] S22. Extract address information and location descriptions related to facility location from the candidate corpus. Perform name merging, entity merging, address completion and place name disambiguation on facility abbreviations, aliases, acronyms, historical names of enterprises, local place names, fuzzy place names or cross-language place name expressions in the text. Generate candidate coordinates and / or candidate spatial ranges based on place name matching, address parsing and spatial relationship reasoning.
[0022] S23. Based on attribute information, address information, location description, and candidate coordinates and / or candidate spatial range, the information is uniformly organized to form a structured information record of potential facilities.
[0023] Furthermore: In S2, the large language model parsing module includes a text input unit, a semantic analysis unit, an entity extraction unit, a location parsing unit, and a result organization unit;
[0024] The text input unit is used to receive a set of candidate corpora.
[0025] The semantic analysis unit is used to identify semantic information related to the target facility in the candidate corpus;
[0026] The entity extraction unit is used to extract facility name, facility type, parent company, parent project, administrative division, and operational status;
[0027] The location resolution unit is used to extract address information and location description, and generate candidate coordinates and / or candidate spatial ranges;
[0028] The results organization unit is used to organize the extracted results into a structured record of potential facility information.
[0029] Furthermore, S3 includes the following sub-steps:
[0030] S31. Input the structured information of potential facilities into the spatial location and auxiliary comparison module, and use network map services or open map services to perform geocoding, coordinate lookup or spatial range delineation on the address information, place name information and spatial location description in the structured information of potential facilities to obtain map location results.
[0031] S32. Based on the map location results, the potential facility information is compared with the energy infrastructure database. The comparison includes similarity of facility names, consistency of facility types, consistency of enterprise affiliation, and proximity of locations. Combined with map label names, road relationships, distribution of surrounding industrial facilities, and geographical references in the network map, the coordinates of potential facilities are reversed and confirmed to obtain the database comparison results.
[0032] S33. Perform spatial correlation analysis between the candidate coordinates and / or candidate spatial range of potential facilities and the hyperspectral satellite remote sensing methane plume point source or anomalous region to obtain remote sensing spatial correlation results. The spatial correlation analysis includes proximity matching, buffer matching and spatial overlay analysis.
[0033] S34. Based on the map location results, remote sensing spatial correlation results, database comparison results, and source information, output the potential methane emission facility result set.
[0034] A geographic information extraction system for methane emission facilities based on a large language model includes:
[0035] The module for targeted acquisition of publicly available online information and construction of candidate corpus is used to determine the range of target facility types, construct a search term list and generate a search task set, and to retrieve, automatically collect, extract, deduplicate and clean publicly available online information from multiple sources to form a candidate corpus set.
[0036] The large language model facility entity extraction and location parsing module is used to perform semantic understanding, entity recognition, attribute extraction and location parsing on the candidate corpus set to form a potential facility structured information record including candidate coordinates and / or candidate spatial range;
[0037] The spatial location, auxiliary comparison and result output module is used to geocode, look up coordinates or delineate the spatial range of the structured information records of potential facilities, and to perform auxiliary comparisons with hyperspectral satellite remote sensing methane plume observation results, international energy infrastructure database and online map information, and output a set of potential methane emission facilities results.
[0038] The beneficial effects of this invention are as follows:
[0039] (1) This invention targets methane emission facilities in energy industries such as coal mines and oil and gas. It organizes facility type terms, industry terms, place names, location descriptions and enterprise names into a retrieval task set according to the combination retrieval rules. It also automatically collects, extracts and preprocesses text from multiple sources of public information such as government announcements, enterprise public information, news reports, bidding announcements, environmental impact assessment documents, industry website information and public map markings through a web page information targeted acquisition program. This can gather facility clues from scattered sources into a candidate corpus set, reducing the repetitive work of manual item-by-item retrieval and sorting.
[0040] (2) This invention uses a large language model parsing module to perform semantic understanding, entity recognition, attribute extraction and location parsing on facility names, facility types, affiliated enterprises, address information, spatial location descriptions and operating status in candidate corpora. It also combines context to merge, complete and disambiguate facility abbreviations, aliases, historical enterprise names, local place names, ambiguous place names, cross-language expressions and relative location descriptions. This can convert unstructured public text into potential structured facility information records with unified fields, and alleviate the difficulty of information organization caused by inconsistent facility name expressions, incomplete address descriptions and ambiguous location descriptions.
[0041] (3) The present invention uses a spatial location and auxiliary comparison module to record the structured information of potential facilities and input it into the map service for geocoding, coordinate lookup and candidate spatial range delineation. It also combines the GEM series energy infrastructure database, network map annotation, road relationships, distribution of surrounding industrial facilities and methane plume observation results for auxiliary comparison, which can reduce the risk of mislocation caused by relying on a single text source or a single geocoding result.
[0042] (4) For relative location descriptions such as “east side of the mining area”, “near the road”, and “along the pipeline” where the coordinates are difficult to determine directly, the present invention can first form a candidate spatial range, and then combine map features and methane plume anomalies to perform spatial correlation analysis, output candidate coordinates, candidate range, source information and auxiliary comparison results, providing basic data support for methane emission accounting, remote sensing attribution analysis, screening of key emission sources and manual verification. Attached Figure Description
[0043] Figure 1 This is a flowchart of the method for extracting geographic information of methane emission facilities based on a large language model according to the present invention.
[0044] Figure 2 This is a schematic diagram of the geographic information extraction system for methane emission facilities based on a large language model according to the present invention.
[0045] Figure 3 This is a schematic diagram of the candidate spatial range of coal mine facilities generated based on publicly available text and the results of large language model parsing.
[0046] Figure 4 The candidate locations of coal mine facilities after map-assisted verification.
[0047] Figure 5 This is a schematic diagram of the methane plume observation results in the target area.
[0048] Figure 6 This is a schematic diagram showing the overlay analysis of methane plume observation results with map features and candidate facility locations. Detailed Implementation
[0049] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0050] Example 1:
[0051] like Figure 1As shown, in one embodiment of the present invention, the method for extracting geographic information of methane emission facilities based on a large language model includes the following steps:
[0052] S1. Determine the target facility type range based on methane emission facilities in the energy industry, construct a search term list and generate combined search rules to form a search task set, conduct targeted retrieval and crawling of publicly available information from multiple sources based on the search task set, preprocess the extracted text content, and generate a candidate corpus set.
[0053] S2. Input the candidate corpus into the large language model parsing module to perform semantic understanding, entity recognition, attribute extraction and location parsing, extract facility name, facility type, address information and spatial location description, and combine with context information to perform name merging, address completion and location reasoning to form a potential facility structured information record including candidate coordinates and / or candidate spatial range.
[0054] S3. Record the structured information of potential facilities and input it into the spatial location and auxiliary comparison module. Use network map services or open map services to perform geocoding, coordinate lookup or spatial range delineation, and combine hyperspectral satellite remote sensing methane plume observation results, international energy infrastructure database and network map information for auxiliary comparison, and output a set of potential methane emission facilities results.
[0055] S1 includes the following steps:
[0056] S11. For methane emission facilities in the energy industry, the target facility types are determined according to the needs of methane emission monitoring and regulation. The target facility types include ventilation shafts, gas extraction pumping stations, gas power plants and related auxiliary facilities in the coal mining industry, as well as gas gathering stations, compressor stations, processing plants, well sites and combined stations in the oil and gas industry.
[0057] S12. Construct a search term list based on the target facility type, combine the search term list according to preset rules to generate combined search rules for targeted acquisition of publicly available information on the Internet, and organize the target facility type, search term list and combined search rules to form a search task set;
[0058] S13. Perform targeted retrieval and crawling of publicly available information from multiple sources of the network based on the retrieval task set, and generate publicly available web page information, including government announcements, publicly available corporate information, news reports, bidding documents, environmental impact assessment documents, industry website information and publicly available map labeling information;
[0059] S14. Extract text from publicly available information on web pages, preprocess the extracted text content, and generate a candidate corpus.
[0060] S11 specifically refers to:
[0061] Based on the needs of methane emission monitoring, emission accounting, and regulatory analysis, the scope of methane emission facility types targeted by this invention is determined. The target facilities mainly include those closely related to methane emissions in the coal mining and oil and gas industries.
[0062] In practice, the types of target facilities can be added to, subtracted from, refined, or reorganized based on different countries, regions, or application tasks to form a set of search objects adapted to specific scenarios. By pre-defining the scope of facility types, the relevance and effectiveness of subsequent online information retrieval can be improved.
[0063] The facilities in the coal mining industry include, but are not limited to: ventilation shafts, gas extraction pumping stations, gas power plants, and auxiliary facilities related to gas extraction, transportation, and emission; the facilities in the oil and gas industry include, but are not limited to: gas gathering stations, compressor stations, processing plants, well sites, combined stations, and facilities related to natural gas extraction, transportation, and processing. For potential differences in facility terminology found in publicly available information from different regions, corresponding common industry expressions can also be retained for use as the basis for subsequent search terminology construction.
[0064] In S12, after identifying the target facility type, a search term list is constructed for targeted acquisition of publicly available information on the Internet. The search term list is used to describe the facility itself, industry background, geographical scope, and common location clues in the text. The search term list includes facility type terms, industry terminology terms, place names, location description terms, and company name terms.
[0065] Among them, facility type terms are used to characterize the standard names, common alternative names, abbreviations, acronyms, and corresponding Chinese and English expressions of different facilities; industry terminology terms are used to characterize professional terms related to facility construction, operation, expansion, renovation, shutdown, emissions, and related engineering activities; place names include the names of geographical units such as countries, states, provinces, cities, counties, townships, mining areas, oil fields, gas fields, and industrial parks; location descriptive terms are words that indicate spatial and distance relationships, including "located in," "located in," "situated on," "east side," "west side," "nearby," "adjacent to," and "approximately ... kilometers away"; and company name terms include the full name, abbreviation, and historical name of the company.
[0066] After constructing the search terminology, keywords of different categories are combined according to preset rules to generate combined search rules for targeted acquisition of publicly available online information. The combined search rules can take the form of "place name + facility type," "company name + facility type," "place name + industry term + facility type," and "company name + location description + facility type." In this embodiment, search expressions such as "a certain region + compressor station," "a certain oil field + natural gas processing plant," "a certain coal mine + ventilation shaft," and "a certain company + gas power plant" can be generated. For foreign regions or non-Chinese-speaking areas, corresponding search expressions in English or other languages can be further constructed to expand the scope of candidate corpus acquisition.
[0067] In specific implementation, the preset rules include: using place names such as administrative divisions, mining areas, oil fields, gas fields, or industrial parks as constraints, and combining them with facility type terms such as ventilation wells, return air wells, gas extraction pumping stations, gas power plants, gas gathering stations, gas compressor stations, processing plants, well sites, and combined stations; using the full name, abbreviation, or historical name of the enterprise as the main item, and combining it with facility type terms or industry terms; and using location descriptions such as "located in," "located on the east side," "nearby," "adjacent to," and "approximately [number] kilometers away" as location clues, and combining them with enterprise names, place names, or facility type terms.
[0068] After the completed facility type set, search term list, and combined search rules are organized and unified, a search task set is formed, which serves as the input for subsequent targeted retrieval and automated collection of publicly available online information.
[0069] In S13, targeted retrieval and automated collection of publicly available information from multiple sources are performed based on the retrieval task set. In practice, webpage information targeting acquisition programs, search engine result parsing programs, site-specific search interfaces, or other methods of obtaining publicly available information can be used to perform item-by-item retrieval and page acquisition from different information sources. For each retrieval result, the webpage title, text content, publication time, source URL, publishing entity, and supplementary information related to the target facility can be further extracted. For map-based information, the label name, label category, address field, and corresponding location description information can also be recorded simultaneously; for announcements, bidding instructions, environmental impact assessment documents, project lists, news reports, and tabular pages, their original paragraph, list, or table structure can be retained.
[0070] During the retrieval and automated data collection phase, webpage content and attachments containing facility type terms, place names, company names, or location descriptions should be prioritized for retention. Results that are clearly irrelevant, duplicated, or severely lacking in information should be initially removed. For overseas regions or cross-language scenarios, simultaneous searches can be performed based on the corresponding language search expressions to expand the scope of access to overseas facility information.
[0071] In S14, the text extraction objects include HTML webpage text, announcement content, news report text, bidding instructions, PDF attachments, Word attachments, environmental impact assessment documents, table-type pages, project lists, map POI label names, map POI categories, map POI address fields and related descriptive information, while retaining webpage titles, source URLs, publication time, publishing entity, attachment names and other auxiliary information that can characterize the source attributes.
[0072] After text extraction, the collected results undergo basic organization and preprocessing. Preprocessing includes: removing duplicate information based on source URL, title similarity, and text repetition rate; removing webpage tags, scripts, and navigation bar text; deleting invalid symbols; correcting garbled characters; standardizing line breaks and spaces; separating titles and body text; removing irrelevant paragraphs; removing duplicate paragraphs; uniformly converting PDF / Word / table text; and standardizing map POI fields. Simultaneously, the collected results are uniformly numbered, and an original information record table is established. This original information record table includes information number, search keywords, information source, publication time, webpage title, original text, supplementary descriptions, source URL, and collection time.
[0073] After basic processing, a candidate corpus is formed. The candidate corpus includes publicly available text content related to the target facility, as well as its source information, time information, and structural information. It serves as input data for subsequent large language models to perform facility name recognition, facility type determination, address information extraction, and location parsing.
[0074] S2 includes the following steps:
[0075] S21. Input the candidate corpus into the large language model parsing module, perform semantic analysis on the candidate corpus, and extract attribute information, including facility name, facility type, affiliated enterprise, affiliated project, administrative division, and operating status.
[0076] S22. Extract address information and location descriptions related to facility location from the candidate corpus. Perform name merging, entity merging, address completion and place name disambiguation on facility abbreviations, aliases, acronyms, historical names of enterprises, local place names, fuzzy place names or cross-language place name expressions in the text. Generate candidate coordinates and / or candidate spatial ranges based on place name matching, address parsing and spatial relationship reasoning.
[0077] S23. Based on attribute information, address information, location description, and candidate coordinates and / or candidate spatial range, the information is uniformly organized to form a structured information record of potential facilities.
[0078] In S2, the large language model parsing module is used to perform semantic understanding, entity recognition, attribute extraction, position description parsing, and result organization on candidate corpora. The large language model parsing module includes a text input unit, a semantic analysis unit, an entity extraction unit, a position parsing unit, and a result organization unit.
[0079] The text input unit is used to receive a set of candidate corpora.
[0080] The semantic analysis unit is used to identify semantic information related to the target facility in the candidate corpus;
[0081] The entity extraction unit is used to extract facility name, facility type, parent company, parent project, administrative division, and operational status;
[0082] The location resolution unit is used to extract address information and location description, and generate candidate coordinates and / or candidate spatial ranges;
[0083] The results organization unit is used to organize the extracted results into a structured record of potential facility information.
[0084] In practical implementation, the large language model parsing module can call existing general-purpose large language model services or locally deployed large language models. This invention does not limit the specific model name, model parameters, or training method. The large language model is used in this application as a semantic parsing tool. It works in conjunction with prompt templates, field constraints, structured output rules, and result organization rules to adapt short texts, long texts, list-type texts, table-type texts, and news report-type texts to improve the consistency and completeness of information extraction.
[0085] The prompt template includes a task description, a list of target facility types, a list of fields to be extracted, output format constraints, source evidence citation rules, and uncertain field marking rules; the structured output rules can require the model to output facility name, facility type, affiliated enterprise, affiliated project, administrative division, address information, location description, candidate coordinates and / or candidate spatial range, source link, extraction confidence, and remarks in Excel spreadsheets, field tables, JSON, or other structured data formats.
[0086] S21 specifically refers to:
[0087] The candidate corpus is input into the large language model parsing module, where the semantic analysis unit performs semantic recognition on the text content based on preset prompting rules and extracts core entity information related to the target facility. Core entity information includes at least the facility name, facility type, address information, and spatial location description, and may further include the parent company, project, administrative division, operational status, and other auxiliary attribute information.
[0088] For cases where the standard facility name and type are directly given in the text, the corresponding fields can be extracted directly. For cases where only the project name, facility abbreviation, company name, or partial facility description appears, the information in the context can be used for completion and judgment. For facility name abbreviations, aliases, acronyms, and different expressions that often appear in public texts, they can be identified and merged using a large language model combined with contextual semantics. For the full name, abbreviation, and historical name of a company, they can also be uniformly merged into the same entity.
[0089] Different extraction methods can be used for different text structures. For short texts, facility names, facility types, and location descriptions can be directly identified; for long texts, facility-related information can be extracted by combining paragraph context; for list-type and table-type texts, facility names, enterprise entities, address information, and affiliated attributes can be identified by combining adjacent field relationships; for news report-type texts, semantic completion can be performed by combining publication time, publishing entity, reporting subject, and contextual content.
[0090] In S22, address information and location descriptions related to the facility location are extracted from the candidate corpus. These include standard administrative or regional names such as provinces, cities, counties, townships, mining areas, oil fields, gas fields, and industrial parks, as well as relative location descriptions such as “located in,” “at,” “located in,” “about … kilometers away,” “near,” “east side,” and “west side.”
[0091] For information with a clear address or a clear place name level, candidate location results can be generated through place name matching, address parsing, or coordinate retrieval. For information with only partial place names, vague place names, or relative location descriptions, address completion and location inference can be performed by combining administrative divisions, place name references, publishing entities, source URLs, reporting areas, and contextual semantics. For example, when the text only mentions the name of a township, a mining area, an oil field, or an industrial site, a more complete address level can be completed by combining its superior administrative divisions and source information. When the text contains descriptions such as "east of a mining area," "approximately several kilometers from a certain road," or "adjacent to an industrial park," corresponding candidate location results can be generated by combining place name references and directional distance relationships.
[0092] For situations where facilities have the same name, different place names, old and new place names coexist, or there are spelling differences across languages in publicly available texts, place name disambiguation and entity merging can be performed by combining contextual information. When multiple candidate location results are generated for the same facility, they can be sorted according to address completeness, place name matching degree, clarity of location description, and source reliability, and one or more candidate coordinates and / or candidate spatial ranges with the most reasonable location can be retained.
[0093] In S23, the structured information record of potential facilities should at least include the facility name, facility type, affiliated enterprise, affiliated project, administrative division, address information, location description, candidate coordinates and / or candidate spatial range, operating status, source text number, source link, extraction confidence level, and result remarks. In specific implementation, information extraction markers, location resolution markers, source reliability markers, database comparison markers, or remote sensing spatial association markers can be added to the structured information record of potential facilities to facilitate subsequent spatial location, auxiliary comparison, and result output. After this step, the candidate corpus can be converted into structured information records of potential facilities, providing basic data for subsequent map service geocoding, database-assisted comparison, and remote sensing spatial association analysis.
[0094] S3 includes the following steps:
[0095] S31. Input the structured information of potential facilities into the spatial location and auxiliary comparison module, and use network map services or open map services to perform geocoding, coordinate lookup or spatial range delineation on the address information, place name information and spatial location description in the structured information of potential facilities to obtain map location results.
[0096] S32. Based on the map location results, the potential facility information is compared with the energy infrastructure database. The comparison includes similarity of facility names, consistency of facility types, consistency of enterprise affiliation, and proximity of locations. Combined with map label names, road relationships, distribution of surrounding industrial facilities, and geographical references in the network map, the coordinates of potential facilities are reversed and confirmed to obtain the database comparison results.
[0097] S33. Perform spatial correlation analysis between the candidate coordinates and / or candidate spatial range of potential facilities and the hyperspectral satellite remote sensing methane plume point source or anomalous region to obtain remote sensing spatial correlation results. The spatial correlation analysis includes proximity matching, buffer matching and spatial overlay analysis.
[0098] S34. Based on the map location results, remote sensing spatial correlation results, database comparison results, and source information, output the potential methane emission facility result set.
[0099] In S31, the spatial location and auxiliary comparison module receives structured information records of potential facilities and performs map location, spatial verification, and multi-source information comparison on the address information, place name information, spatial location description, and candidate location results. The spatial location and auxiliary comparison module includes a data input unit, a map location unit, a candidate range determination unit, an energy infrastructure database comparison unit, a remote sensing methane plume spatial correlation unit, and a result organization unit.
[0100] The data input unit is used to receive the facility name, facility type, affiliated enterprise, administrative division, address information, location description, candidate coordinates and / or candidate spatial range from the structured information record of potential facilities;
[0101] The map placement unit is used to call online map services or open map services to perform geocoding, coordinate lookup and map retrieval on address information, place name information and spatial location description to obtain one or more map candidate locations.
[0102] The candidate range determination unit is used to generate or adjust the candidate spatial range for location descriptions that cannot be directly converted into a single coordinate, such as "east side of the mining area", "near the road", "along the pipeline", and "adjacent industrial park". It combines geographical references, directional relationships, distance relationships, road relationships and the distribution of surrounding features.
[0103] The energy infrastructure database comparison unit is used to compare candidate locations on the map with the energy infrastructure database published by Global Energy Monitor or other publicly available energy facility data. The comparison includes similarity of facility names, consistency of facility types, consistency of enterprise affiliation, and spatial proximity.
[0104] The remote sensing methane plume spatial correlation unit is used to perform proximity matching, buffer matching, or spatial overlay analysis on candidate locations or candidate spatial ranges on the map with self-built methane plume inversion results, publicly available methane plume products, or plume point sources and anomalous areas provided by third-party methane emission observation platforms.
[0105] The results organization unit is used to integrate map location results, candidate range determination results, energy infrastructure database comparison results, remote sensing methane plume spatial correlation results, and original source information to form a record of potential methane emission facility results, including candidate coordinates, candidate spatial range, and auxiliary comparison information.
[0106] The spatial location and auxiliary comparison module can be implemented using existing geographic information processing programs, web map service interfaces, spatial database query programs, and spatial analysis programs. This invention is not limited to specific map platforms, database products, or spatial analysis software. The improvement of this module lies in connecting map location, energy infrastructure database comparison, and methane plume spatial correlation according to a unified data structure, which is used for multi-source auxiliary verification of facility candidate locations generated by large language models.
[0107] In S31, address information, place name information, and spatial location descriptions in the structured information records of potential facilities are geocoded, have their coordinates reversed, or their spatial range delineated. For facility records with clear addresses, clear place name levels, or direct location capabilities, corresponding candidate coordinates can be generated using online map services or open map services. For facility records with only local place names, relative positional relationships, or range descriptions, spatial location can be achieved by combining map label names, surrounding road relationships, industrial facility distribution, and relevant geographic references, and corresponding candidate spatial ranges can be generated. Online map services or open map services can include map platforms such as Google Maps, Tianditu, and Aowei Interactive Map, which have capabilities for place name retrieval, address resolution, remote sensing image browsing, coordinate reverse lookup, and feature-assisted interpretation.
[0108] When there are multiple candidate locations for the same facility, one or more reasonable candidate coordinates and / or candidate spatial ranges can be retained based on address completeness, place name matching degree, clarity of location description and consistency of map search results.
[0109] In S32, after spatial placement is completed, the structured information of potential facilities is recorded and compared with the energy infrastructure database. The comparison includes similarity in facility names, consistency in facility types, consistency in enterprise affiliation, and proximity in location.
[0110] The energy infrastructure database includes a series of databases published by the global energy monitoring platform Global Energy Monitor (GEM), such as the Global Coal Mine Tracker (GCMT), the Global Oil and Gas Extraction Tracker (GOGET), the Global Gas Infrastructure Tracker (GGIT), and the Global Methane Emitters Tracker (GMET). In addition, it also incorporates databases or spatial data products such as the Global Fuel Exploitation Inventory (GFEI), publicly available map data, or self-built energy facility inventories.
[0111] For potential facilities for which corresponding records can be found in the above databases and whose names, types, and locations are relatively consistent, the corresponding database information can be retained as supplementary supporting information; for records with some fields that are consistent but have differences in names, translations, or coordinate deviations, further comparison can be made by combining the original text, contextual information, and map placement results; for facilities for which there are no corresponding records in the database but whose textual evidence is sufficient, they can also be retained as potential results.
[0112] In S33, spatial correlation analysis is used to determine whether there is a spatial proximity or overlap between the potential facility location and the methane emission anomaly. The hyperspectral satellite remote sensing methane plume observation results can be derived from self-built hyperspectral methane plume inversion results, publicly available methane plume products, or third-party methane emission observation platforms, such as Carbon Mapper data products. Specifically, the coordinates or spatial range of the candidate facility can be matched with the methane plume point source or anomaly region identified by the satellite through proximity matching, buffer matching, or spatial overlay analysis. If there is a methane plume anomaly around the potential facility that matches its spatial location, the result can be retained as additional spatial evidence; if no obvious correlation is found, the record of the potential facility can still be retained, and further processing can be carried out in conjunction with map location results, text sources, and database comparison results.
[0113] For situations where only a large area can be determined in the text but it is difficult to pinpoint a single facility, spatial delineation can be performed first based on candidate areas, and then a matching analysis within the range can be conducted in conjunction with remote sensing plume anomaly areas.
[0114] In S34, the methane emission facility result set includes facility name, facility type, industry, company, administrative division, address information, location description, candidate coordinates and / or candidate spatial range, information source, map location results, database comparison results, and remote sensing spatial association results. For the same facility appearing with multiple names, locations, or attribute descriptions from different sources, the relevant results can be retained side-by-side or organized as a master record, with other information saved as supplementary explanations. The final output includes a potential facility list, coordinate points, candidate ranges, attribute information table, and related explanatory information, used for subsequent methane emission accounting, remote sensing attribution analysis, screening of key emission sources, and manual verification.
[0115] like Figure 2 As shown, the geographic information extraction system for methane emission facilities based on a large language model includes:
[0116] The module for targeted acquisition of publicly available online information and construction of candidate corpus is used to determine the range of target facility types, construct a search term list and generate a search task set, and to retrieve, automatically collect, extract, deduplicate and clean publicly available online information from multiple sources to form a candidate corpus set.
[0117] The large language model facility entity extraction and location parsing module is used to perform semantic understanding, entity recognition, attribute extraction and location parsing on the candidate corpus set to form a potential facility structured information record including candidate coordinates and / or candidate spatial range;
[0118] The spatial location, auxiliary comparison and result output module is used to geocode, look up coordinates or delineate the spatial range of the structured information records of potential facilities, and to perform auxiliary comparisons with hyperspectral satellite remote sensing methane plume observation results, international energy infrastructure database and online map information, and output a set of potential methane emission facilities results.
[0119] Example 2:
[0120] This embodiment provides a specific implementation case based on Embodiment 1. This embodiment takes ventilation shaft facilities in the coal mining industry as the target object and performs intelligent extraction of geographic information.
[0121] First, based on the characteristics of publicly available texts related to methane emission facilities in the coal mining industry, a search expression for targeted acquisition of publicly available information online is constructed, including combinations such as "coal mine + ventilation shaft," "mine + return air shaft," "enterprise + ventilation shaft project," and "county X + town Y + ventilation shaft." Based on these search expressions, targeted retrieval, automated collection, and text extraction are performed on government announcements, publicly available corporate data, news reports, bidding documents, environmental impact assessment documents, and publicly available map annotations. The obtained results are then deduplicated, cleaned, and fundamentally organized to form a candidate corpus.
[0122] In this embodiment, the candidate corpus contains the following text: Text A1 states "The project construction includes a return air shaft, a ventilation room, and ancillary works"; Text A2 states "The ventilation shaft is located on the east side of a coal mine area in Y Town, X County"; Text A3 states "The construction unit is a mining company, and the project location is within Y Town"; Text A4 is map labeling information, with the label name matching the coal mine name. The above texts retain their titles, main text, publication time, source URL, and supplementary explanatory information, and are input into the large language model parsing module after being uniformly numbered.
[0123] The large language model parsing module first performs semantic recognition on the candidate corpora. After parsing, "return air shaft" in text A1, "ventilation shaft" in text A2, and other expressions related to ventilation engineering are identified as the same type of target facility, and the facility type is determined to be a coal mine ventilation shaft facility. Further, the enterprise to which it belongs is a mining company, and the administrative division information of County X, Town Y is extracted from texts A2 and A3. The location description "east side of the mining area" is extracted from text A2. For different expressions such as "ventilation shaft," "return air shaft," and "ventilation shaft project" appearing in different texts, the large language model merges the names based on the context; for texts that only mention "within Town Y" without providing a complete hierarchical address, the address is completed by combining the mine name, administrative division, and source information from other candidate corpora.
[0124] After entity extraction and address completion, the facility location is analyzed. For clearly extracted place name information such as County X, Town Y, and a coal mine area, corresponding candidate location results are generated. For the relative location description "east of the mine area," it is not directly determined as a single coordinate point, but rather a candidate spatial range is generated based on the mine area's scope and the eastward orientation. If different candidate corpora produce multiple location results, they are sorted according to address completeness, source reliability, text clarity, and place name matching degree, and the candidate coordinates and candidate spatial ranges with the most reasonable location are retained.
[0125] like Figure 3 As shown, the large language model forms a candidate spatial range for coal mine ventilation shaft facilities based on the mine area name, administrative division, and location descriptions such as "east side of the mine area" in the candidate corpus, which is used to limit the subsequent map retrieval and spatial delineation range.
[0126] Subsequently, the parsed address information, place name information, and candidate location results are input into the map service for geocoding, coordinate lookup, and spatial delineation. For results that can directly match County X, Town Y, and a specific coal mine area, the map service returns the corresponding location; for the relative location description "east side of the mine area," a candidate area to the east is formed based on the mine area's boundaries. If the map results show that the candidate area is basically consistent with the road relationships, mine boundaries, and surrounding industrial facility distribution mentioned in the text, the corresponding candidate location result is retained as a potential facility location.
[0127] like Figure 4 As shown, with the assistance of map services and public map annotations, the above-mentioned candidate spatial ranges are subjected to coordinate reverse lookup, verification of surrounding features, and range shrinkage to obtain the candidate locations of coal mine facilities.
[0128] Furthermore, the potential facility locations are compared with energy infrastructure databases and hyperspectral satellite remote sensing methane plume observations. If a record corresponding to the coal mine or its affiliated enterprise exists in the energy infrastructure database, and its name, type, or location relationship matches the candidate result, this database information is retained as supplementary supporting information. If an anomalous methane plume region is located near the candidate location, the remote sensing result is retained as additional spatial evidence. Even if no completely corresponding database record is found or no obvious plume anomaly is observed, the potential facility result can still be retained and further processed in conjunction with textual sources and map location results.
[0129] like Figure 5 and Figure 6 As shown, the hyperspectral satellite remote sensing observation results of methane plumes in this area are used as spatial evidence and overlaid with map features, road relationships, mining area boundaries and candidate facility locations to help confirm the spatial correlation between plume anomalies and target facilities.
[0130] Through the above processing, a potential result record for the ventilation shaft facility is ultimately generated. This result record includes the facility name, facility type, affiliated company, administrative division, address information, location description, candidate coordinates and candidate spatial range, source text number, map location results, and auxiliary comparison results. This result record can serve as the basis for subsequent methane emission accounting, remote sensing attribution analysis, screening of key emission sources, and manual verification.
[0131] Example 3:
[0132] This embodiment provides a specific implementation case based on Embodiment 1. This embodiment takes the compressor station facilities of an oil and gas field enterprise as the target object and performs intelligent extraction of geographic information.
[0133] First, based on the characteristics of publicly available texts related to methane emission facilities in the oil and gas industry, a search expression for targeted acquisition of publicly available information online is constructed, including combinations such as "a certain oil and gas field enterprise + compressor station," "a certain oil field + compressor station," "a certain company + natural gas station," and "county A + township B + compressor station." Based on these search expressions, targeted retrieval, automated collection, and text extraction are performed on publicly available enterprise information, news reports, bidding documents, environmental impact assessment documents, project announcements, and publicly available map annotations. The obtained results are then deduplicated, cleaned, and fundamentally organized to form a candidate corpus.
[0134] In this embodiment, the candidate corpus contains the following text content: Text B1 states "The project construction includes compressor units, station process equipment, and supporting ancillary works"; Text B2 states "The compressor station is located in Township B of County A and is an important station in the natural gas transmission system of a certain oil and gas field enterprise"; Text B3 states "The construction unit is a branch of a certain oil and gas field, and the station site is located along a certain gas pipeline"; Text B4 states "The station is located near the intersection of a certain gathering and transmission trunk line and a county road"; Text E is map labeling information, with the label name corresponding to the name of the oil and gas block or station to which the enterprise belongs. The above texts retain their titles, main text, publication time, source URL, and supplementary explanatory information, and are input into the large language model parsing module after being uniformly numbered.
[0135] The large language model parsing module first performs semantic recognition on the candidate corpora. After parsing, expressions such as "compressor unit" and "station process unit" in text B1, "compressor station" in text B2, and "natural gas transmission system station" in text B3 are identified as the same type of target facility, and the facility type is determined to be a compressor station facility. Further, the entity to which the facility belongs is extracted from texts B2 and B3 as an oil and gas field company or its subsidiary; the administrative division information of county A and township B is extracted from text B2; and the location description "located near the intersection of a certain gathering and transmission trunk line and a county road" is extracted from text B4. For different expressions such as "compressor station," "natural gas station," and "station project" appearing in different texts, the large language model performs name merging and entity merging based on the context; for cases where only the company name, oil and gas block name, or gas transmission system name appears in the text, but the complete station address information is not directly provided, the address is completed by combining the administrative division, station attributes, and location descriptions from other candidate corpora.
[0136] After entity extraction and address completion, the facility location is analyzed. For clearly extracted place name information such as County A, Township B, and areas along the pipeline, corresponding candidate location results are generated. For relative location descriptions such as "located along a pipeline," "near a county road," or "near an intersection area," they are not directly determined as a single coordinate point, but rather a candidate spatial range is formed by combining the location of the gas pipeline, road, and surrounding industrial sites. If different candidate corpora produce multiple location results, they are sorted according to address completeness, source reliability, clarity of text description, consistency of enterprise entity, and degree of matching of location description, and the candidate coordinates and candidate spatial ranges with more reasonable locations are retained.
[0137] Subsequently, the parsed address information, place name information, and candidate location results are input into the map service for geocoding, coordinate lookup, and spatial delineation. For results that can directly match County A, Township B, and related pipeline areas, the map service returns the corresponding location. For descriptions such as "near intersection areas" or "stations along the pipeline," candidate areas are formed by combining road relationships, station distribution characteristics, surrounding industrial facilities, and geographical references. If the map results show that the candidate area is basically consistent with the pipeline direction, road relationships, and the location of the enterprise's block mentioned in the text, the corresponding candidate location result is retained as a potential facility location.
[0138] Furthermore, the potential facility locations are compared with energy infrastructure databases and hyperspectral satellite remote sensing methane plume observations. If a record corresponding to the oil and gas field enterprise, related block, or natural gas transmission system exists in the energy infrastructure database, and its name, type, or location relationship matches the candidate result, the database information is retained as supplementary supporting information. If an anomalous methane plume area is present near the candidate location, the remote sensing result is retained as additional spatial evidence. Even if no completely corresponding database record is found or no obvious plume anomaly is observed, the potential facility result can still be retained and further processed in conjunction with text sources, enterprise information, and map location results.
[0139] Through the above processing, a potential result record for the compressor station facility is ultimately generated. This result record includes the facility name, facility type, parent company, administrative division, address information, location description, candidate coordinates and candidate spatial range, source text number, map location results, and auxiliary comparison results. This result record can serve as the basis for subsequent methane emission accounting, remote sensing attribution analysis, screening of key emission sources, and manual verification.
[0140] In the description of this invention, the above are merely preferred embodiments and are not intended to limit the scope of protection of this invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for extracting geographic information of methane emission facilities based on a large language model, characterized in that, Includes the following steps: S1. Determine the target facility type range based on methane emission facilities in the energy industry, construct a search term list and generate combined search rules to form a search task set, conduct targeted retrieval and crawling of publicly available information from multiple sources based on the search task set, preprocess the extracted text content, and generate a candidate corpus set. S2. Input the candidate corpus into the large language model parsing module to perform semantic understanding, entity recognition, attribute extraction and location parsing, extract facility name, facility type, address information and spatial location description, and combine with context information to perform name merging, address completion and location reasoning to form a potential facility structured information record including candidate coordinates and / or candidate spatial range. S3. Record the structured information of potential facilities and input it into the spatial location and auxiliary comparison module. Use network map services or open map services to perform geocoding, coordinate lookup or spatial range delineation, and combine hyperspectral satellite remote sensing methane plume observation results, international energy infrastructure database and network map information for auxiliary comparison, and output a set of potential methane emission facilities results. S3 includes the following steps: S31. Input the structured information of potential facilities into the spatial location and auxiliary comparison module, and use network map services or open map services to perform geocoding, coordinate lookup or spatial range delineation on the address information, place name information and spatial location description in the structured information of potential facilities to obtain map location results. S32. Based on the map location results, the potential facility information is compared with the energy infrastructure database. The comparison includes similarity of facility names, consistency of facility types, consistency of enterprise affiliation, and proximity of locations. Combined with map label names, road relationships, distribution of surrounding industrial facilities, and geographical references in the network map, the coordinates of potential facilities are reversed and confirmed to obtain the database comparison results. S33. Perform spatial correlation analysis between the candidate coordinates and / or candidate spatial range of potential facilities and the hyperspectral satellite remote sensing methane plume point source or anomalous region to obtain remote sensing spatial correlation results. The spatial correlation analysis includes proximity matching, buffer matching and spatial overlay analysis. S34. Based on the map location results, remote sensing spatial correlation results, database comparison results, and source information, output the potential methane emission facility result set.
2. The method for extracting geographic information of methane emission facilities based on a large language model according to claim 1, characterized in that, S1 includes the following steps: S11. For methane emission facilities in the energy industry, the target facility types are determined according to the needs of methane emission monitoring and regulation. The target facility types include ventilation shafts, gas extraction pumping stations, gas power plants and related auxiliary facilities in the coal mining industry, as well as gas gathering stations, compressor stations, processing plants, well sites and combined stations in the oil and gas industry. S12. Construct a search term list based on the target facility type, combine the search term list according to preset rules to generate combined search rules for targeted acquisition of publicly available information on the Internet, and organize the target facility type, search term list and combined search rules to form a search task set; S13. Perform targeted retrieval and crawling of publicly available information from multiple sources of the network based on the retrieval task set, and generate publicly available web page information, including government announcements, publicly available corporate information, news reports, bidding documents, environmental impact assessment documents, industry website information and publicly available map labeling information; S14. Extract text from publicly available information on web pages, preprocess the extracted text content, and generate a candidate corpus.
3. The method for extracting geographic information of methane emission facilities based on a large language model according to claim 2, characterized in that, In S12, the search term list includes facility type terms, industry terminology terms, place names, location description terms, and company name terms; Among them, facility type terms are used to characterize the standard names, common alternative names, abbreviations, acronyms, and corresponding Chinese and English expressions of different facilities; industry terminology terms are used to characterize professional terms related to facility construction, operation, expansion, renovation, shutdown, emissions, and related engineering activities; place names include the names of geographical units such as countries, states, provinces, cities, counties, townships, mining areas, oil fields, gas fields, and industrial parks; location descriptive terms are words indicating spatial and distance relationships, including "located in", "located in", "situated on", "east", "west", "nearby", "adjacent to", and "approximately ... kilometers away"; enterprise name terms include the full name, abbreviation, and historical name of the enterprise. The combined search rules can take the form of "place name + facility type", "company name + facility type", "place name + industry term + facility type" and "company name + location description + facility type".
4. The method for extracting geographic information of methane emission facilities based on a large language model according to claim 1, characterized in that, S2 includes the following steps: S21. Input the candidate corpus into the large language model parsing module, perform semantic analysis on the candidate corpus, and extract attribute information, including facility name, facility type, affiliated enterprise, affiliated project, administrative division, and operating status. S22. Extract address information and location descriptions related to facility location from the candidate corpus. Perform name merging, entity merging, address completion and place name disambiguation on facility abbreviations, aliases, acronyms, historical names of enterprises, local place names, fuzzy place names or cross-language place name expressions in the text. Generate candidate coordinates and / or candidate spatial ranges based on place name matching, address parsing and spatial relationship reasoning. S23. Based on attribute information, address information, location description, and candidate coordinates and / or candidate spatial range, the information is uniformly organized to form a structured information record of potential facilities.
5. The method for extracting geographic information of methane emission facilities based on a large language model according to claim 4, characterized in that, In S2, the large language model parsing module includes a text input unit, a semantic analysis unit, an entity extraction unit, a location parsing unit, and a result organization unit; The text input unit is used to receive a set of candidate corpora. The semantic analysis unit is used to identify semantic information related to the target facility in the candidate corpus; The entity extraction unit is used to extract facility name, facility type, parent company, parent project, administrative division, and operational status; The location resolution unit is used to extract address information and location description, and generate candidate coordinates and / or candidate spatial ranges; The results organization unit is used to organize the extracted results into a structured record of potential facility information.
6. A geographic information extraction system for methane emission facilities based on a large language model, employing the geographic information extraction method for methane emission facilities based on a large language model as described in any one of claims 1 to 5, characterized in that, The system includes: The module for targeted acquisition of publicly available online information and construction of candidate corpus is used to determine the range of target facility types, construct a search term list and generate a search task set, and to retrieve, automatically collect, extract, deduplicate and clean publicly available online information from multiple sources to form a candidate corpus set. The large language model facility entity extraction and location parsing module is used to perform semantic understanding, entity recognition, attribute extraction and location parsing on the candidate corpus set to form a potential facility structured information record including candidate coordinates and / or candidate spatial range; The spatial location, auxiliary comparison and result output module is used to geocode, look up coordinates or delineate the spatial range of the structured information records of potential facilities, and to perform auxiliary comparisons with hyperspectral satellite remote sensing methane plume observation results, international energy infrastructure database and online map information, and output a set of potential methane emission facilities results.
Citation Information
Patent Citations
Methane parameter estimation method, system, equipment, medium and product
CN120404620A
Intelligent question and answer method and device based on multi-source ESG knowledge base and agent reasoning
CN122240792A