Underground space entity identification method and system based on open source spatio-temporal data
By using network crawler, BiLSTM-CRF and BERT models for text preprocessing in underground space information processing, combining spatiotemporal clustering and graph neural network to complete spatiotemporal attributes, the inefficiency problem of underground space entity recognition and spatiotemporal information extraction is solved, and efficient dynamic recognition and intelligent processing of open source data is achieved.
Patent Information
- Application Number
- CN202510626880.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art has problems in the processing of underground space information, insufficient data openness, insufficient semantic understanding, weak spatial and fusion capabilities and low system integration. It is difficult to achieve dynamic identification of underground space entities and efficient extraction of spatial and temporal information, especially in the automatic extraction of complex spatial and temporal features implicitly in open source data.
The underground space entity recognition method based on open source spatiotemporal data is adopted, multimodal data is collected through network crawlers, text preprocessing is performed in combination with BiLSTM-CRF model and BERT model, entity-spatiotemporal structure is constructed by using the BERT+CRF model, and the missing spatiotemporal attribute information is completed through spatiotemporal clustering algorithm and graph neural network.
It realizes efficient identification of underground space entities in open source data and dynamic extraction of space-time information, expands data acquisition boundaries and update capabilities, improves the intelligence level of information processing, supports high-quality dynamic modeling and trend analysis, and improves the operating efficiency of the system.
Smart Images

Figure CN120493929A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information perception and intelligent processing technology, and in particular to a method and system for identifying underground space entities based on open source spatiotemporal data. Background Art
[0002] Currently, the acquisition and management of underground space information primarily relies on traditional geographic information system modeling, remote sensing image interpretation, geological survey data integration, and administrative reporting mechanisms. This technology has some applicability in describing static spatial structures and archiving data during the construction phase, demonstrating particularly stable data support capabilities in mappable areas such as underground pipeline networks, tunneling, and subway lines. However, these methods generally rely on closed data sources and periodic manual update processes, making it difficult to achieve dynamic perception, near-real-time monitoring, and automated deduction of underground space changes. Consequently, they suffer from significant deficiencies in the timeliness and flexibility of urban operations.
[0003] At the same time, the rise of natural language processing (NLP) technology has opened up new possibilities for mining unstructured underground spatial data. Some existing studies have attempted to apply rule-based named entity recognition (NER) techniques to structured text corpora, such as construction bulletins or environmental assessment reports. However, these approaches rely heavily on domain vocabularies and sentence templates, making them difficult to adapt to the complex, ambiguous, and highly erratic language found in open source text corpora. Their semantic parsing capabilities are limited, robustness is poor, and generalization capabilities are insufficient. In particular, recognition rates drop significantly when dealing with non-standard expressions in open contexts such as social media and news information.
[0004] More critically, most existing methods focus solely on identifying "entities" themselves, lacking the ability to collaboratively model entities across temporal dimensions (e.g., start date, lifespan) and spatial dimensions (e.g., latitude and longitude, administrative divisions). Time and location information in text is often non-standardized, context-dependent, or even implicitly expressed, making it difficult for rule-based or shallow model-based methods to accurately parse and map entities. Consequently, the unified extraction and integration of entity information and its spatiotemporal attributes is difficult, hindering the ability to support high-quality dynamic modeling, trend analysis, and identification of spatiotemporal evolution.
[0005] Furthermore, text related to underground spaces often contains ambiguous descriptions, semantic gaps, and incomplete information. For example, expressions like "recently constructed tunnel" and "this section of underground space is closed" often lack clear time, location, or entity boundaries. Traditional methods in such scenarios often rely on manual intervention or outright data discarding. They lack the ability to intelligently complete and infer information through mechanisms such as contextual semantics, similarity pattern transfer, and graph structure reasoning, resulting in significant drawbacks in data utilization and information integrity.
[0006] Finally, from the perspective of overall system architecture, existing underground space information processing systems often exhibit decentralized processes, independent components, and a low degree of automation. The lack of unified data flow scheduling and intelligent collaboration mechanisms across modules such as collection, processing, analysis, and storage limits system efficiency, making it difficult to meet the high-frequency processing and real-time response requirements of massive data volumes. This severely restricts their adoption and value realization in application scenarios such as smart cities, emergency management, and dynamic planning.
[0007] In summary, existing underground space information processing technologies have significant shortcomings in key dimensions such as data openness, depth of semantic understanding, spatiotemporal fusion capabilities, model intelligence, and system integration. Therefore, there is an urgent need for an intelligent system framework that integrates open-source data acquisition, deep semantic analysis, spatiotemporal attribute reasoning, and dynamic behavior modeling. This framework aims to build a semantic recognition and deduction system for the entire lifecycle of underground space for unstructured information, providing solid technical support for spatial governance and intelligent decision-making in future cities. Summary of the Invention
[0008] The present invention aims to solve the problems of inefficiency and inaccuracy in the existing technology of underground space entity recognition and spatiotemporal information extraction, especially in the automatic extraction of complex spatiotemporal features implicit in open source data. A method and system for underground space entity recognition based on open source spatiotemporal data is proposed. By comprehensively applying natural language processing technology, spatiotemporal data analysis algorithms and deep learning models, it can efficiently extract entities related to underground space and their spatiotemporal information from open source Internet data such as news social networks.
[0009] In order to achieve the above purpose, the technical solutions adopted are:
[0010] The present invention provides a method for identifying underground space entities based on open source spatiotemporal data, comprising the following steps:
[0011] Step 1: Collect multimodal spatiotemporal data containing underground space information from multiple open source platforms through web crawlers and API interfaces;
[0012] Step 2: Standardize and preprocess the original text, including denoising, word segmentation, part-of-speech tagging, and syntactic analysis, and combine the BiLSTM-CRF model with the BERT model to improve semantic recognition capabilities;
[0013] Step 3: Identify underground space-related entities by integrating the BERT+CRF model with domain rules;
[0014] Step 4: Based on the spatiotemporal clustering algorithm and graph neural network, the missing spatiotemporal attribute information is completed to construct the entity-spatiotemporal structure.
[0015] According to the underground space entity recognition method based on open source spatiotemporal data of the present invention, further, the multi-source open source platform in step 1 includes news social networks, blogs and forums, and the multimodal spatiotemporal data includes text, image and video data.
[0016] According to the underground space entity recognition method based on open source spatiotemporal data of the present invention, further, the standardized preprocessing in step 2 specifically includes:
[0017] Feed the text into the BERT encoder to obtain the context vector representation of each word;
[0018] Input the vector representation into the BiLSTM model to obtain bidirectional context features;
[0019] The label sequence is modeled through CRF to output the optimal word segmentation and part-of-speech sequence labels.
[0020] According to the method for identifying underground space entities based on open source spatiotemporal data of the present invention, further, in the standardization preprocessing, the method further includes:
[0021] Perform syntactic dependency analysis on the segmented text to extract the core verbs, subject-verb-object relationships, and modifying structures in the sentence, and generate a word dependency graph. The syntactic dependency analysis is implemented using a Transformer- or Tree-LSTM-based syntactic analyzer.
[0022] The word dependency graph is input into step 3 of underground space related entity recognition to assist in determining the semantic role and boundary of the named entity.
[0023] According to the underground space entity recognition method based on open source spatiotemporal data of the present invention, further, the underground space-related entities in step 3 include underground facilities, tunnels, underground pipe networks and underground buildings; the underground space-related entity recognition is achieved by:
[0024] Input the token sequence into the BERT model and output the context vector representation of each word;
[0025] A state transition diagram is constructed based on the vector representation through the CRF module, the optimal label sequence is decoded, and the entity boundaries and categories are determined.
[0026] According to the underground space entity recognition method based on open source spatiotemporal data of the present invention, step 3 further includes:
[0027] Design a set of entity labels specific to the underground space industry to provide a domain-adapted semantic framework for the model, enabling model training and prediction.
[0028] Post-process the model results through the rule engine and use the term dictionary to align and fill in the gaps in the results;
[0029] Merge entities with the same semantics to generate a unique Entity-ID.
[0030] According to the underground space entity recognition method based on open source spatiotemporal data of the present invention, further, the spatiotemporal attribute information in step 4 includes a timestamp and geographic coordinates, and the spatiotemporal clustering algorithm adopts K-means or DBSCAN algorithm.
[0031] According to the underground space entity recognition method based on open source spatiotemporal data of the present invention, further, the missing spatiotemporal attribute information in step 4 is supplemented by the following method:
[0032] Extract the original time and space information associated with the identified underground space entities and preliminarily label them through regular matching and NER model;
[0033] Unify the formatting of non-standardized time phrases and use standard time representation; call the geocoding interface to parse spatial information into latitude and longitude coordinates and administrative division codes;
[0034] Use spatiotemporal clustering algorithms or DBSCAN algorithms to group similar spatiotemporal entities;
[0035] For entities that are not clearly labeled, a support vector regression model is first used to perform regression prediction based on the entity context and semantically similar entities. Then, an entity-entity graph is constructed, with nodes as entities and edges as co-occurrence or context similarity. Graph neural networks are used to propagate and complete node attributes.
[0036] Furthermore, the present invention also provides an underground space entity recognition system based on open source spatiotemporal data, which is used to implement the above-mentioned underground space entity recognition method based on open source spatiotemporal data, including:
[0037] The data acquisition module is used to collect multimodal spatiotemporal data containing underground space information from multiple open source platforms through web crawlers and API interfaces;
[0038] The text preprocessing module is used to perform standardized preprocessing on the original text, including denoising, word segmentation, part-of-speech tagging, and syntactic analysis. It combines the BiLSTM-CRF model with the BERT model to improve semantic recognition capabilities.
[0039] The entity recognition module is used to identify underground space-related entities by integrating the BERT+CRF model with domain rules;
[0040] The spatiotemporal information completion module is used to complete missing spatiotemporal attribute information based on spatiotemporal clustering algorithms and graph neural networks, and to construct entity-spatiotemporal structures.
[0041] The beneficial effects achieved by adopting the above technical solution are:
[0042] (1) The paradigm of information sources has shifted from closed to open. Traditional underground space information acquisition relies on government systems, professional surveys, or closed databases, with long data update cycles and limited coverage. This invention uses open source Internet media as its information base, achieving a fundamental transformation of data sources from "static closed" to "dynamic open". Through the automated integration of news, social platforms, blogs, and other content, it greatly expands the collection boundaries and update capabilities of underground space information.
[0043] (2) The intelligent upgrade of information processing methods from rule-driven to model-driven. This invention embeds artificial intelligence technologies such as deep learning, graph neural networks, and natural language processing into the traditional information extraction process, breaking through the traditional model that relies on keyword and template matching. It achieves accurate understanding and extraction of complex semantic structures such as fuzzy language, polysemous expressions, and contextual dependencies, and promotes text information processing in the underground space field into the era of "semantic intelligence."
[0044] (3) The dimension of information structure is expanded from planar entities to spatiotemporal entities. This invention not only identifies underground space-related entities in text, but also further anchors the entities in the spatiotemporal dimension, constructing a structured data model of "semantics + time + space" in a three-in-one manner. This makes underground space information traceable, evolvable, and has a visualization foundation, truly realizing the transformation from "static recognition" to "dynamic mapping." BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings of the embodiments of the present invention. The drawings are only used to illustrate some embodiments of the present invention, but not to limit all embodiments of the present invention thereto.
[0046] Figure 1 1 is a flow chart of an underground space entity recognition method based on open source spatiotemporal data according to an embodiment of the present invention;
[0047] Figure 2 It is a structural block diagram of an underground space entity recognition system based on open source spatiotemporal data according to an embodiment of the present invention. DETAILED DESCRIPTION
[0048] The following will be combined with the accompanying drawings of specific embodiments of the present invention to clearly and completely describe the exemplary embodiments of the present invention. Unless otherwise defined, technical or scientific terms used in the present invention should be given the common meanings understood by people with ordinary skills in the relevant field.
[0049] like Figure 1As shown, this embodiment discloses a method for identifying underground spatial entities based on open-source spatiotemporal data. This method builds a fusion model framework that integrates natural language processing algorithms, spatiotemporal data analysis algorithms, and deep learning algorithms to automatically extract underground spatial entities and their spatiotemporal information from open-source data sources such as news and social networks. This method achieves intelligent identification of underground spatial entities and dynamic activity analysis through four steps.
[0050] Step S1: Collect multimodal spatiotemporal data containing underground space information from a multi-source open source platform through a web crawler and an API interface.
[0051] This step automatically collects underground space-related data from open-source internet platforms such as news social networks, blogs, and forums. Efficient web crawler technology and API interfaces, combined with deep crawling and dynamic data acquisition methods, ensure high-quality data acquisition from multiple channels, covering content in various formats such as text, images, and video. To improve the ability to handle large-scale datasets, a distributed computing framework is utilized for data storage and processing, effectively ensuring the timeliness, comprehensiveness, and accuracy of the data, providing rich spatiotemporal data support for subsequent analysis and entity recognition.
[0052]
[0053] Among them, Data(t) represents the underground space related data obtained in the time interval [t1, t2], and f(t) is the spatiotemporal characteristic function related to the underground space entity.
[0054] The specific steps for open source spatiotemporal data collection are as follows:
[0055] ①Keyword library initialization and update
[0056] A keyword library encompassing underground space terminology (e.g., "subway construction," "underground passage," "tunnel collapse," "underground pipeline network construction") is built, supporting both manual maintenance and automatic expansion. The system periodically uses NLP models to perform semantic analysis on new media content, dynamically adding high-frequency new terms.
[0057] ②Site source configuration and distribution scheduling
[0058] During the platform startup phase, multiple open source data sources are loaded through configuration files, the crawling cycle, depth level and data format requirements are set according to the site type, and the task scheduler is used to implement distributed asynchronous crawling task delivery.
[0059] ③Web page structure analysis and deep crawling
[0060] Separate collection strategies are designed for structured and unstructured pages: the structured platform directly extracts text fields, time, user information, media links, etc. through API or DOM positioning; the unstructured page content is captured through regular expressions and HTML parsers, supporting simulated clicks, page turning and comment capture.
[0061] ④Content type identification and pre-filtering
[0062] The media type (text, pictures, videos, etc.) of the original collected content is identified, and coarse filtering is performed based on content summary generation and keyword matching to remove redundant data irrelevant to the underground space and reduce subsequent processing pressure.
[0063] ⑤ Preliminary indexing of spatiotemporal information
[0064] By using place name dictionaries and time regularization rules, geographic entities and time expressions can be quickly extracted from the original text, providing candidate inputs for subsequent NER and spatiotemporal completion modules.
[0065] ⑥Data standardization and storage
[0066] The extracted data is encapsulated into a unified data structure, including fields: data source, crawling time, content summary, media type, preliminary location information, URL, etc.
[0067] Step S2: Perform standardized preprocessing on the original text, including denoising, word segmentation, part-of-speech tagging, and syntactic analysis, and combine the BiLSTM-CRF model with the BERT model to improve semantic recognition capabilities.
[0068] The core task of this step is to clean and process the acquired raw text data, including denoising, word segmentation, part-of-speech tagging, and syntactic analysis. To improve the accuracy of word segmentation and entity recognition, advanced deep learning models such as the BiLSTM-CRF model and the BERT model are used. BiLSTM-CRF combines a bidirectional long short-term memory network and conditional random fields to efficiently process lexical boundaries and contextual information in Chinese text. The BERT pre-trained model further improves the accuracy of named entity recognition (NER), enabling precise identification of multiple categories of entities related to underground space, such as underground buildings, facilities, and tunnels.
[0069]
[0070] Among them, P(y|x) represents the conditional probability of entity tag sequence y under text sequence x, φ(y i ,x i ) is from x i to y i The feature mapping function of , Y is the set of all possible label sequences.
[0071] The specific steps for standardization preprocessing of raw text data are as follows:
[0072] ① Word segmentation and part-of-speech tagging
[0073] The BERT encoder is introduced to embed the enhanced BiLSTM-CRF model to perform Chinese word segmentation and part-of-speech tagging on the text. The system process is as follows: the text is input into the BERT encoder to obtain the context vector representation of each word Will Input the BiLSTM model to obtain bidirectional context features. Use CRF to model the label sequence (such as B-MISC, I-LOC, O) and output the optimal word segmentation and part-of-speech sequence labels.
[0074] ②Syntactic dependency analysis
[0075] A Transformer- or Tree-LSTM-based syntactic analyzer is used to analyze the word segmentation results for dependency syntactic structure, extracting dependency edges such as verb centers, subject-verb-object, and modifiers, and constructing a word dependency graph. This word dependency graph is used to assist in the subsequent identification of underground space-related entities and semantic role labeling.
[0076] ③ Pre-screening of named entity candidate recognition
[0077] Perform the first round of entity recognition on the text to identify basic entities such as places, organizations, facilities, and time, and mark their start and end locations and types to form a preliminary entity candidate set for the next step.
[0078] The word segmentation and part-of-speech tagging in this step provide support for accurate positioning and semantic understanding of entity recognition in step S3:
[0079] Precisely locate entity boundaries: Word segmentation breaks text into independent units (tokens), enabling more precise determination of the boundaries of underground space entities in step S3. When identifying the entity "underground mall," accurate word segmentation allows for clear identification of "underground" and "mall" as a whole, avoiding misclassification as two separate, unrelated entities and improving entity recognition accuracy.
[0080] Assisting semantic understanding: Part-of-speech tagging provides part-of-speech information for each word segment, which helps the model in step S3 better understand the role and semantic relationship of vocabulary in a sentence. Knowing whether a word is a noun, verb, or adjective allows the model to more accurately determine its relationship with other words, thereby more precisely identifying entities. In the sentence "The tunnel connects two areas," "tunnel" is a noun. Through part-of-speech tagging, the model can more clearly identify its role as an entity in the sentence, rather than confusing it with the action of connection, thereby improving its semantic understanding of underground space entities.
[0081] Step S3: Accurately identify underground space-related entities by integrating the BERT+CRF model with domain rules.
[0082] This step accurately identifies named entities closely related to underground space (such as underground facilities, tunnels, underground pipe networks, underground structures, underground shopping malls, etc.) through the combination of deep learning and traditional rule engines. The BERT model is used to extract contextual information in the text, and combined with CRF for precise annotation, which significantly improves the accuracy, recall rate and robustness of entity recognition in the professional context of underground space. This method is particularly suitable for processing underground space entity information with complex contextual dependencies. In addition, combined with a domain-specific rule engine, it can be optimized according to the unique vocabulary related to underground space (such as tunnels, underground pipe networks, underground facilities, etc.), thereby further improving the accuracy of recognition and the robustness of the system.
[0083] First, the token sequence is input into the BERT model, and the context vector representation h of each word is output. i , capturing the deep semantic relationship between words. Then, the CRF module is used to i Sequence constructs a state transition diagram and decodes the optimal BIO or BIOES tag sequence {y1,y2,...y n}, thereby determining the entity boundaries and categories. The joint training objective of the model is to maximize the following log-likelihood function:
[0084]
[0085] Where s(X,Y) is the sum of BERT output and CRF state transition score.
[0086] In addition to the above, this step also includes:
[0087] ① Design of labeling system for underground space industry
[0088] In response to the specialized nature of the underground space sector, a domain-specific entity tag set was established, enabling model training and prediction based on annotation specifications. This tagging system combines corpus annotation semantic specifications with a database of industry literature terminology to enhance model generalization and recognition accuracy.
[0089] ② Rule engine post-processing optimization
[0090] To improve the integrity and consistency of domain entity recognition, a professionally constructed rule engine was introduced to post-process the model results. The constructed underground space entity terminology dictionary was used to align and fill in the model results; semantic merging was performed on scenes that continuously hit multiple related entities; context clues were used to determine the true referent of the pronoun; confidence weights were assigned to the model recognition results and the rule supplementation results respectively, and finally the optimal entity set was generated by fusion.
[0091] ③Entity fusion and unique coding
[0092] When semantically identical entities appear multiple times in the same text or the same data source, the system merges the entities through semantic vector similarity and position overlap detection to generate a unique Entity-ID, enabling consistent tracking across sentences and paragraphs.
[0093] Step S4: Based on the spatiotemporal clustering algorithm and graph neural network (GNN), the missing spatiotemporal attribute information is completed to construct the entity-spatiotemporal structure.
[0094] This step is based on the spatiotemporal data analysis algorithm to extract the relevant spatiotemporal information from the identified underground space entities, including timestamps, geographic coordinates, etc. This step uses the spatiotemporal association analysis algorithm, combined with spatiotemporal K-means clustering and graph-based spatiotemporal association analysis, to conduct a detailed analysis of the distribution of underground space entities in different time and space dimensions, thereby revealing the spatiotemporal evolution laws of underground space. In addition, for missing or incomplete spatiotemporal data, graph neural networks (GNN) and regression analysis (such as support vector regression SVR) are used to complete spatiotemporal data to ensure data integrity and high-precision inference. The spatiotemporal clustering process can be expressed as:
[0095]
[0096] Among them, J(θ) is the cost function of clustering, μ ci is the center point of the i-th point, λ is the regularization parameter, θ i is the spatiotemporal data feature, χ i is the data point.
[0097] This step aims to annotate the spatiotemporal attributes of identified underground entities, constructing a three-dimensional "entity-time-space" structure. By parsing temporal expressions in natural language and extracting geographic information, combined with clustering and machine learning models, missing or ambiguous spatiotemporal data can be supplemented and predicted, ensuring that all entities have complete and usable spatiotemporal labels, laying the foundation for dynamic monitoring and evolutionary modeling.
[0098] ① Joint time-place parsing and standardization of underground space semantic entities
[0099] After identifying underground space entities, the system further conducts in-depth mining and structural processing of expressions related to spatiotemporal attributes in the text. First, by combining regular expressions, domain templates, and contextual semantic matching mechanisms, it identifies temporal and spatial text fragments from the entity's sentence. Temporal information includes absolute and relative time, while spatial information primarily encompasses administrative divisions, place names, streets, and building names. The system uniformly formats the identified time phrases, using a standard time representation, and performs semantic reasoning and contextual inversion on relative fuzzy time. Simultaneously, the spatial information recognition module integrates a place name NER model with a place name dictionary to identify location fragments such as administrative regions and infrastructure names. It then uses a geocoding interface to parse these fragments into precise latitude and longitude coordinates and administrative division codes, forming structured geographic attribute annotations. This stage achieves preliminary spatiotemporal anchoring of underground space entities in natural language by constructing a ternary association of "text-time-place," providing complete raw input for subsequent clustering analysis and missing-completion.
[0100] ② Spatiotemporal K-means cluster analysis
[0101] Perform cluster analysis on the initially acquired entity temporal and spatial data to identify entity groups in the same spatiotemporal cluster for pattern induction and anomaly detection. Use K-means or DBSCAN algorithms for clustering, output cluster centers and boundaries, and determine whether entities belong to known spatial clusters or temporally dense segments. Construct a spatiotemporal vector representation of the entity:
[0102]
[0103] where time i Convert the standardized timestamp into a numerical value, lon i is longitude, lat i is the latitude.
[0104] ③ Intelligent completion of missing information (SVR+GNN)
[0105] For entities whose complete spatiotemporal attributes cannot be directly parsed, the following method is used for prediction and completion: first, a support vector regression (SVR) model is used to perform regression prediction based on the entity context, semantically similar entities, and the time distribution of the paragraph in which the sentence is located. Then, an entity-entity graph is constructed, with nodes as entities and edges as co-occurrence or context similarity. A graph neural network (GNN) is used to propagate and complete node attributes.
[0106] Here is an example of completing missing spatiotemporal attribute information:
[0107] Scenario: An "underground utility corridor" is identified in news reports, but the text only mentions its location in Chaoyang District and recent completion, lacking its specific latitude, longitude, and completion date. SVR (Support Vector Regression) and GNN (Graph Neural Network) are needed to fill in the missing spatiotemporal information.
[0108] Step 1: Input data
[0109] Identified entities and partial information
[0110] {
[0111] Chaoyang District's underground integrated pipeline corridor was recently completed.
[0112] "entity":"underground integrated pipeline corridor",
[0113] "type":"FACILITY",
[0114] "partial_info":{
[0115] "location":"Chaoyang District", / / only administrative district, no latitude and longitude
[0116] "time":"recent" / / fuzzy time, no specific date
[0117] }
[0118] }
[0119] Step ②: Spatiotemporal information completion process
[0120] (1) SVR (Support Vector Regression) prediction
[0121] Input features:
[0122] Text context (such as "Chaoyang District" and "completion"); historical data of similar entities (such as the latitude and longitude and completion time of other pipeline corridors); administrative division statistical information (such as the average coordinates of Chaoyang District's infrastructure).
[0123] Predicted latitude and longitude:
[0124] The model learns from historical data:
[0125] The pipeline corridors in Chaoyang District are mostly concentrated around (116.48, 39.95); "Completion" is often associated with dates in the last three months.
[0126] Output:
[0127] predicted_lonlat=(116.4832,39.9524)#SVR predicted latitude and longitude
[0128] (2) GNN (Graph Neural Network) Verification and Correction
[0129] Build an entity relationship diagram:
[0130] Node: other known “underground integrated pipeline corridor” entities (such as the Haidian District pipeline corridor and the Dongcheng District pipeline corridor);
[0131] Edge: An association between entities with similar spatial distance (<10km) or similar time (construction at the same time).
[0132] GNN completion:
[0133] Correct the SVR prediction results through node information propagation:
[0134] If the coordinates of adjacent pipeline corridors are concentrated in (116.47±0.01, 39.95±0.01), the coordinates of the Chaoyang District pipeline corridor will be adjusted to this range.
[0135] If most of the adjacent pipeline corridors are completed in 2023-10, the "short term" will be revised to 2023-10-15.
[0136] Final output:
[0137] {
[0138] "entity":"underground integrated pipeline corridor",
[0139] "location":{
[0140] "admin":"Chaoyang District",
[0141] "lonlat":[116.475,39.953] / / GNN corrected coordinates
[0142] },
[0143] "time":{
[0144] "text":"recent",
[0145] "timestamp":"2023-10-15" / / Specific date inferred by GNN
[0146] }
[0147] }
[0148] Corresponding to the above method, this embodiment discloses an underground space entity recognition system based on open source spatiotemporal data, such as Figure 2 Shown, including:
[0149] The data acquisition module is used to collect multimodal spatiotemporal data containing underground space information from multiple source open source platforms through web crawlers and API interfaces.
[0150] The text preprocessing module is used to perform standardized preprocessing on raw text, including denoising, word segmentation, part-of-speech tagging, and syntactic analysis. It combines the BiLSTM-CRF model with the BERT model to improve semantic recognition capabilities.
[0151] The entity recognition module is used to identify underground space-related entities by integrating the BERT+CRF model with domain rules.
[0152] The spatiotemporal information completion module is used to complete missing spatiotemporal attribute information based on spatiotemporal clustering algorithms and graph neural networks, and to construct entity-spatiotemporal structures.
[0153] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A method for identifying underground space entities based on open source spatiotemporal data, characterized in that: The following steps are involved: Step 1: Collect multimodal spatiotemporal data containing underground space information from multiple open source platforms through web crawlers and API interfaces; Step 2: Standardize and preprocess the original text, including denoising, word segmentation, part-of-speech tagging, and syntactic analysis, and combine the BiLSTM-CRF model with the BERT model to improve semantic recognition capabilities; Step 3: Identify underground space-related entities by integrating the BERT+CRF model with domain rules; Step 4: Based on the spatiotemporal clustering algorithm and graph neural network, the missing spatiotemporal attribute information is completed to construct the entity-spatiotemporal structure.
2. The underground space entity recognition method based on open source spatiotemporal data according to claim 1 is characterized in that: The multi-source open source platform in step 1 includes news social networks, blogs and forums, and the multimodal spatiotemporal data includes text, image and video data.
3. The underground space entity recognition method based on open source spatiotemporal data according to claim 1 is characterized in that: The standardized preprocessing in step 2 specifically includes: Feed the text into the BERT encoder to obtain the context vector representation of each word; Input the vector representation into the BiLSTM model to obtain bidirectional context features; The label sequence is modeled through CRF to output the optimal word segmentation and part-of-speech sequence labels.
4. The underground space entity recognition method based on open source spatiotemporal data according to claim 3 is characterized in that: The standardization preprocessing also includes: Perform syntactic dependency analysis on the segmented text to extract the core verbs, subject-verb-object relationships, and modifying structures in the sentence, and generate a word dependency graph. The syntactic dependency analysis is implemented using a Transformer- or Tree-LSTM-based syntactic analyzer. The word dependency graph is input into step 3 of underground space related entity recognition to assist in determining the semantic role and boundary of the named entity.
5. The underground space entity recognition method based on open source spatiotemporal data according to claim 1 is characterized in that: The underground space-related entities mentioned in step 3 include underground facilities, tunnels, underground pipe networks, and underground buildings. The identification of underground space-related entities is achieved through the following methods: Input the token sequence into the BERT model and output the context vector representation of each word; A state transition diagram is constructed based on the vector representation through the CRF module, the optimal label sequence is decoded, and the entity boundaries and categories are determined.
6. The underground space entity recognition method based on open source spatiotemporal data according to claim 1 is characterized in that: Step 3 also includes: Design a set of entity labels specific to the underground space industry to provide a domain-adapted semantic framework for the model, enabling model training and prediction. Post-process the model results through the rule engine and use the term dictionary to align and fill in the gaps in the results; Merge entities with the same semantics to generate a unique Entity-ID.
7. The underground space entity recognition method based on open source spatiotemporal data according to claim 1 is characterized in that: The spatiotemporal attribute information in step 4 includes a timestamp and geographic coordinates, and the spatiotemporal clustering algorithm adopts a K-means or DBSCAN algorithm.
8. The underground space entity recognition method based on open source spatiotemporal data according to claim 7 is characterized in that: The missing spatiotemporal attribute information in step 4 is completed by the following method: Extract the original time and space information associated with the identified underground space entities and preliminarily label them through regular matching and NER model; Unify the formatting of non-standardized time phrases and use standard time representation; call the geocoding interface to parse spatial information into latitude and longitude coordinates and administrative division codes; Use spatiotemporal clustering algorithms or DBSCAN algorithms to group similar spatiotemporal entities; For entities that are not clearly labeled, a support vector regression model is first used to perform regression prediction based on the entity context and semantically similar entities. Then, an entity-entity graph is constructed, with nodes as entities and edges as co-occurrence or context similarity. Graph neural networks are used to propagate and complete node attributes.
9. An underground space entity recognition system based on open source spatiotemporal data, characterized by: The method for realizing underground space entity recognition based on open source spatiotemporal data according to any one of claims 1 to 8 comprises: The data acquisition module is used to collect multimodal spatiotemporal data containing underground space information from multiple open source platforms through web crawlers and API interfaces; The text preprocessing module is used to perform standardized preprocessing on the original text, including denoising, word segmentation, part-of-speech tagging, and syntactic analysis. It combines the BiLSTM-CRF model with the BERT model to improve semantic recognition capabilities. The entity recognition module is used to identify underground space-related entities by integrating the BERT+CRF model with domain rules; The spatiotemporal information completion module is used to complete missing spatiotemporal attribute information based on spatiotemporal clustering algorithms and graph neural networks, and to construct entity-spatiotemporal structures.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.