A three-dimensional spatio-temporal geographic entity intelligent extraction method based on a large language model and related devices

CN122176246BActive Publication Date: 2026-08-21HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610652772.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-08-21
Estimated Expiration
2046-05-13

AI Technical Summary

Technical Problem

[0003]然而,现有的技术存在显著缺陷:一是接收机载LiDAR数据进行建筑物三维重建时,受屋顶形态多样、点云密度低且分布不均等因素影响,建筑物提取难度大;二是现有公开的符合CityGML 3.0标准的三维城市模型普遍缺少建筑年份、翻新状态等语义属性,并且,实现异构非结构化数据向CityGML与3DCityDB的映射的门槛高

Benefits of technology

本申请提供了一种基于大语言模型的三维时空地理实体智能提取方法及相关装置,通过提取目标地理实体对应的点云子集并生成点云稀疏张量,解决了原始点云数据量大、冗余信息多的问题。通过将点云稀疏张量输入包含点云条件编码器以及相互独立的第一解码分支和第二解码分支的自回归生成模型,生成三维顶点序列和网格面序列,并对其进行几何有效性与建筑逻辑合规性的联合校验,解决了传统三维重建方法中因点云质量问题导致的模型精度不高、结构不合理等难题。通过基于大语言模型的语义富集智能体从多源异构数据中提取语义属性,并进行基于语义属性几何约束信息的三维多边形网格模型结构校验以及基于三维多边形网格模型几何边界的语义属性空间匹配结果语义数据校验,解决了现有三维城市模型语义属性缺失以及异构数据映射门槛高的问题,实现了语义属性的自动提取与双向校验,保证了语义属性与几何模型的一致性和准确性,最终得到符合标准规范的三维城市模型文件,为城市数字孪生、智慧城市规划等领域提供了高质量的空间数据底座。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122176246B_ABST
    Figure CN122176246B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional space-time geographic entity intelligent extraction method based on a large language model and related devices, relates to the fields of three-dimensional geographic information systems and smart city digital twin technologies, and comprises the following steps: extracting a point cloud subset corresponding to a target geographic entity, and generating a point cloud sparse tensor; inputting the point cloud sparse tensor into an autoregressive generation model to generate a three-dimensional vertex and a mesh surface sequence; jointly checking the three-dimensional vertex and the mesh surface sequence to obtain a three-dimensional polygon mesh model of the target geographic entity; based on a semantic enrichment intelligent agent of the large language model, extracting semantic attributes associated with the three-dimensional polygon mesh model from multi-source heterogeneous data; based on geometric constraint information in the semantic attributes, checking the three-dimensional polygon mesh model; based on a geometric boundary of the three-dimensional polygon mesh model, performing semantic data checking on the semantic attributes; and storing the three-dimensional polygon mesh model and the semantic attributes in a spatial database after bidirectional checking to obtain a three-dimensional city model file conforming to standard specifications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of three-dimensional geographic information systems and smart city digital twin technology, and in particular to a method and related apparatus for intelligent extraction of three-dimensional spatiotemporal geographic entities based on a large language model. Background Technology

[0002] With the large-scale implementation of smart city and digital twin technologies, semantic 3D city models have become the core spatial data foundation for urban planning, traffic simulation, and other fields. 3D building models are the core spatial carriers for smart city and digital twin construction. The construction of 3D city models conforming to the CityGML specification mainly includes two core steps: geometric reconstruction and semantic enrichment.

[0003] However, existing technologies have significant drawbacks: First, when reconstructing buildings using airborne LiDAR data, building extraction is challenging due to factors such as diverse roof shapes, low and uneven point cloud density. Second, publicly available 3D city models conforming to the CityGML 3.0 standard generally lack semantic attributes such as building age and renovation status. Furthermore, the mapping of heterogeneous unstructured data to CityGML and 3DCityDB presents a high barrier to entry. Therefore, a highly automated, high-fidelity, and cross-modal self-correcting intelligent extraction solution for 3D spatiotemporal geographic entities is urgently needed. Summary of the Invention

[0004] The purpose of this application is to provide a method and related apparatus for intelligent extraction of three-dimensional spatiotemporal geographic entities based on a large language model, which can realize the generation of geometric structures from point clouds and achieve semantic enrichment, and can be applied to fields such as urban digital twins and smart city planning.

[0005] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a method for intelligent extraction of three-dimensional spatiotemporal geographic entities based on a large language model, including: Based on the original point cloud data of the region to be processed, extract the point cloud subsets corresponding to the target geographic entities and generate point cloud sparse tensors.

[0006] Obtain an autoregressive generative model; the autoregressive generative model includes a point cloud conditional encoder, and a first decoding branch and a second decoding branch that are independent of each other.

[0007] The point cloud sparse tensor is input into the autoregressive generation model to generate a three-dimensional vertex sequence and a mesh surface sequence.

[0008] The geometric validity and architectural logic compliance of the three-dimensional vertex sequence and mesh surface sequence are jointly verified to obtain a three-dimensional polygon mesh model of the target geographic entity.

[0009] A semantic enrichment agent based on a large language model extracts semantic attributes associated with the three-dimensional polygonal mesh model from multi-source heterogeneous data.

[0010] Based on the geometric constraint information in the semantic attributes, the structure of the three-dimensional polygonal mesh model is verified; and based on the geometric boundary of the three-dimensional polygonal mesh model, the spatial matching results of the semantic attributes are verified for semantic data.

[0011] The bidirectionally validated 3D polygonal mesh model and its semantic attributes are stored in a spatial database to obtain a 3D city model file that conforms to the standard specifications.

[0012] In a second aspect, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the intelligent extraction method for three-dimensional spatiotemporal geographic entities based on a large language model as described above.

[0013] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the intelligent extraction method for three-dimensional spatiotemporal geographic entities based on a large language model as described above.

[0014] Fourthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the intelligent extraction method for three-dimensional spatiotemporal geographic entities based on a large language model as described above.

[0015] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a method and related apparatus for intelligent extraction of 3D spatiotemporal geographic entities based on a large language model. By extracting a subset of point clouds corresponding to the target geographic entity and generating a sparse point cloud tensor, it solves the problems of large original point cloud data volume and excessive redundant information. By inputting the sparse point cloud tensor into an autoregressive generative model containing a point cloud conditional encoder and independent first and second decoding branches, it generates a 3D vertex sequence and a mesh surface sequence. Joint verification of geometric validity and architectural logic compliance is then performed, solving the problems of low model accuracy and unreasonable structure caused by point cloud quality issues in traditional 3D reconstruction methods. Furthermore, a semantic enrichment agent based on a large language model extracts semantic attributes from multi-source heterogeneous data. Structural verification of the 3D polygonal mesh model based on semantic attribute geometric constraints and semantic data verification of the semantic attribute spatial matching results based on the geometric boundaries of the 3D polygonal mesh model are performed. This solves the problems of missing semantic attributes in existing 3D city models and high barriers to heterogeneous data mapping. It achieves automatic extraction and bidirectional verification of semantic attributes, ensuring the consistency and accuracy of semantic attributes and geometric models. Finally, it yields a 3D city model file that conforms to standards and specifications, providing a high-quality spatial data foundation for fields such as urban digital twins and smart city planning. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is an application environment diagram of a three-dimensional spatiotemporal geographic entity intelligent extraction method based on a large language model, according to one embodiment of this application.

[0018] Figure 2 This is a flowchart illustrating a method for intelligent extraction of three-dimensional spatiotemporal geographic entities based on a large language model, provided in one embodiment of this application.

[0019] Figure 3 This is a diagram of the three-dimensional geographic entity autoregressive generation architecture provided in an embodiment of this application.

[0020] Figure 4 This is a diagram illustrating the multi-agent construction framework provided in one embodiment of this application.

[0021] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0023] Traditional geometric reconstruction methods suffer from significant error accumulation and are highly dependent on initial plane segmentation, making them prone to reconstruction failure in scenarios with complex roofs, vegetation occlusion, and uneven point cloud density. Extraction of unstructured data is inefficient and highly reliant on manual intervention. Architectural semantic information is largely scattered across PDF documents and OpenStreetMap interfaces, while existing technologies lack the ability to automatically parse multimodal heterogeneous data. Furthermore, mapping attributes to CityGML and 3DCityDB requires a high level of expertise, leading to low operational efficiency. The lack of a bidirectional consistency verification mechanism between geometric and semantic information means that current processes cannot utilize semantic information (such as known building floors) to constrain and correct geometric reconstruction errors, nor can they use geometric information to verify semantic data deviations, resulting in inconsistencies between geometric and semantic information in the final database. Additionally, interaction with 3D spatiotemporal databases presents technical barriers. CityGML and 3DCityDB have complex architectures, and data queries require writing specialized spatial SQL statements, limiting the application of 3D city models to non-professional users.

[0024] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] The intelligent extraction method for 3D spatiotemporal geographic entities based on a large language model provided in this application can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on another server. Terminal 102 can send the raw point cloud data of the area to be processed to server 104. After receiving the raw point cloud data of the area to be processed, server 104 extracts the point cloud subset corresponding to the target geographic entity, generates a point cloud sparse tensor, obtains an autoregressive generation model, which includes a point cloud conditional encoder and independent first and second decoding branches; inputs the point cloud sparse tensor into the autoregressive generation model to generate a 3D vertex sequence and a mesh surface sequence; and performs geometric validity and architectural logic checks on the 3D vertex sequence and mesh surface sequence. Joint compliance verification yields a 3D polygonal mesh model of the target geographic entity. A semantic enrichment agent based on a large language model extracts semantic attributes associated with the 3D polygonal mesh model from multi-source heterogeneous data. Based on the geometric constraint information in the semantic attributes, structural verification is performed on the 3D polygonal mesh model. Furthermore, based on the geometric boundaries of the 3D polygonal mesh model, semantic data verification is performed on the spatial matching results of the semantic attributes. The bidirectionally verified 3D polygonal mesh model and its semantic attributes are stored in a spatial database, resulting in a 3D city model file conforming to standards, which is then fed back to terminal 102. In some embodiments, the intelligent extraction method for 3D spatiotemporal geographic entities based on a large language model can also be implemented independently by server 104 or terminal 102. For example, terminal 102 can directly perform intelligent extraction of 3D spatiotemporal geographic entities based on a large language model on the original point cloud data of the area to be processed, or server 104 can obtain the original point cloud data of the area to be processed from the data storage system and perform intelligent extraction of 3D spatiotemporal geographic entities based on a large language model on the original point cloud data of the area to be processed.

[0026] The terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster composed of multiple servers, or it can be a cloud server.

[0027] In one exemplary embodiment, such as Figure 2As shown, a method for intelligent extraction of 3D spatiotemporal geographic entities based on a large language model is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps 201 to 207. Wherein: Step 201: Based on the original point cloud data of the area to be processed, extract the point cloud subsets corresponding to the target geographic entities and generate point cloud sparse tensors.

[0028] Step 202: Obtain the autoregressive generative model; the autoregressive generative model includes a point cloud conditional encoder, and a first decoding branch and a second decoding branch that are independent of each other.

[0029] Step 203: Input the point cloud sparse tensor into the autoregressive generation model to generate a 3D vertex sequence and a mesh surface sequence.

[0030] Step 204: Perform a joint verification of the geometric validity and architectural logic compliance of the three-dimensional vertex sequence and mesh surface sequence to obtain a three-dimensional polygon mesh model of the target geographic entity.

[0031] Step 205: The semantic enrichment agent based on the large language model extracts semantic attributes associated with the three-dimensional polygonal mesh model from multi-source heterogeneous data; the multi-source heterogeneous data includes at least one of the following: static attribute data in unstructured documents, dynamic attribute data provided by spontaneous geographic information platforms, and spatial positioning data in structured address databases.

[0032] Step 206: Based on the geometric constraint information in the semantic attributes, perform structural verification on the three-dimensional polygonal mesh model; and based on the geometric boundary of the three-dimensional polygonal mesh model, perform semantic data verification on the spatial matching results of the semantic attributes.

[0033] Step 207: Store the bidirectionally validated 3D polygonal mesh model and semantic attributes in the spatial database to obtain a 3D city model file that conforms to the standard specifications.

[0034] In one exemplary embodiment, when performing steps 201-207, the specific steps may be as follows: Step 1: Process the raw point cloud data of the area to be processed. In this embodiment, the raw point cloud of the airborne LiDAR is received, and non-building points such as vegetation and ground are removed according to the classification labels. Point clouds marked as buildings are selected and normalized and multi-scale voxelized according to the ASPRS classification criteria.

[0035] Specifically, coordinate normalization in the normalization process is used to normalize the coordinates of a single building point cloud, normalizing the point cloud coordinates to the unit sphere. The normalization formula is: .

[0036] In the formula, The original point cloud coordinates; The coordinates of the center point of the point cloud; This represents the maximum distance from the point cloud to the center point.

[0037] Multi-scale voxelization is converted into sparse voxelization by constructing a multi-scale voxel mesh with a resolution of 0.05m to complete sparse tensor formatting, which is then input into the subsequent point cloud conditional encoding branch.

[0038] Step 2: Construction of a 3D geographic entity autoregressive generation model. For example... Figure 3 As shown, a shared point cloud conditional encoder (point cloud conditional encoding layer) is constructed to extract point cloud spatial context features. The decoding layer consists of two independent decoding branches: a vertex generation decoding branch (first decoding branch) and a face generation decoding branch (second decoding branch), which generate 3D vertex sequences and face mesh sequences. The first decoding branch is composed of a first Transformer decoder layer; the second decoding branch is composed of a vertex encoder and a second Transformer decoder layer.

[0039] Specifically, model building includes the following steps: Step 2.1 Construction of Point Cloud Conditional Encoding Layer: Sparse 3D convolutional feature extraction: Sparse 3DResNet encoding is implemented using the Minkowski Engine, with the core being the BasicBlock residual module. The forward propagation formula is as follows: .

[0040] .

[0041] .

[0042] .

[0043] In the formula, The input is a sparse tensor of a point cloud, where N is the number of points in the point cloud. The dimension of point cloud coordinate features; , It is a sparse convolution kernel. The kernel size is the convolution kernel size. The convolution stride; This is a Minkowski sparse convolution operation; LeakyReLU activation function is specifically designed for sparse tensors.

[0044] 3D sparse feature tensor The global flattening is a two-dimensional sequence, which is then projected through a linear layer to a fixed embedding dimension of D=256, resulting in the context feature sequence. (S is the length of the feature sequence), which serves as the key (K) and value (V) input for the cross attention of the decoding branch.

[0045] When constructing a sequence autoregressive decoder, such as Figure 3 As shown, an independent design is adopted for the vertex decoder and face decoder. While their structures share the same origin, their parameters are not shared, allowing them to adapt to the different task characteristics of vertex coordinate generation and face index generation, respectively. For the vertex decoder, the input is a serialized vertex coordinate token, and the output is the coordinate classification logits. For the face decoder, the input is a serialized vertex index token, and the output is the query vector for the pointer network. This branch requires an additional vertex feature encoding module to encode vertices into feature vectors for the pointer network to retrieve. Both decoders contain the following common components: Transformer decoder layer: The single-layer structure consists of layer normalization, multi-head self-attention with causal masking, cross-attention, feedforward network, and Dropout. This embodiment sets up 8 Transformer decoder layers, and its forward propagation formula is as follows: .

[0046] .

[0047] .

[0048] .

[0049] .

[0050] .

[0051] .

[0052] .

[0053] .

[0054] In the formula, Input features to the decoder (input point cloud sparse tensor). For normalization layer; Initialize the coefficients for the ReZero residuals; Dropout is for random discarding. is the discard rate; causal is the lower triangular causal mask, which forces the model to only access the generated sequence from 1 to n-1 when predicting the nth token; Spatial context features; FFN is a feedforward network with dimension D→ →D, where D=256 is the embedding dimension. is the dimension of the hidden layer of the feedforward network.

[0055] The formula for calculating multi-head self-attention is as follows: .

[0056] In the formula, These are the query, key, and value vectors, respectively; H=4 represents the number of attention heads. This is a multi-head output projection matrix.

[0057] The formula for calculating single-head self-attention is as follows: .

[0058] In the formula, These are the query, key, and value vectors for the h-th head, respectively.

[0059] When constructing the output header module, it specifically includes vertex output headers and face output headers: Vertex output head: A linear classification layer that maps the decoder's D=256-dimensional hidden features to... Class (corresponding to 8-bit quantized coordinates), output logits vector The coordinate classification probability distribution is converted using Softmax: .

[0060] in Softmax is the temperature coefficient used to control the randomness of sampling.

[0061] Surface output header: Pointer network layer, which displays the features queried by the decoder. With encoded vertex features Perform the dot product operation and output the vertex index probability distribution: .

[0062] in, Let m be the index of the m-th vertex in the face sequence. The query pointer feature vector generated by the face decoder at step m is... This is the encoded context feature vector of the j-th vertex.

[0063] Step 3: Generate the 3D vertex sequence and mesh surface sequence of the building based on the autoregressive model. The steps are as follows: The autoregressive model constructed in step 2 is called to perform post-processing such as Top-p sampling, dequantization, and decentralization to generate a compliant 3D vertex sequence of buildings; then, based on the 3D vertex sequence, the autoregressive generation of the mesh surface sequence is completed through a pointer network.

[0064] Specifically, point cloud feature encoding is performed first: the output of step 1 is input into the point cloud conditional encoder constructed in step 2 to extract spatial context features. .

[0065] Coordinate quantization and sequence modeling: Quantizing the coordinates of the original point cloud Perform 8-bit uniform quantization and map to For integer intervals, the quantization formula is: .

[0066] in, For the number of quantization bits, This is a rounding function. After quantization, it is expanded into a one-dimensional token sequence in the order of z, y, x, and a terminator 's' is added at the end.

[0067] Then, a triple embedding fusion is performed: coordinate embedding (distinguishing x / y / z axes), position embedding (identifying sequence position), and numerical embedding (expressing quantization value) are superimposed on the sequence, and the sum is input into the vertex decoder.

[0068] Autoregressive sampling: Employs a Top-p kernel sampling strategy to generate point sequences for each token, with a maximum length of... (Corresponding to 50 3D vertices), the core logic formula is: .

[0069] .

[0070] .

[0071] .

[0072] .

[0073] .

[0074] In the formula, The generated vertex sequence; Spatial context features; Temperature coefficient; It is a descending sorting function; This is a cumulative summation function; The threshold for Top-p sampling probability; Mask to restore the original order; The token sampled in step n.

[0075] Dequantization and centering: After generation, dequantization is performed to restore the coordinate range to [-1,1], convert it to x, y, z axis order, and centering alignment is performed.

[0076] Then, a sequence of building mesh surfaces is generated: the generated vertex sequence is input into a point cloud conditional encoder to obtain a vertex feature set. Feature set The spatial location and topological association of vertices are encoded, and finally, a face sequence is generated token by token through a pointer network, with a maximum length. .

[0077] Step 4: Joint Verification of Vertex and Face Sequences. The vertex and face sequences generated in Step 3 are jointly verified for geometric validity and architectural logic compliance. An iterative rejection sampling strategy is used to eliminate unreasonable predictions, outputting a 3D polygonal mesh model of the building that conforms to engineering specifications. Specifically, this includes the following checks: Vertex sequence verification: Checks if the generated vertex sequence contains a valid stop token; if not, regenerates the vertex sequence. The verification includes the placement rules of the termination marker s, z-coordinate verification, y-coordinate verification, and x-coordinate verification. The placement rule of the sequence termination marker s means that the sequence must place the stop marker only after the x-coordinate; the z-coordinate verification means that the sequence of z-coordinates must maintain non-decreasing order; the y-coordinate verification means that if the z-coordinate corresponding to a subsequent y-coordinate is equal, then the coordinate should not decrease; the x-coordinate verification means that when the y-coordinate corresponding to a subsequent x-coordinate is equal to the preceding coordinate, the coordinate must strictly increase.

[0078] Initial face sequence verification: Check if the generated face sequence contains a valid stop token. If not, regenerate the face sequence. The face sequence verification rules include face index increment rules, vertex index rules, and vertex uniqueness rules. The face index increment rule means that the first index of the newly generated face must be greater than or equal to the first index of the previous face. The vertex index order rule means that the indices of adjacent vertices within the same face must increment sequentially. The vertex uniqueness rule means that all vertex indices within the same face must be distinct.

[0079] Building geometry verification: Perform three core verifications sequentially: ① Does the ground surface set completely cover the point cloud? ① If the projection of the plane is incomplete, regenerate the vertex sequence; ② Check if all edges of the ground polygon are correctly connected to the perpendicular edges of the wall. If the connection is incorrect, regenerate the face sequence; ③ Check if there are diagonals in the wall structure. If so, regenerate the face sequence; Perform ground projection coverage verification, structural connectivity verification, and diagonal detection.

[0080] The ground projection coverage verification refers to the ground set Projecting the surface in the middle to The coverage area is obtained, and the point cloud is projected onto... , obtain point set ,statistics The number of points falling within the coverage area Calculate coverage: ,like If the condition is met, the test is considered a failure; otherwise, it is considered a success.

[0081] The structural connectivity verification refers to the data from the ground set. and wall collection Extract all edges to form a ground edge set. and wall edge collection For each ground edge, in Find wall edges that share a common vertex and determine if they are perpendicular and whether their horizontal projections are orthogonal to the horizontal projections of the ground edges. Calculate the percentage of successfully matched ground edges; if the percentage is ≤90%, the match is considered a failure; otherwise, it is considered a success.

[0082] The diagonal detection refers to calculating the horizontal modulus for each side of the wall surface. absolute value of the vertical component Calculate their ratio: If there exists any edge that satisfies If it is neither purely horizontal nor purely vertical, then it is determined that a diagonal exists and the operation fails; otherwise, it passes.

[0083] Iteration termination rules: When the face sequence and vertex sequence pass all checks, the final 3D building polygon mesh is output; if the face sequence iteration count reaches the preset limit and still fails to pass the check, the vertex sequence is regenerated; if the vertex sequence iteration count reaches the preset limit and still fails to pass the check, the process is terminated and an error message is output.

[0084] Step 5: Geographic Entity Semantic Extraction Based on a Large Language Model. For the reconstructed 3D mesh, a semantic enrichment agent driven by a large language model is used. The Retrieval Enhancement Generation (RAG) module reads building-related PDF documents to extract static attribute data such as building year, or calls the external OpenStreetMap API to extract dynamic attributes such as current building usage. This completes tasks such as geographic entity spatial matching, semantic attribute extraction, standard mapping, and automated database writing.

[0085] Step 5 describes the semantic extraction of geographic entities based on a large language model, such as... Figure 4As shown, by constructing a semantic enrichment agent, the automatic extraction, standard mapping, and database insertion of semantic attributes are completed. The semantic enrichment agent is a production-grade agent for administrators, possessing the minimum necessary read and write permissions for the 3DCityDB database. Specifically, it includes the following sub-steps: Construction of the semantic enrichment agent: This agent uses an externally invoked general-purpose large language model (such as GPT-4, Tongyi Qianwen, etc.) as its model foundation. The specific construction content is as follows: Core Customized Prompt System Construction: This application constructs a semantic enrichment agent based on a large language model. Through prompt words, it pre-embeds the CityGML 3.0 standard data model, 3DCityDB v5 database table association rules, and standard SQL query examples into the LLM, enabling the LLM to learn the semantic storage specifications of the 3D city model. Specifically, this includes: ① embedding the schema definition, attribute naming conventions, and namespace encoding rules of the CityGML 3.0 standard core module; ② embedding the table structure, field relationships, and spatial SQL syntax specifications of the 3DCityDB v5 spatial database, along with accompanying small-sample examples; ③ a cadastral classification code table for the target area, achieving standardized mapping of building uses.

[0086] Built-in toolchain module development: Four callable tool modules are developed for intelligent agents to achieve end-to-end processing of multi-source heterogeneous data. These modules include: ① RAG retrieval module: Supports accurate attribute extraction from unstructured documents such as PDF and Word documents; ② API function call module: Built-in API call rules for platforms such as OSM Nominatim / Overpass; ③ SQL generation and execution module: Built-in run_sql execution function, supports automatic generation and compliant execution of SQL statements conforming to the 3DCityDB specification. This function automatically generates corresponding SQL statements to implement attribute addition, deletion, and modification operations; ④ User interaction confirmation module: Supports user confirmation and correction of attribute candidate sets and mapping rules, achieving accurate enrichment through human-machine collaboration.

[0087] Hierarchical permission control configuration: Configure database read and write permissions for the intelligent agent, grant only the minimum necessary permissions that the administrator can configure, and completely isolate permissions from the read-only query agent in subsequent step 7 to ensure the data integrity and security of the spatial database.

[0088] The execution flow of the semantic enrichment agent is as follows: 1) Geographic entity spatial matching: Based on the 3D geographic entity geometric mesh generated in step 3, extract the WGS84 coordinates of its bounding box and center point. Use both address matching and spatial topology matching to complete the accurate binding of building entities to multi-source semantic data sources and obtain the multi-source semantic data corresponding to the geographic entity as the basis for subsequent semantic enrichment. The address matching involves the agent calling spatial SQL to query the address of the target building from the ADDRESS and PROPERTY tables in 3DCityDB: `SELECT a.street, a.house_number, a.city FROM feature f JOIN property p ON f.id = p.feature_id JOIN address a ON p.val_address_id = a.id WHERE f.id ={feature_id}`, obtaining the structured address `Addr=(street,house_number,city)` from the structured address database. The spatial topology matching involves the agent calling spatial SQL to extract the geometric boundary of the target building from the geometry_data table in 3DCityDB and converting it to WGS84 coordinates: `SELECT ST_AsText(ST_Transform(ST_Envelope(geometry), 4326)) AS bbox_wgs84 FROM geometry_data WHERE id = {geom_id}`, obtaining the bounding box. .

[0089] 2) Semantic Attribute Extraction: Utilizing the RAG (Relative Aggregator) and tool-calling capabilities of the semantic enrichment agent, semantic attributes are extracted from the matched multi-source heterogeneous data. Specifically, this includes the following three types of data: ① Using the RAG module, key attributes such as construction year and number of floors are extracted from municipal documents in PDF format, generating an attribute candidate set; ② Using the built-in API function call module of the semantic enrichment agent, a structured query request is sent to the OSM Overpass API to obtain the complete tag dataset for the building. The query request uses the Overpass QL language, with the format: plaintext[out:json][timeout:25];way({osm_way_id});>;out tags, returning attributes such as building category and building function in standard JSON format, generating the corresponding attribute candidate set; ③ The extracted attribute candidate set is output to the user, who confirms the final attribute items to be added to the database, and irrelevant attribute data is removed.

[0090] 3) CityGML 3.0 Standard Semantic Mapping: Based on the CityGML 3.0 international standard and the 3DCityDB v5 database schema, standardized semantic mapping is completed for user-confirmed attributes.

[0091] 4) Automated 3DCityDB Database Write: Automatically generates SQL statements conforming to the 3DCityDB v5 specification, queries the unique feature_id of the target geographic entity, generates an INSERT statement to write to the PROPERTY table, and completes the automated storage of semantic attributes, achieving a unique binding between geometric grids and semantic attributes. The semantic enrichment agent, using the set of building attributes confirmed and retained by the user in the previous steps, calls the SQL generation and execution module in the built-in tools to complete the automated write. The automated write operation is based on the WGS84 bounding box coordinates of the geographic entity obtained from the previous spatial matching, generates a query statement according to the table structure and relationships of the 3DCityDB v5 database, and obtains the unique identifier feature_id of the geographic entity after execution. Subsequently, the SQL generation and execution module generates an INSERT statement according to the field meanings and storage formats of the PROPERTY table in the 3DCityDB v5 database, completing the automated semantic write.

[0092] Step 6: Geometric-Semantic Bidirectional Consistency Verification. Based on the extracted semantic attributes, optimize the geometric model, verify and correct deviations such as model height and topology, and use geometric boundaries to perform a secondary verification of the semantic matching results, eliminating spatially misaligned semantic data. The specific implementation process is as follows: When optimizing the geometric model with semantic constraints, the geometric model generated in step 4 is optimized based on the semantic attributes extracted in step 5. Specifically, based on the extracted floor number attribute, the matching degree between the height of the geometric model and the floor number is checked. If the deviation exceeds a preset threshold, the vertex sequence and face sequence are regenerated. Based on the building function attribute, the geometric constraint rules corresponding to the building type are loaded to correct the topological errors of the geometric model. The theoretical building height is calculated based on the floor number attribute and single-floor height, and compared with the actual height of the geometric model. When the height deviation exceeds a preset threshold, the vertex sequence and face sequence are regenerated. Based on the building function attribute, the pre-set geometric structural constraint conditions corresponding to the building type are called to check whether the connection status and face composition of the walls, ground, and other components of the geometric model are reasonable. For problems such as missing surfaces and abnormal structural forms that do not meet the constraint conditions, the vertex and face sequences of the model are directly corrected. The specific thresholds are shown in Table 1. Table 1 Threshold Illustration

[0093] Semantic matching optimization of geometric boundaries: Based on the high-precision geometric boundaries generated in step 4, the spatial matching results of semantic attributes are verified a second time, semantic data with mismatched spatial ranges are removed, and mismatched attribute items are corrected.

[0094] Consistency Verification: A joint verification is performed on the geometric model and semantic attributes after they are imported into the database. This joint verification includes: uniqueness verification, standardization verification, and spatial consistency verification. The uniqueness verification involves the agent driving the LLM to automatically generate a SELECT query statement conforming to the 3DCityDB v5 specification. This query retrieves the FEATURE table, geometry_data table, and PROPERTY attribute table, comparing the feature_id of the target geographic entity with the corresponding values ​​in these tables. The standardization verification involves the agent generating a SELECT query statement based on the CityGML 3.0 standard namespace encoding rules learned by the LLM. This query retrieves the corresponding namespace_id from the PROPERTY attribute table and compares it with the CityGML 3.0 standard namespace encoding to ensure that the imported semantic attributes conform to the CityGML 3.0 international standard. The spatial consistency verification involves the agent driving the LLM to generate a spatial query statement conforming to the PostGIS specification. This query extracts the WGS84 coordinate bounding box of the corresponding geometric model from the geometry_data table and compares it with the source coordinate range obtained during the spatial matching stage, verifying whether the deviation value is within a preset threshold range.

[0095] Step 7: Standardized Storage and Interactive Natural Language Query. This involves importing 3D spatiotemporal geographic entity data into a spatial database, exporting it in the CityGML standard format, and configuring a read-only query agent to enable convenient natural language querying.

[0096] The standardized storage and natural language interactive query mentioned above aim to build a query agent based on a large language model to achieve standardized storage and natural language query of enriched 3D spatiotemporal geographic entity data. The specific steps are as follows: To achieve standardized storage and natural language querying of 3D spatiotemporal geographic entity data, a query agent needs to be constructed. The underlying inference framework of the query agent can share a large language model framework with the semantic enrichment agent. Its core function of read-only querying is clearly defined through system prompts, while hard constraints are set to prohibit it from performing any data modification, writing, or deletion operations. The specific construction process includes three parts: Customized Prompt: The Prompt system comprises three core modules: ① 3DCityDB v5 core table structure and SQL syntax knowledge, along with a few sample examples of natural language to SQL conversion; ② Hard constraints on read-only permissions, explicitly allowing only the generation of SELECT type SQL statements; ③ Query result formatting rules, requiring the output to be user-friendly natural language or tables.

[0097] Lightweight toolchain integration: Only two read-only tools are integrated, with the core effects being: ① Read-only SQL generation and execution module with built-in multi-layer security verification mechanism; ② Query result formatting module: converts the structured data returned by the database into natural language descriptions or structured tables.

[0098] Establish hard isolation for read-only permissions: Achieve triple hard isolation of read-only permissions through database account permission configuration, code-level statement validation, and system-level operation restrictions.

[0099] The following is the execution flow of the query agent: Standardized data storage and export: The 3D geometric mesh model and its bound semantic attributes are uniformly stored in the 3DCityDB v5 spatial database. It supports direct export of model files conforming to the CityGML 3.0 standard for downstream applications such as urban planning and traffic simulation. After the semantic enrichment agent stores the bound semantic attributes in the 3DCityDB v5 spatial database, the query agent uses the built-in read-only query module to call the run_sql_query function to read the data in 3DCityDB v5. Then, relying on the built-in CityGML 3.0 standard export interface of 3DCityDB v5, model files conforming to international standards can be exported to support downstream applications such as urban planning and traffic simulation.

[0100] Interactive Natural Language Query: After the user inputs a natural language query command, the query agent automatically generates the corresponding spatial SQL query statement based on 3DCityDBschema knowledge. The SQL is executed using the `run_sql_query(query)` function, and after verification, the query results are returned to the user in a human-readable format. Alternatively, after the user inputs a natural language query command, the query agent calls the LLM to generate the corresponding spatial SQL SELECT query statement, executes the SQL using the `run_sql_query` function, and then returns the query results to the user in a human-readable format.

[0101] Furthermore, this application compared two mainstream point prediction methods, showing significant improvements in precision, recall, and F1 score, indicating enhanced accuracy in vertex identification. The reduction in chamfer distance (m) further confirms higher accuracy in reconstructing vertex positions, meaning that the reconstructed building surface using this method is geometrically closer to the true shape of the ground. Specific results are shown in Table 2. Table 2 Algorithm Model

[0102] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 5 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores 3D city model files. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a method for intelligent extraction of 3D spatiotemporal geographic entities based on a large language model.

[0103] Those skilled in the art will understand that Figure 5 The structures shown are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0104] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0105] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0106] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of the relevant data are carried out in compliance with the relevant data protection laws and policies of the country where the location is located, and with the authorization granted by the owner of the corresponding device.

[0107] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0108] Compared with the prior art, this application has the following advantages: This application constructs a multi-agent system, which corely comprises two types of units: semantic enrichment agents and query agents. It employs a large language model combined with a hard-access permission mechanism to achieve automated processing of unstructured data and natural language retrieval. This architecture addresses the problems of low efficiency in manual loading of unstructured data and high barriers to mapping to the 3DCityDB spatial database in traditional methods.

[0109] This application proposes a conditional autoregressive generative model based on point cloud space, employing a dual-branch architecture to sequentially generate building vertices and mesh surface sequences. This method reduces the excessive reliance on initial plane fitting in traditional methods and addresses the failure problem of 3D reconstruction in scenarios with complex roof structures and vegetation occlusion.

[0110] This application establishes two related processes: a geometric optimization process based on semantic prior constraints, and a semantic verification process based on geometric boundaries. These processes can utilize extracted attribute information such as the number of floors to reverse-correct height deviations and topological errors in geometric modeling, thus solving the data deviation problem caused by the disconnect between 3D reconstruction and attribute mounting in existing technologies.

[0111] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0112] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0113] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for intelligent extraction of three-dimensional spatiotemporal geographic entities based on a large language model, characterized in that, include: Based on the original point cloud data of the area to be processed, extract the point cloud subsets corresponding to the target geographic entities and generate point cloud sparse tensors. Obtain an autoregressive generative model; The autoregressive generative model includes a point cloud conditional encoder, and a first decoding branch and a second decoding branch that are independent of each other; The point cloud sparse tensor is input into the autoregressive generation model to generate a 3D vertex sequence and a mesh surface sequence. The geometric validity and architectural logic compliance of the three-dimensional vertex sequence and mesh surface sequence are jointly verified to obtain a three-dimensional polygon mesh model of the target geographic entity. A semantic enrichment agent based on a large language model extracts semantic attributes associated with the three-dimensional polygonal mesh model from multi-source heterogeneous data. Based on the geometric constraint information in the semantic attributes, the structure of the three-dimensional polygonal mesh model is verified. Based on the geometric boundaries of the three-dimensional polygonal mesh model, semantic data verification is performed on the spatial matching results of semantic attributes. Based on the extracted floor number attribute, the matching degree between the height of the geometric model and the floor number is checked. If the deviation between the two exceeds the preset threshold, the vertex sequence and face sequence are regenerated. Based on the building's functional attributes, load the corresponding geometric constraint rules for the building type to correct topological errors in the geometric model; The theoretical building height is calculated based on the number of floors and the height of a single floor. This theoretical height is then compared with the actual height of the geometric model. If the height deviation exceeds a preset threshold, the vertex sequence and face sequence are regenerated. Based on the building's functional attributes, the pre-set geometric constraints for the corresponding building type are invoked to check whether the connection status and face composition of the walls and ground components of the geometric model are reasonable. The bidirectionally validated 3D polygonal mesh model and its semantic attributes are stored in a spatial database to obtain a 3D city model file that conforms to the standard specifications.

2. The intelligent extraction method for three-dimensional spatiotemporal geographic entities based on a large language model according to claim 1, characterized in that, In the autoregressive generation model, the point cloud conditional encoder is used to extract the spatial context features of the point cloud sparse tensor, the first decoding branch is used to generate a three-dimensional vertex sequence of the target geographic entity based on the spatial context features, and the second decoding branch is used to generate a mesh surface sequence based on the spatial context features and the three-dimensional vertex sequence.

3. The intelligent extraction method for three-dimensional spatiotemporal geographic entities based on a large language model according to claim 1, characterized in that, The point cloud sparse tensor is input into the autoregressive generation model to generate a 3D vertex sequence and a mesh surface sequence, specifically including: Input the sparse tensor of the point cloud into the point cloud conditional encoder to extract the spatial context features of the sparse tensor of the point cloud. According to the formula The coordinates of the three-dimensional vertices in the spatial context features are uniformly quantized and mapped to... Integer range; in, The number of quantization bits; This is a rounding function. The original point cloud coordinates; Based on the quantized three-dimensional vertex coordinates, a one-dimensional token sequence is determined. The one-dimensional token sequence is triple-embedded and fused, and the fused one-dimensional token sequence is input into the first decoding branch; The first decoding branch uses a Top-p kernel sampling strategy to generate a three-dimensional vertex sequence; The generated 3D vertex sequence is input into the second decoding branch to obtain the mesh surface sequence.

4. The intelligent extraction method for three-dimensional spatiotemporal geographic entities based on a large language model according to claim 3, characterized in that, The first decoding branch consists of a first Transformer decoder layer; the second decoding branch consists of a vertex encoder and a second Transformer decoder layer.

5. The intelligent extraction method for three-dimensional spatiotemporal geographic entities based on a large language model according to claim 1, characterized in that, The point cloud conditional encoder uses the Minkowski Engine to implement sparse 3DResNet encoding; the forward propagation formula of the BasicBlock residual module in the point cloud conditional encoder is: ; ; ; ; in, The input is a sparse tensor of a point cloud, where N is the number of points in the point cloud. The dimension of point cloud coordinate features; , It is a sparse convolution kernel. The kernel size is the convolution kernel size. The convolution stride; This is a Minkowski sparse convolution operation; LeakyReLU activation function is specifically designed for point cloud sparse tensors.

6. The intelligent extraction method for three-dimensional spatiotemporal geographic entities based on a large language model according to claim 1, characterized in that, The joint verification includes vertex sequence verification, face sequence initial verification, and architectural geometry verification.

7. The intelligent extraction method for three-dimensional spatiotemporal geographic entities based on a large language model according to claim 1, characterized in that, The multi-source heterogeneous data includes at least one of the following: static attribute data in unstructured documents, dynamic attribute data provided by spontaneous geographic information platforms, and spatial positioning data in structured address databases.

8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that the processor executes the computer program to implement the intelligent extraction method for three-dimensional spatiotemporal geographic entities based on a large language model according to any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the intelligent extraction method for three-dimensional spatiotemporal geographic entities based on a large language model, as described in any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the intelligent extraction method for three-dimensional spatiotemporal geographic entities based on a large language model, as described in any one of claims 1-7.