A 3D urban spatial data processing method and system for large models

By decomposing urban space into a three-dimensional grid network with dynamically adjusting particle accuracy and constructing a multi-dimensional tensor structure [C, H, W, L], the problem of poor adaptability of three-dimensional data processing and AI big models in the existing technology is solved, and efficient multi-factor correlation and dynamic adaptability are achieved.

CN119849014BActive Publication Date: 2025-06-20SHANGHAI YINGYI URBAN PLANNINGDESIGN CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510331219.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-20
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

The existing technology has significant shortcomings in the adaptability of urban three-dimensional data processing and AI big models, including data dimension fragmentation, inefficient three-dimensional data expression, insufficient multi-factor fusion capability and poor dynamic adaptability.

Method used

By decomposing the urban space into a three-dimensional grid network with dynamically adjusting the particle accuracy, and building a multi-dimensional tensor structure [C, H, W, L] that integrates multi-factor attributes, the problems of poor compatibility with AI large models and three-dimensional urban data, multi-factor fragmentation and insufficient dynamic adaptability are solved.

Benefits of technology

It realizes efficient compatibility between AI models and three-dimensional urban data, improves the semantic correlation ability and dynamic adaptability of multiple elements, and reduces the complexity of data processing and computing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119849014B_ABST
    Figure CN119849014B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional urban spatial data processing method and system for large models, mainly solving problems such as poor compatibility between three-dimensional data and AI models, fragmentation of multiple elements, and limitations of static annotation. Its innovative solutions include: first, dynamic rasterization, which divides the urban space into three-dimensional pixel points according to adjustable granularity and integrates three-dimensional coordinates and multi-channel attributes; second, multi-channel tensors, which support multi-channel One-Hot encoding for classification tasks and three-dimensional bounding box normalization for object detection tasks, and automatically expand the protection scope of elements based on a rule engine; third, rule-driven label expansion, which retains the original value for binary data and uses Min-Max or Z-Score normalization for continuous data to improve the model training efficiency; fourth, channel-wise normalization, which includes a three-dimensional grid division engine, a multi-source element fusion module, a dynamic label encoder, etc., to achieve full-link connection from data collection to AI model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technology of urban spatial data processing, and particularly relates to a method and system for processing three-dimensional urban spatial data for large models. Background Art

[0002] With the breakthrough of AI large models (such as GPT, Vision Transformer, etc.) in the fields of natural language processing and image recognition, their application potential in urban spatial analysis has gradually emerged. However, there are significant deficiencies in the compatibility between existing technologies and AI large models in urban three-dimensional data processing, specifically manifested as follows:

[0003] 1. Data dimension fragmentation: The training and output of AI large models usually rely on single-dimensional data (such as text sequences or two-dimensional images), while the urban space is essentially a complex system containing multi-dimensional information such as geographical coordinates, height, and facility attributes. Existing technologies lack a method for uniformly encoding three-dimensional spatial structures and multi-element attributes, resulting in difficulty for AI large models to understand the three-dimensional correlation of urban spaces.

[0004] 2. Inefficient three-dimensional data representation: Traditional methods describe three-dimensional urban spaces through point clouds, BIM models, or oblique photography data. However, these data formats have high redundancy, fixed granularity, and are not directly compatible with the tensor input structure of AI large models. For example, point cloud data lacks semantic attribute channels, and BIM models are difficult to dynamically adjust precision, severely limiting the model training efficiency.

[0005] 3. Insufficient multi-element fusion ability: Urban spaces contain cross-domain elements such as nature, matter, culture, and economy. Existing technologies usually adopt independent modeling of subsystems (such as GIS systems for processing geographical information and BIM systems for processing building data), resulting in fragmentation of multi-source data. AI large models cannot extract cross-element correlation laws from heterogeneous data and are difficult to support comprehensive decision-making (such as disaster early warning requires simultaneous analysis of geological, building, and pipeline data).

[0006] 4. Poor dynamic adaptability: The boundaries and attributes of urban elements change dynamically over time and scenarios (such as subway construction requires real-time expansion of the protection range). However, existing label encoding methods rely on manual static annotation and cannot automatically adapt to dynamic requirements, resulting in limited model generalization ability. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a method and system for processing three-dimensional urban spatial data for large models, aiming at the deficiencies in the above-mentioned existing technologies. By decomposing urban spaces into three-dimensional grid networks with dynamically adjustable granular precision and constructing multi-dimensional tensors ([C, H, W, L]) that integrate multi-element attributes, the industrial problems of poor compatibility between AI large models and three-dimensional urban data, fragmentation of multi-elements, and insufficient dynamic adaptability are fundamentally solved.

[0008] The first aspect of the present invention discloses a method for processing three-dimensional urban spatial data for large models, including the following steps:

[0009] (a) Divide the target urban space into a three-dimensional grid network according to a preset granularity, and each grid corresponds to a three-dimensional pixel point;

[0010] (b) Determine the types of urban elements intersecting within each three-dimensional pixel point through spatial calculation, and the urban elements include natural elements, material elements, cultural elements, and economic elements;

[0011] (c) Construct a multi-dimensional urban spatial pixel tensor with dimensions [C, H, W, L], where:

[0012] C is the number of channels, and each channel corresponds to a classification of urban elements;

[0013] H, W, and L are the numbers of three-dimensional pixel points in the length, width, and height directions of the urban space respectively;

[0014] (d) Perform label encoding on the pixel tensor according to the task type:

[0015] If it is a classification task, use multi-channel one-hot encoding to mark the category of each three-dimensional pixel point;

[0016] If it is an object detection task, generate normalized detection labels based on the absolute coordinates of the three-dimensional bounding box;

[0017] (e) Perform channel-wise data normalization on the pixel tensor, including retaining binary channel data and performing Min-Max normalization or Z-Score standardization on continuous channel data;

[0018] (f) Generate the urban spatial data required for training the urban spatial analysis model.

[0019] In the above method, the three-dimensional grid granularity in step (a) is 1m×1m×1m.

[0020] In the above method, the classification task label encoding in step (d) further includes dynamic boundary expansion:

[0021] Expand the original boundary of the specified urban element according to a preset rule to generate an expansion area;

[0022] Mark the three-dimensional pixel points within the expansion area as new protected categories;

[0023] The expansion rule is: Based on the original boundary parameters M(Hm, Wm, Lm), expand the distance Δ along the three-dimensional direction to generate the expansion boundary M’(Hm±Δ, Wm±Δ, Lm±Δ).

[0024] For the above method, the expansion distance Δ is dynamically set according to the element type, including:

[0025] Subway protection area: Δ = 50m;

[0026] Municipal pipeline protection area: Δ = 20m;

[0027] Historical building protection area: Δ = 30m;

[0028] The expanded area is marked as an independent category channel and coexists with the original element channel in the label tensor.

[0029] For the above method, the target detection task label encoding in step (d) includes:

[0030] Calculate the absolute coordinates of the three-dimensional bounding box of the target area, including the center point coordinates (x, y, z) and dimensions (dx, dy, dz);

[0031] Normalize the absolute coordinates to relative coordinates, and the formula is:

[0032] x_center = x / H, y_center = y / W, z_center = z / L;

[0033] dx_norm = dx / H, dy_norm = dy / W, dz_norm = dz / L;

[0034] Generate a detection label vector [class_id, x_center, y_center, z_center, dx_norm, dy_norm, dz_norm], where class_id is the category identifier.

[0035] For the above method, the data normalization processing in step (e) includes:

[0036] Retain the original 0 / 1 values for binary channel data;

[0037] For continuous economic element data, use Min - Max normalization, and the formula is:

[0038] x_norm = (x - x_min) / (x_max - x_min);

[0039] Save the normalization parameters of the training set and synchronously apply the same parameters to the test set.

[0040] The second aspect of the present invention discloses an urban spatial data processing system, including:

[0041] A three-dimensional grid division engine for dividing urban space into uniform three-dimensional pixel points;

[0042] A multi-source feature fusion module for generating a pixel tensor containing multi-channel urban features;

[0043] A dynamic label encoder that supports one-hot encoding for classification tasks and bounding box normalization for object detection tasks;

[0044] A per-channel normalization processor for differentially and / or standardizing tensor data according to channel type.

[0045] In the above system, the dynamic label encoder includes a boundary expansion unit for automatically expanding the protection range and updating the label tensor according to the feature type.

[0046] The present invention has the following advantages compared with the prior art:

[0047] 1. Three-dimensional rasterization and multi-channel fusion: Decompose urban space into three-dimensional pixel points with dynamically adjustable granularity (such as 1m³), and each pixel point integrates three-dimensional coordinates (H, W, L) and multi-channel attribute information (such as natural feature status, facility function, cultural label, economic indicator) to form a unified tensor structure [C, H, W, L], which directly adapts to the multi-dimensional input requirements of AI large models.

[0048] 2. Dynamic granularity control: Support flexible adjustment of grid granularity according to task requirements (such as 1m³ for fine facility management and 10m³ for macro planning), achieve a balance between data accuracy and computational efficiency, and overcome the rigidity problem of traditional point clouds or BIM models.

[0049] 3. Multi-factor semantic association: Explicitly encode cross-domain feature attributes through the channel dimension (C), enabling the AI large model to learn the interaction rules between features from a unified data source (such as the impact of subway vibrations on surrounding buildings), and improving the reasoning ability in complex scenarios.

[0050] 4. Automatic label expansion: Dynamically expand the feature boundary based on a rule engine (such as the subway protection range automatically extends 50m with the construction stage), and update the pixel point label in real time, solve the problem of model lag caused by static annotation, and significantly improve the adaptability to dynamic scenarios.

[0051] 5. Industrial-level compatibility: The generated three-dimensional pixel tensor can be directly input into vision large models (such as ViT) and multi-modal large models (such as GPT-4V) for end-to-end training, reducing the threshold for the implementation of AI technology in urban space analysis tasks.

[0052] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Brief Description of the Drawings

[0053] Figure 1 This is a flowchart of the method for processing urban spatial data based on 3D grids according to the present invention.

[0054] Figure 2 This is an example diagram of the initial pixel points of the urban space.

[0055] Figure 3 This is an architecture diagram of the urban spatial data processing system according to the present invention. Detailed Description of the Invention

[0056] As Figure 1 shown, a method for processing 3D urban spatial data for large models includes the following steps:

[0057] Step 1: 3D grid division and data collection (as Figure 2 shown)

[0058] Step 1.1: Define the urban space range and granularity

[0059] Input parameters: Set the 3D dimensions of the target urban space (such as length H = 1000m, width W = 800m, height L = 50m) and the grid granularity (such as 1m×1m×1m);

[0060] Grid generation: Divide the urban space into 1000×800×50 3D pixel points through the spatial coordinate system, and the coordinates of each pixel point are (h, w, l) (h ∈ [0, 1000), w ∈ [0, 800), l ∈ [0, 50));

[0061] Dynamic adjustment: If it is necessary to reduce the computational complexity, the granularity can be adjusted to 5m×5m×5m, and the number of grids is reduced to 200×160×10.

[0062] Step 1.2: Collect multi-source urban element data

[0063] Data sources:

[0064] Natural elements: Terrain elevation, vegetation distribution, hydrological data (GIS system);

[0065] Material elements: Building BIM models, transportation road networks, underground pipelines (CAD drawings);

[0066] Cultural elements: Coordinates of historical buildings, protection levels (database);

[0067] Economic elements: Regional GDP, population density (statistical reports).

[0068] Spatial Computing: Use spatial indexing techniques (such as R-trees) to determine the intersection relationship between each three-dimensional voxel and urban elements.

[0069] Example: If a voxel (h, w, l) intersects with a subway tunnel, mark its material element channel value as 1.

[0070] Step 2: Construct a multi-dimensional voxel tensor

[0071] Step 2.1: Channel Allocation and Data Mapping

[0072] Channel Definition:

[0073] Channel C1 (Natural Elements): Binary Marking (0 / 1 indicates the presence or absence of natural elements);

[0074] Channel C2 (Material Elements): Classification Coding (1 = Transportation Facilities, 2 = Building Complexes, 3 = Municipal Pipelines);

[0075] Channel C3; (Cultural Elements): Protection Level (0 - 5 levels);

[0076] Channel C4 (Economic Elements): Continuous Value (such as GDP density)

[0077] Tensor Generation: Fill the attributes of each voxel into the corresponding channels to generate a four-dimensional tensor Tensor[C, H, W, L], where C = 4, H = 1000, W = 800, L = 50;

[0078] Step 2.2: Data Storage Optimization

[0079] Sparse Storage: Adopt sparse matrix compression for null value regions (such as areas without buildings underground) to reduce storage overhead.

[0080] Step 3: Label Encoding

[0081] 3.1 Classification Task (Semantic Segmentation)

[0082] Step 3.1.1: Define the Class Set

[0083] Preset Classes: Natural Elements (N), Transportation Facilities (T), Building Complexes (B), Subway Protection Range (MP), Historical Buildings (HB).

[0084] Step 3.1.2: Dynamic Boundary Expansion

[0085] Rule Engine:

[0086] Input the original boundary parameters (such as the center of the subway tunnel Hm = 500m, Wm = 400m, Lm = 10m);

[0087] The extended distance Δ = 50m, generating an extended area: H′ ∈ [500 - 50, 500 + 50], W′ ∈ [400 - 50, 400 + 50], L′ ∈ [10 - 5, 10 + 5];

[0088] Traverse the pixel points within the extended area and update the class label from "Traffic Facility (T)" to "Subway Protection Range (MP)".

[0089] Step 3.1.3: Generate One - Hot labels

[0090] Each pixel point generates a 5 - dimensional vector (corresponding to 5 classes), for example, the pixel label of the subway protection range is [0, 0, 0, 1, 0].

[0091] 3.2 Object Detection Task (Region Detection)

[0092] Step 3.2.1: Calculate the 3D bounding box

[0093] Input subway tunnel parameters: center coordinates (500, 400, 10), size (100m, 80m, 10m);

[0094] Normalization calculation:

[0095] x center = 500 / 1000 = 0.5, y center = 400 / 800 = 0.5, z center = 10 / 50 = 0.2;

[0096] dx norm = 100 / 1000 = 0.1, dy norm = 80 / 800 = 0.1, dz norm = 10 / 50 = 0.2;

[0097] Label vector: Generate [4, 0.5, 0.5, 0.2, 0.1, 0.1, 0.2], where 4 is the class ID of "Subway Protection Range".

[0098] Step 4: Data Normalization

[0099] Step 4.1: Channel - by - channel processing

[0100] Binary channels (such as natural elements): Keep the original 0 / 1 values;

[0101] Continuous channels (such as economic elements):

[0102] Calculate the global minimum / maximum values of the training set (such as GDP density x min = 0, x max = 1000);

[0103] Apply Min - Max normalization: ;

[0104] Step 4.2: Parameter Persistence

[0105] Save the normalization parameters of the training set (such as x min , x max ), and apply them synchronously during the inference of the test set.

[0106] Step 5: Model Training and Validation

[0107] Step 5.1: Task Adaptation

[0108] Classification task: Use the 3D U - Net model, input tensor [4, 1000, 800, 50], and output the segmentation result [5, 1000, 800, 50];

[0109] Object detection task: Use 3D Faster R - CNN, input the same tensor, and output the bounding box coordinates and classes.

[0110] Step 5.2: Dynamic Scene Validation

[0111] Test case: Add a subway construction phase, and extend the protection range to Δ = 70m;

[0112] System response: Automatically update the extended area labels, retrain the model, and the F1 - score is increased by 12%.

[0113] Example 1: Dynamic Labeling of Subway Protection Area

[0114] Input: Subway tunnel center (500m, 400m, 10m), original size (100m, 80m, 10m), Δ = 50m;

[0115] Output:

[0116] Classification label: Pixels within the extended area [450 - 550m, 350 - 450m, 5 - 15m] are labeled as "MP";

[0117] Detection label: Normalized bounding box [4, 0.5, 0.5, 0.2, 0.1, 0.1, 0.2].

[0118] Example 2: Multi - factor Analysis of Historical Buildings

[0119] Input: Building center (200m, 300m, 20m), Δ = 30m, cultural protection level 3;

[0120] Output:

[0121] Classification label: Pixels within the extended area [170 - 230m, 270 - 330m, 0 - 50m] are labeled as "HB".

[0122] Economic corridor data: The regional GDP density is normalized from 500 to 0.5.

[0123] The verification of technical effects is as follows:

[0124]

[0125] As Figure 3 shown, the specific implementation of the urban spatial data processing system is described below:

[0126] 1. Overview of System Architecture

[0127] This system consists of the following core modules, supporting the dynamic processing, label encoding, and model adaptation of 3D urban spatial data: 3D grid division engine, multi-source feature fusion module, dynamic label encoder, channel-by-channel normalization processor, and AI model interface layer;

[0128] 2. Detailed Implementation of Modules

[0129] 2.1 3D Grid Division Engine

[0130] Function: Cut the urban space into a 3D pixel network with adjustable granularity.

[0131] Hardware implementation: A parallel computing cluster accelerated by GPU (such as NVIDIA A100).

[0132] Software process:

[0133] 1) Input parameter configuration:

[0134] Spatial range: Absolute coordinates of length (H), width (W), and height (L) (such as H = 1000m, W = 800m, L = 50m);

[0135] Granularity: Preset value (such as 1m³) or dynamic rule (such as "1m³ in the city center area, 10m³ in the suburbs").

[0136] 2) Grid generation algorithm:

[0137] Adopt the spatial hashing algorithm to map the 3D coordinates (x, y, z) to a unique grid ID:

[0138] where Δ is the granularity (such as 1m).

[0139] 3) Dynamic adjustment mechanism:

[0140] If the density of area features in a certain area is detected to exceed the threshold (such as a densely built-up area), the granularity is automatically switched from 10m to 1m.

[0141] Example:

[0142] Input: Urban space of 1000m×800m×50m, granularity 1m³;

[0143] Output: Generate 1000×800×50 = 40000000 three-dimensional pixel points, and each pixel point stores the reference coordinates (h, w, l).

[0144] 2.2 Multi-source Feature Fusion Module

[0145] Function: Map multi-source data such as GIS, BIM, and statistical reports to three-dimensional pixel channels.

[0146] Key technologies: Spatial relational database (such as PostGIS), semantic parsing engine.

[0147] Processing flow:

[0148] 1) Data access:

[0149] Support multiple format inputs:

[0150] Vector data (Shapefile, GeoJSON);

[0151] 3D models (IFC, OBJ);

[0152] Tabular data (CSV, Excel).

[0153] 2) Spatial relationship calculation: Use R-tree index to accelerate spatial intersection judgment;

[0154] 3) Channel mapping rules:

[0155] Natural feature channel: The terrain elevation value is directly filled;

[0156] Cultural feature channel: The historical building protection level is weighted by radius attenuation (such as level 5 in the core area and level 3 in the outer extension area).

[0157] Example:

[0158] Input: Subway BIM model (center line coordinates + radius);

[0159] Output: All pixel points intersecting with the tunnel, and their material feature channels are marked as "1" (transportation facilities).

[0160] 2.3 Dynamic Label Encoder

[0161] Function: Generate classification or detection labels according to the task type, and support rule-driven dynamic expansion.

[0162] Sub-modules:

[0163] Classification Label Unit: One-Hot encoding based on the rule engine;

[0164] Detection Label Unit: 3D bounding box coordinate normalization;

[0165] Dynamic Expansion Engine: Automatically adjust the protection range of elements.

[0166] Implementation details:

[0167] 1) Rule Engine:

[0168] Predefined Extension Rule Library (JSON format): {

[0169] "Subway Protection Area": {

[0170] "Element Type": "Transportation Facility",

[0171] "Expansion Distance": {"H": 50, "W": 50, "L": 5},

[0172] "New Category": "MP"

[0173] },

[0174] "Historical Building Protection Area": {

[0175] "Element Type": "Cultural Element",

[0176] "Expansion Distance": {"H": 30, "W": 30, "L": 10},

[0177] "New Category": "HB"

[0178] }

[0179] }

[0180] Monitor element changes in real time (such as new construction areas) and trigger rule execution.

[0181] 2) Classification Label Generation:

[0182] For each pixel point, assign the final category according to the priority (such as MP > HB > building complex) and generate a One-Hot vector.

[0183] 3) Detection Label Generation:

[0184] Call the 3D minimum bounding box algorithm (such as PCA-AABB) to calculate the bounding box and output the normalized coordinates.

[0185] Example:

[0186] Input: The original subway area is H = 500m, W = 400m, L = 10m, and the expansion rule is Δ = 50m;

[0187] Output: The expanded area label is [450−550m, 350−450m, 5−15m], and the detection box is [0.5, 0.5, 0.2, 0.1, 0.1, 0.2].

[0188] 2.4 Channel normalization processor

[0189] Function: Process data differently according to channel types to improve the stability of model training.

[0190] Processing logic:

[0191] 1) Channel type recognition:

[0192] Binary channels (such as element existence): Skip normalization;

[0193] Continuous channels (such as economic indicators): Automatically select Min - Max or Z - Score.

[0194] 2) Distributed computing:

[0195] Use Apache Spark for parallel computing of channel statistics (such as mean, variance).

[0196] 3) Parameter persistence:

[0197] Store the normalization parameters of the training set (such as the maximum GDP value of 1000) in the Redis database for use during inference.

[0198] Example:

[0199] Input: Original economic channel data [0, 350, 920, 1000];

[0200] Output: After Min - Max normalization [0, 0.35, 0.92, 1.0].

[0201] 2.5 AI model interface layer

[0202] Function: Convert the processed tensor data into the input format of the AI large model.

[0203] Adapter type:

[0204] 1) Visual large model adapter:

[0205] Reshape the tensor [C, H, W, L] into a sequence format supported by ViT (e.g., H×W×L C-dimensional vectors).

[0206] 2) Multimodal large model adapter:

[0207] Concatenate the tensor data with the text description (e.g., "subway protection area") to generate a multimodal input compatible with GPT-4V.

[0208] Example:

[0209] Input: Four-dimensional tensor [4, 1000, 800, 50];

[0210] Output: ViT input sequence of 1000×800×50 = 40,000,000 4-dimensional embedding vectors.

[0211] 3. System dynamic collaboration process

[0212] Scenario: Subway construction causes a change in the protection range

[0213] 1) Data update: The construction party submits new shield tunnel coordinates;

[0214] 2) Grid re-partitioning: The engine switches the granularity to 1m³ in the changed area;

[0215] 3) Label extension: The rule engine automatically extends the MP label range;

[0216] 4) Model hot update: Inject new data into the online model through the interface layer to achieve real-time inference.

[0217] 4. Technical effect verification

[0218]

[0219] Through modular design, rule engine-driven, and distributed computing, this system realizes the full-link automation of three-dimensional urban spatial data from collection to AI model training, and overcomes key technical bottlenecks such as multi-source heterogeneous data fusion, dynamic scenario adaptation, and large model compatibility.

[0220] The above are only the preferred embodiments of the present invention, and do not impose any limitations on the present invention. Any simple modifications, changes, and equivalent structural changes made to the above embodiments according to the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for processing three-dimensional urban space data for a large model, characterized in that: The following steps are involved: (a) dividing the target urban space into a three-dimensional grid network according to a preset granularity, where each grid corresponds to a three-dimensional pixel point; (b) determining the type of urban elements intersecting within each three-dimensional pixel point through spatial calculation, where the urban elements include natural elements, material elements, cultural elements, and economic elements; (c) Construct a multi-dimensional urban space pixel tensor with the dimension [C, H, W, L], where: C is the number of channels, each channel corresponds to a city feature classification; H, W, L are the number of three-dimensional pixels in the length, width, and height of the urban space, respectively; (d) Label encode the pixel tensor according to the task type: if it is a classification task, use multi-channel one-hot encoding to label each three-dimensional pixel point; If it is a target detection task, a normalized detection label is generated based on the absolute coordinates of the three-dimensional bounding box; (e) the pixel tensor is normalized by channel data, including retaining binary channel data and performing Min-Max normalization or Z-Score normalization on continuous channel data; (f) urban space data required for training an urban space analysis model is generated; The classification task label encoding in step (d) further includes dynamic boundary extension: expanding the original boundary of the specified urban element according to preset rules to generate an extended area; marking the three-dimensional pixels in the extended area as a newly added protection category; the extension rule is: based on the original boundary parameters M (Hm, Wm, Lm), expanding the distance Δ along the three-dimensional direction to generate an extended boundary M' (Hm±Δ, Wm±Δ, Lm±Δ).

2. A method for processing three-dimensional urban space data for a large model according to claim 1, characterized in that: The three-dimensional grid particle size in step (a) is 1m×1m×1m.

3. A method for processing three-dimensional urban space data for a large model according to claim 1, characterized in that: The extended distance Δ is dynamically set according to the feature type, including: subway protection area: Δ=50m; municipal pipeline protection area: Δ=20m; historical building protection area: Δ=30m; the extended area is marked as an independent category channel and coexists with the original feature channel in the label tensor.

4. A method for processing three-dimensional urban space data for a large model according to claim 1, characterized in that: The object detection task label encoding in step (d) includes: calculating the absolute coordinates of the three-dimensional bounding box of the target area, including the center point coordinates (x, y, z) and size (dx, dy, dz); normalizing the absolute coordinates to relative coordinates, the formula is: x_center=x / H,y_center=y / W,z_center=z / L; dx_norm=dx / H,dy_norm=dy / W,dz_norm=dz / L; Generate a detection label vector [class_id,x_center,y_center,z_center,dx_norm,dy_norm,dz_norm], where class_id is the class identifier.

5. A method for processing three-dimensional urban space data for a large model according to claim 1, characterized in that: The data normalization processing in step (e) includes: retaining the original 0 / 1 value for binary channel data; using Min-Max normalization for continuous economic factor data, the formula is: x_norm=(x-x_min) / (x_max-x_min); saving the normalization parameters of the training set, and synchronously applying the same parameters to the test set.

6. An urban space data processing system based on the three-dimensional urban space data processing method for large models according to claim 1, characterized in that: include: A 3D grid partitioning engine for segmenting urban space into uniform 3D pixels; a multi-source feature fusion module for generating pixel tensors containing multi-channel urban features; a dynamic label encoder that supports one-hot encoding for classification tasks and bounding box normalization for object detection tasks; and a channel-by-channel normalization processor that differentiates and / or normalizes tensor data by channel type. The dynamic label encoder includes a boundary extension unit for automatically extending the protection range and updating the label tensor according to the feature type.

Citation Information

Patent Citations

  • Point cloud target detection method, system, device and medium

    CN116403062A