Spatial intelligent data expression design and definition mode highly suitable for large model
By reconstructing the physical world into discrete, atomized spatial entities and employing semantic primitives and logical stitching mechanisms, the conflict between determinism and probability, as well as the problem of multimodal data fusion, in generating an interactive digital earth from a large model are resolved, achieving efficient and continuous spatial data representation and generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies for generating interactive digital earths using large models suffer from problems such as conflicts between deterministic structures and probabilistic reasoning, lack of cross-grid semantic logic stitching, conflicts in multimodal observation data attributes, and a lack of unified data containers.
By employing a probabilistic inference mechanism of generative large models, the physical world is reconstructed into discrete, atomized spatial entities. Probabilistic features and semantic primitives are defined, and cross-grid logical continuity and multimodal data fusion are achieved through normalized grid segmentation, logical stitching, and topological linking.
It achieves uncertainty compatibility in the large model generation process, reduces data lexical consumption, ensures spatial logical continuity and data accuracy, and supports efficient fusion of dynamic evolution and multimodal observations.
Smart Images

Figure CN121811246A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer data processing and artificial intelligence technology, and in particular to a spatial intelligent data representation design and definition method that is highly applicable to large models. Background Technology
[0002] With the rapid development of generative artificial intelligence technology, the application of large models has gradually expanded from single natural language processing or two-dimensional image generation to the fields of three-dimensional spatial understanding and digital earth construction. Spatial intelligence requires computers not only to render three-dimensional scenes, but also to understand the semantic relationships, physical attributes and evolution laws between entities within space.
[0003] In the existing field of spatial data processing, the application of large models mainly focuses on retrieving existing data or generating visual representations of specific scenes. However, existing technologies have significant technical shortcomings when facing the new demand of "directly generating an interactive digital earth using large models." First, traditional geographic information data or 3D mesh models are deterministic and rigid geometric expressions, while generative large models output results based on probabilistic inference. There is a conflict between "deterministic structure" and "probabilistic reasoning" at the underlying data logic, resulting in the inability to structurally record or correct the illusions of model output. Second, when processing global-scale spatial data, the context window limitation of large models necessitates physical slicing of space. However, existing technologies lack a semantic logic stitching mechanism across meshes, leading to geometric and logical breaks between cross-boundary entities. Finally, existing data structures lack a definition of the "public-private domain" micro-topology. When multimodal observation data have attribute conflicts at the same spatial location, there is a lack of a unified data container that can integrate multi-source information and express the possibility of existence. Summary of the Invention
[0004] The objective of this application is to provide a spatial intelligent data representation design and definition method that is highly applicable to large-scale models, including: receiving multimodal observation data as input, the multimodal observation data including satellite remote sensing images, ground-view street view images, and radar point cloud data; based on the probabilistic inference mechanism of generative large-scale models, reconstructing the continuous space of the physical world into a discrete set of atomized spatial entities, the atomized spatial entities serving as the smallest independent unit of data representation; defining probabilistic features for each atomized spatial entity, the probabilistic features including an existence confidence level reflecting the possibility of the entity's actual existence in the physical world, the existence confidence level being quantized into a numerical range field between zero and one, used to accommodate the uncertain output during the large-scale model generation process; constructing the geometric shape of the atomized spatial entities using semantic primitives, discarding non-semantic mesh patches, and using parameterized vector combinations or geometric parameter sets to describe the centroid coordinates, orientation angle vectors, and spatial contours of the entities; generating a spatial entity objectification representation model containing the probabilistic features and semantic primitives, serving as the structured foundation data for digital earth generation.
[0005] By adopting the above technical solution, an existence confidence field is defined for each entity, quantifying the real existence of the physical world into a numerical range of 0 to 1, thereby compatibility with the uncertainty in the large model generation process at the data level; at the same time, the high word consumption of grid patches is abandoned, and instead semantic primitives (such as parameterized vector groups or geometric combinations) are used to construct geometric shapes, so that spatial data can be understood, generated and stored by large models with extremely low word cost.
[0006] Optionally, to address the issues of parallel generation and logical continuity of global-scale spatial data under the constraints of large model context windows, the design and definition methods further include: defining a normalized grid segmentation strategy to divide global space into regular grid units adapted to the memory and lexical constraints of single inference in large models; based on the regular grid units, the large model is triggered in parallel to generate local atomized spatial entities within each grid; a logical connection and fusion mechanism for cross-domain entities is established, whereby when linear or planar atomized spatial entities cross grid boundaries, a logical stitching operation is performed by identifying the semantic primitive features and attribute fingerprints at the breakpoints, recombining entity fragments located in different grid units into complete entities with globally unique identifiers, ensuring spatial logical continuity across grids.
[0007] By adopting the above technical solution, parallel reasoning is achieved by dividing the global space into regular units that adapt to the model's memory. More importantly, a logical stitching mechanism is established at the grid boundary. By identifying the semantic primitive features and attribute fingerprints at the breakpoints, the physically cut entity fragments are logically reorganized into globally unique complete entities, ensuring the continuity of spatial logic.
[0008] Optionally, the step of constructing the geometric form of the atomized spatial entity using semantic primitives specifically includes: for linear entities, defining them using parametric primitives containing a centerline vector group and width attributes; for volumetric entities, using a set or more parametric combinations of cuboids and ellipsoids for fitting expression, and replacing vertex data by storing the axis length, rotation angle and combination logic of the geometric body to reduce the lexical consumption of data expression.
[0009] By adopting the above technical solutions, parametric definitions of centerline vector groups and width attributes are used for linear entities, and parametric combination fitting of cuboids or ellipsoids is used for volumetric entities. Instead of relying on massive vertex data, the shape is described by storing axis lengths, rotation angles and combination logic, which significantly reduces the data volume and improves the efficiency of generating complex 3D scenes from large models.
[0010] Optionally, the logical connection and fusion mechanism further includes: generating soft connection anchors at the grid boundary, wherein the soft connection anchors carry the semantic category and geometric trend information of the entity; using the semantic reasoning capability of the large model, matching and determining the soft connection anchors of adjacent grids; when determined to be the same entity, fusing the probabilistic features of both ends, and updating the geometric parameters of the entity to eliminate physical cracks.
[0011] By adopting the above technical solution, a fusion strategy based on soft connection anchor points is proposed. Anchor points carrying semantic categories and geometric trends are generated at the segmentation boundary. The semantic reasoning capability of the large model is used to match and determine the anchor points of adjacent grids. Once it is confirmed that they belong to the same entity, their probabilistic features are fused and the geometric parameters are corrected, thereby eliminating physical gaps and ensuring the integrity of the data.
[0012] Optionally, the probabilistic features also include the definition of the entity's lifecycle dimension: pre-setting fields for establishment time, current state time, and estimated disappearance time in the data structure; based on the timestamp differences of multimodal observation data, the dynamic evolution process of the entity is inferred through a large model, and the probabilistic features are used to express the confidence of the entity's state in different time dimensions.
[0013] By adopting the above technical solution, a lifecycle dimension is introduced into the data structure. By pre-setting fields for establishment time, state time, and predicted disappearance time, and combining the timestamp differences of multimodal data, the large model can deduce the dynamic change process of entities and express the state confidence at different time points using probabilistic features, thus realizing the leap from static map to dynamic digital twin.
[0014] Optionally, to address the issues of missing micro-connectivity and attribute conflicts of spatial entities under multimodal observation, the design and definition method further includes: defining entity topological links, which are used to describe the logical interconnection and spatial orientation relationships between different atomized spatial entities; constructing a cross-domain micro-topological network, explicitly defining the connection nodes between private domain entities and public domain entities, whereby the connection nodes, as a special type of atomized spatial entity, are used to bridge the internal road network of a closed area with the public road network, forming a continuous topological path that can be used for navigation inference; and when observation data from different modalities generate attribute conflicts for the same area, using the logical consistency constraints of the entity topological links to correct the semantic category of the entity or adjust the existence confidence.
[0015] By adopting the above technical solution, a cross-domain micro-topology network was constructed. By defining entity topology links, the connection nodes between private and public entities were clarified, thus eliminating the blind spots of traditional navigation data. When multimodal data conflict, the logical consistency constraints of the topology network are used to correct the semantic category of entities or adjust the existence confidence, thereby improving the accuracy of the data.
[0016] Optionally, the entity topology link includes hierarchical membership and spatial inclusion relationships: establishing a parent-child entity hierarchy, defining functional areas as parent entities, defining facilities within the areas as child entities, and child entities inheriting the spatial attributes of the parent entities; defining a three-dimensional stacking relationship, using spatial orientation semantics to describe the positional order of entities in the vertical direction, and distinguishing above-ground, ground-level, and underground entities through logical identifiers rather than pure geometric coordinates.
[0017] By adopting the above technical solution, the hierarchy and spatial relationship between entities are refined, a parent-child entity hierarchy structure and three-dimensional stacking relationship are established, and entities above ground, on the ground and underground are distinguished by logical identifiers rather than simple geometric coordinates, so that large models can understand complex spatial semantics.
[0018] Optionally, the correction of the semantic category of the entity or the adjustment of the existence confidence specifically includes: if the entity type identified by the visual modality data contradicts the spatial logic deduced from the entity topology link, then the existence confidence weight of the entity generated based on the visual modality is reduced; based on the connectivity rules of the topology network, the missing entities in the observation blind zone are automatically inferred and filled using a large model, and given a probabilistic feature identifier based on the deduction.
[0019] By adopting the above technical solution and using a weighted correction mechanism based on topological logic, when the visual modality recognition result contradicts the spatial logic, the confidence of the visual weight is automatically reduced, and the topological rules are used to fill the observation blind zone, giving the inferred entities a specific probability identifier, thereby ensuring the logical self-consistency of the spatial data.
[0020] Optionally, it also includes: defining a visual attribute mapping field in the atomized spatial entity, the field being associated with side facade texture features and physical material parameters; filtering entities based on the existence confidence, and converting entities with probability values higher than a preset threshold into 3D rendering instructions through their semantic primitives and visual attribute mapping fields to achieve structured reconstruction of the digital earth.
[0021] By adopting the above technical solution, a mapping channel from abstract data to visual presentation is established. Visual attribute mapping fields are defined in atomic entities, and texture and material parameters are associated. After filtering low-probability entities based on existence confidence, high-confidence semantic primitives are converted into 3D rendering instructions, thus realizing efficient visualization reconstruction of structured data.
[0022] The second objective of this application is to provide a spatial intelligent data generation system, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the aforementioned highly applicable large-scale model spatial intelligent data expression design and definition method. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the spatial entity construction method based on multimodal input and probabilistic inference in this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0025] like Figure 1 As shown in the figure, this application discloses a spatial entity construction method based on multimodal input and probabilistic inference, which includes the following steps.
[0026] S01: Receives multimodal observation data as input, including satellite remote sensing images, ground-view street view images, and radar point cloud data.
[0027] Understandably, this step involves the preprocessing and standardization cleaning of massive amounts of heterogeneous data. Specifically, data sources include, but are not limited to, high-resolution series optical satellite remote sensing images, synthetic aperture radar point cloud data, and ground-view street scene images collected by urban-level vehicle-mounted mobile measurement systems. When acquiring optical satellite remote sensing data, the system automatically filters image slices with clear physical environments, atmospheric transparency greater than a threshold, and low cloud cover. The sampling frequency is set to once a day to ensure timeliness. At the same time, radiometric calibration parameters are used to convert the raw digital quantization values into apparent reflectance at the top of the atmosphere, and orthorectification is performed based on a rational polynomial coefficient model to ensure accurate correspondence between pixel coordinates and geographic coordinates. For radar point cloud data, the backscattering coefficient is extracted to help identify building outlines under cloud cover. For ground-view street scene data, the camera pose is solved using a motion reconstruction structure algorithm to map two-dimensional image pixels to a three-dimensional spatial coordinate system. These multimodal data are feature-aligned through a multimodal encoder based on a transformer architecture, mapping observation data of different resolutions and perspectives into a unified high-dimensional tensor sequence, which serves as the input context for generative large models.
[0028] S02: Based on the probabilistic inference mechanism of generative large models, the continuous space of the physical world is reconstructed into a discrete set of atomized spatial entities, with the atomized spatial entities serving as the smallest independent unit for data representation.
[0029] Understandably, traditional geographic information system (GIS) data structures are often based on the concept of layers, storing roads, buildings, and water systems separately, lacking logical connections between objects. In contrast, the "atomic spatial entity" in this embodiment refers to the smallest data unit with independent semantics, geometry, and attributes, such as an independent residential building, a road connecting two intersections, or a specific bus stop. The system uses a spatially intelligent large model that has been fine-tuned by instructions to perform end-to-end decoding of the input feature tensor. The model no longer outputs semantically meaningless pixel masks, but directly generates a stream of structured objects representing entities. For example, when the model identifies a building area in an image, it does not output a set of discrete triangular faces, but instantiates a "building" type entity object and assigns it a globally unique universal identifier. This process simulates how humans perceive the physical world, that is, abstracting specific "object" concepts from chaotic visual signals.
[0030] S03: Define probabilistic features for each atomized spatial entity. The probabilistic features include an existence confidence level that reflects the possibility that the entity actually exists in the physical world. The existence confidence level is quantized into a numerical range field between zero and one to accommodate the uncertain output in the large model generation process.
[0031] Due to noise, occlusion, and resolution limitations in observational data, the inference of the objective world by large models is essentially a probabilistic distribution process. Therefore, in the data structure design, each entity object is required to include a field called "existence confidence," which is a double-precision floating-point number with a value strictly limited to between zero and one, and the unit is a percentage normalized value. The specific value of this field is calculated by the normalized exponential function of the large model's output layer, reflecting the model's confidence in the actual existence of the entity in the current spatiotemporal environment. For example, when processing a blurry nighttime remote sensing image, the model may identify an object that appears to be a "temporary shed," assigning it a specific type and a 65% confidence level, while inferring another possibility as a "truck," with a 35% confidence level. This probabilistic expression allows the system to retain uncertainty information in subsequent processing, rather than prematurely performing binarization truncation, thus providing a mathematical basis for subsequent multi-source fusion and temporal correction. In addition, this probability field also supports Bayesian updates, and the probability value will dynamically converge as observational data accumulates.
[0032] S04: Use semantic primitives to construct the geometric shape of atomized spatial entities, abandon non-semantic mesh patches, and use parametric vector combinations or geometric parameter sets to describe the centroid coordinates, orientation angle vectors and spatial contours of the entities.
[0033] S05: Generate a spatial entity objectification representation model containing probabilistic features and semantic primitives as the structured foundation data for digital earth generation.
[0034] Understandably, in scenarios involving the generation of large models, storing mesh models containing tens of thousands of vertices using traditional general-purpose 3D model formats would consume enormous amounts of lexical resources, leading to rapid overflow of the context window. Therefore, this embodiment employs "parametric semantic primitives" to describe geometry. For buildings, the system defines basic geometric shapes such as "cuboids," "prisms," "cylinders," and "spheres" as primitives. A complex building is expressed as a Boolean combination of these primitives. For example, a house with a pointed roof is described as a combination of cuboids containing length, width, and height parameters and pyramids containing base and height parameters, and its spatial position and combination operation logic are recorded. For road entities, a "centerline control point array" plus a "road width function" is used for expression. This expression method compresses the amount of geometric data by several orders of magnitude, enabling large models to generate city models covering several square kilometers in a single interactive command. Furthermore, the generated models are inherently editable, completely solving the computational bottleneck and storage redundancy problems in the expression of 3D spatial data in generative artificial intelligence.
[0035] To address the challenges of parallel generation and logical continuity of global-scale spatial data under the context window constraints of large models, the design and definition methods include: implementing a normalized grid partitioning strategy. First, a global index grid system based on Mercator projection or spherical geometry is established, discretizing the Earth's surface into regular grid units. Considering the current limitations on the maximum number of terms per inference in mainstream generative large models and the peak memory usage, this embodiment sets the standard physical size of the grid to 300 meters by 300 meters. This size is chosen because a 300-meter-scale urban block typically contains a suitable number of major building entities and corresponding road network topology. The corresponding parameterized entity description text volume is precisely within the optimal performance range for large model inference. This avoids information loss or increased illusion caused by excessively large grids, as well as a surge in interface call overhead caused by excessively small grids. Each grid is assigned a unique hash index based on the space-filling curve, ensuring that spatially adjacent grids are as close as possible in storage address to optimize subsequent parallel reading efficiency.
[0036] Based on regular grid cells, the system generates local atomized spatial entities within each grid using a large model in parallel. A distributed task scheduling cluster is constructed, dividing the global region to be generated into independent task queues. Each computing node receives grid coordinates and the corresponding multimodal observation data slice, independently calling the large model inference instance. The input prompts for the large model include not only the current observation image but also the geocoding and neighborhood environment description of the grid, guiding the model to generate all entity objects within that grid. The standard output format is a structured object sequence, containing the identifier, type, probability value, and geometric parameters in the local coordinate system for all entities within the grid. This strategy allows the construction of the global digital earth to be horizontally scaled using cloud computing resources.
[0037] A logical connection and fusion mechanism for cross-domain entities is established. When linear or planar entities physically cross two or more grid boundaries, they are treated as truncated fragments by the large model in a single grid inference. To restore their integrity, the system designs a post-processing logical stitching process. This process first traverses all grid boundaries to identify all geometrically accessible entity fragments. Then, it extracts the "semantic primitive features" and attribute fingerprints at the breakpoints and uses a graph matching algorithm based on semantic similarity to calculate the probability that two breakpoints on adjacent grid boundaries belong to the same entity. For example, if there is a road of a specific level on the right boundary of grid A and a road of the same level on the corresponding position on the left boundary of grid B, and the angle between the geometric tangent directions is less than a preset threshold, the system determines that they are the same entity. Once the match is successful, the system performs a logical merging operation to generate a new global parent entity identifier, smoothly interpolates and connects the geometric parameters of the two fragments, and marks the original two local entities as child nodes of the parent entity, thereby eliminating the physical cracks caused by grid cutting at the logical level.
[0038] In this embodiment, the algorithm details of cross-mesh fusion are further refined. The large model's capabilities are utilized for soft connection processing. During the mesh splitting and independent generation stages, when the large model detects an entity extending to the mesh boundary, it is instructed to generate a special soft connection anchor point at the boundary intersection. This soft connection anchor point is not a data structure rich in semantic information, but contains an anchor point identifier, boundary intersection coordinates, a tangential vector indicating the extension direction, entity category, and cross-sectional feature fingerprint. These soft connection anchor points are stored in the metadata index at the mesh edge for rapid retrieval in subsequent steps. Upon entering the fusion stage, the system launches a dedicated boundary stitching inference engine to read all soft connection anchor points on the shared boundary between two adjacent meshes, constructing a bipartite graph matching problem. For two anchor points that are close in distance and whose tangential vector angle is less than a threshold, the system considers them as candidate matching pairs. Since geometric matching alone is often unreliable, the system introduces the semantic reasoning capability of a large model, inputting the semantic descriptions of the two candidate anchor points and the contextual information of their respective entities into the similarity model. When the matching score exceeds a set threshold, the system determines that the two are the same entity. At this point, a fusion operation is performed: First, the probabilistic features of both ends are fused, usually using Bayes' theorem to update the entity's existence confidence; second, the geometric parameters are updated, and the coordinates at the connection point are interpolated and corrected using a spline curve smoothing algorithm to eliminate small displacements or angular deviations caused by independent generation, ultimately generating a continuous and smooth entity object that spans the grid boundary.
[0039] In this embodiment, for linear entities such as roads, railways, pipelines, and rivers, the system abandons the traditional polyline or strip mesh representation method and instead adopts a parameterized definition method of centerline vector groups plus cross-sectional attributes. Specifically, the road entity data structure output by the large model includes an array of centerline control points, storing the three-dimensional coordinates of several key control points, and a set of parameters describing the curve interpolation method. At the same time, a width profile attribute is defined, which is a function description that varies along the centerline, such as the starting width, ending width, and gradient method, or a segmented description. In addition, for complex road entities, a lane structure field is also included, using string encoding to describe the lane composition. This representation method not only compresses the amount of data required to store a curved road from thousands of vertices to dozens of parameters, but more importantly, it preserves the topological skeleton and functional semantics of the road, allowing subsequent artificial intelligence navigation algorithms to directly read the number of lanes and width for capacity calculation without performing complex geometric analysis.
[0040] For solid entities such as buildings, bridge piers, and tree canopies, the system uses a simplified version of the concept of constructing solid geometry for fitting and representation. When a large model perceives a building, it decomposes it into a combination of several basic geometric shapes. The system predefines a geometric primitive library containing standard shapes such as cuboids, cylinders, ellipsoids, and prisms. For a typical residential building, the model might output the parameters of a main cuboid, including the origin coordinates, length, width, height, and rotation angle. For more complex structures, such as a commercial center with a podium, the model outputs a list of primitives and a combinational logic tree. The core advantage of this representation method is its semantic editability. When the building height needs to be adjusted, only the dimensional parameters need to be modified, without recalculating the coordinates of thousands of triangle vertices. At the same time, by storing the axis lengths, rotation angles, and centroid coordinates of the geometric shapes instead of discrete vertex data, the network transmission bandwidth and the video memory pressure of edge rendering are greatly reduced.
[0041] It is understandable that this embodiment also introduces an entity lifecycle dimension. In this embodiment, when defining the data structure of atomized spatial entities, in addition to spatial and semantic attributes, a time dimension field is specifically pre-defined, including establishment time, current state update time, and estimated disappearance time. These fields are not just simple timestamp records, but conclusions drawn from in-depth analysis of multimodal and multi-temporal observation data by a large model. In specific operation, the system receives satellite image sequences of the same area at different time points, and the large model compares the time series images to identify the change patterns of ground features. For example, if the model observes that a certain coordinate point is flat at the first time point, has a foundation at the second time point, and a building is topped out at the third time point, the model will write the establishment time as the second time point in the generated entity and set the probability field as a function that increases with time.
[0042] Understandably, based on the large model's understanding of urban planning documents, building life cycle patterns, and demolition signs in imagery, the model can make probabilistic predictions about the future state of entities. For example, when the model identifies demolition-related words on a building wall in the latest street view, or discovers heavy machinery features indicating demolition operations in satellite imagery, the large model will set the entity's predicted disappearance time to a recent time window and reduce its confidence in its presence in future time slices. This full life cycle representation makes Digital Earth a dynamic system with timeline backtracking and extrapolation capabilities, supporting advanced application scenarios such as urban evolution simulation and post-disaster damage assessment.
[0043] This application embodiment also constructs a micro-topological network and resolves multimodal conflicts. In this embodiment, entity topological links are defined, which is a logical graph structure independent of geometric data. When generating entities, the large model synchronously outputs relational triples between entities, such as "entity A connects to entity B" or "entity C is located inside entity D". The system uses a graph database to store these relationships, focusing on constructing a cross-domain micro-topological network to bridge the gap between the public and private domains. In traditional maps, these two are often not connected. In this embodiment, the large model is specially trained to identify connecting nodes, such as community gates and underground parking garage entrance gates. These connecting nodes are defined as special atomic entities with bridging properties. For example, a gate entity logically connects both the external municipal road entity and the internal community main road entity. In this way, the navigation algorithm can calculate the complete path from the city's main road to the community's underground parking garage, realizing a continuous topology at the micro-scale.
[0044] Furthermore, topological logic is used to resolve multimodal attribute conflicts. When a satellite image shows an area as an impermeable road surface, while a street view image shows that the area is blocked by a wall, an attribute conflict occurs. In this case, the system calls the entity topology link for logical verification. If the topology network shows that the road segment is closed at both ends and located inside a closed community entity, and the connection node is normally closed, the large model will determine that the road segment is an internal road rather than a public road based on the principle of logical consistency, and correct its semantic category or adjust its drivability confidence. That is, logical relationships have a higher weight than single visual features in conflict determination, thereby ensuring that the generated spatial data is logically self-consistent.
[0045] Furthermore, this embodiment establishes a parent-child entity hierarchy, defining multi-level container entities, such as a hierarchical chain from city to room. In the data structure, each entity contains a parent entity identifier and a list of child entity identifiers. This tree structure allows attribute inheritance. When the large model is generated, a macroscopic parent entity is generated first, and then child entities are generated within the constraints of the parent entity, thus ensuring the rationality of the spatial layout. A three-dimensional stacking relationship is defined to solve the problem of vertical spatial layering in complex urban environments. This embodiment introduces semantic stacking descriptions, defining logical predicates such as "above," "below," and "on the surface." For example, for an underground parking lot entity, in addition to its geometric coordinates being underground, it also explicitly includes a stacking relationship attribute with the above-ground buildings. This stacking relationship based on logical identifiers enables the large model to construct a clear vertical topology when processing multi-layered spaces. Even when the geometric coordinates slightly overlap due to measurement errors, the logical relationship can still ensure the correct parsing of the spatial structure and clearly distinguish between above-ground, ground-level, and underground entity systems.
[0046] It is understandable that in the process of multimodal fusion, visual modalities are often limited by lighting, occlusion, or camouflage, which may lead to incorrect identification. This application's embodiment designs a set of logical weighted correction mechanisms. Suppose that the visual model identifies a certain place as a body of water with high confidence, but the topology reasoning module finds that the area is a necessary passage connecting two busy intersections and belongs to the lower level entity of an overpass. At this time, the system detects a logical paradox. According to the preset rule base or the common sense reasoning ability of the large model, the weight of the logical modal is dynamically increased, the system automatically reduces the confidence based on visual generation, and corrects it to a wet road surface, thereby avoiding serious semantic errors. Furthermore, the connectivity rules of the topology network are used to fill the observation blind spots. In areas where satellite images are obscured by the shadows of tall buildings or cannot be covered by street view images, this embodiment uses the generation ability of the large model to perform logical completion. If the topology network shows that a road enters the blind spot and there is a corresponding road extending out on the other side of the blind spot, the large model will infer that there must be a connecting road segment in the blind spot. The system will automatically generate a deduced entity to complete the geometry of this missing road and assign it a special probabilistic feature identifier.
[0047] In this embodiment, although atomized spatial entities are highly abstract parameterized data, end users need to see realistic 3D scenes. Therefore, a visual attribute mapping field is defined in the entity data structure. This field does not directly store texture images, but rather stores texture semantic tags and material parameters. For example, for a building, the field content includes texture style, hue, and physical material parameters such as roughness and metallicity. These tags correspond to the pre-built procedural material library in the client rendering engine. The rendering process is as follows: First, based on the viewpoint distance and hardware performance, the entities in the scene are... The system employs a filtering mechanism based on existence confidence. A dynamic threshold is set, and entities with confidence levels below this threshold are removed from the rendering process to purify the scene and improve performance. Secondly, for entities that pass the filtering, the rendering engine reads their semantic primitive parameters and generates geometric meshes in real-time – this step is calculated instantaneously. Finally, based on the visual attribute mapping fields, the corresponding texture shader is called from the local material library and applied to the geometric surface. Through this decoupled "data-instruction-generation" model, the system can reconstruct a detailed and realistic digital earth scene on the terminal device with extremely low data transmission volume.
[0048] In this application embodiment, a spatial intelligent data generation system is also disclosed. Its physical architecture includes a distributed memory and a high-performance processor cluster. The memory adopts a hierarchical architecture: hot data is stored in a high-speed solid-state drive array to meet the throughput requirements of large models for massive tensor data; cold data is stored in a distributed object storage system. In addition, the memory also hosts a compiled computer program, including a multimodal coding module, a large model inference engine based on a converter architecture, a grid scheduling middleware, and a topology construction module. The processor part adopts a heterogeneous computing architecture, with a central processing unit cluster responsible for the execution of grid partitioning, task scheduling, logical edge joining, and post-processing fusion algorithms; a graphics processing unit cluster is specifically responsible for loading generative large models and executing deep neural network inference calculations. A high-speed interconnection network based on remote direct data access is deployed within the system to ensure low-latency synchronization of parameters between different computing nodes. By executing the above-mentioned computer program, the system can automatically extract information from the input raw observation data, and through a series of pipeline operations such as probabilistic inference, semantic construction, and grid stitching, finally output a standardized spatial intelligent data product containing probabilistic features and topological logic, realizing an intelligent mapping from the physical world to the digital twin world.
[0049] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0050] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0051] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0052] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0053] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0054] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.
Claims
1. A spatial intelligent data representation design and definition method highly applicable to large models, characterized in that, include: The system receives multimodal observation data as input, including satellite remote sensing images, ground-view street view images, and radar point cloud data. Based on the probabilistic inference mechanism of generative large models, the continuous space of the physical world is reconstructed into a discrete set of atomized spatial entities, which serve as the smallest independent unit for data representation. For each of the atomic spatial entities, a probabilistic feature is defined, which includes an existence confidence level reflecting the possibility that the entity actually exists in the physical world. The existence confidence level is quantized into a numerical range field between zero and one to accommodate the uncertain output in the large model generation process. The geometric shape of the atomized spatial entity is constructed using semantic primitives, abandoning non-semantic mesh patches, and using parameterized vector combinations or geometric parameter sets to describe the centroid coordinates, orientation angle vectors and spatial contours of the entity. A spatial entity objectification representation model containing the probabilistic features and semantic primitives is generated as the structured foundation data for digital earth generation.
2. The spatial intelligent data representation design and definition method for a highly applicable large-scale model according to claim 1, characterized in that, To address the challenges of parallel generation and logical continuity of global-scale spatial data under the constraint of large model context windows, the design and definition methods also include: Define a normalized grid partitioning strategy to divide the global space into regular grid units that are adapted to the memory and word constraints of single inference of large models; Based on the regular grid cells, the large model is triggered in parallel to generate local atomized spatial entities within each grid. Establish a logical connection and fusion mechanism for cross-domain entities. When linear or planar atomized spatial entities cross grid boundaries, logical stitching operations are performed by identifying semantic primitive features and attribute fingerprints at breakpoints. Entity fragments located in different grid cells are recombined into complete entities with globally unique identifiers, ensuring the spatial logical continuity across grids.
3. The spatial intelligent data representation design and definition method for a highly applicable large-scale model according to claim 2, characterized in that, The specific methods for constructing the geometric form of the atomized spatial entity using semantic primitives include: For linear entities, a parameterized primitive containing a centerline vector group and a width attribute is used for definition; For solid entities, a parameterized combination of one or more sets of cuboids and ellipsoids is used for fitting representation. By storing the axis length, rotation angle and combination logic of the geometry to replace vertex data, the amount of word consumption in data representation is reduced.
4. The spatial intelligent data representation design and definition method for a highly applicable large-scale model according to claim 2, characterized in that, The logical connection and fusion mechanism also includes: Soft connection anchors are generated at the grid boundaries, and the soft connection anchors carry the semantic category and geometric trend information of the entities; By leveraging the semantic reasoning capabilities of large models, soft connection anchor points of adjacent grids are matched and determined. When they are determined to be the same entity, the probabilistic features of both ends are fused and the geometric parameters of the entity are updated to eliminate physical cracks.
5. The spatial intelligent data representation design and definition method for a highly applicable large-scale model according to claim 1, characterized in that, The probabilistic features also include the definition of the entity's lifecycle dimension: Pre-defined fields for creation time, current status time, and estimated disappearance time in the data structure; Based on the timestamp differences of multimodal observation data, the dynamic evolution process of entities is inferred through a large model, and the probabilistic features are used to express the state confidence of entities in different time dimensions.
6. The spatial intelligent data representation design and definition method for a highly applicable large-scale model according to claim 1, characterized in that, To address the issues of missing microscopic connectivity and attribute conflicts in spatial entities under multimodal observation, the design and definition methods also include: Define entity topology links, which are used to describe the logical interconnection and spatial orientation relationship between different atomized spatial entities; Construct a cross-domain micro-topology network and clearly define the connection nodes between private domain entities and public domain entities. The connection nodes are a special type of atomized spatial entity used to bridge the internal road network of a closed area with the public road network, forming a continuous topological path that can be used for navigation reasoning. When observation data from different modalities cause attribute conflicts in the same region, the semantic category of the entity or the existence confidence is adjusted by utilizing the logical consistency constraints of the entity topology links.
7. The spatial intelligent data representation design and definition method for a highly applicable large-scale model according to claim 6, characterized in that, The entity topology links include hierarchical membership relationships and spatial inclusion relationships: Establish a parent-child entity hierarchy, define the functional area as the parent entity, define the facilities within the area as child entities, and the child entities inherit the spatial attributes of the parent entity; Define three-dimensional layering relationships, use spatial orientation semantics to describe the positional order of entities in the vertical direction, and distinguish above-ground, ground-level and underground entities through logical identifiers rather than pure geometric coordinates.
8. The spatial intelligent data representation design and definition method for a highly applicable large model according to claim 6, characterized in that, The specific inclusion of correcting the semantic category of the entity or adjusting the existence confidence includes: If the entity type identified by the visual modality data contradicts the spatial logic deduced from the entity topology link, then the existence confidence weight of that entity generated based on the visual modality is reduced. Based on the connectivity rules of the topological network, the missing entities in the observation blind zone are automatically inferred and filled using a large model, and assigned probabilistic feature labels based on inference.
9. The spatial intelligent data representation design and definition method for a highly applicable large-scale model according to any one of claims 1-8, characterized in that, Also includes: Define a visual attribute mapping field in the atomized spatial entity, the field being associated with the side facade texture features and physical material parameters; Entities are filtered based on the existence confidence level. Entities with probability values higher than a preset threshold are converted into 3D rendering instructions through their semantic primitives and visual attribute mapping fields, thereby realizing the structured reconstruction of the digital earth.
10. A spatial intelligent data generation system, characterized in that, include: Memory, used to store computer programs; A processor, used to execute the computer program to implement the spatial intelligent data representation design and definition method of the highly applicable large model as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Method for constructing large model for artificial general spatiotemporal intelligence
CA3265620A1
Entity-based network space map expression method, space-time entity visualization method and system
CN119537715A
Multi-source data processing system for geographic information big data
CN120353874A
Building three-dimensional model lightweight design method and system based on artificial intelligence
CN120429937A
3D modeling method based on digital twin cities
CN120612425A
Cited By
Multi-modal large model space sensing method based on 3D scene graph driving
CN121996993A