Large Language Model-Driven Parametric 3D Modeling Method, System, Storage Medium and Program Product for Geographic Scenes

Through the large language modeling parametric three-dimensional modeling method of geographic scenarios, we build a progressive knowledge graph and perform implicit correlation reasoning, solving the problem of time-consuming and lack of semantics in 3D geographic scenario modeling, and achieving an efficient and intelligent three-dimensional modeling process.

CN119445016BActive Publication Date: 2025-07-01ARMOR ACADEMY OF CHINESE PEOPLES LIBERATION ARMY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411565925.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-05
Publication Date
2025-07-01
Estimated Expiration
2044-11-05

AI Technical Summary

Technical Problem

The existing 3D geographic scenario modeling methods are time-consuming in the data acquisition and processing process, and the generated models lack semantic information, affecting spatial query, analysis and intelligent decision-making; parameterized modeling has problems of degree of intelligence and low efficiency.

Method used

The parametric three-dimensional modeling method of geographic scenarios driven by large language models is adopted. By constructing a progressive three-dimensional scene modeling knowledge graph, combining the inference context screening mechanism of spatial autocorrelation and information entropy, the implicit correlation reasoning rules between entities and entities guided by thinking chains are used to update the knowledge graph, and the three-dimensional model generation method is generated through the primitive model and the three-dimensional model modeling object semantics.

Benefits of technology

The semantic richness and consistency of three-dimensional geographic scenario modeling is improved, making the generated models and real environment more realistic, improving the modeling efficiency and intelligence, and providing semantic information to support spatial query, analysis and intelligent decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119445016B_ABST
    Figure CN119445016B_ABST
Patent Text Reader

Abstract

The present invention discloses a large language model-driven parametric three-dimensional modeling method, system, storage medium and program product for geographical scenes, belonging to the field of geographic information science, and solves the problems that the geographical 3D scene reconstruction method is time-consuming in the data acquisition and processing process, and the generated models and scenes lack semantic information. The present invention constructs a progressive three-dimensional scene modeling knowledge graph for geographical scenes; constructs an inference context screening mechanism based on spatial autocorrelation and information entropy based on the geographical entity space and the progressive three-dimensional scene modeling knowledge graph to obtain prompt texts for LLMs inference; constructs an implicit association inference rule for entities and between entities guided by a chain of thought based on the prompt texts for LLMs inference to update the progressive three-dimensional scene modeling knowledge graph; performs parametric three-dimensional modeling of geographical scenes based on the three-dimensional model generation method based on primitive models and modeling object semantics and the updated progressive three-dimensional scene modeling knowledge graph. The present invention is used for parametric three-dimensional modeling of geographical scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] A large language model-driven parametric 3D modeling method, system, storage medium and program product for parametric 3D modeling of geographical scenes, belonging to the field of geographic information science. Background Art

[0002] Three-dimensional (3D) geographical scene modeling is an important research direction in the cross-field of geographic information science and computer graphics. Different from ordinary 3D scenes, 3D geographical scenes are digital abstractions and expressions of the real geographical environment, so they need to maintain spatial, semantic and temporal consistency with the real environment. 3D geographical scenes play a key role in applications such as digital twin cities, smart cities and virtual geographical environments, providing intuitive and interactive digital space decision-making support for fields such as urban planning, emergency management, environmental monitoring and navigation services (Biljecki et al., 2015; Dang et al., 2022). However, due to the complexity, diversity and dynamics of the geographical environment, there are still many technical challenges in constructing and maintaining large-scale realistic 3D geographical scene models (Ghamisi et al., 2019).

[0003] Currently, there are mainly two methods for 3D geographical scene modeling: 3D reconstruction method and parametric modeling. The 3D reconstruction method is a technology that directly constructs a three-dimensional model from data collected from the real world (such as images, point clouds). Common techniques include Structure from Motion (SfM) and Multi-View Stereo (MVS) based on images, and LiDAR reconstruction based on laser scanning. The main advantage of this method is that it can generate highly realistic and accurate models, especially suitable for capturing details of complex geographical environments, with high geometric accuracy and visual realism. However, 3D reconstruction also has some limitations. First, 3D reconstruction is time-consuming in the process of data acquisition and processing, especially for the reconstruction of large-scale scenes. More importantly, the generated models and scenes often lack semantic information. Semantic information is the core element of 3D geographical scenes and GIS. It not only endows geometric models with rich attributes and meanings, but also is the key foundation for realizing spatial queries, spatial analysis, and intelligent decision-making. Parametric modeling is a method of generating 3D models based on the abstraction and generalization of the real world by setting parameters and rules. The core of this method lies in transforming complex geographical entities and phenomena into a series of semantic parameters and rules. First, by analyzing the data characteristics of the real world, key semantic attributes and relationships are extracted; then, based on these semantic information, modeling parameters and rules are formulated. Common techniques include Grammar-based Modeling and Procedural Modeling, etc. These methods essentially encode semantic knowledge into the modeling process. The advantage of this method is that it can quickly generate large-scale complex scenes, can generate multiple variants by adjusting parameters, and the models themselves have rich semantic information. However, this method has limitations in accurately restoring scene details.

[0004] The fundamental reason for the above limitations is that parametric modeling attempts to describe complex and diverse real environments through limited and general parameters and rules, which makes parametric modeling have to balance between generality and particularity. For example, methods such as extracting modeling information from more data, increasing parameters and rules, and manual adjustment can make the 3D scene more detailed, but will lead to the loss of the original efficiency advantage of parametric modeling. Generally speaking, the key to this limitation lies in the balance between efficiency and accuracy. Without changing the basic paradigm of parametric modeling, improving the degree of intelligence is the most effective way to solve this limitation. Intelligent parametric modeling can not only more effectively integrate and analyze the data used for modeling, extract more valuable information, but also adaptively change the modeling strategy according to the particularity of the modeling object.

[0005] To sum up, the existing 3D geographical scene modeling methods have the following technical problems:

[0006] 1. The 3D reconstruction method is time-consuming in the data acquisition and processing process, and the generated models and scenes lack semantic information, which is not conducive to spatial query, spatial analysis, intelligent decision-making, etc.;

[0007] 2. There are problems with low intelligence level, efficiency or accuracy in parametric modeling. Summary of the Invention

[0008] The purpose of the present invention is to provide a large language model-driven parametric three-dimensional modeling method, system, storage medium and program product for geographical scenes, to solve the problems that the 3D reconstruction method is time-consuming in the data acquisition and processing process, and the generated models and scenes lack semantic information, which is not conducive to spatial query, spatial analysis, intelligent decision-making, etc.; and there are problems with low intelligence level, efficiency or accuracy in parametric modeling.

[0009] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0010] A large language model-driven parametric three-dimensional modeling method for geographical scenes, comprising the following steps:

[0011] Step 1, construct a progressive three-dimensional scene modeling knowledge graph of the geographical scene;

[0012] Step 2, construct an inference context screening mechanism based on spatial autocorrelation and information entropy based on the geographical entity space and the progressive three-dimensional scene modeling knowledge graph to obtain prompt texts for LLMs inference;

[0013] Step 3, update the progressive three-dimensional scene modeling knowledge graph based on the implicit association inference rules between entities guided by the chain of thought based on the prompt texts for LLMs inference;

[0014] Step 4, perform parametric three-dimensional modeling of the geographical scene based on the three-dimensional model generation method based on the primitive model and the semantics of the modeling object and the updated progressive three-dimensional scene modeling knowledge graph.

[0015] Further, the specific steps of the said Step 1 are:

[0016] Express and abstract the scene from three dimensions of geography, geometry, and rendering of the geographical scene to achieve the construction of a progressive knowledge graph from geographical semantics to rendering data; that is, through a semantic mapping mechanism, associate and transform different granularity semantics from the three dimensions of geography, geometry, and rendering of the geographical scene, and finally obtain an associated multi-dimensional semantic graph, that is, obtain a progressive three-dimensional scene modeling knowledge graph. Among them, the associated multi-dimensional semantic graph includes an explicitly associated geographical semantic knowledge graph, a geometric semantic knowledge graph, and a rendering semantic knowledge graph. The geographical knowledge graph is used to abstract and conceptualize geographical information in the real world. The geometric semantic knowledge graph is used to further refine the geographical semantics to support the specific geometric modeling process. The rendering semantic knowledge graph is used to further abstract on the basis of geometric semantics to support high-fidelity visual rendering. Explicit association means directly establishing a deterministic mapping between different-dimensional semantics through the entity connection of the knowledge graph.

[0017] Further, the specific steps of step 2 are as follows:

[0018] Taking the spatial position of the geographical entity to which the object to be inferred belongs as the center, calculate the spatial autocorrelation under the spatial distance threshold of the geographical entity using Moran's I index, and select the distance with the largest absolute value of the spatial autocorrelation as the final spatial reasoning range. Among them, the formula for Moran's I index is:

[0019]

[0020]

[0021] Among them, represents the value of Moran's I index, n represents the total number of geographical entity spatial objects whose spatial distance from the geographical entity of the object to be inferred is less than D k of the geographical entity spatial object, D k represents the kth geographical entity spatial distance threshold, x i and x j are the attribute values of the objects to be inferred i and j respectively, is the average value of the attribute values, w ij is the geographical entity spatial weight between the objects to be inferred i and j, d is the interval distance between geographical entities in space, is a non-zero positive integer;

[0022] Based on the obtained spatial reasoning range, taking the entities in the knowledge graph to which the object to be inferred belongs in the progressive three-dimensional scene modeling knowledge graph as the center, forming sub-networks with different hop counts, and using the number of relationships and the number of attributes existing in the network as the information richness, and then using information entropy to calculate the information richness of each sub-network. When the information entropy reaches the threshold H sWhen the number of hops is reached, the text descriptions of the parameters of all entities within the number of hops and the relationships between entities are used as the prompt text for the inference of LLMs, and are represented in the form of a graph as a local subgraph G q , the formula for calculating the information entropy within different hops when taking a certain entity as the center is:

[0023]

[0024] where H k (v) represents the information entropy of the sub-network composed of different hop counts when taking a certain entity as the center, and N h (v) represents the set of all nodes with a hop count of h when the entity in the knowledge graph to be inferred is used as the central node v. A(N h (v)) represents the set of attributes in the sub-network with the central node v and a hop count of h, and R(N h (v) represents the set of relationships in the sub-network with the central node v and a hop count of h. P(a i ) and P(r i ) are the probabilities of the attribute a i and the relationship r i in the sets A(N h (v)) and R(N h (v), respectively.

[0025] Furthermore, the specific steps of step 3 are as follows:

[0026] Step 3.1: Convert the local subgraph G q into a context description text C in natural language form. q , and use the Embedding technology and the external knowledge base K to convert the context description text C q into vectors, and then calculate the distance between the vectors through cosine similarity. Take the context description text corresponding to the vector with the closest distance as the external knowledge base K q . Finally, according to the given query entity e q and the query relationship r q retrieve the external knowledge base K q to obtain the target context description text to construct the thought chain prompt P. Among them, the thought chain prompt P includes query descriptions, knowledge descriptions, and inference step prompts. Among them, the external knowledge base K q is a knowledge base of text descriptions about the modeling scenario, including the query entity e q and the query relationship r q and the corresponding context description text C q ;

[0027] Step 3.2: Input the constructed thought chain prompt P into the pre-trained large language model for implicit association inference to generate the answer relationship ra and the new entity e a , add the new entity e a and the new relationship r a to the progressive 3D scene modeling knowledge graph G, and at the same time update the representations of entities and relationships, that is, the updated progressive 3D scene modeling knowledge graph is obtained.

[0028] A computer system includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method for large language model-driven parametric 3D modeling of geographical scenes.

[0029] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the method for large language model-driven parametric 3D modeling of geographical scenes are implemented.

[0030] A computer program product includes a computer program. When the computer program is executed by a processor, the steps of the method for large language model-driven parametric 3D modeling of geographical scenes are implemented.

[0031] Compared with the prior art, the advantages of the present invention are as follows:

[0032] First, the present invention proposes a progressive 3D scene modeling knowledge graph, which realizes a unified semantic expression from abstract to concrete by integrating semantic information in three dimensions of geography, geometry, and rendering. This method of multi-dimensional semantic fusion improves the semantic richness and consistency of 3D geographical scene modeling, making the 3D modeled scene more realistic compared to the real environment;

[0033] Second, the present invention designs an inference context screening mechanism based on spatial autocorrelation and information entropy, which compresses the original data submitted to the large language model within 128k, enabling the large language model to perform implicit associative reasoning in a larger-scale knowledge graph;

[0034] Third, the present invention introduces large language model reasoning based on chain-of-thought prompting. Combining with the constructed knowledge graph, the reasoning accuracy of the Qwen2-75b large language model in the modeling knowledge graph is increased from 82% to 94%;

[0035] Fourth, the 3D model generation method based on primitive models and the semantics of modeling objects proposed by the present invention combines the reasoning results of the large language model with parametric modeling technology, saving the time for manually constructing the knowledge graph and the time for 3D model modeling in traditional modeling, greatly improving the degree of intelligence and modeling efficiency, and the generated models and scenes have semantic information, which is conducive to spatial query, spatial analysis, intelligent decision-making, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0037] Figure 1 Schematic diagram of three knowledge graphs and relationship definitions in the present invention;

[0038] Figure 2 Schematic diagram of the screening of context entities and subgraphs in the present invention;

[0039] Figure 3 Schematic diagram of the Prompt template for implicit relationship reasoning in the present invention;

[0040] Figure 4 Schematic diagram of the Prompt template for new entity reasoning in the present invention;

[0041] Figure 5 Schematic diagram of the primitive model example in the present invention. Detailed implementation manners

[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0043] Construct a progressive three-dimensional scene modeling knowledge graph for the geographical scene:

[0044] 3D geographic scene modeling is the process of abstracting the geographical environment of the real world into a digital model that can be recognized and rendered by a computer. To achieve intelligent modeling of 3D geographic scenes, a knowledge graph construction with progressive multi-dimensional semantic fusion is proposed. The core is to formally express and abstract the scene from three dimensions: geographical semantics, geometric semantics, and rendering semantics, and through a semantic mapping mechanism, associate and transform different granularity semantics from the geographical, geometric, and rendering dimensions of the geographic scene, and finally obtain a related multi-dimensional semantic graph, that is, a progressive 3D scene modeling knowledge graph. Geographical semantics is used to abstract and conceptualize geographical information in the real world to form a geographical semantic knowledge graph. Then, geometric semantics further refines the geographical semantics to generate a geometric semantic knowledge graph to support the specific geometric modeling process. Finally, rendering semantics is further abstracted based on geometric semantics to form a rendering semantic knowledge graph to support high-fidelity visual rendering.

[0045] The geographical semantic knowledge graph focuses on the ontological definition and instantiation description of geographical entity types, attributes, and spatio-temporal relationships, etc., to form a structured semantic network reflecting the geographical world. For example, in the geographical semantic knowledge graph, a tree can be represented as a geographical entity, and its attributes include tree species, height, crown width, etc. The geometric semantic knowledge graph focuses on depicting the geometric elements, structural relationships, and constraint rules of the scene to form a semantic representation reflecting the geometric characteristics of the scene. In the geometric semantic knowledge graph, this tree can be represented as a geometric model, and its geometric elements include the tree trunk (cylinder) and the tree crown (sphere or cone), the structural relationship is that the tree trunk supports the tree crown, and the constraint rule is that the height of the tree crown does not exceed the total height of the tree. The rendering semantic knowledge graph focuses on the semantic organization of visual attributes such as the material, texture, and lighting of the scene to form a parametric representation supporting high-fidelity rendering. In the rendering semantic graph, the geometric model of the tree can be given the material of the bark, the texture of the leaves, and the lighting parameters are set to achieve a realistic rendering effect.

[0046] These three semantic knowledge graphs are both relatively independent and logically related. From geographical semantics to geometric semantics and then to rendering semantics, it reflects a hierarchical expression from abstract to concrete, from concept to instance. At the same time, based on the semantic mapping mechanism, any semantic knowledge graph can be associated with other graphs to achieve cross-dimensional semantic conversion and reasoning. The associations between multi-dimensional semantic graphs include two forms: explicit association and implicit association. Explicit association directly establishes a deterministic mapping between different-dimensional semantics through the entity connection of the knowledge graph, such as the one-to-one correspondence between geographical entities and geometric models. Implicit association is a potential uncertain association between different-dimensional semantics that may be established by intelligent reasoning mechanisms such as rule reasoning, path reasoning, or the large language model used in this case, such as the dependency relationship between geographical attributes and rendering parameters. Through the combination of explicit and implicit associations, multi-dimensional semantic graphs can co-evolve dynamically, continuously enriching and improving the semantic representation of the scene model. For example, given a geographical entity and its attributes, its geometric model can be directly generated through explicit association, and appropriate rendering parameters can be inferred through implicit association. Conversely, given a geometric model, its geographical entity type can also be determined through explicit association, and its geographical attributes can be inferred through implicit association. The interaction and cooperation between multi-dimensional semantic graphs enhance the semantic consistency and completeness of the scene model.

[0047] The most ideal situation is that the three semantic knowledge graphs can be converted into each other. Considering that the final visual effect is determined by the rendering semantics, the rendering semantics need to be strictly set according to different 3D engines and cannot be customized by users or developers. Each entity, attribute, and relationship should have a corresponding implementation method in the 3D engine (the present invention provides a new implementation method). On the contrary, the geographical semantic knowledge graph can be as rich and diverse as possible and is not restricted. A richer and larger geographical semantic knowledge graph helps LLMs capture more detailed details, making the modeling results more accurate. The geometric semantic knowledge graph is in between. The geometric semantic knowledge graph is a feature extraction of the geographical knowledge graph, and the rendering semantic knowledge graph is the implementation of the geometric semantic knowledge graph. This also means that the semantic knowledge graph is independent of the specific 3D engine, which facilitates the implementation of this method on different platforms and in different environments. The entity and relationship definitions of the three knowledge graphs are as Figure 1 shown.

[0048] In addition, for explicitly defined relationships, corresponding mapping methods need to be provided during the modeling process. For example, if the green attribute of grass in geographical semantics is connected to the ARGB value in rendering semantics through explicit definition, corresponding tools or methods need to be provided. The inference of implicit relationships is the rule update of the implicit association inference between entities guided by the chain of thought construction of the prompt text for inference based on LLMs introduced later in the progressive 3D scene modeling knowledge graph of geographical scenes. Therefore, the initial construction of the progressive 3D scene modeling knowledge graph of geographical scenes only has explicit associations, and implicit associations are added through subsequent mining.

[0049] Based on the geographical entity space and the progressive 3D scene modeling knowledge graph, a reasoning context screening mechanism based on spatial autocorrelation and information entropy is constructed to obtain the prompt text for LLMs inference:

[0050] When using a large language model (LLM) for knowledge inference, it is crucial to select appropriate context information. The quality and quantity of context information directly affect the accuracy and efficiency of inference. Excessive information will lead to redundancy and interference, making it difficult for the LLM to grasp the key points; while insufficient information may result in inaccurate or incomplete inference results. Therefore, a method is needed to balance the breadth and depth of context information to optimize the inference performance of the LLM. Considering that the object to be inferred has corresponding instances in the knowledge graph and also corresponding objects in the geographical entity space, this case evaluates whether other objects and their relationships need to be included in the inference context from two dimensions: the geographical entity space and the 3D scene modeling knowledge graph space.

[0051] Spatial autocorrelation refers to the correlation between the similarity of objects in space in a certain attribute and their spatial positions. Spatial autocorrelation follows the "first law of distance" (Tobler's first law of geography), that is, the closer the objects in space are, the more similar their attributes are; conversely, the farther the objects in space are, the greater the difference in their attributes. Spatial autocorrelation is used to measure the correlation degree between the object to be inferred and its surrounding objects in the selection of inference context. When performing 3D geographical scene modeling, the size range of the space is uncertain. To select the most suitable objects in the geographical entity space dimension as the inference basis, by calculating the spatial autocorrelation under different spatial distance thresholds with the object to be inferred (such as the RGB value of the leaves of a tree) as the center, the distance with the largest absolute value of spatial autocorrelation is selected as the final spatial inference range.

[0052] The Moran's I index is used to quantify spatial autocorrelation. For different distance thresholds Among them, the definition formula of the Moran's I index is:

[0053]

[0054] Among them, represents the Moran's l index value, n represents the total number of geographical entity spatial objects whose spatial distance from the geographical entity being inferred is less than D k , D k represents the k-th geographical entity spatial distance threshold, x i and x j are the attribute values of the inferred objects i and j respectively, is the average value of the attribute values, wij is the geographical entity spatial weight between the inferred objects i and j, d is the interval distance between geographical entities, is a non-zero positive integer.

[0055] The value range of the Moran's l index is [-1, 1], corresponding to perfect negative correlation, no correlation, and perfect positive correlation respectively. Considering that both negative and positive correlations may contribute to the reasoning of LLMs, we choose the D corresponding to the Moran's l index with the largest absolute value k as the best. By calculating the Moran's l index at different spatial distances, the spatial autocorrelation between the inferred object and the objects within different ranges around it can be quantified. However, this is only the correlation in geographical space, and there are also other related objects or attributes corresponding to these filtered objects in the knowledge graph, and these objects and attributes may also affect the inferred object. In the network structure of the knowledge graph, the concept of 'hop' also applies. The hop count in the knowledge graph refers to the number of edges required to reach other entities from the central entity. By different hop counts, starting from a certain central entity, sub-networks of different scales can be constructed. This gradually expanding process will include more and more entities and relationships, thus forming neighborhoods of different ranges. The information content of the networks composed of different neighborhoods is different. Considering that the information content of the entities in the knowledge graph is different, in order to balance this difference, taking a certain entity in the knowledge graph as the center, different hop counts are used to form sub-networks, and the number of relationships and attributes existing in the network is used as the information richness. The information entropy is used to calculate the information richness of each sub-network. The information entropy is an index that measures the uncertainty or amount of information, reflecting the degree of disorder or uncertainty of a random variable. As Figure 2 shown, taking a certain entity as the center, calculate the information entropy H k (v) of the sub-networks formed by different hops. When the information entropy reaches the threshold H s , the parameter texts of all entities within this hop count and the relationships between entities are used as the prompt text for the reasoning of LLMs, and are represented in the form of a graph as a local sub-graph G q . The formula for calculating the information entropy of different hops of a certain entity is:

[0056]

[0057] Among them, H k (v) represents the information entropy of the sub-network composed of different hop counts when centered on a certain entity, and N h (v) represents the set of all nodes with hop count h when the entity in the knowledge graph to be inferred is used as the central node v, and A(N h (v)) represents the set of attributes in the sub-network composed of the central node v and hop count h, and R(N h (v) represents the set of relationships in the sub-network composed of the central node v and hop count h, and P(a i ) and P(r i ) are the probabilities of attribute a i and relationship r i in the sets A(N h (v)) and R(N h (v) respectively. The complete context screening process is as Figure 2 shown.

[0058] D k 's value determines the sampling accuracy of the spatial distance. The smaller the D k value, the more reasonable the selected spatial objects are, but an overly large D k value is the opposite. An overly large H s value will result in a small advantage of reasonable context, and an overly small value will result in a reduction of useful information selected. Therefore, the setting of these two values depends on the requirements in practice and needs to be adjusted according to parameters such as the density of geographical objects, spatial scale, and the number of knowledge graph entities.

[0059] Updating the Progressive 3D Scene Modeling Knowledge Graph with Implicit Association Inference Rules Guided by the Chain of Thought in Prompt Text Construction for LLMs Inference: The Chain of Thought (CoT) is an emerging large language model inference paradigm that aims to improve the interpretability and accuracy of inference by decomposing complex inference tasks into a series of intermediate steps and guiding the model to perform step-by-step reasoning. Different from traditional end-to-end inference methods, CoT inference introduces a structured reasoning path, simulating the step-by-step thinking process of humans, enabling the model to more effectively utilize existing knowledge for inference. In this case, the content of the inference includes implicit relationship inference and new entity inference. Implicit relationship inference is based on the existing entities and relationships in the knowledge graph to infer potential relationships, and new entity inference is to insert new entities into the knowledge graph according to the external knowledge base. Specifically:

[0060] Convert the local subgraph G q into the context description text C in natural language formq , the context description text C is converted into a vector by using the Embedding technology and the external knowledge base K q , and then the distance between the vectors is calculated by cosine similarity, and the context description text corresponding to the vector with the closest distance is used as the external knowledge base K q , finally, according to the given query entity e q and the query relation r q the external knowledge base K is retrieved q to obtain the target context description text to construct the chain of thought prompt P. Among them, the chain of thought prompt P includes query description, knowledge description and reasoning step prompt. For example, when reasoning about the color of tree leaves, the query description is: "Please reason about the color of tree leaves in winter." The knowledge description is: "Tree leaves contain chlorophyll and can carry out photosynthesis. In autumn, the temperature drops and the sunlight decreases, and the tree stops producing chlorophyll." The reasoning step prompt is: Step 1: Analyze the main influencing factors of the color of tree leaves; Step 2: Consider the color of tree leaves in winter and the reasons; Step 3: Analyze the principle of the color change of tree leaves in winter; Step 4: Infer the color difference between the leaves of deciduous trees and evergreen trees in winter; Step 5: Summarize the RGB values of the colors in winter.

[0061] Among them, the external knowledge base K q is a knowledge base of text descriptions about the modeling scenario, including the query entity e q and the query relation r q and the corresponding context description text C q ;

[0062] The constructed chain of thought prompt P is input into the pre-trained large language model for implicit association reasoning to generate the answer relation r a and the new entity e a , the new entity e a and the new relation r a are added to the progressive three-dimensional scene modeling knowledge graph G, and at the same time, the representations of the entities and relations are updated, that is, the updated progressive three-dimensional scene modeling knowledge graph is obtained.

[0063] The new knowledge graph is to reason about the name, attributes and explicit relations of the new entity e a according to the external knowledge base K by guiding the pre-trained large language model LLM in multiple steps, and the implicit relation about e a , that is, the answer relation r a . And insert it into the progressive three-dimensional scene modeling knowledge graph of the constructed geographical scene, such as Figure 4 .

[0064] A 3D model generation method based on primitive models and the semantics of modeling objects and a parameterized 3D geographical scene modeling is carried out using the updated progressive 3D scene modeling knowledge graph:

[0065] In this case, all 3D models are realized by modifying the attributes of primitive models (cubes, irregular geometric bodies, roads, trees). A simple example is shown as Figure 5 follows.

[0066] The attributes of primitive models include: position, size, color, texture, etc. These attributes are also mapped to the geometric semantic knowledge graph and can also be associated with other nodes or attributes through explicit or implicit relationships. Some of the primitive models are complete 3D models imported externally, such as models composed of multiple components. By adding appropriate modification functions to this model, parameter control of different components can be achieved. While some primitive models do not rely on external modeling software and are directly generated by control functions, such as double-lane roads and irregular buildings. For double-lane roads, a series of 3D points are input, and roads are generated according to attributes such as width, length, and color. For irregular buildings, a series of 3D points are also input as the shape of the building, and then the appearance of the 3D building is constructed through floor height, number of floors, etc.

[0067] This case aims at 3D geographical scene modeling and proposes a parameterized 3D modeling method driven by a large language model. First, features are extracted from the perspectives of geographical entities, geometric information, and rendering engines, and three knowledge graphs are designed, and a progressive 3D scene modeling knowledge graph expression from geography to rendering is proposed. The semantics and parameters of modeling objects are managed through the progressive 3D scene modeling knowledge graph. Then, a large language model based on Chain-of-Thought (CoT) prompting is introduced to perform commonsense and multi-step reasoning on the knowledge graph. To avoid the problem of loss of reasoning focus caused by the context limitation of the large language model, a reasoning context screening mechanism based on spatial autocorrelation and information entropy is used to constrain the context of reasoning content. Finally, the method of controlling primitive models with parameters (i.e., the 3D model generation method of the semantics of modeling objects) is adopted to map the spatial knowledge obtained by reasoning into the geometric, texture, and semantic information of the 3D geographical scene model (i.e., the updated progressive 3D scene modeling knowledge graph), realizing the automated construction from knowledge to model. Compared with traditional methods, this case significantly improves the intelligence level and efficiency of the modeling process, reduces manual intervention, and has advantages in semantic completeness and realism.

Claims

1. A parameterized 3D modeling method for geographic scenes driven by a large language model, characterized in that: The steps include: Step 1: Construct a knowledge graph of progressive 3D scene modeling of geographic scenes; the specific steps are: The scene is expressed and abstracted from the three dimensions of geography, geometry and rendering of the geographic scene to realize the construction of a progressive knowledge graph from geographic semantics to rendering data; that is, through the semantic mapping mechanism, different granularity semantics are associated and converted from the three dimensions of geography, geometry and rendering of the geographic scene, and finally an associated multi-dimensional semantic graph is obtained, that is, a progressive three-dimensional scene modeling knowledge graph is obtained, wherein the associated multi-dimensional semantic graph includes an explicitly associated geographic semantic knowledge graph, a geometric semantic knowledge graph and a rendering semantic knowledge graph. The geographic knowledge graph is used to abstract and conceptualize geographic information in the real world, the geometric semantic knowledge graph is used to further refine the geographic semantics to support the specific geometric modeling process, and the rendering semantic knowledge graph is used to further abstract on the basis of geometric semantics to support high-realistic visual rendering. Explicit association refers to the direct establishment of deterministic mappings between semantics of different dimensions through the entity connection of the knowledge graph; Step 2: Based on the geographic entity space and progressive 3D scene modeling knowledge graph, a reasoning context screening mechanism based on spatial autocorrelation and information entropy is constructed to obtain prompt text for LLMs reasoning; Step 3: Based on the prompt text of LLMs reasoning, the implicit association reasoning rules between entities guided by the thinking chain are constructed to update the progressive 3D scene modeling knowledge graph; the specific steps are: Step 3.1: Sub-graph G q Context description text C converted into natural language form q , using Embedding technology and external knowledge base K to describe the context text C q Convert it into a vector, and then calculate the distance between the vectors by cosine similarity, and use the context description text corresponding to the vector with the closest distance as the external knowledge base K q ,Finally, according to the given query entity e q and query relation r q Retrieve external knowledge base K q Get the target context description text to construct a thought chain prompt P, where the thought chain prompt P includes query description, knowledge description and reasoning step prompts, where the external knowledge base K q is a knowledge base of textual descriptions of modeling scenarios, including query entities e q and query relation r q and the corresponding context description text C q ; Step 3.2: Input the constructed thought chain prompt P into the pre-trained large language model for implicit association reasoning to generate the answer relation r a and the new entity e a , the new entity e a and new relationships a Add it to the progressive 3D scene modeling knowledge graph G, and update the representation of entities and relationships at the same time, that is, obtain the updated progressive 3D scene modeling knowledge graph; Step 4: Perform parametric geographic scene 3D modeling based on the 3D model generation method based on the primitive model and modeling object semantics and the updated progressive 3D scene modeling knowledge graph.

2. The method for parameterized 3D modeling of geographic scenes driven by a large language model according to claim 1, characterized in that: The specific steps of step 2 are: Taking the spatial location of the geographic entity to which the inferred object belongs as the center, the Moran's I index is used to calculate the spatial autocorrelation under the spatial distance threshold of the geographic entity, and the distance with the largest absolute value of the spatial autocorrelation is selected as the final spatial reasoning range. The formula of the Moran's I index is: in, Represents the Moran's I index value, n means that the spatial distance from the geographic entity of the inferred object is less than D k The total number of geographic entity spatial objects, D k represents the k-th geographic entity spatial distance threshold, x i and x j are the attribute values ​​of the inferred objects i and j, is the average value of the attribute, w ij is the spatial weight of the geographic entity between the inferred objects i and j, d is the interval distance between the geographic entity spaces, is a non-zero positive integer; Based on the obtained spatial reasoning scope, the entity in the knowledge graph to which the inferred object belongs in the progressive 3D scene modeling knowledge graph is taken as the center, and sub-networks are formed with different hop numbers. The number of relationships and attributes in the network is used as the information richness. Then, the information entropy is used to calculate the information richness of each sub-network. When the information entropy reaches the threshold H s When the text description of the parameters of all entities within the hop number and the relationship between entities is used as the prompt text of LLMs reasoning, it is represented as a local subgraph G in the form of a graph. q , the formula for calculating the information entropy in different hops with a certain entity as the center is: Among them, H k (v) represents the information entropy of sub-networks with different hop numbers when a certain entity is the center, N h (v) represents the set of all nodes with a hop count of h when the entity in the knowledge graph to which the object to be inferred belongs is the central node v. A(N h (v)) represents the attribute set in the sub-network with a central node of v and a hop count of h. h (v) represents the set of relations in a subnetwork with a central node of v and a hop count of h. i ) and P(r i ) are attributes a i and the relationship i In the set A(N h (v)) and R(N h The probability in (v).

3. A computer system comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 2.

4. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 2 are implemented.

5. A computer program product, comprising a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 2 are implemented.