Method and device for constructing city-level 3D model
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-08-11
AI Technical Summary
[0009]本发明提供一种城市级3D模型构建方法和装置,用以解决现有技术中城市级3D模型的数据成本较高、可扩展性较低、生成的建筑几何形状视觉失真的问题
[0025]本发明还提供一种非暂态计算机可读存储介质,其上存储有计算机程序,该计算机程序被处理器执行时实现如上述任一种所述城市级3D模型构建方法。
Smart Images

Figure CN122550818A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a method and apparatus for constructing city-level 3D models. Background Technology
[0002] The generation of high-quality 3D (3D) worlds represents a key research frontier with profound strategic implications for immersive media, large-scale simulation, and the development of embodied intelligence world models. As highly complex mega-systems, modern cities possess extremely complex spatial layouts and high component heterogeneity, including diverse architectural forms, complex road network topologies, and a vast array of fine-grained surface elements. Therefore, the ability to simulate or reconstruct entire 3D cities is indispensable for applications with enormous industrial potential, such as urban planning simulation, autonomous driving, and embodied AI (Artificial Intelligence) algorithm training. However, the creation of complex city-level 3D scenes remains a highly manual, labor-intensive process, not only extremely costly and time-consuming but also a significant technical bottleneck for achieving larger-scale, higher-quality environmental simulations, urgently requiring the development of fully automated, high-quality solutions.
[0003] In existing technologies, 3D representation methods, visual generative models, and visual language models are generally used to construct city-level 3D scenes. However, when constructing city-level 3D scenes, the number of independent objects contained in the 3D scene increases exponentially, significantly increasing the representational difficulty of 3D representation methods such as NeRF (Neural Radiance Fields). Visual generative models are essentially pixel- or voxel-level renderers, unable to truly understand the spatial relationships or physical constraints in complex urban spatial layout instructions. This results in instructions guiding 3D generation often being limited to plain text, ignoring the richer context in multimodal inputs. While visual language models (VLMs) possess powerful multimodal perception and high-dimensional cognitive capabilities, they are limited by their autoregressive text output characteristics, unable to directly render continuous latent spaces or high-fidelity 3D meshes.
[0004] Therefore, due to the technical bottlenecks of the aforementioned general underlying technology model when facing city-level scale, the following defects exist when constructing city-level 3D models.
[0005] (1) Existing methods for generating urban geometry models are limited by technology and implementation. They often adopt simple rule-based methods, lack understanding of multimodal environmental information, and generate overly simple building geometry, resulting in visual distortion.
[0006] (2) The 3D representation-based method does not have the scalability for city scale.
[0007] (3) The method of relying on structured data sources is highly dependent on the corresponding data. Such data is difficult to obtain in general real city-level large-scale application scenarios, resulting in high construction costs.
[0008] Therefore, how to provide a cost-effective, high-fidelity, and scalable method for building city-level 3D models is a technical problem that urgently needs to be solved. Summary of the Invention
[0009] This invention provides a method and apparatus for constructing city-level 3D models, which solves the problems of high data cost, low scalability, and visual distortion of generated building geometry in existing city-level 3D models.
[0010] This invention provides a method for constructing a city-level 3D model, comprising the following steps.
[0011] Based on the geographic coordinates of the target area, determine the structured geospatial data and street view panoramic image corresponding to the target area; The structured geospatial data and the street view panoramic image are input into the visual generation model to obtain the initial two-dimensional building representation output by the visual generation model; the visual generation model is obtained by cognitive alignment training based on multimodal geographic sample data; the cognitive alignment is used to distill the multimodal semantic understanding ability of the visual language model into the visual generation model; A quality assessment is performed based on the initial two-dimensional architectural representation to obtain the target two-dimensional architectural representation. Generate textured 3D models corresponding to different elements in the target two-dimensional architectural representation; Based on the structured geospatial data, the texture 3D models corresponding to different elements are spatially assembled to obtain a city-level 3D model.
[0012] According to the city-level 3D model construction method provided by the present invention, the visual language model includes a first visual language model and a second visual language model; The visual generation model is trained based on the following steps: The multimodal geographic sample data is input into the first visual language model for cognitive distillation to obtain multimodal training data pairs; the multimodal training data pairs are used to characterize the correlation between geospatial geometric features, local geographic images and complete two-dimensional building representations; Based on the multimodal training data, the initial visual generation model is fine-tuned under supervision to obtain the first visual generation model; The intermediate image generated by the first visual generation model is input into the second visual language model for quality evaluation to obtain the scalar reward value corresponding to the intermediate image; the scalar reward value is used to characterize the degree of deviation between the potential pixel evolution trajectory corresponding to the intermediate image and the preset semantic constraints; Based on the scalar reward value, the first visual generation model is semantically aligned and trained to obtain the trained visual generation model.
[0013] According to the city-level 3D model construction method provided by the present invention, the visual language model further includes a third visual language model; The quality assessment based on the initial two-dimensional building representation to obtain the target two-dimensional building representation includes: Obtain the initial building image generated by the visual generation model based on the initial two-dimensional building representation; The initial building image is input into the third visual language model for quality assessment, resulting in a quality assessment result corresponding to the initial building image; the quality assessment result includes a quality score and image correction suggestions. If the quality score is greater than or equal to a first preset threshold, the architectural representation corresponding to the initial architectural image is determined as the target two-dimensional architectural representation. If the quality score is less than a first preset threshold, a correction prompt word corresponding to the visual generation model is determined based on the image correction opinion; the correction prompt word is used to instruct the visual generation model to generate a corrected building image; the iteration stops when the quality score of the corrected building image is greater than or equal to the first preset threshold, and the building representation corresponding to the finally obtained corrected building image is determined as the target two-dimensional building representation.
[0014] According to the city-level 3D model construction method provided by the present invention, the step of generating textured 3D models corresponding to different elements in the target two-dimensional architectural representation includes: The target two-dimensional architectural representation is input into a 3D large-scale model to obtain a textureless 3D mesh output by the 3D large-scale model. The textureless 3D mesh is input into a large texture synthesis model to obtain initial textured 3D models corresponding to different elements output by the large texture synthesis model; the large texture synthesis model is used to perform physically based rendering texture mapping on the textureless 3D mesh. By removing interfering artifacts from the initial texture 3D models corresponding to different elements, texture 3D models corresponding to different elements are obtained.
[0015] According to the city-level 3D model construction method provided by the present invention, the step of spatially assembling textured 3D models corresponding to different elements based on the structured geospatial data to obtain a city-level 3D model includes: Rendering is performed based on the structured geospatial data to obtain the initial 3D support corresponding to the target area; Based on the underlying map coordinates corresponding to different elements, the texture 3D models corresponding to different elements are integrated into the initial 3D support; Adjust the spatial transformation parameters of each texture 3D model in the initial 3D scaffold to obtain the initial city-level 3D model; Fine-grained city element models were retrieved from the 3D material library; The fine-grained city element model is placed in the initial city-level 3D model to obtain the city-level 3D model.
[0016] According to the city-level 3D model construction method provided by the present invention, the visual language model further includes a fourth visual language model; The step of placing the fine-grained city element model into the initial city-level 3D model to obtain the city-level 3D model includes: Based on a preset rule placement mechanism and a preset spatial location, the fine-grained urban element model is deployed in the initial city-level 3D model to obtain the city-level 3D model. The preset rule placement mechanism is used to determine the specific category spacing of the fine-grained urban element model deployed along the road boundaries in the initial city-level 3D model based on the road geometry and lane type in the structured geospatial data. The preset spatial location is obtained by the fourth visual language model performing semantic analysis on the street view panoramic image.
[0017] According to the city-level 3D model construction method provided by the present invention, the spatial transformation parameters include: scaling ratio and spatial orientation.
[0018] According to the city-level 3D model construction method provided by the present invention, the step of determining the structured geospatial data and street view panoramic image corresponding to the target area based on the geographic coordinates of the target area includes: Based on the geographic coordinates of the target area, the OSM API is invoked to determine the structured geospatial data corresponding to the target area; Based on the geographic coordinates of the target area, the map API is invoked to determine the corresponding street view panoramic image of the target area.
[0019] According to the city-level 3D model construction method provided by the present invention, the step of calling the map API to determine the street view panoramic image corresponding to the target area includes: Call the map API to determine the initial street view image corresponding to the target area; If there are interfering elements in the initial street view image, the initial street view image is input into the target detection model to remove the interfering elements, thereby obtaining the street view panoramic image corresponding to the target area. If there are no interfering elements in the initial street view image, the initial street view image is determined as the panoramic street view image corresponding to the target area.
[0020] According to the city-level 3D model construction method provided by the present invention, the method further includes: Based on the structured geospatial data, the city-level 3D model is integrated into a micro-traffic simulator for simulation, resulting in a virtual city ecological model.
[0021] According to the city-level 3D model construction method provided by the present invention, the quality score includes: semantic rationality score, structural integrity score, and aesthetic appearance score.
[0022] The present invention also provides a city-level 3D model building device, comprising the following modules.
[0023] The data perception module is used to determine the structured geospatial data and street view panoramic image corresponding to the target area based on the geographic coordinates of the target area. The cognitive alignment module is used to input the structured geospatial data and the street view panoramic image into the visual generation model to obtain the initial two-dimensional building representation output by the visual generation model; the visual generation model is obtained by cognitive alignment training based on multimodal geographic sample data; the cognitive alignment is used to distill the multimodal semantic understanding capability of the visual language model into the visual generation model; The quality assessment module is used to perform quality assessment based on the initial two-dimensional building representation to obtain the target two-dimensional building representation; The generation module is used to generate texture 3D models corresponding to different elements in the target two-dimensional architectural representation; The assembly module is used to spatially assemble the textured 3D models corresponding to different elements based on the structured geospatial data to obtain a city-level 3D model.
[0024] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the city-level 3D model construction method as described above.
[0025] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the city-level 3D model construction method as described above.
[0026] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the city-level 3D model construction method as described above.
[0027] The present invention provides a method and apparatus for constructing city-level 3D models. It extracts structured geospatial data and street view panoramic images corresponding to the target area using the geographical coordinates of the target area. The structured geospatial data and street view panoramic images are then input into a visual generation model after knowledge distillation using a visual language model, resulting in a cognitively aligned initial two-dimensional architectural representation. By evaluating the quality of this initial two-dimensional architectural representation, a target two-dimensional architectural representation is obtained. Textured 3D models corresponding to different elements in the target two-dimensional architectural representation are generated. Based on the structured geospatial data, the textured 3D models of different elements are spatially assembled to obtain a spatially and cognitively aligned city-level 3D model. In this invention, the use of open-source software APIs to obtain publicly available structured geospatial data and street view panoramic images corresponding to the target area effectively reduces the data acquisition threshold and construction cost for city-level 3D modeling. By training cognitive alignment using multimodal geographic sample data, the multimodal semantic understanding capabilities of the visual language model are distilled into the visual generative model. This model understands the geometric relationships and physical constraints within urban spatial layouts, generating structurally sound and visually realistic initial 2D architectural representations without human intervention. Furthermore, the quality of these initial 2D architectural representations is assessed, automatically filtering and optimizing the generated results to ensure high-fidelity alignment between the final target 2D architectural representation and the complex and ever-changing visual features of the real world. The generated textured 3D model is spatially assembled using structured geospatial data, ensuring that the final city-level 3D model achieves both spatial and cognitive alignment. This allows for flexible adaptation to 3D city modeling needs of varying scales and complexities, and also possesses excellent city-level scalability. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0029] Figure 1 This is a flowchart illustrating the city-level 3D model construction method provided in this embodiment of the invention.
[0030] Figure 2 This is a comparative diagram showing the model generation effects of different model building methods provided in the embodiments of the present invention.
[0031] Figure 3 This is a comparative diagram showing the architectural model generation effects of different model construction methods provided in the embodiments of the present invention.
[0032] Figure 4 This is a schematic diagram of the structure of the city-level 3D model building device provided in an embodiment of the present invention.
[0033] Figure 5 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0035] To address the problems of high data cost, low scalability, and visual distortion of generated building geometry in existing city-level 3D models, this invention provides a method for constructing city-level 3D models. Figure 1 This is a flowchart illustrating the city-level 3D model construction method provided in this embodiment of the invention, as shown below. Figure 1 As shown, the method includes steps 110 to 150.
[0036] Step 110: Based on the geographic coordinates of the target area, determine the structured geospatial data and street view panoramic image corresponding to the target area.
[0037] Specifically, the electronic device receives the geographic coordinates of the target area as input data and extracts structured geospatial data and street view panoramic images corresponding to the target area by calling multiple open-source software APIs (Application Programming Interfaces) in parallel. The structured geospatial data may include: building outline data, road network data, land use data, vegetation distribution data, and public service facility data. Specifically, the building outline data may include building base area and height information; the road network data may include road type, number of lanes, and traffic direction; the land use data may include the boundaries of parks, water bodies, and squares; the vegetation distribution data may include tree locations and green space areas; and the public service facility data may include the locations of streetlights, trash cans, and bus stops. The street view panoramic image can be a real-world image captured and stitched together by a street view acquisition device along the street. This image may include visual details of elements such as building facades, road surfaces, vegetation, traffic signs, pedestrians, and vehicles in the real world.
[0038] Optionally, the geographic coordinates are used to define the geographic extent of the target area, and the geographic coordinates may include at least one of the following: the central latitude and longitude of the target area, the latitude and longitude range of the target area, and the administrative division code.
[0039] Optionally, the structured geospatial data may include: building outline data, road network data, land use data, vegetation distribution data, and public service facility data. The building outline data may include information such as building base area and height. The road network data may include road type, number of lanes, and direction of travel. The land use data may include the area boundaries of parks, water bodies, and squares. The vegetation distribution data may include tree locations and green space areas. The public service facility data may include the locations of streetlights, trash cans, and bus stops. This embodiment of the invention does not limit these aspects.
[0040] Step 120: Input the structured geospatial data and the street view panoramic image into the visual generation model to obtain the initial two-dimensional building representation output by the visual generation model; the visual generation model is obtained by cognitive alignment training based on multimodal geographic sample data; the cognitive alignment is used to distill the multimodal semantic understanding ability of the visual language model into the visual generation model.
[0041] Specifically, before executing step 120, multimodal geographic sample data from different regions is acquired. This multimodal geographic sample data can be structured geospatial sample data and street view panoramic sample images extracted from different regions by calling multiple open-source software APIs. The initial visual generation model is then trained using this multimodal geographic sample data to achieve cognitive alignment, distilling the high-dimensional and abstract multimodal semantic understanding capabilities of the visual language model into the trained visual generation model.
[0042] While panoramic street view images provide rich information, they are limited by the viewing angle and cannot capture the complete three-dimensional space and volumetric features of buildings. Furthermore, transient obstacles often cause occlusion. Therefore, in this embodiment of the invention, inspired by the human cognitive ability to form a complete mental image from incomplete sensory data, after training the visual generation model, the spatial span, rough geometric structure, and geographical volume of buildings in structured geospatial data are used as strong prior cues. These, along with the panoramic street view images, are input into the trained visual generation model. This model then performs in-depth analysis, repair, and complete reconstruction of urban elements such as buildings in visual space. This results in an implicit alignment between the output initial two-dimensional architectural representation and the physical rules, geometric relationships, and semantic concepts understood by the visual language model. For example, the initial two-dimensional architectural representation contains rich spatial, structural, and textural information and conforms to rules such as buildings being perpendicular to the ground and windows being neatly arranged. This overcomes the cognitive blindness inherent in pure visual generation models and improves the visual fidelity of the initial two-dimensional architectural representation.
[0043] Step 130: Perform a quality assessment based on the initial two-dimensional building representation to obtain the target two-dimensional building representation.
[0044] Specifically, after determining the initial two-dimensional architectural representation, the visual language model is invoked to perform a multi-dimensional quality assessment of the architectural image corresponding to the initial two-dimensional architectural representation. If the quality assessment result does not meet the quality assessment requirements, the visual generation model is controlled to redraw the image until the redrawn image meets the quality assessment requirements or the maximum number of iterations is reached, at which point the iteration stops. The target two-dimensional architectural representation is then determined based on the final image, thus blocking the propagation of accumulated errors and achieving dynamic alignment of cognitive abilities without manual annotation.
[0045] Step 140: Generate texture 3D models corresponding to different elements in the target two-dimensional architectural representation.
[0046] Specifically, after determining the target two-dimensional building representation, 3D tools are called to generate textured 3D models that can be directly used by the computer graphics pipeline, and different elements correspond to a single textured 3D model. These elements can include different urban elements such as buildings, vegetation, roads, and water bodies.
[0047] Step 150: Based on the structured geospatial data, spatially assemble the texture 3D models corresponding to different elements to obtain a city-level 3D model.
[0048] Specifically, after determining the texture 3D model corresponding to different elements, the spatial layout of different elements in the real world is determined based on structured geospatial data, and they are precisely placed and assembled according to their respective spatial layouts to obtain a complete city-level 3D model.
[0049] Furthermore, before executing step 110, the electronic device receives a sequence of subtask instructions from the planning module. This sequence includes the execution order of instructions for each step from 110 to 150, the content of the subtask instructions, and the invocation of models and tools. The electronic device executes each subtask instruction in sequence to trigger subsequent steps 110 to 150. The planning module is a hardware device independent of the electronic device, but it is communicatively connected to the electronic device. After receiving the user's input task to build a city-level 3D model, the planning module invokes a large language model to decompose the task into five deeply collaborative sub-stages: Perception, Imagination, Reflection, 3D Generation, and Scene Design. Each sub-stage corresponds to its own sub-tasks: the Perception stage corresponds to step 110, the Imagination stage to step 120, the Reflection stage to step 130, the 3D Generation stage to step 140, and the Scene Design stage to step 150. This modular task decomposition architecture effectively decouples the inherently heterogeneous sub-tasks in the generation of complex 3D worlds, significantly enhancing the flexibility and controllability of the entire construction process.
[0050] The city-level 3D model construction method provided in this invention extracts structured geospatial data and street view panoramic images corresponding to the target area using the geographic coordinates of the target area. The structured geospatial data and street view panoramic images are then input into a visual generation model after knowledge distillation using a visual language model to obtain a cognitively aligned initial two-dimensional architectural representation. By evaluating the quality of the initial two-dimensional architectural representation, a target two-dimensional architectural representation is obtained. Texture 3D models corresponding to different elements in the target two-dimensional architectural representation are generated, and the texture 3D models of different elements are spatially assembled based on the structured geospatial data to obtain a spatially and cognitively aligned city-level 3D model. In this invention, the structured geospatial data and street view panoramic images corresponding to the target area are obtained by calling open-source software APIs, effectively reducing the data acquisition threshold and construction cost of city-level 3D modeling. By training cognitive alignment using multimodal geographic sample data, the multimodal semantic understanding capabilities of the visual language model are distilled into the visual generative model. This model understands the geometric relationships and physical constraints within urban spatial layouts, generating structurally sound and visually realistic initial 2D architectural representations without human intervention. Furthermore, the quality of these initial 2D architectural representations is assessed, automatically filtering and optimizing the generated results to ensure high-fidelity alignment between the final target 2D architectural representation and the complex and ever-changing visual features of the real world. The generated textured 3D model is spatially assembled using structured geospatial data, ensuring that the final city-level 3D model achieves both spatial and cognitive alignment. This allows for flexible adaptation to 3D city modeling needs of varying scales and complexities, and also possesses excellent city-level scalability.
[0051] In one embodiment, determining the structured geospatial data and street view panoramic image corresponding to the target area based on the geographic coordinates of the target area includes: Based on the geographic coordinates of the target area, the OSM API is invoked to determine the structured geospatial data corresponding to the target area; Based on the geographic coordinates of the target area, the map API is invoked to determine the corresponding street view panoramic image of the target area.
[0052] Specifically, after determining the geographic coordinates of the target area, the OSM (OpenStreetMap) API and online map API are invoked in parallel. OSM is an open-source geospatial database that provides structured geospatial data globally. Geographic data in OSM is constructed using three spatial elements: points, lines / or relationships, and attribute information is labeled with key-value pairs, resulting in a high degree of data structure. The electronic device extracts structured geospatial data from OSM, including the geometry, spatial span, and semantic labels of geographic elements within the target area, through API requests, facilitating subsequent steps. Online maps can include Baidu Maps, Gaode Maps, etc. The electronic device also obtains multi-view panoramic street view images of the target area through API requests, providing rich details for subsequent construction of real-world textures, facade details, and materials.
[0053] In this embodiment of the invention, data is obtained by calling open-source software APIs, which greatly reduces the threshold for data acquisition and the cost of building the system.
[0054] In one embodiment, calling the map API to determine the street view panoramic image corresponding to the target area includes: Call the map API to determine the initial street view image corresponding to the target area; If there are interfering elements in the initial street view image, the initial street view image is input into the target detection model to remove the interfering elements, thereby obtaining the street view panoramic image corresponding to the target area. If there are no interfering elements in the initial street view image, the initial street view image is determined as the panoramic street view image corresponding to the target area.
[0055] Specifically, to obtain a high-quality, unobstructed image source, after calling the map API to obtain the initial street view image corresponding to the target area, it is determined whether there are interfering elements such as irrelevant vegetation, construction sites, or vehicles in the initial street view image that hinder the subsequent image generation process. If so, the electronic device calls the object detection model to accurately segment building instances and other elements in the initial street view image, identify and extract the interfering elements, and input the processed initial street view image into the visual language model for image quality scoring. The processed initial street view image with the highest image quality score is determined as the high-clarity panoramic street view image. If no such elements are found, the initial street view image is determined as the panoramic street view image.
[0056] In one embodiment, the visual language model includes a first visual language model and a second visual language model; The visual generation model is trained based on the following steps: The multimodal geographic sample data is input into the first visual language model for cognitive distillation to obtain multimodal training data pairs; the multimodal training data pairs are used to characterize the correlation between geospatial geometric features, local geographic images and complete two-dimensional building representations; Based on the multimodal training data, the initial visual generation model is fine-tuned under supervision to obtain the first visual generation model; The intermediate image generated by the first visual generation model is input into the second visual language model for quality evaluation to obtain the scalar reward value corresponding to the intermediate image; the scalar reward value is used to characterize the degree of deviation between the potential pixel evolution trajectory corresponding to the intermediate image and the preset semantic constraints; Based on the scalar reward value, the first visual generation model is semantically aligned and trained to obtain the trained visual generation model.
[0057] Specifically, this invention provides a two-stage cognitive alignment training architecture, which aims to implicitly distill the high-dimensional multimodal understanding capabilities of the visual language model into the latent space of the visual generative model without changing the underlying network architecture of the initial visual model or forcing its output language logic. In this cognitive alignment training architecture, the first stage is a supervised fine-tuning-based static cognitive distillation, specifically including: using a first visual language model with strong perception and reasoning capabilities as a fully autonomous data purifier, automatically parsing the three-dimensional geometric contours in incomplete street view panoramic sample images and structured geospatial sample data collected from the real world, synthesizing a multimodal training data pair containing geospatial geometric features, local geographic images, and complete two-dimensional building representations. This multimodal training data pair is a high-quality structured data pair. The initial visual generative model is then supervised fine-tuned using this multimodal training data pair. After training, a first visual generative model is obtained. In the latent space of this first visual generative model, a static pixel mapping relationship from abstract geometric priors to concrete three-dimensional spatial features is initially established. The second stage involves deep semantic alignment based on AI-driven feedback reinforcement learning. Specifically, to prevent the final visual generation model from falling into rote memorization and overfitting of the training data distribution, the second visual language model is deployed as a reward model with multi-dimensional, multi-modal scoring capabilities, constructing a dynamic feedback loop for reinforcement learning. While the first visual generation model performs denoising evolution in the latent space, the second visual language model, based on physical laws and semantic logic, evaluates the quality of the intermediate image output by the first visual generation model in real time, outputting a scalar reward value corresponding to the intermediate image. If the scalar reward value is too low, it indicates that the latent pixel evolution trajectory corresponding to the intermediate image deviates from the high-dimensional preset semantic constraints; if the scalar reward value is high, it indicates that the latent pixel evolution trajectory approaches the high-dimensional preset semantic constraints. This scalar reward value is fed back to the first visual generation model, causing it to update its cross-attention weights and denoising trajectory. This dynamic gradient backpropagation mechanism forces the first visual generative model to strictly align its continuous denoised trajectory with the discrete semantic space defined by the visual language model without outputting any text parsing, until training is complete and a visual generative model is obtained.
[0058] It should be noted that while the first and second visual language models have the same model structure and similar parameters, their functions differ. The first visual language model acts as a data refiner, used to parse structured geospatial sample data and street view panoramic sample images to determine multimodal training data pairs representing geospatial geometric features and the correlation between local geographic images and complete two-dimensional building representations. The second visual language model, on the other hand, serves as a reward model with multi-dimensional, multimodal scoring capabilities, constructing a dynamic feedback loop of reinforcement learning to score the intermediate images generated by the first visual generation model.
[0059] In one embodiment, the visual language model further includes a third visual language model; The quality assessment based on the initial two-dimensional building representation to obtain the target two-dimensional building representation includes: Obtain the initial building image generated by the visual generation model based on the initial two-dimensional building representation; The initial building image is input into the third visual language model for quality assessment, resulting in a quality assessment result corresponding to the initial building image; the quality assessment result includes a quality score and image correction suggestions. If the quality score is greater than or equal to a first preset threshold, the architectural representation corresponding to the initial architectural image is determined as the target two-dimensional architectural representation. If the quality score is less than a first preset threshold, a correction prompt word corresponding to the visual generation model is determined based on the image correction opinion; the correction prompt word is used to instruct the visual generation model to generate a corrected building image; the iteration stops when the quality score of the corrected building image is greater than or equal to the first preset threshold, and the building representation corresponding to the finally obtained corrected building image is determined as the target two-dimensional building representation.
[0060] Specifically, after training the visual generation model, the electronic device inputs image generation constraints into the visual generation model. For example, the image generation constraints can be: the generated building must perfectly match the rough shape and footprint proportion of the 3D model extracted by OSM, while transferring realistic facade materials and window details from real street scene photos, and forcing the generation of a clean image isolated from the background from a drone perspective at a 45-degree angle.
[0061] The visual generation model, based on the aforementioned image generation constraints, transforms the initial two-dimensional architectural representation into a visualized initial architectural image. This initial architectural image is then input into a third visual language model, which is instructed to ignore general image ambiguity. According to a pre-defined quality assessment guideline, a multi-dimensional quality assessment is performed on the initial architectural image, yielding a quality assessment result. This quality assessment structure includes quality scores for each dimension and image correction suggestions. The quality score for each dimension is compared with a first pre-defined threshold for that dimension. If the quality score for any dimension is greater than or equal to the first pre-defined threshold, it indicates that the image quality for that dimension is acceptable. At this point, the final target two-dimensional architectural representation can be determined based on the initial architectural image. If the quality score for any dimension is less than the first preset threshold, it indicates that the image quality for that dimension is unqualified. In this case, image correction suggestions are embedded into the visual generation model's prompt word template to obtain correction prompt words. These prompt words are then input into the visual generation model, forcing it to enter an iterative refinement process. The model is compelled to regenerate the corrected architectural image based on the structural weaknesses and improvement guidelines indicated in the correction prompt words. The third visual language model then re-evaluates the quality of the corrected architectural image. If it passes the evaluation, the iteration stops; otherwise, the above steps are repeated, forcing the visual generation model to redraw the image. Iteration continues until the corrected architectural image meets the quality requirements or the maximum number of iterations is reached. The target two-dimensional architectural representation is then determined based on the final corrected architectural image. This target two-dimensional architectural representation, evaluated through the above iterations, contains rich structural and appearance information.
[0062] It should be noted that the first preset threshold is a threshold value for quality assessment standards of different dimensions, and the first preset threshold is different for different dimensions.
[0063] It should be noted that the quality score includes: semantic rationality score, structural integrity score, and aesthetic appearance score. For example, the semantic rationality score is used to characterize the overlap accuracy between the initial architectural image and the 3D footprint model in the structured geospatial data; the structural integrity score is used to characterize the physical structural soundness of the building in the initial architectural image; and the aesthetic appearance score is used to characterize the seamless integration of texture and reference style in the initial architectural image. This embodiment of the invention does not limit this aspect.
[0064] It should be noted that this third visual language model has the same model structure and similar parameters as the first and second visual language models, but its function differs. This third visual language model serves as an automated quality critic, used to conduct multi-dimensional reviews of images generated by the visual generative model from different evaluation dimensions.
[0065] In this embodiment of the invention, through quality assessment and iteration between the third visual language model and the visual generation model, a system-level self-correction closed loop without human intervention is formed, ultimately achieving dynamic alignment of cognitive abilities without manual annotation.
[0066] In one embodiment, generating textured 3D models corresponding to different elements in the target two-dimensional architectural representation includes: The target two-dimensional architectural representation is input into a 3D large-scale model to obtain a textureless 3D mesh output by the 3D large-scale model. The textureless 3D mesh is input into a large texture synthesis model to obtain initial textured 3D models corresponding to different elements output by the large texture synthesis model; the large texture synthesis model is used to perform physically based rendering texture mapping on the textureless 3D mesh. By removing interfering artifacts from the initial texture 3D models corresponding to different elements, texture 3D models corresponding to different elements are obtained.
[0067] Specifically, after determining the target 2D architectural representation, a large 3D composite model is invoked. This representation is input into the large 3D composite model to generate a textureless 3D mesh with correct topology. This textureless 3D mesh can be a white model mesh composed of vertex coordinates and connectivity relationships, and it does not contain any color or material information. The large 3D composite model can be a Hunyuan3D-DiT-v2.1 large model.
[0068] Next, the textureless 3D mesh is input into a large texture compositing model. This model then performs physically based texture mapping on different elements within the textureless 3D mesh, resulting in an initial textured 3D model with a realistic visual appearance for each element. This large texture compositing model can be a Hunyuan3D-Paint-v2.1 model.
[0069] Considering that the textured 3D models of various elements may contain interfering artifacts, electronic devices can input the initial textured 3D model into 3D manipulation software for operations such as moving, rotating, scaling, and spatial clipping. This precisely removes interfering artifacts such as redundant ground planes and geometrically irregular fragments, ultimately resulting in an extremely clean and high-fidelity textured 3D model. This 3D manipulation software can be Blender.
[0070] In one embodiment, the spatial assembly of textured 3D models corresponding to different elements based on the structured geospatial data to obtain a city-level 3D model includes: Rendering is performed based on the structured geospatial data to obtain the initial 3D support corresponding to the target area; Based on the underlying map coordinates corresponding to different elements, the texture 3D models corresponding to different elements are integrated into the initial 3D support; Adjust the spatial transformation parameters of each texture 3D model in the initial 3D scaffold to obtain the initial city-level 3D model; Fine-grained city element models were retrieved from the 3D material library; The fine-grained city element model is placed in the initial city-level 3D model to obtain the city-level 3D model.
[0071] Specifically, after determining the textured 3D models of different elements, a low-fidelity initial 3D scaffold is rendered based on building outlines, road lines, and vegetation distribution points in structured geospatial data. This initial 3D scaffold serves as the spatial reference frame for the final city-level 3D model, encompassing the placement area for all textured 3D models. Then, all textured 3D models are traversed, and based on the underlying map coordinates corresponding to different elements, the textured 3D models corresponding to each element are integrated into the corresponding area of the underlying map coordinates within the initial 3D scaffold. Repeating this process completes the initial integration of the textured 3D models. The spatial layout of the initially integrated textured 3D models is relatively coarse. At this point, the electronic device can calculate the spatial transformation parameters of each textured 3D model, which may include scaling and spatial orientation. The scaling ratio of each textured 3D model is calculated using an alignment algorithm; this scaling ratio is the ratio of the bounding volume of each textured 3D model to the target volume. The spatial orientation of each textured 3D model is determined by exhaustively enumerating rotation angles; this spatial orientation maximizes the overlap between the model's bottom surface and the actual building plots projected onto the ground. Based on the scaling ratio of each textured 3D model, all textured 3D models are uniformly scaled. The spatial orientation of each textured 3D model is then adjusted to determine its final spatial pose within the initial 3D scaffold, resulting in a preliminary assembled initial city-level 3D model. Next, to enrich scene details, fine-grained city element models corresponding to different scene types are retrieved from a pre-built 3D asset library. These models include streetlights, traffic signs, trash cans, and trees. These fine-grained city element models are then placed in their corresponding positions within the initial city-level 3D model, resulting in a spatially aligned city-level 3D model. This city-level 3D model is a static model.
[0072] It should be noted that the 3D asset library includes a large number of standardized and high-quality fine-grained city element models, which can be categorized and stored according to scene type. You can perform hierarchical searches by scene type and city element name to determine the fine-grained city element models to be placed in the initial city-level 3D model.
[0073] In one embodiment, the visual language model further includes a fourth visual language model; The step of placing the fine-grained city element model into the initial city-level 3D model to obtain the city-level 3D model includes: Based on a preset rule placement mechanism and a preset spatial location, the fine-grained urban element model is deployed in the initial city-level 3D model to obtain the city-level 3D model. The preset rule placement mechanism is used to determine the specific category spacing of the fine-grained urban element model deployed along the road boundaries in the initial city-level 3D model based on the road geometry and lane type in the structured geospatial data. The preset spatial location is obtained by the fourth visual language model performing semantic analysis on the street view panoramic image.
[0074] Specifically, after retrieving the fine-grained urban element models, a dual-track placement mechanism is used to place them in their corresponding positions within the initial city-level 3D model. This dual-track placement mechanism includes a preset rule placement mechanism and a model-assisted placement mechanism. The preset rule placement mechanism is suitable for placing fine-grained urban element models with clear engineering specifications and high repetition, such as those corresponding to urban elements like streetlights and trash cans. Based on the road geometry and lane type in the structured geospatial data, specific category spacing for placing the fine-grained urban element models is determined. The road geometry can include the curvature of straight segments and curved segments, and the lane type can include urban arterial roads and secondary arterial roads. After determining the specific category spacing, the models are aligned and placed along road boundaries in the initial city-level 3D model according to the specific category spacing. The model-assisted placement mechanism is suitable for placing fine-grained urban element models that require contextual semantic inference, such as those corresponding to urban elements like traffic lights or signs. A panoramic street view image is input into a fourth visual language model. Semantic analysis is performed using this model to determine the most likely pre-defined spatial location of a fine-grained urban element model at the intersection. This fine-grained urban element model is then placed at this pre-defined spatial location. Repeating this dual-track placement mechanism yields a city-level 3D model with added urban details.
[0075] It should be noted that the fourth visual language model has the same model structure and similar model parameters as the first visual language model, but the model functions are different.
[0076] In one embodiment, the method further includes: Based on the structured geospatial data, the city-level 3D model is integrated with a micro-traffic simulator to obtain a virtual city ecological model.
[0077] Specifically, after constructing a static city-level 3D model, a refined lane-level topology is obtained based on structured geospatial data, and road segments in the city-level 3D model are connected. The city-level 3D model is then integrated with a micro traffic simulator, and dynamic traffic flow of people and vehicles is dynamically injected to obtain a virtual city ecosystem model that is highly aligned in both space and semantics and can operate dynamically.
[0078] In this embodiment of the invention, the city-level 3D model constructed according to this invention is compared with the generation effect of existing model construction methods such as SGAM, CityDreamer, Syncity, CityCraft, and Urbanworld. Figure 2 This is a comparative diagram showing the model generation effects of different model building methods provided in the embodiments of the present invention, such as... Figure 2 As shown, the city-level 3D model generated in this embodiment of the invention has the highest fidelity to the real world. Figure 3 This is a comparative diagram showing the architectural model generation effects of different model building methods provided in the embodiments of the present invention, such as... Figure 3 As shown, the architectural model generated in this embodiment of the invention has the most realistic effect.
[0079] The city-level 3D model building device provided by the present invention is described below. The city-level 3D model building device described below and the city-level 3D model building method described above can be referred to in correspondence.
[0080] The present invention also provides a city-level 3D model building device. Figure 4 This is a schematic diagram of the structure of the city-level 3D model building device provided in an embodiment of the present invention, as shown below. Figure 4 As shown, the city-level 3D model building device 400 includes: a data perception module 410, a cognitive alignment module 420, a quality assessment module 430, a generation module 440, and an assembly module 450.
[0081] The data sensing module 410 is used to determine the structured geospatial data and street view panoramic image corresponding to the target area based on the geographic coordinates of the target area. The cognitive alignment module 420 is used to input the structured geospatial data and the street view panoramic image into the visual generation model to obtain the initial two-dimensional building representation output by the visual generation model; the visual generation model is obtained by cognitive alignment training based on multimodal geographic sample data; the cognitive alignment is used to distill the multimodal semantic understanding ability of the visual language model into the visual generation model; Quality assessment module 430 is used to perform quality assessment based on the initial two-dimensional building representation to obtain the target two-dimensional building representation; The generation module 440 is used to generate texture 3D models corresponding to different elements in the target two-dimensional architectural representation; Assembly module 450 is used to spatially assemble texture 3D models corresponding to different elements based on the structured geospatial data to obtain city-level 3D models.
[0082] The city-level 3D model building device provided in this invention extracts structured geospatial data and street view panoramic images corresponding to the target area using the geographic coordinates of the target area. The structured geospatial data and street view panoramic images are then input into a visual generation model after knowledge distillation using a visual language model to obtain a cognitively aligned initial two-dimensional architectural representation. By evaluating the quality of the initial two-dimensional architectural representation, a target two-dimensional architectural representation is obtained. Texture 3D models corresponding to different elements in the target two-dimensional architectural representation are generated, and the texture 3D models of different elements are spatially assembled based on the structured geospatial data to obtain a spatially and cognitively aligned city-level 3D model. In this invention, the structured geospatial data and street view panoramic images corresponding to the target area are obtained by calling open-source software APIs, effectively reducing the data acquisition threshold and construction cost of city-level 3D modeling. By training cognitive alignment using multimodal geographic sample data, the multimodal semantic understanding capabilities of the visual language model are distilled into the visual generative model. This model understands the geometric relationships and physical constraints within urban spatial layouts, generating structurally sound and visually realistic initial 2D architectural representations without human intervention. Furthermore, the quality of these initial 2D architectural representations is assessed, automatically filtering and optimizing the generated results to ensure high-fidelity alignment between the final target 2D architectural representation and the complex and ever-changing visual features of the real world. The generated textured 3D model is spatially assembled using structured geospatial data, ensuring that the final city-level 3D model achieves both spatial and cognitive alignment. This allows for flexible adaptation to 3D city modeling needs of varying scales and complexities, and also possesses excellent city-level scalability.
[0083] Optionally, the data sensing module 410 is specifically used for: Based on the geographic coordinates of the target area, the OSM API is invoked to determine the structured geospatial data corresponding to the target area; Based on the geographic coordinates of the target area, the map API is invoked to determine the corresponding street view panoramic image of the target area.
[0084] Optionally, the data sensing module 410 is specifically used for: Call the map API to determine the initial street view image corresponding to the target area; If there are interfering elements in the initial street view image, the initial street view image is input into the target detection model to remove the interfering elements, thereby obtaining the street view panoramic image corresponding to the target area. If there are no interfering elements in the initial street view image, the initial street view image is determined as the panoramic street view image corresponding to the target area.
[0085] Optionally, the visual language model includes a first visual language model and a second visual language model.
[0086] Optionally, the city-level 3D model building device 400 includes a training module, which is specifically used for: The multimodal geographic sample data is input into the first visual language model for cognitive distillation to obtain multimodal training data pairs; the multimodal training data pairs are used to characterize the correlation between geospatial geometric features, local geographic images and complete two-dimensional building representations; Based on the multimodal training data, the initial visual generation model is fine-tuned under supervision to obtain the first visual generation model; The intermediate image generated by the first visual generation model is input into the second visual language model for quality evaluation to obtain the scalar reward value corresponding to the intermediate image; the scalar reward value is used to characterize the degree of deviation between the potential pixel evolution trajectory corresponding to the intermediate image and the preset semantic constraints; Based on the scalar reward value, the first visual generation model is semantically aligned and trained to obtain the trained visual generation model.
[0087] Optionally, the visual language model may further include a third visual language model.
[0088] Optionally, the quality assessment module 430 is specifically used for: Obtain the initial building image generated by the visual generation model based on the initial two-dimensional building representation; The initial building image is input into the third visual language model for quality assessment, resulting in a quality assessment result corresponding to the initial building image; the quality assessment result includes a quality score and image correction suggestions. If the quality score is greater than or equal to a first preset threshold, the architectural representation corresponding to the initial architectural image is determined as the target two-dimensional architectural representation. If the quality score is less than a first preset threshold, a correction prompt word corresponding to the visual generation model is determined based on the image correction opinion; the correction prompt word is used to instruct the visual generation model to generate a corrected building image; the iteration stops when the quality score of the corrected building image is greater than or equal to the first preset threshold, and the building representation corresponding to the finally obtained corrected building image is determined as the target two-dimensional building representation.
[0089] Optionally, the quality score includes: semantic rationality score, structural integrity score, and aesthetic appearance score.
[0090] Optionally, the generation module 440 is specifically used for: The target two-dimensional architectural representation is input into a 3D large-scale model to obtain a textureless 3D mesh output by the 3D large-scale model. The textureless 3D mesh is input into a large texture synthesis model to obtain initial textured 3D models corresponding to different elements output by the large texture synthesis model; the large texture synthesis model is used to perform physically based rendering texture mapping on the textureless 3D mesh. By removing interfering artifacts from the initial texture 3D models corresponding to different elements, texture 3D models corresponding to different elements are obtained.
[0091] Optionally, the assembly module 450 is specifically used for: Rendering is performed based on the structured geospatial data to obtain the initial 3D support corresponding to the target area; Based on the underlying map coordinates corresponding to different elements, the texture 3D models corresponding to different elements are integrated into the initial 3D support; Adjust the spatial transformation parameters of each texture 3D model in the initial 3D scaffold to obtain the initial city-level 3D model; Fine-grained city element models were retrieved from the 3D material library; The fine-grained city element model is placed in the initial city-level 3D model to obtain the city-level 3D model.
[0092] Optionally, the visual language model may further include a fourth visual language model.
[0093] Optionally, the assembly module 450 is specifically used for: Based on a preset rule placement mechanism and a preset spatial location, the fine-grained urban element model is deployed in the initial city-level 3D model to obtain the city-level 3D model. The preset rule placement mechanism is used to determine the specific category spacing of the fine-grained urban element model deployed along the road boundaries in the initial city-level 3D model based on the road geometry and lane type in the structured geospatial data. The preset spatial location is obtained by the fourth visual language model performing semantic analysis on the street view panoramic image.
[0094] Optionally, the spatial transformation parameters include: scaling ratio and spatial orientation.
[0095] Optionally, the city-level 3D model building device 400 includes a simulation module, which is specifically used for: Based on the structured geospatial data, the city-level 3D model is integrated into a micro-traffic simulator for simulation, resulting in a virtual city ecological model.
[0096] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present invention, such as... Figure 5 As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communications bus 540. The processor 510 can call logical instructions in the memory 530 to execute a city-level 3D model construction method. This method includes: determining structured geospatial data and a street view panoramic image corresponding to the target area based on its geographic coordinates; inputting the structured geospatial data and the street view panoramic image into a visual generation model to obtain an initial two-dimensional building representation output by the visual generation model; the visual generation model is obtained through cognitive alignment training based on multimodal geographic sample data; the cognitive alignment is used to distill the multimodal semantic understanding capability of the visual language model into the visual generation model; performing quality assessment based on the initial two-dimensional building representation to obtain a target two-dimensional building representation; generating textured 3D models corresponding to different elements in the target two-dimensional building representation; and spatially assembling the textured 3D models corresponding to different elements based on the structured geospatial data to obtain a city-level 3D model.
[0097] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0098] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the city-level 3D model construction method provided by the above methods. The method includes: determining structured geospatial data and street view panoramic images corresponding to the target area based on the geographic coordinates of the target area; inputting the structured geospatial data and the street view panoramic images into a visual generation model to obtain an initial two-dimensional building representation output by the visual generation model; the visual generation model is obtained by cognitive alignment training based on multimodal geographic sample data; the cognitive alignment is used to distill the multimodal semantic understanding capability of the visual language model into the visual generation model; performing quality assessment based on the initial two-dimensional building representation to obtain a target two-dimensional building representation; generating texture 3D models corresponding to different elements in the target two-dimensional building representation; and spatially assembling the texture 3D models corresponding to different elements based on the structured geospatial data to obtain a city-level 3D model.
[0099] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a city-level 3D model construction method provided by the above methods. This method includes: determining structured geospatial data and a street view panoramic image corresponding to the target area based on the geographic coordinates of the target area; inputting the structured geospatial data and the street view panoramic image into a visual generation model to obtain an initial two-dimensional architectural representation output by the visual generation model; the visual generation model is obtained through cognitive alignment training based on multimodal geographic sample data; the cognitive alignment is used to distill the multimodal semantic understanding capability of a visual language model into the visual generation model; performing a quality assessment based on the initial two-dimensional architectural representation to obtain a target two-dimensional architectural representation; generating textured 3D models corresponding to different elements in the target two-dimensional architectural representation; and spatially assembling the textured 3D models corresponding to different elements based on the structured geospatial data to obtain a city-level 3D model.
[0100] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing an urban-level 3D model, characterized in that, include: Based on the geographic coordinates of the target area, determine the structured geospatial data and street view panoramic image corresponding to the target area; The structured geospatial data and the street view panoramic image are input into the visual generation model to obtain the initial two-dimensional building representation output by the visual generation model; the visual generation model is obtained by cognitive alignment training based on multimodal geographic sample data; The cognitive alignment is used to distill the multimodal semantic understanding capabilities of the visual language model into the visual generative model; A quality assessment is performed based on the initial two-dimensional architectural representation to obtain the target two-dimensional architectural representation. Generate textured 3D models corresponding to different elements in the target two-dimensional architectural representation; Based on the structured geospatial data, the texture 3D models corresponding to different elements are spatially assembled to obtain a city-level 3D model.
2. The urban level 3D model construction method of claim 1, wherein, The visual language model includes a first visual language model and a second visual language model; The visual generation model is trained based on the following steps: The multimodal geographic sample data is input into the first visual language model for cognitive distillation to obtain multimodal training data pairs; the multimodal training data pairs are used to characterize the correlation between geospatial geometric features, local geographic images and complete two-dimensional building representations; Based on the multimodal training data, the initial visual generation model is fine-tuned under supervision to obtain the first visual generation model; The intermediate image generated by the first visual generation model is input into the second visual language model for quality evaluation, and the scalar reward value corresponding to the intermediate image is obtained. The scalar reward value is used to characterize the degree of deviation between the potential pixel evolution trajectory corresponding to the intermediate state image and the preset semantic constraints. Based on the scalar reward value, the first visual generation model is semantically aligned and trained to obtain the trained visual generation model.
3. The urban level 3D model construction method of claim 1, wherein, The visual language model also includes a third visual language model; The quality assessment based on the initial two-dimensional building representation to obtain the target two-dimensional building representation includes: Obtain the initial building image generated by the visual generation model based on the initial two-dimensional building representation; The initial building image is input into the third visual language model for quality assessment, and the quality assessment result corresponding to the initial building image is obtained; the quality assessment result includes a quality score and image correction suggestions. If the quality score is greater than or equal to a first preset threshold, the architectural representation corresponding to the initial architectural image is determined as the target two-dimensional architectural representation. If the quality score is less than a first preset threshold, a correction prompt word corresponding to the visual generation model is determined based on the image correction opinion; the correction prompt word is used to instruct the visual generation model to generate a corrected building image; the iteration stops when the quality score of the corrected building image is greater than or equal to the first preset threshold, and the building representation corresponding to the finally obtained corrected building image is determined as the target two-dimensional building representation.
4. The city-scale 3D model construction method of claim 1, wherein, The process of generating textured 3D models corresponding to different elements in the target two-dimensional architectural representation includes: The target two-dimensional architectural representation is input into a 3D large-scale model to obtain a textureless 3D mesh output by the 3D large-scale model. The textureless 3D mesh is input into a large texture synthesis model to obtain initial textured 3D models corresponding to different elements output by the large texture synthesis model; the large texture synthesis model is used to perform physically based rendering texture mapping on the textureless 3D mesh. By removing interfering artifacts from the initial texture 3D models corresponding to different elements, texture 3D models corresponding to different elements are obtained.
5. The city-scale 3D model construction method of claim 1, wherein, The process of spatially assembling textured 3D models corresponding to different elements based on the structured geospatial data to obtain a city-level 3D model includes: Rendering is performed based on the structured geospatial data to obtain the initial 3D support corresponding to the target area; Based on the underlying map coordinates corresponding to different elements, the texture 3D models corresponding to different elements are integrated into the initial 3D support; Adjust the spatial transformation parameters of each texture 3D model in the initial 3D scaffold to obtain the initial city-level 3D model; Fine-grained city element models were retrieved from the 3D material library; The fine-grained city element model is placed in the initial city-level 3D model to obtain the city-level 3D model.
6. The city-level 3D model construction method according to claim 5, characterized in that, The visual language model also includes a fourth visual language model; The step of placing the fine-grained city element model into the initial city-level 3D model to obtain the city-level 3D model includes: Based on a preset rule placement mechanism and a preset spatial location, the fine-grained urban element model is deployed in the initial city-level 3D model to obtain the city-level 3D model. The preset rule placement mechanism is used to determine the specific category spacing of the fine-grained urban element model deployed along the road boundaries in the initial city-level 3D model based on the road geometry and lane type in the structured geospatial data. The preset spatial location is obtained by the fourth visual language model performing semantic analysis on the street view panoramic image.
7. The urban level 3D model construction method of claim 5, wherein, The spatial transformation parameters include: scaling ratio and spatial orientation.
8. The city-scale 3D model construction method of claim 1, wherein, The determination of structured geospatial data and street view panoramic images corresponding to the target area based on the geographic coordinates of the target area includes: Based on the geographic coordinates of the target area, the OSM API is invoked to determine the structured geospatial data corresponding to the target area; Based on the geographic coordinates of the target area, the map API is invoked to determine the corresponding panoramic street view image of the target area.
9. The urban level 3D model construction method of claim 8, wherein, The step of calling the map API to determine the street view panoramic image corresponding to the target area includes: Call the map API to determine the initial street view image corresponding to the target area; If there are interfering elements in the initial street view image, the initial street view image is input into the target detection model to remove the interfering elements, thereby obtaining the street view panoramic image corresponding to the target area. If there are no interfering elements in the initial street view image, the initial street view image is determined as the panoramic street view image corresponding to the target area.
10. The city-scale 3D model construction method of claim 1, wherein, The method further includes: Based on the structured geospatial data, the city-level 3D model is integrated into a micro-traffic simulator for simulation, resulting in a virtual city ecological model.
11. The urban level 3D model construction method of claim 3, wherein, The quality score includes: semantic rationality score, structural integrity score, and aesthetic appearance score.
12. An urban level 3D model construction apparatus, characterized by comprising: include: The data perception module is used to determine the structured geospatial data and street view panoramic image corresponding to the target area based on the geographic coordinates of the target area. The cognitive alignment module is used to input the structured geospatial data and the street view panoramic image into the visual generation model to obtain the initial two-dimensional building representation output by the visual generation model; the visual generation model is obtained by cognitive alignment training based on multimodal geographic sample data; The cognitive alignment is used to distill the multimodal semantic understanding capabilities of the visual language model into the visual generative model; The quality assessment module is used to perform quality assessment based on the initial two-dimensional building representation to obtain the target two-dimensional building representation; The generation module is used to generate texture 3D models corresponding to different elements in the target two-dimensional architectural representation; The assembly module is used to spatially assemble the textured 3D models corresponding to different elements based on the structured geospatial data to obtain a city-level 3D model.