Urban planning design method based on live-action three-dimensional simulation system

By integrating multi-source data and using deep learning models, we have achieved deep integration of 3D simulation, GIS analysis and VR interaction in urban planning and design. This solves the problem of insufficient integration in existing technologies, improves planning accuracy and response efficiency, and supports real-time iterative optimization and user participation.

CN121328124APending Publication Date: 2026-01-13QINGDAO URBAN PLANNING & DESIGN INST

Patent Information

Application Number
CN202511497629.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

In existing urban planning and design methods, the integration of 3D simulation, GIS data analysis, and VR interactive technology is insufficient, resulting in large deviations in planning accuracy and high response delays, making it impossible to achieve real-time dynamic response and comprehensive optimization.

Method used

By collecting multi-source urban data, using deep learning models for cross-modal feature extraction and alignment, generating fused feature vectors, constructing a dynamic 3D urban model, embedding a GIS spatial analysis module, injecting VR interactive logic, simulating user behavior feedback, optimizing planning parameters, realizing multi-technology collaborative decision-making, and outputting a visual report.

Benefits of technology

It achieves deep integration of 3D simulation, GIS spatial analysis and VR interaction, shortens response latency to the millisecond level, improves planning accuracy and public participation, optimizes the iterative decision-making process, reduces the risk of planning disputes, and enhances visualization intuitiveness and scheme adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121328124A_ABST
    Figure CN121328124A_ABST
Patent Text Reader

Abstract

The invention discloses an urban planning design method based on a live-action three-dimensional simulation system, relates to the technical field of urban planning, and is used for solving the problem of insufficient technology fusion depth in an existing method. Firstly, an initial data set is formed by collecting multi-source city data; performing cross-modal feature extraction and alignment based on a deep learning model to generate a fusion feature vector; constructing a dynamic three-dimensional city model by using the fusion feature vector, and embedding a GIS spatial analysis module to perform preliminary layout simulation; vR interaction logic is injected into the model, user behavior feedback is simulated, and environment influence indexes are calculated; according to simulation results and index feedback optimization planning parameters, multi-technology collaborative decision making is achieved; outputting an optimization scheme and generating a visual report, and supporting real-time iterative updating; according to the method, the technology fusion depth is remarkably improved, the response delay is shortened to millisecond level, the planning precision is improved, the environmental influence evaluation deviation is reduced, and efficient and sustainable urban planning practice is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of urban planning, in particular to a city planning design method based on a real scene three-dimensional simulation system. BACKGROUND

[0002] With the acceleration of urbanization and the deepening of smart city construction, real scene three-dimensional simulation technology is increasingly widely used in urban planning and design, and is an important tool for the sustainable development of modern cities. Through the integration of geographic information systems, three-dimensional simulation technology and virtual reality, comprehensive evaluation and optimization of urban spatial layout, functional zoning and environmental impact are achieved. A three-dimensional modeling method based on real scene data is usually used to digitally reconstruct urban areas. High-precision models are generated through aerial image acquisition and laser scanning, and planning software is used to simulate building layout, traffic flow and landscape effects. This method improves the visualization and interactivity of planning and design to some extent, especially for preliminary scheme evaluation of large-scale urban projects, which can reduce errors in manual mapping and support multi-scheme comparison. However, existing urban planning and design methods often treat three-dimensional simulation, GIS data analysis and VR interaction as independent modules, with shallow integration, limited to data superimposition or simple interface calls, resulting in insufficient depth interaction between different technologies and inability to achieve real-time dynamic response and comprehensive optimization. For example, in patent document CN117974912A (City planning real scene three-dimensional simulation system), although a modeling system based on deep learning is proposed, the scheme only achieves static data exchange when integrating GIS programmatic path optimization and VR light pollution simulation, lacks deep algorithm integration, and causes large planning precision deviations. Similarly, patent document WO2020192027A1 (Embedded city design scene simulation method) uses embedded simulation, but does not solve the problem of deep integration of three-dimensional models and real-time environmental data, resulting in low efficiency and high response delay in multi-modal data processing. These defects not only reduce the accuracy and reliability of planning and design, but also may cause environmental assessment distortion and resource waste, especially in high-density urban areas with dynamic changes. The above-mentioned defects of existing technologies result in low efficiency and simulation results that deviate from reality in the urban planning and design process, and an improved method that can improve the depth of technology integration is urgently needed to achieve more accurate and efficient real scene three-dimensional simulation and planning optimization.

[0003] A set of pre-set compliance thresholds are compared with any constraint parameters; Generate event log entries containing object identification, interaction time, constraint items and pose, and write back the scene state.

[0004] In a preferred embodiment, it includes: The environmental indicators include light and glare distribution, noise intensity estimation and flood sensitivity. Preset a second threshold and compare it with the target area ratio or peak value of any indicator; Automatically generate a candidate set of directional adjustments for at least one of the parameters: reflectivity, building height, building spacing, or land density.

[0005] In a preferred embodiment, based on simulation results and indicator feedback, planning parameters are optimized, and the layout scheme is adjusted through an iterative algorithm to achieve multi-technology collaborative decision-making, including: Select several parameters to be adjusted according to the priority queue; perform single-round fine-tuning with a fixed micro-step size and recalculate the core indicators in real time; preset the convergence threshold or any hard constraint and compare it with the comprehensive score; maintain the Pareto non-dominated solution set throughout the iteration process and retain candidate solutions that pass the constraint consistency check.

[0006] In a preferred embodiment, the optimized urban planning design scheme is output and a visualization report is generated, supporting real-time iterative updates, including: verifying each strong constraint clause on building spacing, sunshine duration, ventilation assessment, ecological green space ratio and traffic accessibility; checking the version consistency of the 3D model, spatial analysis results and interactive modification records; and generating standardized 3D model files and a list of structured indicators.

[0007] In a preferred embodiment, the report generation employs template-driven feature assembly and incremental updates, including: Use a template engine to arrange text descriptions, indicator charts, and multi-view scene screenshots into a page format; Embed the scheme number, data snapshot fingerprint, and camera pose in the report metadata; Summary of the Invention

[0008] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides an urban planning and design method based on a real-scene 3D simulation system, thereby solving the problem of insufficient technological integration in existing methods.

[0009] (II) Technical Solution To achieve the goal of addressing insufficient depth of technology integration mentioned in the background section, the present invention provides the following technical solution: A method for urban planning and design based on a real-scene 3D simulation system includes: S1. Collect multi-source urban data, including real-scene 3D images, GIS geographic information and VR interactive feedback data, to form an initial dataset; S2. Based on a deep learning model, cross-modal feature extraction and alignment are performed on the initial dataset to generate a fused feature vector, thereby achieving deep integration of the data layer. S3. Construct a dynamic three-dimensional city model using the fused feature vector and embed a GIS spatial analysis module for preliminary layout simulation. S4. Inject VR interaction logic into the dynamic three-dimensional model, simulate user behavior feedback, and calculate environmental impact indicators. S5. According to the simulation results and indicator feedback, optimize the planning parameters, adjust the layout scheme through iterative algorithms, and realize multi-technology collaborative decision-making. S6. Output the optimized urban planning and design scheme, and generate a visual report to support real-time iterative updates.

[0010] In a preferred embodiment, multi-source city data is collected, including real three-dimensional images, GIS geographic information, and VR interaction feedback data, forming an initial data set, including: Configure the sampling period and timestamp for real three-dimensional images, GIS geographic information, and VR interaction feedback respectively; Process the data source, determine the trigger for mirror source switching and differential re-collection; Perform incremental re-collection on spatial fragments that have been verified and confirmed as changed by metadata; Establish a one-to-one mapping relationship based on data source identification and spatial coordinates and generate a traceable index table and collection log.

[0011] In a preferred embodiment, cross-modal feature extraction and alignment are performed on the initial data set based on a deep learning model to generate a fused feature vector, realizing deep integration at the data layer, including: Deep integration of the initial data set using cross-modal feature alignment mechanism; Extract sub-vectors from different modalities, normalize them, and then perform weighted fusion according to learnable weights; Configure learnable weight coefficients for each sub-vector; Pre-set alignment threshold and compare with cross-modal similarity; The fused feature vector is accompanied by three types of meta-information: modal weight, unified time label, and spatial range.

[0012] In a preferred embodiment, a dynamic three-dimensional city model is constructed using the fused feature vector, and a GIS spatial analysis module is embedded for preliminary layout simulation, including: The construction of the dynamic three-dimensional city model uses multi-level detail (LOD) management and maintains grid topological relationships; Trigger LOD level down when the calculation or rendering resource occupancy reaches the preset resource threshold; Prioritize simplifying detailed triangular facets while maintaining building boundary outlines and road network topology unchanged; Embed the GIS spatial analysis module into the scene pipeline in an interface manner; A closed loop of local reconstruction and spatial analysis on the same data snapshot is performed according to the chain of trigger - writeback - retrigger.

[0013] In a preferred embodiment, VR interaction logic is injected into the dynamic 3D model, simulates user behavior feedback, and calculates environmental impact indicators, including: VR interaction employs ray hit detection and constrained drag-and-drop logic; Slope, land use constraints, and boundary buffer space attributes are read for candidate object placement locations; A set of compliance thresholds is preset and compared with any constraint parameter; Event log entries containing object identification, interaction time, constraint items, and pose are generated and written back to the scene state.

[0014] In a preferred embodiment, including: Environmental indicators include lighting and glare distribution, noise intensity estimation, and flood sensitivity; A second threshold is preset and compared with the target area proportion or peak value of any indicator; Automatically generate a candidate set of directional adjustments for at least one of the parameters of reflectivity, building height, building spacing, or land use density.

[0015] In a preferred embodiment, according to the simulation results and indicator feedback, optimize the planning parameters, adjust the layout scheme through an iterative algorithm, and realize multi-technology collaborative decision-making, including: Select a number of parameters to be adjusted according to the priority queue; Single round fine-tuning is performed with a fixed micro-step size, and the core indicators are recalculated immediately; A convergence threshold or any hard constraint is preset and compared with the comprehensive score; Maintain a set of Pareto non-dominated solutions throughout the entire iteration process, and retain candidate schemes that pass the constraint consistency check.

[0016] In a preferred embodiment, output the optimized urban planning and design scheme, and generate a visual report to support real-time iterative updates, including: Verify each item of the strong constraint clauses of building spacing, sunshine duration, ventilation evaluation, ecological green land proportion, and traffic accessibility; Perform version consistency checks on the 3D model, spatial analysis results, and interaction modification records; Generate standardized 3D model files and structured indicator lists.

[0017] In a preferred embodiment, report generation employs template-driven element assembly and incremental updates, including: Use a template engine to page layout text descriptions, indicator charts, and multi-perspective scene screenshots; Embedding scheme number, data snapshot fingerprint and camera pose in report metadata; Subsequent fine-tuning of the same scheme only generates incremental pages and merges into the original report, while outputting a machine-readable parameter tree file.

[0018] In a preferred embodiment, comprising: Coordinate precision downgrading and regional desensitization of external shared space positions; Implementing anonymization and identification hashing on user feedback; Enable end-to-end encryption at transmission and storage stages, and maintain tamper-proof iterative update logs.

[0019] Compared with the prior art, the present application provides a city planning design method based on a real scene three-dimensional simulation system, which has the following beneficial effects: 1. The present application uses a deep learning model to generate a fusion vector dimension of 512 through cross-modal feature extraction and alignment, realizes deep integration at the data layer, converts heterogeneous data such as visual details of real scene images, spatial attributes of GIS, and user preferences of VR into a unified representation, realizes deep integration of three-dimensional simulation, GIS spatial analysis, and VR interaction, shortens the response delay from seconds to milliseconds, improves the planning accuracy, and in complex urban scenes, the fusion vector supports real-time semantic consistency, solves the deviation problem of existing independent modules, can fuse multiple source data in real time, avoids data silos, significantly improves the depth of technology fusion and planning accuracy, and solves the problem of insufficient interaction caused by shallow fusion.

[0020] 2. The present application optimizes parameters according to simulation results and index feedback, adjusts the layout through a feedback loop, realizes multi-technology collaborative decision-making, compared with existing manual adjustment, shortens the iteration time, improves the scheme optimization score, and supports users to modify in real time through VR feedback, integrates human factors, thereby reducing the planning dispute risk and improving the public participation, optimizing the iterative decision-making process, solving the problems of multi-scheme comparison and insufficient user participation.

[0021] 3. The present application outputs an optimized scheme and a visual report, realizes dynamic maintenance, reduces the update response time, ensures the survival of the fittest through A / B testing, and the report includes 3D screenshots, index change graphs and SWOT analysis, improves the visual intuitiveness, improves the user understanding rate, supports cloud collaboration, reduces the maintenance cost, guarantees the sustainability and adaptability of the scheme in long-term planning, thereby enhancing the output and update mechanism, and solving the limitations of planning scheme presentation and maintenance. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1A flowchart of a city planning and design method based on a real scene three-dimensional simulation system according to the present application. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0024] Embodiment: Figure 1 The present application provides a city planning and design method based on a real scene three-dimensional simulation system, which comprises the following steps: S1, collecting multi-source city data, including real scene three-dimensional images, GIS geographic information and VR interactive feedback data, to form an initial data set; S2, performing cross-modal feature extraction and alignment on the initial data set based on a deep learning model to generate a fusion feature vector, and realizing deep integration at the data layer; S3, constructing a dynamic three-dimensional city model using the fusion feature vector, and embedding a GIS spatial analysis module to perform preliminary layout simulation; S4, injecting VR interactive logic into the dynamic three-dimensional model, simulating user behavior feedback, and calculating environmental impact indicators; S5, optimizing planning parameters according to the simulation results and indicator feedback, adjusting the layout scheme through an iterative algorithm, and realizing multi-technology collaborative decision-making; S6, outputting the optimized city planning and design scheme, and generating a visual report to support real-time iterative updating.

[0025] S1, collecting multi-source city data, including real scene three-dimensional images, GIS geographic information and VR interactive feedback data, to form an initial data set, which is specifically implemented as: In the city planning and design process based on multi-modal deep fusion, first, multi-source city data is collected, including real scene three-dimensional images, GIS geographic information and VR interactive feedback data, to form an initial data set, To achieve comprehensive perception and multi-dimensional input of urban environment, and ensure the accuracy of subsequent cross-modal feature extraction and alignment; The multi-source data includes real scene three-dimensional images captured by unmanned aerial vehicle aerial photography, which are used to capture urban building and landscape details, GIS geographic information extracted from database, which is used to provide spatial coordinates and land use attributes, and VR interaction feedback data from user equipment records, which is used to integrate planning preferences and behavior simulation. The selection of data sources is based on the complexity and dynamic factors of urban planning, for example, in high-density urban areas, real scene images can reflect the building density, while GIS information supplements the land zoning constraints, and VR feedback enhances user participation; The collection frequency is set to 1Hz to balance real-time and data processing load, and the frequency setting is based on the statistical distribution of urban change rate to ensure that small dynamics are captured without generating redundant data. The entire collection process is performed on the integrated platform of the planning system, and the collection delay is controlled within 50ms. The delay threshold is set according to the requirements of the system end-to-end response time; Real scene three-dimensional images are collected using unmanned aerial vehicles; The specific collection process is as follows: First, the unmanned aerial vehicle performs aerial photography through the pre-set flight path to collect real scene three-dimensional image data; After collection, the three-dimensional model is reconstructed through the SfM algorithm, and the model precision error is controlled to be less than 5cm. The precision calculation formula is:

[0026] Wherein, the number of control points is >10, and an OBJ format file is generated, with additional metadata such as shooting timestamp and geographic reference. Image data provides a visual basis for alignment with GIS information; Geographic information is extracted from the GIS database. The GIS database queries the planning area boundary through the API interface to extract geographic information, including land use map and slope distribution, wherein the slope calculation formula is:

[0027] Wherein, the data format is GeoJSON, and the extraction interval is 1s to support real-time update; If the database response timeout is >2s, switch to the standby mirror server to ensure data integrity and link VR feedback to provide spatial constraints for GIS information to verify the planning feasibility of user interaction; The user interaction feedback is recorded by the VR device. First, the VR device is deployed in the planning interaction room. The user behavior, such as the scheme preference and modification feedback, is collected through the handle and head tracker. The data source is the user input event log, and the recording format is JSON array. The collection frequency is 1 Hz, which ensures the capture of fine-grained interaction. The feedback data is transmitted to the platform through the USB interface, with the user ID and session time attached. The VR feedback supplements the subjective input, and is fused with the real scene image and GIS information to form a multi-dimensional initial data set for feature extraction in S2, ensuring the comprehensiveness and user orientation of the planning design.

[0028] S2, cross-modal feature extraction and alignment of the initial data set based on a deep learning model to generate a fusion feature vector, realizing deep integration at the data layer, specifically: In the urban planning and design process based on multi-modal deep fusion, cross-modal feature extraction and alignment of the initial data set formed in S1 step are performed to generate a fusion feature vector, realizing deep integration at the data layer to eliminate the heterogeneity between different data sources and ensure the accuracy and consistency of subsequent three-dimensional model construction. The selection of data sources is based on the multi-dimensional needs of urban planning, such as real scene images providing visual details, GIS information supplementing spatial attributes, and VR feedback incorporating user subjective input. The extraction frequency is uniformly set to 1 Hz to match the collection rhythm. The frequency setting is based on the statistical distribution of urban dynamic changes to ensure timely feature capture without computational redundancy. The entire extraction and alignment process is performed in the deep learning module of the planning system, which is configured as a high-performance GPU server with a processing delay controlled within 100 ms. The delay threshold is set based on the requirement of the system end-to-end response time, with a total delay not exceeding 200 ms. Through the Transformer model and contrastive learning mechanism, deep fusion from multi-source data to a unified vector is realized. The Transformer model extracts features from the initial data set. First, the real scene three-dimensional image is converted into point cloud embedding, the GIS geographic information is converted into spatial embedding, and the VR interaction feedback is converted into sequence embedding. The extraction process is calculated through the self-attention mechanism, with the formula as follows:

[0029] The dimension of the query key vector is 64. Multi-head parallel processing is used to process different modal subspaces to generate a preliminary feature vector, which is normalized by LayerNorm to stabilize the output and ensure that heterogeneous data is converted into a unified representation. The modal is aligned through contrastive learning. The contrastive learning uses a CLIP-style loss function to align the feature vectors of different modalities. First, a positive and negative sample pair is constructed. The positive sample is the different modal data of the same planning area, and the negative sample is the random data of other areas. The alignment process calculates the similarity matrix, with the formula as follows:

[0030] The dimension of the query key vector is 64; multi-head parallel processing of different modal subspaces generates a preliminary feature vector, and is normalized by LayerNorm to stabilize the output and ensure that heterogeneous data is converted into a unified representation; Align the modalities through contrastive learning. Contrastive learning uses a CLIP-style loss function to align the feature vectors of different modalities. First, construct a positive and negative sample pair. The positive sample is different modal data in the same planning area, and the negative sample is random data in other areas. The alignment process calculates a similarity matrix, with the formula:

[0031] Cosine similarity, loss function is cross entropy, formula is:

[0032]

[0033] The quantity supports the semantic consistency of multi-source data, and the fusion vector is used as input to improve the accuracy of three-dimensional simulation; Generate a fusion feature vector to realize deep integration at the data layer. The fusion process weights the contributions of different modalities through a gating mechanism, with the fusion formula:

[0034]

[0035] Data resampling ensures comprehensive integration of planning and design.

[0036] S3, use the fusion feature vector to construct a dynamic three-dimensional city model and embed a GIS spatial analysis module for preliminary layout simulation, with the specific implementation being: In the process of urban planning and design based on multi-modal deep fusion, the fusion feature vector generated in step S2 is used to construct a dynamic three-dimensional city model and embed a GIS spatial analysis module for preliminary layout simulation, to realize digital reconstruction of urban space and preliminary evaluation of the planning scheme, ensuring that the model dynamicity and analysis accuracy; the fusion feature vector includes point cloud coordinates and texture information, 512-dimensional real scene three-dimensional image visual details, 512-dimensional GIS geographic information spatial attributes and behavior sequence and scheme modification records, and 512-dimensional VR interaction feedback user preferences; the data source is selected based on the multi-level needs of urban planning, for example, visual details provide architectural forms, spatial attributes supplement functional zoning, and user preferences enhance human factors; the construction frequency is uniformly set to 1 per minute to match the planning iteration rhythm, ensure that the model captures progressive adjustments without generating computational redundancy, and control the processing delay within 500 ms to realize deep integration from the feature vector to the simulation output; Constructing a dynamic three-dimensional urban model using a fusion feature vector: First, extract key components such as point cloud data and topological relationships from the fusion feature vector, generate an initial mesh model through the Mesh component of the Unity engine, and according to the planning precision test, when the mesh resolution is 0.1 m, the detail retention is > 95%, and when the mesh resolution is < 0.05 m, the rendering overhead increases by 50%, then set the mesh resolution to 0.1 m; In the construction, particle systems are used to simulate dynamic elements, the density is adjusted according to the user preferences of VR feedback, and the fusion vector is bound through scripts to update the model attributes in real time, such as building height and texture coordinates obtained through the texture coordinate formula:

[0037] The spatial analysis module is loaded into the Unity scene, binds the dynamic model through the API interface, analyzes the spatial attributes, extracts the model Z-axis height difference and GIS coordinate horizontal distance, calculates the slope through the slope calculation formula, and the formula is:

[0038]

[0039]

[0040] and outputs the preliminary indicators as a JSON report, including coverage, light pollution distribution map, and sustainability score calculated through the sustainability score calculation formula, and the sustainability score calculation formula is:

[0041] To maintain the model's timeliness and continuity, the controller establishes a routine comparison with external geographic data sources. When the hash difference between a new geographic information update and the local cache exceeds a set threshold, local reconstruction is triggered only for the changed spatial areas, ensuring that an update is completed within approximately one second. This achieves rapid change capture driven by differences and minimizes the recalculation range, aligning with mainstream geographic information change detection approaches and enabling hash-driven change identification to reduce the cost of full comparison. In terms of reconstruction and material...

[0042]

[0043]

[0044]

[0045] S4. Inject VR interaction logic into the dynamic 3D model to simulate user behavior feedback and calculate environmental impact indicators. The specific implementation is as follows: Functional enhancements are made to the dynamic 3D city model generated by S3. The model uses a building grid with a precision of 0.1 meters as the display medium, overlaid with a spatial attribute layer (slope distribution, land use type) injected by geographic information data, and uses preliminary layout parameters (such as the interval setting of building height and spacing) as the starting point for interaction. Within the virtual reality interaction module, the interaction toolkit and script system of the engine are used to bind the "grab-move-look-attention confirmation" interaction events to the simulation kernel. User behavior feedback is written back to the model at a frequency of "once per scheme iteration" to ensure that the interaction rhythm is synchronized with the simulation step size without lag, thereby realizing the visualization and real-time modification of the scheme. In terms of environmental impact assessment, the system calculates the correlation between illumination and glare in the frame loop and uses industry-standard light environment simulation and glare evaluation processes to generate light pollution distribution maps and threshold determinations, serving as the quantitative basis for linkage optimization. The system runs on a high-performance virtual reality workstation, with the graphics processing unit providing real-time ray tracing and high-resolution rendering capabilities. The eye-tracking accuracy of the headset is 0.5°, which can be used for gaze point sampling and controller-free interface interaction. To suppress dizziness and maintain immersion, and with end-to-end processing latency controlled within a 50-millisecond budget, the system coordinates with the Unity engine's scripting system and environmental simulation components to achieve logical continuity from model enhancement to indicator output. The optimized dynamic 3D model, packaged with resources, is loaded into the target scene (to reduce the initial package size and runtime memory, and facilitate on-demand loading across different platforms). Then, an interactive script derived from the script base class is attached to the root node of the object hierarchy to uniformly receive and distribute input from virtual reality devices. During initialization, the script parses external fusion feature vectors and maps them to adjustable model attributes (such as building spacing). During runtime, it uses the agreed-upon input mapping to switch modes using controller buttons, adjust the viewpoint using head posture, and move objects using dragging, while utilizing box colliders to implement lightweight collision constraints. To ensure stability, the script incorporates built-in exception handling and rollback branches, logging and restoring to the default model in case of mesh or hierarchy anomalies. Without altering the existing content creation and operation mechanisms, the "on-demand resource loading—root node injection—input mapping—attribute linkage—exception rollback" process is chained into a reproducible injection pipeline, enabling instant conversion from static models to interactive objects while maintaining end-to-end auditability and scalability. Once the virtual reality script receives user input, it first performs hit detection on interactive objects in the scene using raycasting (in dragging scenes, a ray emanating from the camera or controller is used, with an infinite length, and layer masks are used to only hit building layers to reduce false positives). Based on this, it locks the manipulated object and the hit point. When the object is dragged, the system calculates the target's new position in real time according to a formula: The system calculates the displacement changes caused by the input and performs parallel verification of placement constraints. A slope threshold of 30° is set, and the current slope is read via the geographic information module. If the slope exceeds the threshold, a vibration pulse is immediately sent back to the handle, and placement is rejected. Upon successful verification, the system enters a state-fixed state and triggers dual-channel feedback. On one hand, the selected building is visually highlighted; on the other hand, a prompt sound effect is played at key operation points during scheme switching. User preferences are logged, and this feedback chain is executed in each frame loop to ensure interactive continuity. A conflict handling branch is also set: if the input causes an out-of-bounds or violation of scene boundaries, an "invalid position" warning immediately pops up on the user interface canvas, and the object state is rolled back to the previous valid stack frame. Furthermore, user-modified data generated by simulated feedback is used as input for re-evaluation of indicators such as light pollution, ensuring user-oriented iteration of the planning process. In the updated model state, the system prioritizes establishing solar source parameters. The azimuth angle is obtained from the astronomical database according to the target date (e.g., summer solstice tilt angle 23.5°), and the light intensity is set to 1000 W / m² based on the standard solar constant corrected for latitude 30°. Then, it iterates through the outer surface of the model (number of triangular facets > 1,000,000), calculating the reflected light intensity facet by face. The reflected light intensity is obtained using the light intensity formula:

[0046] In S4, the calculated light pollution values ​​are used as raster input. The rendering pipeline is invoked to generate a distribution map, and boundary detection is first performed on high-risk areas. The extracted boundaries are then overlaid and marked with 2-pixel line widths. Subsequently, the resulting map is projected onto the VR view in a screen-space overlay mode, enabling user interactive query capabilities. After generation, the map is stored on disk in PNG format, and metadata fields are written to support evidence collection and traceability, including the calculation of timestamps and average risk values. The formula for the average risk value is:

[0047] If an anomaly occurs during the distribution map generation process (such as memory overflow), the fault-tolerant block process will be automatically activated, splitting the rendering task into 4 blocks, each 1024×1024, and then stitching them together without loss after completion to ensure stable and reliable output. The front-end integrates gesture recognition, triggering model scaling with pinch gestures and writing gesture events and view transformation parameters into an interaction queue. To support collaborative solutions, multi-user synchronization is enabled, and sequential merging of editing conflicts is achieved through session identifiers and timestamps. In the back-end log layer, user behavior is measured for preferences, and preference scores are calculated using a formula: The weights are the sum of 0.6 for scheme type and 0.4 for location, thus forming a traceable human-computer interaction profile; the environmental assessment is extended to the noise dimension and calculated according to the noise formula, which is:

[0048] The solution enables collaborative decision-making across multiple technologies, and its specific implementation is as follows: In the urban planning and design process based on multimodal deep fusion, the decision module performs real-time feedback analysis on the preliminary layout simulation results and environmental impact indicators generated in step S4. The simulation results include the preliminary layout scheme of the 3D model and indicator feedback. The data sources are selected based on the comprehensive needs of urban planning. Coverage rate is used to reflect land use efficiency, light pollution index is used to assess environmental impact, and sustainability score is used as an overall evaluation that integrates human and ecological factors. In the feedback loop, the system adjusts the planning parameters according to the above indicators and continuously corrects the layout scheme through iterative algorithms to achieve multi-technology collaborative decision-making. The basis is the historical average adjustment rate of 0.2 schemes / min. The optimization frequency is uniformly set to once per iteration cycle to ensure that the optimization can capture incremental improvements and avoid decision redundancy caused by high-frequency invalid oscillations. This achieves synchronous convergence of layout and indicators, ensures the sustainability and feasibility of the scheme, and significantly improves the overall design quality and efficiency. Based on the simulation results and indicator feedback analysis of planning parameters, firstly, the JSON report generated in step S4 is read and parsed to obtain building layout coordinates, traffic accessibility indicators, and environmental impact indicators. Simultaneously, a structured table evaluated by the GIS module is received as indicator feedback input, including flood risk percentage and ventilation rate. These multi-source indicators are then compared item by item with the preset planning targets, where the coverage target is >70%, and the sustainability score is >80 points. The sustainability score is calculated using the following formula:

[0049] Calculate the compliance level of the current plan and locate the source of deviation; when the coverage deviation is >5%, it is automatically marked as an inefficient area, and threshold verification is performed in combination with the historical planning case library. If any key indicator deviation exceeds the threshold, a parameter correction list for the execution unit is generated, and the indicator feedback is used as input to ensure the pertinence of the analysis and provide a data basis for adjustment, avoiding blind optimization. To prepare for adjusting the layout scheme by optimizing planning parameters, the candidate planning parameters are first quantified and sorted based on the analysis list. Then, the top three parameters are selected as the focus of optimization using a priority calculation formula, which is as follows: The influence weights consist of coverage (0.4), risk (0.3), and matching (0.3). Subsequently, under the three-tiered constraints of "compliance-feasibility-controllability" and in accordance with urban regulations and environmental standards, the parameter adjustment range is defined as follows: height 20–50m, spacing >5m, and density 0.5–1.0. Simultaneously, spatial constraints are loaded from the geographic information module, imposing a density upper limit on slope-sensitive areas. The rule "high-density buildings are prohibited in areas with a slope of less than 30°" is solidified into a geometric mask to ensure that the constraints in the preparation stage are consistent with the actual terrain. All preparation data is stored in the database in a temporary cache and metadata is attached at the entry level so that it can be used as standardized input for the optimizer in subsequent iterations. This supports incremental convergence and backoff control of the scheme, supports incremental optimization of the scheme, and ensures logical continuity. The layout scheme was adjusted to achieve initial optimization. Starting with parameter preparation, candidate schemes were fine-tuned one by one in a predetermined order, prioritizing low-risk parameters. "Increasing building spacing" was a typical operation, using a step size of 1 m and a test cycle of 5 times. After each adjustment, core indicators were recalculated immediately, and the coverage rate was calculated using the coverage rate calculation formula. The deviation was compared with the baseline. If the target deviation decreased by more than 2%, the adjustment was retained as the starting point for the next round. The overall optimization ran within a "comparison-feedback" closed loop, with a maximum cycle of 10 times and a termination condition (deviation <3%) to control convergence and computational costs. After solution space convergence, a set of 5 optimized schemes was generated based on the Pareto front principle, balancing multiple objectives such as coverage indicators and risk constraints. Geographic information constraints and immersive feedback were incorporated simultaneously throughout the process. User preferences in the virtual reality environment were weighted at 0.3 in the comprehensive score, ensuring integrated technical and humanistic decision-making. The final output was based on engineering deliverability, providing clear technical boundaries and auditable evidence for subsequent in-depth optimization and review. To achieve collaborative decision-making across multiple technologies, based on a predetermined adjustment plan, the system synchronously integrates 3D models, geographic information analysis, and virtual reality interaction, and calculates a comprehensive score using a unified scoring logic. The formula is as follows: In this process, user satisfaction is converted from VR rating 0-10 to 0-1. During the decision-making stage, based on the threshold corresponding to a 90% success rate in historical planning, candidate solutions with scores >85 are prioritized. If the threshold is not met, iterative rollback is triggered for re-evaluation. The collaborative mechanism is reflected in cross-module data exchange. For example, the coordinates of the 3D model are pushed to the GIS end for flood risk query, and the risk results are fed back to the VR side to dynamically adjust the scene and user interaction. The final output is a reviewable solution report, which includes 3D screenshots of key scenes and indicator tables, realizing measurable and traceable decision-making driven by technology integration. To further expand the optimization process, we first perform a solar radiation analysis according to the solar radiation calculation formula, where the formula is: Under latitude of 30°, daily sunshine hours were calculated using astronomical database data, and candidates with less than 6 hours of sunshine per day were screened out using GB 50033-2013 as a hard constraint. Simultaneously, wind field assessments were conducted, and a ventilation map was output using a CFD model. A ventilation rate threshold >80% was used as the qualification criterion. Based on this, layout iterations were triggered, and a gradient descent variant was used to fine-tune and verify key parameters. For example, the impact of a 2% reduction in coverage and a 5% reduction in risk when the building spacing increased by 0.5m was recorded. Iteration automatically terminated when the score converged to less than 1% change. The comprehensive evaluation of the optimization scheme incorporated a sustainability sub-indicator, focusing on verifying the green space ratio using the green space ratio calculation formula, and setting a green space ratio threshold of 30%. If the calculated green space ratio value was lower than the threshold, the planning parameters were forcibly adjusted to increase the green space. The calculation formula is as follows: In terms of multi-technology collaboration, the GIS module performs buffer queries to undertake spatial compliance verification, the VR terminal supports users to drag and drop modifications and write back the model and scores in real time, the decision-making level retains the preferred solution with a fusion score threshold of >85%, and the final report automatically exports 3D screenshots and Excel index tables, and attaches a change trajectory map generated by Matplotlib. It completes a verifiable and quantifiable optimization loop within a ten-second time limit, ensuring the logical continuity and engineering feasibility from multi-source vector fusion to layout finalization. Furthermore, as a supplementary implementation method for multi-technology collaborative decision-making, a risk prediction submodule can be integrated into the platform: based on ≥100 historical cases, a quantitative assessment of the long-term impact of light pollution can be performed, calculated using the following formula: When the predicted risk value is greater than 500 lux* per year, it is judged as high-risk. The system automatically generates adjustment suggestions and lowers the reflectivity parameter by 0.1 in the candidate solutions to prioritize the elimination of high-risk sources. At the human-machine collaboration level, user feedback adaptive weights are introduced. If VR satisfaction is less than 0.8, the corresponding subjective indicator weight is increased by 0.1 to ensure that the evaluation results take into account both engineering objectivity and humanistic orientation. The solution output is simultaneously designed for system integration and visualization verification, and provides API interfaces for external business systems to call. When generating reports, Unity screenshots are called to ensure perceptual consistency between key sections and bird's-eye views. To control complexity and timeliness, when the number of decision loops is greater than 10, a model simplification strategy is triggered, reducing geometric details to LOD level 3, reducing the number of faces by 40%, and reducing the total time of this round of optimization to less than 2 seconds. This process not only enhances the robustness of decision-making but also provides a data-driven, multi-technology collaborative foundation for urban planning. It ensures the logical integrity from feedback to optimization and achieves refined optimization of planning parameters through the analysis of simulation results, preparation of parameter optimization, adjustment of layout schemes, and multi-technology collaborative decision-making, thereby supporting sustainable urban planning and design.

[0050] S6. Output the optimized urban planning design scheme and generate a visualization report, supporting real-time iterative updates. Specific implementation details are as follows: In the process of urban planning and design based on multimodal deep fusion, the system completes the finalization of the scheme, report generation, and real-time iterative integration within the same visualization module, based on the optimized layout obtained by S5. First, based on the previous multimodal simulation and indicator feedback, the consistency of building layout parameters (height, spacing, density), traffic network configuration, and green space distribution is checked. Then, a threshold rule is used to impose strong constraints on key indicators (such as building spacing > 5m, traffic accessibility > 85%, sustainability score > 80 points), so that high-density areas prioritize ventilation and sunlight, and suburban areas prioritize the proportion of green space. The verified elements are then integrated into the urban spatial base map. Based on the weights of three categories of results—"ventilation rate, accessibility rate, and sustainability"—a decision-oriented solution view and element list are generated. According to the statistical distribution of planning iterations, the output frequency is set to once per optimization cycle to capture the final adjustment and avoid redundancy. At the implementation level, the report and view are rendered and arranged by a high-performance graphics server, ensuring that the processing latency within the module is <1s and the total end-to-end latency is ≤3s. The final results are also provided to the outside world through file export and application interface to achieve deep integration from solution output to iteration. The optimized urban planning design scheme is output. First, after completing the multi-objective optimization of S5, the highest-scoring scheme is selected from the Pareto scheme set according to the scoring criteria. Then, schemes with scores >85 are selected using a weighted scoring formula, where the formula is: The parameter list of the proposed solution is used as input for legal compliance review to ensure that it meets the requirements of existing urban regulations. Then, standardized output is organized to generate OBJ 3D model files containing grid, texture and coordinate information to support subsequent tracking. Once the output is completed, an integrity check is performed. If any item fails to meet the standard, a rollback process is triggered and the process returns to S5 for one iteration until the check is passed. This ensures the continuity from decision-making to presentation and provides verifiable data and model carriers for visualization reports and archival preservation, avoiding output deviations. The system generates a visualization report. After obtaining the final solution, it automatically compiles the report elements, first generating a text block summarizing the solution, then visualizing the core indicators of coverage, ventilation, and light pollution using bar charts. The 3D engine's camera captures scene screenshots from multiple angles and poses, and multi-sampling anti-aliasing (MSAA) is enabled during the rendering stage. 4× to improve edge quality; then combine text, charts and 3D screenshots into a PDF document, and use iTextSharp to complete page management and export, and embed hyperlinks to external OBJ models in the page, so that users can interact and browse in the 3D viewer after clicking. At the same time, key feedback, such as light pollution heat map, is highlighted on the indicator visualization page, and the change trajectory page is retained to compare parameters and results before and after iteration, intuitively showing the optimization benefits such as "coverage increased from 60% to 75%". If the exported file size exceeds 50MB, the process automatically performs lossy compression on the embedded images to ensure transmission and archiving. The entire report retains the ventilation score calculation formula, and generates a hyperlink directory and bookmarks during export to support subsequent iteration updates and cross-document navigation. The overall generation time is controlled within 10 seconds, and a feedback form is embedded in the report for users to input new preferences. Supports real-time iterative updates. After the solution and report are output, the "parameter update" capability is enabled through the API interface. The client submits change requests in JSON format. The server first validates the input boundaries; if the validation passes, the S5 iteration process is triggered, and during the iteration, parameters are updated in JSON format. According to the data, incremental reports are generated only for pages that have changed, and a PDF tool is used to merge the incremental pages with the original reports. To ensure timely interaction, the system maintains a two-way long connection with the front end via WebSocket. The user end displays the visual model exported by Unity WebGL in real time through a browser plugin. When the update frequency exceeds 5 times per minute, priority queuing and rate limiting are enabled to ensure service stability and deterministic response. This mechanism forms a closed loop in S6, with adopted parameters and results fed back to S1 as input for the next round of data collection and simulation, forming a dynamic cycle of planning. The output scheme is further expanded by integrating a sustainability assessment extension when outputting the optimization scheme. On the one hand, carbon emissions are assessed through the carbon emission calculation formula, which is: The volume is calculated using a 3D model, the material density is taken as 7.8t / m³ for steel, and the emission factor is 0.5t CO2 / t. The results are compared with the carbon accounting boundaries and management requirements defined in the "Building Carbon Emission Calculation Standard". The threshold for the scheme is set as emissions <100t / year to achieve pre-compliance interception and traceability verification. On the other hand, a unified and exchangeable XML meta-file containing a parameter tree is output, which is convenient for direct access by external software. For engineering workflows, the LandXML paradigm can be used to achieve seamless import with AutoCAD series, reducing cross-platform information loss and review costs. At the same time, seismic risk extension is included: during the scheme verification stage, the physics engine is called to simulate swaying, and the coupled response of the structure and site is solved according to rigid body dynamics to generate a risk heat map, so as to achieve intuitive interpretation in line with the seismic intensity classification system and support multi-dimensional planning. The extension for generating visual reports first utilizes Unity's Timeline feature to create a 30-second animation of the solution's evolution (e.g., a gradual increase in scale from low to high) during report generation, ensuring key nodes are presented sequentially along the timeline. The video is then exported and embedded as an attachment to the report, allowing reviewers to directly play the solution demonstration within the same document. To support management decisions, the report simultaneously includes a SWOT analysis (Strengths: High Coverage; Weaknesses: High Cost; Opportunities: Sustainability; Threats: Light Pollution), and the generated Excel statistical table is embedded in the PDF, ensuring integrated presentation of numerical values ​​and charts. It also maintains the corresponding annotations between the aforementioned core evaluation indicators and animation frames, facilitating a traceable link between visual perception and quantification. Upon completion, the report is shared via email or cloud storage, significantly improving interactivity and comprehension efficiency. The extension supports real-time iterative updates by using object storage and enabling version management at the cloud collaboration layer. Every solution change (graphics, parameters, report) creates a traceable history, allowing for easy restoration to any version. This introduces a Git-like version control paradigm into engineering collaboration, ensuring... It ensures consistency and rollbackability in multi-person parallel editing, branch merging, and change auditing, thus providing a compliance foundation for synchronous editing across teams and terminals. The terminal side provides a mobile application to achieve real-time feedback across multiple terminals. Each update cycle performs A / B testing by default (simulating two schemes in parallel, with a test time of <1 minute) and sets a decision threshold. When the voting support rate is >60%, the winning scheme is adopted. If conflict feedback occurs, a weighted voting process is initiated, and the results are calculated based on user roles to ensure a balance between efficiency and fairness, ensuring the real-time nature and collaboration of iterations, and supporting large-scale planning. To ensure the security of the solution, the processing of personal information complies with the General Data Protection Regulation (GDPR), and AES-256 encryption is enabled throughout the entire process. User feedback is collected anonymously, and geographic coordinates are anonymized during the output phase to reduce location identifiability and the risk of re-identification without affecting macro-analysis. The report generation uses the Jinja2 template engine to quickly render a professional appearance with configurable layouts, and the generation time is constrained to within 10 seconds. On the runtime side, an iterative update log is established using SQLite, and the following information is recorded in a standardized manner: time, user ID, and changed parameters. ACID-level transactions ensure write consistency and power-off reliability, meeting audit traceability and auditing requirements, and ensuring the logical integrity from optimization to application. Through solution output, visual report generation, and real-time iteration support, the final presentation and continuous improvement of the planning design are achieved. The detailed parameters and process selections of this implementation are based on actual planning specifications and user needs, ensuring the reliability and scalability of the method and providing a closed-loop guarantee for the overall method, supporting efficient urban planning practices.

[0051] This embodiment's solution generates multi-source initial data by collecting real-world 3D imagery, geographic information, and virtual interactive feedback on a unified platform. Subsequently, deep data fusion is achieved through cross-modal feature extraction and alignment, which is then input into a dynamic 3D city model. Geospatial analysis is embedded in the model to complete preliminary layout and compliance verification, and quantifiable assessments are provided based on scenario indicators such as wind environment, sunshine, light environment, and waterlogging. In the virtual interaction phase, planners are allowed to make real-time adjustments to building height, spacing, density, and road organization in an immersive manner. The system simultaneously recalculates key indicators and records user preferences. Based on indicator feedback, parameter optimization and multi-scheme comparison are performed to form the optimal scheme that meets regulatory constraints and comprehensive scoring requirements. Finally, the scheme finalization and visualization report generation are completed in the same module, supporting integrated presentation of 3D screenshots and indicator tables, and enabling incremental updates and version traceability. The entire process is equipped with data encryption, coordinate anonymization, anonymous feedback, and log auditing to ensure compliance, security, and verifiability, thereby achieving rapid iteration and implementation of urban planning and design through engineering and scalable methods.

[0052] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wireless or wired transmission; wired transmission methods include optical fiber, twisted pair, coaxial cable, etc.; wireless transmission includes infrared, microwave, etc. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center containing one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0053] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0054] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0055] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0056] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0057] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0058] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0059] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for urban planning and design based on a real-scene 3D simulation system, characterized in that, include: S1. Collect multi-source urban data, including real-scene 3D images, GIS geographic information and VR interactive feedback data, to form an initial dataset; S2. Based on a deep learning model, cross-modal feature extraction and alignment are performed on the initial dataset to generate a fused feature vector, thereby achieving deep integration of the data layer. S3. Construct a dynamic 3D city model using fused feature vectors and embed it into a GIS spatial analysis module for preliminary layout simulation; S4. Inject VR interaction logic into the dynamic 3D model, simulate user behavior feedback, and calculate environmental impact indicators. S5. Based on simulation results and indicator feedback, optimize planning parameters, adjust layout schemes through iterative algorithms, and achieve multi-technology collaborative decision-making; S6 outputs the optimized urban planning design scheme and generates a visual report, supporting real-time iterative updates.

2. The urban planning and design method based on a real-scene 3D simulation system according to claim 1, characterized in that, Collect multi-source urban data, including real-scene 3D imagery, GIS geographic information, and VR interactive feedback data, to form an initial dataset, including: Configure sampling periods and timestamps for real-scene 3D images, GIS geographic information, and VR interactive feedback respectively; Process the data source to determine whether to trigger mirror source switching and differential re-collection; Incremental data retrieval will be performed on spatial fragments that have been verified as changed by Hahe metadata. A one-to-one mapping relationship is established between the data source identifier and spatial coordinates, and a traceable index table and collection log are generated.

3. The urban planning and design method based on a real-scene 3D simulation system according to claim 1, characterized in that, Based on a deep learning model, cross-modal feature extraction and alignment are performed on the initial dataset to generate a fused feature vector, achieving deep integration of the data layer, including: A cross-modal feature alignment mechanism is used to deeply integrate the initial dataset; Sub-vectors are extracted from different modalities, normalized, and then weighted and fused according to learnable weights; Configure learnable weight coefficients for each subvector; A preset alignment threshold is set and compared with cross-modal similarity; The feature vector is integrated with three types of meta-information: modal weights, unified time labels, and spatial range.

4. The urban planning and design method based on a real-scene 3D simulation system according to claim 1, characterized in that, A dynamic 3D city model is constructed using fused feature vectors, and a preliminary layout simulation is performed by embedding a GIS spatial analysis module, including: The construction of dynamic 3D city models employs Level of Detail (LOD) management and maintains grid topology relationships; When computation or rendering resource consumption reaches a preset resource threshold, the LOD level is downgraded. Prioritize simplifying detailed triangular facets while maintaining the building boundary outline and road network topology unchanged; Embed the GIS spatial analysis module into the scene pipeline via an interface; Closed-loop updates of local reconstruction and spatial analysis are performed on the same data snapshot according to the trigger-write-re-trigger link.

5. The urban planning and design method based on a real-scene 3D simulation system according to claim 1, characterized in that, Inject VR interaction logic into a dynamic 3D model to simulate user behavior feedback and calculate environmental impact indicators, including: VR interaction employs ray-based hit detection and constrained drag-and-drop logic; Read the slope, land use constraints, and boundary buffer space attributes of the proposed placement location for the candidate object; Preset a set of compliance thresholds and compare them with any constraint parameter; Generate event log entries containing object identifiers, interaction times, constraints, and poses, and write them back to the scene state.

6. The urban planning and design method based on a real-scene 3D simulation system according to claim 5, characterized in that, include: Environmental indicators include light and glare distribution, noise intensity estimation, and flood sensitivity; Preset a second threshold and compare it with the target area ratio or peak value of any indicator; Automatically generate a candidate set of directional adjustments for at least one of the parameters: reflectivity, building height, building spacing, or land density.

7. The urban planning and design method based on a real-scene 3D simulation system according to claim 1, characterized in that, Based on simulation results and indicator feedback, planning parameters are optimized, and layout schemes are adjusted through iterative algorithms to achieve multi-technology collaborative decision-making, including: Select several parameters to be adjusted according to a priority queue; Make single-round fine adjustments with a fixed micro-step size and recalculate core indicators in real time; Set a convergence threshold or any hard constraint and compare it with the comprehensive score; Throughout the iteration process, maintain the Pareto non-dominated solution set and retain candidate solutions that pass the constraint consistency check.

8. The urban planning and design method based on a real-scene 3D simulation system according to claim 1, characterized in that, Output optimized urban planning design schemes and generate visual reports, supporting real-time iterative updates, including: Each of the strongly constrained clauses regarding building spacing, sunshine duration, ventilation assessment, proportion of ecological green space, and traffic accessibility was verified. Perform version consistency checks on the 3D model, spatial analysis results, and interactive modification records; Generate standardized 3D model files and a list of structured indicators.

9. A city planning and design method based on a real-scene 3D simulation system according to claim 8, characterized in that, The report generation utilizes template-driven feature assembly and incremental updates, including: The template engine is used to arrange text descriptions, indicator charts, and multi-view scene screenshots into a page format; the scheme number, data snapshot fingerprint, and camera pose are embedded in the report metadata; Subsequent fine-tuning of the same scheme only generates incremental pages and merges them into the original report, while outputting a machine-readable parameter tree file.

10. The urban planning and design method based on a real-scene 3D simulation system according to claim 8, characterized in that, include: The coordinate accuracy of external shared spatial locations is reduced and regionalized desensitization is performed. User feedback is anonymized and hashed. End-to-end encryption is enabled during transmission and storage, and an immutable iterative update log is maintained.

Citation Information

Patent Citations

  • Urban planning live-action three-dimensional simulation system

    CN117974912A

  • Embedded city design scene simulation method and system

    WO2020192027A1

Cited By

  • VR platform control method and system for synchronous interaction of multiple education terminals

    CN122018701A

  • Design method and program product of canyon low tower cable-stayed aqueduct based on landscape optimization

    CN122154048A

  • Village planning real scene three-dimensional data acquisition method, device and equipment and storage medium

    CN122391511A