Urban skyline real scene visual touch interaction optimization method based on multi-scenario simulation

Through multi-agent deep reinforcement learning and mixed reality technology, a city skyline optimization plan is automatically generated, which solves the problems of single observation perspective and low adjustment efficiency in city skyline design, and realizes efficient and accurate city skyline optimization.

CN119670180BActive Publication Date: 2025-10-10SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411750754.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-10-10
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

Existing technologies in urban skyline optimization design have problems such as a single observation perspective, poor results, low adjustment efficiency, and inconsistent judgment standards, and lack multi-agent artificial intelligence applications in three-dimensional space.

Method used

Building data is collected through a multi-angle tilt camera and GPS positioning module, and a three-dimensional real-scene model is generated in combination with a lidar scanner. Multi-agent deep reinforcement learning methods are used to record user adjustment behaviors, automatically generate optimization plans, and perform interactive optimization through mixed reality glasses and touch screens.

Benefits of technology

It realizes the automation and intelligent optimization of the city skyline scheme, improves the design efficiency and accuracy, reduces the workload of planners, and enhances the accuracy of visual effects and the immediacy of adjustments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119670180B_ABST
    Figure CN119670180B_ABST
Patent Text Reader

Abstract

The application provides a kind of city skyline real scene visual touch interaction optimization method based on multi-scenario simulation, including collecting city target area three-dimensional real scene data;Three-dimensional real scene data is uploaded to digital twin visualization platform, and the skyline lightweight real scene model of city target area is reconstructed;The behavior and instruction information of user adjusting skyline model are recorded and translated one by one using wearable device, and are summarized as preference training data;Each building monomer in skyline model is defined as an agent, and the action and reward function of agent are defined, and the skyline optimization problem of skyline model is solved using multi-agent deep reinforcement learning;Determine whether loss function and reward function meet the requirements;Save the output scheme, output to holographic sand table and run display, and make secondary accurate adjustment.The application can generate city skyline optimization scheme more in line with user demand based on user adjusting skyline model behavior data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of urban planning, and in particular to a method for optimizing real-scene visual touch interaction of a city skyline based on multi-scenario simulation. Background Art

[0002] On the one hand, the city skyline is a crucial window into the cityscape, and its planning and design are key tasks in urban design. Currently, the optimization of city skylines is primarily accomplished by planners observing and manually adjusting three-dimensional models of urban buildings on a computer screen. This process presents challenges such as limited spatial cognition, a significant discrepancy between the skyline display and actual perception, and a time-consuming and labor-intensive adjustment process. Planners also face inconsistent judgment criteria when optimizing different skyline schemes. On the other hand, with the development of artificial intelligence (AI), the methods and presentation of urban design are constantly evolving. Currently, reinforcement learning, as a representative AI technology, has been applied in urban design, demonstrating capabilities exceeding those of human planners. However, reinforcement learning is currently primarily applied to two-dimensional urban design, lacking applications in three-dimensional space and lacking integration with other AI technologies such as multi-agent systems.

[0003] A real-world visual touch interactive optimization method for city skylines based on multi-scenario simulation is of great significance. It addresses several issues inherent in traditional city skyline optimization, including a single observation perspective, poor observation quality, low adjustment efficiency, and inconsistent judgment criteria. It enables automated optimization of city skylines based on multiple scenarios, allowing planners to efficiently and accurately adjust their city skyline plans. Summary of the Invention

[0004] To address the shortcomings mentioned in the above background technology, the purpose of the present invention is to provide a real-life visual touch interaction optimization method for city skylines based on multi-scenario simulation. By recording and summarizing the user's actions and instructions for adjusting the city skyline model, a multi-agent deep reinforcement learning method is used to automatically generate and display city skyline optimization solutions.

[0005] The purpose of the present invention can be achieved through the following technical solutions:

[0006] The method for optimizing the visual touch interaction of a city skyline based on multi-scenario simulation includes the following steps:

[0007] Step (1): Oblique photography data of buildings in the urban target area are collected by a drone equipped with a multi-eye oblique camera and a GPS positioning module, and three-dimensional morphological data of buildings in the urban target area are collected by a lidar scanner. Three-dimensional real scene data of the urban target area are obtained through data registration and fusion.

[0008] Step (2): Upload the real scene data to the digital twin visualization platform, reconstruct the three-dimensional model of the urban target area, perform monomer processing on the three-dimensional model of the urban target area based on the convolutional neural network, and use the digital Li Sheng fusion rendering engine to reconstruct the lightweight real scene model of the urban target area skyline.

[0009] Step (3): The lightweight real-life model of the urban target area skyline is further divided into near-ground buildings, landmark high-rise buildings, and matrix buildings based on the geometric position, density, normal characteristics of the point cloud clusters and their position from the skyline model viewpoint; the mixed reality glasses equipped with an eye movement detection module and the handheld touch device equipped with an inertial sensor are connected to the digital twin visualization platform to record the user's touch behavior and command information for adjusting the skyline model; the user behavior and command information are translated one by one and summarized as the monitoring data of the device during adjustment, the parameter changes of the model before and after adjustment, and the characteristic parameters of the model after adjustment.

[0010] Step (4): Multi-agent deep reinforcement learning (RIAL) is used to solve the skyline optimization problem of the lightweight real-life model of the urban target area skyline. According to the monitoring data of the equipment during adjustment, each building unit in the near-ground building, landmark high-rise building and matrix building is defined as an agent, and the action and reward function of the agent are defined.

[0011] Step (5): During the training iteration process, record the changes in the loss function and judge whether the loss function converges to below 0.1 and whether the reward function converges to the optimal solution. When both of the above conditions are met, stop the iteration and output the city skyline optimization plan; otherwise, clear the experience replay buffer, restart the interaction between the agent and the environment, and loop the training process until the termination condition is met.

[0012] Step (6): Save the output plan, output it to the holographic sandbox that stores the current lightweight real-life model of the city's target area skyline, run the display, compare the city skyline optimization plan with the current skyline, and make secondary interactive precise adjustments.

[0013] Furthermore, the step (1) includes the following steps:

[0014] Using drones equipped with multi-lens oblique cameras and GPS positioning modules, we acquire oblique photography data of buildings in target urban areas with geographic coordinates, including top and facade images. We also use vehicle-mounted and handheld LiDAR scanners equipped with GPS positioning modules to acquire three-dimensional morphological data of buildings in target urban areas, including point cloud data of building heights and outlines. By registering and fusing the oblique photography data with the laser scanning point cloud data, we obtain three-dimensional real-world data of the target urban area.

[0015] Furthermore, the step (2) includes the following steps:

[0016] The three-dimensional real scene data of the urban target area obtained in step (1) is input into the EasyV digital twin visualization platform to reconstruct the three-dimensional model of the urban target area. The three-dimensional model of the urban target area is singulated by the image recognition algorithm based on the convolutional neural network to generate a lightweight real scene model of the current urban target area skyline that can adjust each building separately. The model is then rendered and displayed by the EasyTwi n digital twin fusion rendering engine.

[0017] Furthermore, the step (3) includes the following steps:

[0018] S31. In the EasyV digital twin visualization platform, use the preset target detection algorithm to further divide the building units into near-ground buildings, landmark high-rise buildings, and matrix buildings, record their building heights, facade areas, and base area values, and link them to the corresponding attribute tables;

[0019] The preset target detection algorithm performs division based on the geometric position, density, normal features of the building point cloud clusters and their position from the skyline model viewpoint.

[0020] S32. Connecting MR mixed reality glasses equipped with an eye movement detection module and a handheld touch controller equipped with an inertial sensor to the EasyV digital twin visualization platform to record user touch actions and command information for adjusting the skyline model. The command information includes the following adjustment settings: "remove," "add," "increase," "decrease," and "setback along the vertical line of the building" for near-ground buildings, landmark high-rise buildings, and base buildings.

[0021] S33. After the user completes the touch adjustment of the skyline model of the target urban area, the platform interprets the user's behavior and instruction information one by one and summarizes it into the monitoring data of the device during the adjustment, the parameter changes of the model before and after the adjustment, and the characteristic parameters of the adjusted model, which specifically include:

[0022]

[0023]

[0024] Furthermore, the step (4) includes the following steps:

[0025] S41. A multi-agent deep reinforcement learning (RIAL) algorithm is used to solve the skyline lightweight real-world model optimization problem. Each building in the near-ground building, landmark high-rise building, and matrix building is defined as an agent. At each time t, the local state O of the observed environmental system is the characteristic parameters of other agents within a 1km buffer zone around the geometric center of gravity of the agent. The agent then sends an action from an action set A to the environment. The actions include "remove," "add," "increase," "decrease," and "indent" the facade along the vertical line of the building for the near-ground building, landmark high-rise building, and matrix building.

[0026] S42. In a cloud computing platform with a computing power of 10 TFLOPS, import the user operation training data recorded in step (3), and fit the reward function to evaluate the current state and the touch adjustment result of the skyline model;

[0027] The reward function is defined by integrating the parameter changes of the model before and after adjustment and the similarity between the parameters of the adjusted model and the user training data. The specific formula is:

[0028] R=λ·MinR1+μ·MaxR2

[0029] Among them, MinR1 minimizes the parameter changes of the model before and after adjustment, and MaxR2 maximizes the similarity between the parameters of the adjusted model and the user training data; λ and μ are adjustment parameters used to balance the impact of changes in the two parameters, so that the reward function can better reflect the effect of model adjustment.

[0030] Furthermore, the step (5) includes the following steps:

[0031] S51. During the training iteration process, monitor the changes in the loss function during learning. When the loss function converges to below 0.1 and the reward function converges to the optimal solution, stop the iteration and proceed to the next step of judgment. The loss function uses the policy gradient method to maximize the cumulative reward. The specific formula is:

[0032]

[0033] Among them, L(θ) is the loss function of the model, the trajectory τ is the state sequence experienced by the agent when performing a series of actions in the environment, and R(τ) is the cumulative reward of the agent on the trajectory τ, which represents the total reward obtained by the agent after performing a series of actions. represents the gradient with respect to the parameter θ, and π(a|s;θ) is the action probability distribution output by the agent's policy network, representing the probability of selecting each possible action a given a state s. The optimal decision-making strategy is achieved by guiding the agent to make appropriate action choices in different states.

[0034] If the termination condition is not met, the experience replay buffer is cleared and the interaction between the agent and the environment is restarted to collect new experience data and update the agent's policy network parameters. The training process is repeated until the termination condition is met.

[0035] Furthermore, the step (6) includes the following steps:

[0036] The city skyline scenario optimization solution that meets the termination conditions in step (5) is saved, and the three-dimensional building model of the optimization solution is output to a holographic sandbox that stores a lightweight real-life model of the current city target area skyline. The three-dimensional models of the current city skyline and the city skyline optimization solution are displayed separately, and the characteristic parameters of the current city skyline and the optimization solution model are displayed in a table. The user compares the current city skyline and the city skyline optimization solution through MR mixed reality glasses, and performs secondary precise adjustments to the city skyline scenario optimization solution model through a handheld touch device.

[0037] Beneficial effects of the present invention:

[0038] 1. This invention solves the skyline optimization problem based on user operation datasets by applying a multi-agent deep reinforcement learning method. It can quickly and automatically generate city skyline optimization plans that meet user needs, thus realizing the automation and intelligence of skyline plan optimization, reducing the workload of planners in manually adjusting skyline plans, and improving the efficiency of urban design work.

[0039] 2. This invention records and translates the user's behavior and instruction data for adjusting the skyline model one by one, combines the parameter changes of the model before and after adjustment with the characteristic parameters of the adjusted model, and converts the random user behavior into accurate characteristic parameters. This avoids the problems of inconsistent judgment standards and high arbitrariness in traditional skyline optimization work, and improves the scientific nature and accuracy of skyline optimization work.

[0040] 3. This invention uses MR mixed reality glasses to display a real-life model of the city skyline to users, which can more accurately restore the visual effects of the skyline optimization plan. Users can adjust the optimization plan through a handheld touch device in all-round real-time perception and view a real-time comparison of the skyline model's characteristic parameters before and after the plan adjustment, thereby improving the immediacy of the plan adjustment. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 Flow chart of the method of the present invention;

[0042] Figure 2 Schematic diagram of touch interaction according to the present invention. DETAILED DESCRIPTION

[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0044] A method for optimizing the visual touch interaction of real-life city skylines based on multi-scenario simulation, such as Figure 1-2 As shown, the following steps are included:

[0045] First, drones equipped with multi-lens oblique cameras and GPS positioning modules will acquire oblique photography data of buildings in target urban areas, along with geographic coordinates. Vehicle-mounted and handheld LiDAR scanners equipped with GPS positioning modules will also acquire 3D morphological data of buildings in target urban areas, along with geographic coordinates. The oblique photography data will be aligned and fused with the laser scanning point cloud data to obtain 3D real-world data of the target urban area.

[0046] The building oblique photography data includes a top image and a facade image of the building.

[0047] The three-dimensional building shape data includes point cloud data of building height and building outline.

[0048] 2. Input the 3D real-life data of the urban target area obtained in step 1 into the EasyV digital twin visualization platform to reconstruct a 3D model of the urban target area. The 3D model of the urban target area is singulated using an image recognition algorithm based on a convolutional neural network to generate a lightweight real-life model of the current urban target area skyline that can adjust each building separately. The model is then rendered and displayed using the EasyTwi n digital twin fusion rendering engine.

[0049] 3. In the EasyV digital twin visualization platform, a pre-set target detection algorithm is used to further divide building units into near-ground buildings, landmark high-rise buildings, and matrix buildings. The building height, facade area, and base area values ​​are recorded and linked to the corresponding attribute table. MR mixed reality glasses equipped with an eye movement detection module and a handheld touch controller equipped with an inertial sensor are connected to the EasyV digital twin visualization platform to record the user's touch-sensitive behavior and command information when adjusting the skyline model. After the user completes the touch adjustment of the skyline model of the target urban area, the platform translates the user's behavior and command information one by one and summarizes them into the monitoring data of the device during the adjustment, the parameter changes of the model before and after the adjustment, and the characteristic parameters of the adjusted model.

[0050] The optional adjustment settings of the command information include: "remove", "add", "increase height", "reduce height", and "indent facade along the vertical line of the building" for buildings close to the ground, landmark high-rise buildings, and base buildings.

[0051] The preset target detection algorithm performs division based on the geometric position, density, normal features of the building point cloud clusters and their position from the skyline model viewpoint.

[0052] Fourth, we used multi-agent deep reinforcement learning (RIAL) to solve the problem of optimizing a lightweight, realistic skyline model. We defined each building unit—near-ground buildings, landmark high-rise buildings, and matrix buildings—as an agent. At each time t, we observed the local state O of the environmental system as the characteristic parameters of other agents within a 1km buffer zone around the geometric center of gravity of the agent. We then sent an action from our action set A to the environment. We imported the user action training data recorded in step 3 into a cloud computing platform with 10 TFLOPS of computing power, and fitted a reward function to evaluate the current state and the touch adjustments to the skyline model.

[0053] The actions include: "removal", "addition", "height increase", "height reduction", and "facade indentation along the vertical line of the building" of the near-ground building, the landmark high-rise building and the base building;

[0054] The reward function is defined by integrating the parameter changes of the model before and after adjustment and the similarity between the parameters of the adjusted model and the user training data. The specific formula is:

[0055] R=λ·MinR1+μ·MaxR2

[0056] Among them, MinR1 minimizes the parameter changes of the model before and after adjustment, and MaxR2 maximizes the similarity between the parameters of the adjusted model and the user training data; λ and μ are adjustment parameters used to balance the impact of changes in the two parameters, so that the reward function can better reflect the effect of model adjustment.

[0057] During training iterations, monitor the changes in the loss function. When the loss function converges to below 0.1 and the reward function converges to the optimal solution, stop the iteration and proceed to the next step. If the termination condition is not met, clear the experience replay buffer, restart the interaction between the agent and the environment, collect new experience data, and update the agent's policy network parameters. Repeat the training process until the termination condition is met.

[0058] The loss function uses the policy gradient method to maximize the cumulative reward. The specific formula is:

[0059]

[0060] Among them, L(θ) is the loss function of the model, the trajectory τ is the state sequence experienced by the agent when performing a series of actions in the environment, and R(τ) is the cumulative reward of the agent on the trajectory τ, which represents the total reward obtained by the agent after performing a series of actions. represents the gradient with respect to the parameter θ, and π(a|s;θ) is the action probability distribution output by the agent's policy network, representing the probability of selecting each possible action a given a state s. The optimal decision-making strategy is achieved by guiding the agent to make appropriate action choices in different states.

[0061] 6. Save the city skyline scenario optimization plan that meets the termination criteria from step 5 and export the optimized building 3D model to a holographic sandbox containing a lightweight, real-world model of the existing city skyline and the optimized skyline plan. The 3D models of the existing city skyline and the optimized skyline plan are displayed separately, with the characteristic parameters of the existing and optimized city skyline models presented in a table. The user compares the existing city skyline with the optimized skyline plan using MR mixed reality glasses and makes secondary, precise adjustments to the optimized skyline plan model using a handheld touchpad.

[0062] Example

[0063] The technical solution of the present invention will be described in detail below by taking a certain urban area as an example.

[0064] (1) Urban 3D real-scene oblique photography and scanning data acquisition, specifically including:

[0065] (1.1) Obtain oblique photography data of buildings in urban target areas with geographic coordinate information using a drone equipped with a multi-lens oblique camera and a GPS positioning module, including top and facade images of the buildings.

[0066] (1.2) Using a vehicle-mounted LiDAR scanner and a handheld LiDAR scanner equipped with a GPS positioning module, obtain the three-dimensional morphological data of buildings in the urban target area with geographic coordinate information, including point cloud data of building heights and building outlines.

[0067] (1.3) Align and fuse the oblique photography image data with the laser scanning point cloud data to obtain three-dimensional real-life data of the urban target area.

[0068] (2) Generation of lightweight real-life model of city skyline, including:

[0069] (2.1) Input the three-dimensional real scene data of the urban target area obtained in step (1) into the EasyV digital twin visualization platform to reconstruct the three-dimensional model of the urban target area;

[0070] (2.2) The three-dimensional model of the urban target area is processed individually through an image recognition algorithm based on a convolutional neural network. A lightweight real-life model of the current urban target area skyline is generated, which can be adjusted for each building separately. The model is then rendered and displayed using the EasyTwi n digital twin fusion rendering engine.

[0071] (3) User touch adjustment and preference training data output, specifically including:

[0072] (3.1) In the EasyV digital twin visualization platform, a preset target detection algorithm is used to further divide building units into near-ground buildings, landmark high-rise buildings, and matrix buildings, and their building heights, facade areas, and base area values ​​are recorded and linked to the corresponding attribute tables; the preset target detection algorithm is based on the geometric position, density, normal characteristics of the building unit point cloud clusters and their position from the skyline model viewpoint.

[0073] (3.2) Connecting MR mixed reality glasses equipped with an eye movement detection module and a handheld touch controller equipped with an inertial sensor to the EasyV digital twin visualization platform to record user touch actions and command information for adjusting the skyline model. The available adjustment settings include: "remove," "add," "increase," "decrease," and "indent" the facade along the vertical line of the building for near-ground buildings, landmark high-rise buildings, and base buildings.

[0074] (3.3) After the user completes the touch adjustment of the skyline model of the target urban area, the platform interprets the user's actions and instructions one by one and summarizes them into the monitoring data of the device during the adjustment, the parameter changes of the model before and after the adjustment, and the characteristic parameters of the adjusted model, which specifically include:

[0075]

[0076]

[0077] (4) Multi-agent deep reinforcement learning simulates the skyline adjustment process, specifically including:

[0078] (4.1) Multi-agent deep reinforcement learning (RIAL) is used to solve the skyline lightweight real-world model optimization problem. Each building in the near-ground building, landmark high-rise building, and matrix building is defined as an agent. At each time t, the local state O of the observed environmental system is the characteristic parameters of other agents within a 1km buffer zone around the geometric center of gravity of the agent. The agent then sends an action from the action set A to the environment, specifically including "remove", "add", "increase", "decrease", and "retract the facade along the vertical line of the building" for the near-ground building, landmark high-rise building, and matrix building.

[0079] (4.2) On a cloud computing platform with 10 TFLOPS of computing power, import the user operation training data recorded in step 3 and fit the reward function to evaluate the current state and the touch adjustment results of the skyline model. This is defined by integrating the parameter changes of the model before and after adjustment and the similarity between the parameters of the adjusted model and the user training data. The specific formula is:

[0080] R=λ·MinR1+μ·MaxR2

[0081] Among them, MinR1 minimizes the parameter changes of the model before and after adjustment, and MaxR2 maximizes the similarity between the parameters of the adjusted model and the user training data; λ and μ are adjustment parameters used to balance the impact of changes in the two parameters, so that the reward function can better reflect the effect of model adjustment.

[0082] (5) Iterative training of the city skyline optimization solution, including:

[0083] (5.1) During the training iteration process, the changes in the loss function during learning are monitored. When the loss function converges to below 0.1 and the reward function converges to the optimal solution, the iteration is stopped and the next step is judged. The loss function uses the policy gradient method to maximize the cumulative reward. The specific formula is:

[0084]

[0085] Among them, L(θ) is the loss function of the model, the trajectory τ is the state sequence experienced by the agent when performing a series of actions in the environment, and R(τ) is the cumulative reward of the agent on the trajectory τ, which represents the total reward obtained by the agent after performing a series of actions. represents the gradient with respect to the parameter θ, and π(a|s;θ) is the action probability distribution output by the agent's policy network, representing the probability of selecting each possible action a given a state s. The optimal decision-making strategy is achieved by guiding the agent to make appropriate action choices in different states.

[0086] (5.2) If the termination condition is not met, the experience replay buffer is cleared and the agent-environment interaction is restarted to collect new experience data and update the agent's policy network parameters. The training process is repeated until the termination condition is met.

[0087] (6) Output of the three-dimensional model of the city skyline optimization plan, including:

[0088] (6.1) Save the city skyline scenario optimization plan that meets the termination conditions in step 5, and output the three-dimensional building model of the optimization plan to a holographic sand table that stores a lightweight real-life model of the current city target area skyline, displaying the three-dimensional models of the current city skyline and the city skyline optimization plan respectively, and displaying the characteristic parameters of the current city model and the optimization plan model in a table.

[0089] (6.2) The user compares the current city skyline with the optimized city skyline scenario through MR mixed reality glasses, and uses a handheld touch device to make secondary precise adjustments to the city skyline scenario optimization scenario model.

Claims

1. A method for optimizing the visual touch interaction of real-life city skylines based on multi-scenario simulation, characterized in that: include: Step (1): using a drone equipped with a multi-lens oblique camera and a GPS positioning module to collect oblique photography data of buildings in the urban target area, using a lidar scanner to collect three-dimensional morphological data of buildings in the urban target area, and obtaining three-dimensional real scene data of the urban target area through data registration and fusion; Step (2): Upload the real scene data to the digital twin visualization platform, reconstruct the three-dimensional model of the urban target area, perform monomer processing on the three-dimensional model of the urban target area based on the convolutional neural network, and use the digital twin fusion rendering engine to reconstruct a lightweight real scene model of the urban target area skyline; Step (3): The lightweight real-life model of the urban target area skyline is further divided into near-ground buildings, landmark high-rise buildings, and matrix buildings based on the geometric position, density, normal characteristics of the point cloud clusters and their position from the skyline model viewpoint; the mixed reality glasses equipped with an eye movement detection module and the handheld touch device equipped with an inertial sensor are connected to the digital twin visualization platform to record the user's touch-sensitive behavior and command information for adjusting the skyline model; the optional adjustment settings of the command information include: "remove", "add", "increase height", "decrease height", and "indent facade along the vertical line of the building" for near-ground buildings, landmark high-rise buildings, and matrix buildings; the user behavior and command information are translated one by one and summarized as the monitoring data of the equipment during adjustment, the parameter changes of the model before and after adjustment, and the characteristic parameters of the model after adjustment; Step (4): Multi-agent deep reinforcement learning is used to solve the skyline optimization problem of the lightweight real-life model of the urban target area skyline. Each building unit in the near-ground building, landmark high-rise building and matrix building is defined as an agent based on the monitoring data of the equipment during adjustment, and the action and reward function of the agent are defined; Step (5): During the training iteration, record the changes in the loss function. When both the loss function converges to below 0.1 and the reward function converges to the optimal solution, stop the iteration and output the city skyline optimization solution. Otherwise, the experience replay buffer is cleared, the interaction between the agent and the environment is restarted, and the training process is repeated until the termination condition is met; Step (6): Save the output plan, output it to the holographic sandbox that stores the current lightweight real-life model of the city's target area skyline, run it for display, compare the city skyline optimization plan with the current skyline, and make secondary interactive precise adjustments; The step (4) comprises the following steps: S41. Use multi-agent deep reinforcement learning to solve the skyline lightweight real-world model optimization problem. Define each building unit among the near-ground buildings, landmark high-rise buildings, and matrix buildings as an agent. At each time t, observe the local state O of the environmental system as the characteristic parameters of other agents within a 1km buffer zone around the geometric center of gravity of the agent, and send an action from the action set A to the environment. The actions include: "remove", "add", "increase height", "decrease height", and "retract facade along the vertical line of the building" for the near-ground buildings, landmark high-rise buildings, and matrix buildings. S42. In a cloud computing platform with a computing power of 10 TFLOPS, import the user operation training data recorded in step (3), and fit the reward function to evaluate the current state and the touch adjustment result of the skyline model; The reward function is defined by integrating the parameter changes of the model before and after adjustment and the similarity between the parameters of the adjusted model and the user training data. The specific formula is: R=λ·MinR1+μ·MaxR2 Among them, MinR1 minimizes the parameter changes of the model before and after adjustment, and MaxR2 maximizes the similarity between the parameters of the adjusted model and the user training data; λ and μ are adjustment parameters used to balance the impact of the changes in the two parameters so that the reward function can better reflect the effect of the model adjustment.

2. The method for optimizing the city skyline real-scene visual touch interaction based on multi-scenario simulation according to claim 1 is characterized in that: The step (1) comprises the following steps: Oblique photography data of buildings in the urban target area with geographic coordinate information is obtained by using a drone equipped with a multi-eye oblique camera and a GPS positioning module, including the top image and facade image of the building; three-dimensional morphological data of buildings in the urban target area with geographic coordinate information is obtained by using a vehicle-mounted lidar scanner and a handheld lidar scanner with a GPS positioning module, including point cloud data of building height and building outline; the oblique photography image data is aligned and fused with the laser scanning point cloud data to obtain three-dimensional real-scene data of the urban target area.

3. The method for optimizing the city skyline real-scene visual touch interaction based on multi-scenario simulation according to claim 2 is characterized in that: The step (2) comprises the following steps: The three-dimensional real scene data of the urban target area obtained in step (1) is input into the EasyV digital twin visualization platform to reconstruct the three-dimensional model of the urban target area. The three-dimensional model of the urban target area is singulated by the image recognition algorithm based on the convolutional neural network to generate a lightweight real scene model of the current urban target area skyline that can adjust each building separately. The model is then rendered and displayed by the EasyTwin digital twin fusion rendering engine.

4. The method for optimizing the city skyline real-scene visual touch interaction based on multi-scenario simulation according to claim 3 is characterized in that: The step (3) comprises the following steps: S31. In the EasyV digital twin visualization platform, use the preset target detection algorithm to further divide the building units into near-ground buildings, landmark high-rise buildings, and matrix buildings, and record their building heights, facade areas, and base area values; The preset target detection algorithm is divided based on the geometric position, density, normal characteristics of the building point cloud clusters and their position from the skyline model viewpoint; S32. Connect the MR mixed reality glasses equipped with an eye movement detection module and the handheld touch controller equipped with an inertial sensor to the EasyV digital twin visualization platform to record the user's touch actions and command information for adjusting the skyline model; S33. After the user completes the touch adjustment of the skyline model of the target urban area, the platform translates the user behavior and instruction information one by one, and summarizes it into the monitoring data of the device during the adjustment, the parameter changes of the model before and after the adjustment, and the characteristic parameters of the model after the adjustment.

5. The method for optimizing the city skyline real-scene visual touch interaction based on multi-scenario simulation according to claim 4 is characterized in that: The step (5) comprises the following steps: S51. During the training iteration process, monitor the changes in the loss function during learning. When the loss function converges to below 0.1 and the reward function converges to the optimal solution, stop the iteration and proceed to the next step of judgment. The loss function uses the policy gradient method to maximize the cumulative reward. The specific formula is: Among them, L(θ) is the loss function of the model, the trajectory τ is the state sequence experienced by the agent when performing a series of actions in the environment, and R(τ) is the cumulative reward of the agent on the trajectory τ, which represents the total reward obtained by the agent after performing a series of actions. represents the gradient of the parameter θ, and π(a|s;θ) is the action probability distribution output by the agent's policy network, which represents the probability of selecting each possible action a under a given state s. The optimal decision-making strategy is achieved by guiding the agent to make appropriate action choices under different states. S52. If the termination condition is not met, clear the experience replay buffer, restart the interaction between the agent and the environment, collect new experience data, and update the policy network parameters of the agent; the training process is executed in a loop until the termination condition is met.

6. The method for optimizing the city skyline real-scene visual touch interaction based on multi-scenario simulation according to claim 5 is characterized in that: The step (6) comprises the following steps: The city skyline scenario optimization scheme that meets the termination conditions in step (5) is saved, and the three-dimensional building model of the optimization scheme is output to a holographic sandbox storing a lightweight real-life model of the current city target area skyline, and the three-dimensional models of the current city skyline and the city skyline optimization scheme are displayed respectively, and the characteristic parameters of the models of the current city and the optimization scheme are displayed in a table; the user compares the current city skyline and the city skyline optimization scheme through MR mixed reality glasses, and performs secondary precise adjustments to the city skyline scenario optimization scheme model through a handheld touch device.

Citation Information

Patent Citations

  • Building design scene automatic generation method and system based on artificial intelligence

    CN118940364A

  • Dynamic interactive simulation method for recognition and planning of urban viewing corridor

    US20220309200A1