Urban sustainable development planning decision optimization system based on data driving
By designing a data-driven urban sustainable development planning decision optimization system, the shortcomings of multi-source data integration, dynamic optimization and multi-objective coordination in the existing technology are solved, and efficient urban resource management and sustainable development planning are achieved.
Patent Information
- Application Number
- CN202510044453.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-11
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing urban planning system has shortcomings in multi-source data integration, dynamic optimization, real-time feedback and multi-objective coordination, resulting in bias in optimization results and insufficiency in the system.
A data-driven urban sustainable development planning decision optimization system is designed, and through collaborative work of multiple modules, a full-process closed-loop dynamic optimization from data acquisition to strategy optimization and feedback is realized. The system includes data acquisition and preprocessing, differential dynamic game modeling, multi-agent reinforcement learning optimization and system feedback modules.
It realizes accurate integration and dynamic optimization of multi-source data, improves the accuracy and applicability of the optimization model, can quickly respond to changes in urban operations, and achieves multi-party interest balance and efficient resource allocation.
Smart Images

Figure CN120013271A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of smart city planning and management, and in particular to a data-driven urban sustainable development planning decision optimization system. Background Art
[0002] With the rapid advancement of urbanization, urban planning plays an important role in improving resource utilization efficiency, alleviating environmental pressure, and promoting social equity. Especially under the framework of smart cities and sustainable development, how to dynamically optimize urban operation based on large-scale, multi-dimensional data has become a major technical challenge in global urban management and planning. At present, urban planning and decision-making optimization systems have been widely studied and applied, but there are still technical bottlenecks in many aspects.
[0003] In the existing technology, most urban planning systems rely on traditional data analysis and optimization models. These systems are usually based on a single data source, such as traffic flow or energy consumption, and ignore the correlation and spatiotemporal dynamic characteristics between multi-source data. At the same time, the collection and integration of multi-source data are often limited by the diversity of data collection equipment, the inconsistency of collection frequency, and the inconsistency of data format. These problems make it impossible for existing systems to fully and dynamically describe the complex state of urban operation, which restricts the accuracy and applicability of optimization models.
[0004] In addition, existing urban planning technologies mainly use static optimization methods, such as linear programming, genetic algorithms, or particle swarm optimization algorithms. These methods have certain application value in solving single-objective optimization problems, but their limitations gradually emerge in complex dynamic scenarios in urban planning. Static optimization models are usually based on preset parameters and lack the ability to effectively handle dynamic multi-objective conflicts. For example, the government, enterprises, and the public have different or even conflicting goals in terms of resource allocation, environmental protection, etc. Traditional optimization methods are difficult to achieve a dynamic balance between the interests of multiple parties, and the optimization results show deviations in practical applications.
[0005] At the same time, the lack of real-time response and feedback mechanisms in existing technologies further limits the effectiveness of urban planning systems. Traditional systems are mostly based on offline modeling and historical data analysis, and optimization schemes cannot quickly respond to real-time changes in urban operating conditions. Especially in emergencies such as traffic congestion and fluctuations in energy demand, existing technologies lack the ability to adjust dynamically, resulting in delayed optimization results and reducing the practicality of the system. In addition, the lack of a closed-loop feedback mechanism makes it difficult to accurately monitor and adjust the actual effects of the optimization strategy after execution, and the adaptability and reliability of the model are significantly reduced.
[0006] Therefore, the present invention proposes a data-driven urban sustainable development planning decision optimization system to address the deficiencies of the prior art. Summary of the invention
[0007] In view of the problems of the existing technology, such as insufficient multi-source data integration capability, lack of dynamic optimization capability, delayed real-time feedback response, and insufficient multi-objective coordination capability, this paper proposes a data-driven urban sustainable development planning decision optimization system. Through the collaborative work of multiple modules, the system realizes the closed-loop dynamic optimization of the entire process from data collection to strategy optimization and then to execution feedback, thereby achieving a balance of interests among multiple parties and efficient allocation of resources.
[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions: a data-driven urban sustainable development planning decision optimization system, the system comprising:
[0009] The data collection and preprocessing module is used to collect real-time multi-source urban operation data, including traffic flow, energy consumption, air quality, and land utilization rate, and to perform data cleaning, dimension reduction, and formatting to generate urban state variables;
[0010] The differential dynamic game modeling module is used to construct the state dynamic evolution equation, instant benefit function and cumulative utility function based on the city state variables and multi-party control strategy variables, and establish a dynamic game model involving multiple parties;
[0011] The decision-making optimization module is used to optimize the high-dimensional strategy space of the dynamic game model based on multi-agent reinforcement learning technology to solve the optimal dynamic strategy of each participant;
[0012] The system feedback module is used to apply the optimization strategy to urban operations, monitor the optimization results in real time, and dynamically update the game model and optimization strategy based on the monitoring data.
[0013] Preferably, the data acquisition and preprocessing module includes:
[0014] Sensor networks and IoT devices to collect dynamic data such as traffic flow, energy consumption, and pollutant emissions;
[0015] Data fusion unit, used to uniformly transmit multi-source data to the data center;
[0016] The data preprocessing unit is used to perform data cleaning to remove outliers, data dimension reduction to simplify data complexity, and formatting to generate structured urban status data for modeling.
[0017] Preferably, the differential dynamic game modeling module is used to establish a game model involving multiple parties, wherein:
[0018] Urban state variables are used to describe the urban operation status of traffic flow, energy consumption, and air quality;
[0019] Multiple parties include the government, enterprises and the public, who influence the city state by controlling variables;
[0020] The state dynamics evolution equation is used to describe the changes of urban state variables over time, and its changes are jointly determined by the control variables and the internal dynamics of the system;
[0021] The immediate benefit function is used to reflect the interest objectives of each participant in the current state, including the positive contribution of the state to utility and the negative impact of control variables.
[0022] Preferably, the cumulative utility function is used to describe the target optimization results of the participants in the entire planning period, and is the integration result of the immediate benefit function over time. The immediate benefit function includes state utility, control variable cost and the joint utility of state and control variables.
[0023] Preferably, the optimization decision module is based on multi-agent reinforcement learning technology and includes:
[0024] Construct a reinforcement learning framework, model the game participants as intelligent agents, use the city state variables as the input of the intelligent agents, and use the control variables as the output of the intelligent agents;
[0025] Define the immediate benefit as the agent's reward function, and use the participant's immediate benefit function as part of the reinforcement learning reward function;
[0026] The strategies of intelligent agents are trained through deep reinforcement learning algorithms, so that the strategies of each intelligent agent gradually converge to a dynamic Nash equilibrium.
[0027] Preferably, the optimization decision module uses a deep Q-learning algorithm or a policy gradient optimization method to optimize the agent strategy, and a joint training method based on reinforcement learning enables multiple agents to alternately update strategies while sharing state information, and finally converge to the Nash equilibrium solution of the dynamic game.
[0028] Preferably, the system feedback module includes:
[0029] The optimization strategy execution unit is used to execute the control measures of urban operation according to the optimal strategy generated by the optimization decision module;
[0030] Feedback monitoring unit, used to monitor the changes in urban status after the optimization strategy is executed in real time;
[0031] The model updating unit is used to adjust the state evolution equation and profit function in the dynamic game model based on feedback data.
[0032] Preferably, the model updating unit dynamically adjusts the game model based on the feedback monitoring data, and updates the state dynamic evolution equation parameters and the weight factor of the instant benefit function in the game model to adapt to the real-time changing city operation status.
[0033] Preferably, the optimization decision module is able to make trade-offs for multi-objective conflict issues, including the government's social equity goal, the enterprise's profit maximization goal, and the public's quality of life improvement goal, and optimize the balance between the goals through the cumulative utility function in the dynamic game model.
[0034] Preferably, a data-driven urban sustainable development planning decision optimization method comprises the following steps:
[0035] S1. Data collection and preprocessing:
[0036] Collect multi-source data on the city’s operating status, including traffic flow, energy consumption, air quality, and land use, through urban sensor networks and IoT devices;
[0037] The collected multi-source data are cleaned to remove outliers, dimension reduction is performed to simplify data complexity, and formatting is performed to generate structured urban state variables for modeling;
[0038] S2. Construction of dynamic game model:
[0039] Based on the city state variables and multi-party control variables, the city state dynamic evolution equation is established to describe the dynamic changes of the city state over time;
[0040] Define an immediate benefit function to reflect the benefits and costs of urban states and control strategies to multiple parties;
[0041] Construct a cumulative utility function to describe the optimization objectives of multiple parties during the entire planning period and establish a differential dynamic game model involving multiple parties;
[0042] S3. Optimization strategy solution:
[0043] The differential dynamic game model is transformed into a multi-agent reinforcement learning problem, and each participant is modeled as an agent;
[0044] The city state variable is used as the input of the intelligent agent, the control variable is used as the output of the intelligent agent, and the immediate benefit function is used as the reward of the intelligent agent;
[0045] Based on the deep reinforcement learning algorithm, the agent strategy is trained so that each agent's strategy gradually converges to the Nash equilibrium of the dynamic game;
[0046] S4. Optimize strategy execution and feedback:
[0047] Control the city's operating status according to optimization strategies, including adjusting traffic signal cycles, optimizing energy distribution and implementing pollution control measures;
[0048] Monitor changes in city status after the optimization strategy is implemented, and verify the optimization effect through feedback data;
[0049] The state evolution equation and profit function in the dynamic game model are adjusted according to the feedback data, and the optimization strategy is updated to adapt to the real-time changes in urban operations.
[0050] The present invention provides a data-driven urban sustainable development planning decision optimization system, which has the following beneficial effects:
[0051] 1. The present invention adopts a technical solution based on multi-source data collection and differential dynamic game modeling, which can accurately integrate multi-dimensional data such as traffic flow, energy consumption, air quality and land utilization rate, and achieve the technical effect of accurately describing the dynamic operation status of the city and capturing the interactive relationship of multiple interests. Compared with the technical solutions of the prior art with a single data source, lack of dynamic description and difficulty in modeling complex multi-party relationships, the present invention solves the problem of insufficient data dimension coverage and lack of dynamicity caused by static model, and provides a clear and rigorous mathematical foundation for subsequent optimization decisions.
[0052] 2. The present invention significantly improves the optimization efficiency for high-dimensional control strategy space by combining multi-agent reinforcement learning and dynamic game solving methods, achieving the technical effect of quickly solving dynamic equilibrium solutions and realizing real-time adjustment of strategies. Compared with the technical solutions in the prior art that use traditional genetic algorithms or particle swarm optimization methods and cannot cope with multi-party strategy interactions in dynamic games, the present invention overcomes the bottleneck of algorithm computational complexity in high-dimensional strategy scenarios and the problem of poor adaptability to dynamic multi-objective conflicts, and can also achieve continuous optimization decisions of all participants in a changing environment.
[0053] 3. The present invention adopts a closed-loop feedback control mechanism, which significantly improves the adaptability and optimization ability of the system by real-time monitoring of the execution effect of the optimization strategy and dynamically adjusting the game model, achieving the technical effect of quickly responding to changes in urban operations and improving the actual optimization effect. Compared with the technical solutions in the prior art that lack real-time feedback, lag in optimization results and are difficult to adjust, the present invention solves the problem of separation between feedback data and model updates, and can continuously optimize strategy generation by dynamically updating evolution equations and weight parameters to ensure the system's agile response to emergencies (such as traffic jams or abnormal energy demand).
[0054] 4. The present invention constructs a complete system integrating multi-source data drive, differential dynamic game modeling, reinforcement learning strategy optimization and feedback adjustment, achieving the technical effect of comprehensive coordination of multi-objective conflicts in the city and dynamic optimization of resource allocation. Compared with the technical solutions in the prior art that rely on single-objective optimization or independent operation of each module, the present invention solves the problem of being unable to balance the multi-objective needs of economic benefits, social equity and environmental protection, and at the same time realizes the global optimization of complex urban systems through the collaborative operation of modules, providing an innovative technical path and a more efficient solution for realizing urban sustainable development planning. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 Optimize the system architecture for data-driven urban sustainable development planning decisions;
[0056] Figure 2 Flowchart of the optimization method for data-driven urban sustainable development planning decision-making. DETAILED DESCRIPTION
[0057] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0058] Please refer to the attached Figure 1 The embodiment of the present invention provides a data-driven urban sustainable development planning decision optimization system. The various modules of the system of the present invention are described in detail below.
[0059] The data collection and preprocessing module is used to collect real-time multi-source urban operation data, including traffic flow, energy consumption, air quality, and land utilization rate, and to perform data cleaning, dimension reduction, and formatting to generate urban state variables.
[0060] First, the main goal of this step is to provide basic data support for subsequent urban planning optimization and ensure the accuracy and availability of the data. The optimization system of the present invention needs to obtain multi-source data of the city's operating status in real time and perform necessary cleaning and preprocessing operations on these data. The collection and processing of these data directly determines the construction effect of the subsequent dynamic game model and the accuracy of the optimization algorithm. Therefore, this step needs to pay special attention to the integrity, authenticity and adaptability of the data processing method.
[0061] Generally speaking, urban operation data comes from a wide range of sources, including but not limited to traffic flow monitoring equipment, air quality monitoring stations, energy consumption meters, remote sensing data collection devices, and Internet of Things (IoT) sensor networks. These data usually have problems such as inconsistent formats, noise interference, and different collection frequencies. Therefore, in order to ensure the effectiveness of subsequent analysis, a set of standardized pre-processing processes is required.
[0062] Specifically, in this embodiment, the data acquisition module obtains multi-source dynamic data of the city through sensor networks and Internet of Things devices. Among them, traffic flow data is provided by roadside cameras, signal light detectors and vehicle networking devices; air quality data is collected by distributed air monitoring stations; energy consumption data is obtained in real time through smart meters of the power supply system; and land utilization rate data can be extracted through high-resolution remote sensing satellite imaging.
[0063] As an option, an edge computing architecture can be used for data transmission and integration. In this implementation, data is pre-processed and formatted by distributed computing nodes before being transmitted to a central data center. This approach can significantly reduce data transmission latency while reducing bandwidth usage.
[0064] In general, the collected raw data may contain outliers, missing values, and redundant information. In this embodiment, the data cleaning module performs preliminary processing on the raw data. Specifically, outliers are detected and eliminated by statistical methods, such as using the 3σ principle to identify traffic flow data or air pollutant concentration data that are beyond a reasonable range. In some embodiments, regression prediction can be performed in combination with historical data to complete missing values. For example, energy consumption data from the same time period in history is used to construct a time series model to predict missing values.
[0065] In one possible implementation, in order to reduce data dimensions and improve modeling efficiency, the principal component analysis (PCA) method is used to reduce the data dimension. Through this method, multiple related variables such as traffic flow, energy consumption, and air pollution can be integrated into a few principal components to facilitate the subsequent construction and optimization calculation of the dynamic game model.
[0066] In one possible implementation, in order to reduce data dimensions and improve modeling efficiency, the principal component analysis (PCA) method is used to reduce the data dimension. Through this method, multiple related variables such as traffic flow, energy consumption, and air pollution can be integrated into a few principal components to facilitate the subsequent construction and optimization calculation of the dynamic game model.
[0067] In addition, formatting is an important part of data preprocessing. In this embodiment, all cleaned and dimension-reduced data are uniformly converted into standardized structured data to form state variables x(t). These state variables include but are not limited to:
[0068] x1(t): traffic flow (unit: vehicles / hour), describing the real-time traffic conditions of the urban road network;
[0069] x2(t): energy consumption (unit: kilowatt-hour), reflecting the current energy usage such as electricity and gas;
[0070] x3(t): Air quality index (AQI), which measures the degree of environmental pollution;
[0071] x4(t): Land utilization rate (unit: percentage), used to describe the utilization of urban space resources.
[0072] Generally speaking, the units and ranges of state variables need to be consistent with the input requirements of the dynamic game model to ensure that the physical meaning of the values is clear and conducive to the stability of the model calculation.
[0073] In this embodiment, in order to meet the real-time requirements of the subsequent model, a dynamic data update mechanism based on stream processing is also implemented. Specifically, the cleaned data will be regularly updated to the state variable data pool at fixed time intervals (such as every minute). Through this real-time dynamic update method, the subsequent optimization module can obtain the latest city operation status, thereby improving the timeliness of strategy optimization.
[0074] In order to further improve the reliability of data collection and processing, blockchain technology can also be introduced in some embodiments to encrypt and verify the collected data. The source, timestamp and integrity check information of the data are recorded in a distributed ledger to ensure the tamper-proof nature of the data during transmission and storage.
[0075] In summary, this step forms state variables that meet the input requirements of the dynamic game model through multi-source data collection, cleaning, dimensionality reduction and formatting, providing an accurate and efficient data basis for the subsequent optimization process. The clear design of the formula part and data features ensures the rigor and reproducibility of the technology of the present invention.
[0076] Differential dynamic game modeling module, used to construct state dynamic evolution equations, instant benefit functions and cumulative utility functions based on city state variables and multi-party control strategy variables, and to establish a dynamic game model involving multiple parties
[0077] After data collection and preprocessing, the city state variables have been standardized into structured data. These state variables are directly input into the differential dynamic game modeling module to construct the dynamic evolution process of the urban system and the immediate benefits and objective functions of each participant. The core of this module is to describe the dynamic interaction relationship between the parties in the operation of the city through mathematical modeling, clarify the impact of their control strategies on the overall system state, and provide a rigorous foundation for subsequent optimization solutions. In this embodiment, the differential dynamic game modeling module is based on the quintuple Γ = (N, x(t), u i (t),f(x,u),J i ) constructs a game model of the urban system. Here, N represents the set of participants, including the government, enterprises, and the public; x(t) is the urban state variable, which is an abstract description of the system operation state; u i (t) is the control strategy variable of the participant; f(x,u) is the state dynamic evolution equation; J i is the cumulative utility function.
[0078] In general, the city state variable x(t) includes but is not limited to the following dimensions: traffic flow x1(t) (unit: vehicle / hour), energy consumption x2(t) (unit: kilowatt-hour), air quality index x3(t) and land utilization rate x4(t) (unit: percentage). These variables constitute a multidimensional vector that fully reflects the core elements of the city's operating status.
[0079] As an alternative, the strategy variables u of each participant i (t) can be defined according to its role. For example, the government’s policy variables include traffic light cycle adjustment u 信号 (t) and energy tax rate u 税率 (t). The enterprise’s strategic variables may involve production emission control u 排放 (t) and investment allocation u 投资 (t). The public’s strategy variable may be the probability of choosing different travel modes u 出行 (t).
[0080] Specifically, the dynamic evolution of the city state is described by the following differential equation:
[0081]
[0082] in, The time derivative of the state variable reflects the speed at which the system's operating state changes over time. The function f(x,u) describes how the state variable is affected by the control strategy u and the system's internal dynamics. For example:
[0083]
[0084] The traffic flow x1(t) is affected by the traffic light cycle u 信号 (t) and the current traffic density x1(t). Here, β1 and γ1 are the signal control coefficient and the traffic self-enhancement coefficient, respectively.
[0085] In one possible implementation, the instantaneous profit function g i (x,u) is defined as the benefit or cost of each participant in the current state. Its mathematical expression is:
[0086] g i (x,u)=α i f1(x)+β i f2(u)+γ i f3(x,u),
[0087] in:
[0088] f1(x) is the state utility function, which is used to quantify the positive or negative impact of the city state on the participants.
[0089] For example:
[0090]
[0091] It represents the resources consumed by the government in implementing traffic signal control and adjusting tax rates.
[0092] f3(x,u) is the joint utility between the state and the strategy, which is used to represent the interaction between the two.
[0093] In general, the specific form of the immediate benefit function will be adjusted according to the roles and goals of the participants. For example, the immediate benefit function of a company may be more inclined to maximize profits, while the immediate benefit function of the public may focus on improving the quality of life.
[0094] Cumulative utility function J i is the integral of the instantaneous payoff function over the entire planning period:
[0095]
[0096] where t0 and t f are the start and end time of the planning period respectively. The cumulative utility function defines the global optimization goal of the participants and is the core basis for solving the subsequent optimization module.
[0097] In order to adapt to the complex urban environment and real-time changes, in this embodiment, the instantaneous benefit function and state dynamic evolution equation of the participants can be dynamically adjusted according to the feedback data. For example, sudden changes in air quality may trigger the recalibration of relevant parameters in the model to ensure the accuracy of the game model.
[0098] Through the above mathematical modeling, this module establishes a clear and rigorous multi-party dynamic game model, which provides basic support for subsequent optimization. Its core feature is that it not only takes into account the dynamic and multidimensional nature of the urban system, but also reflects the conflict and synergy between the goals of multiple parties. Through this module, the system can capture the dynamic interaction behaviors of all parties in the complex and ever-changing urban operation.
[0099] The decision-making optimization module is used to optimize the high-dimensional strategy space of the dynamic game model based on multi-agent reinforcement learning technology to solve the optimal dynamic strategy of each participant.
[0100] After completing the differential dynamic game modeling, the state dynamic evolution equation and cumulative utility function of the participants have been established. In order to achieve system optimization, it is necessary to solve the optimal strategy of the participants for the game model. This module is based on the multi-agent reinforcement learning method to solve the complex high-dimensional strategy optimization problem in the dynamic game model and realize the dynamic equilibrium strategy generation under multi-party conflict of interest. Closely connected with the aforementioned modeling module, this module directly uses the mathematical description of the game model and combines reinforcement learning technology to complete the optimization task.
[0101] In this embodiment, the optimization decision module models each participant as an intelligent agent. The city state variable x(t) is used as the input of the intelligent agent, and the control strategy u i (t) is the output of the agent, the immediate benefit function g i (x,u) is used to define the reward function of the agent. The goal of each agent is to optimize its strategy through learning so that its cumulative utility function J i To maximize. In general, the multi-agent reinforcement learning framework includes the definition of state space, action space and reward function. Among them, the state space is composed of the city state variable x(t), which reflects the real-time operation status of the city. The action space corresponds to the control strategy of the agent, such as the government's traffic signal regulation or the enterprise's resource allocation decision. The reward function is directly defined by the immediate benefit function of the participants.
[0102] Specifically, this module uses deep reinforcement learning (DRL) technology to optimize the strategy. In some embodiments, the deep QN (DQN) algorithm is used to process discrete strategy problems. DON approximates the action value function Q through a neural network. i (x,u), the update formula is:
[0103]
[0104] Among them, (x,u): action value function, which represents the expected benefit of the agent when it chooses action u in state x; g i(x,u): instant profit function; V i (x′): state value function, which represents the maximum cumulative benefit of the future state x′; γ: discount factor, which is used to balance immediate benefits and long-term benefits.
[0105] In another possible implementation, a policy gradient method (such as PPO or DDPG) is used to optimize the continuous policy. Specifically, through the policy network π i (u i |x) Parameterize the agent’s strategy, the goal is to maximize the expected cumulative reward:
[0106]
[0107] The gradient calculation formula for policy update is:
[0108]
[0109] where θ i are the parameters of the policy network.
[0110] As an option, this module introduces a joint training mechanism to handle the interactive learning problem of multiple agents. By sharing the environment state x(t), all agents train their policy networks simultaneously. In some embodiments, in order to improve the training efficiency and stability, the opponent modeling technique (Opponent Modeling) is used. This method predicts the strategies u of other agents. -i (t), adjust its own strategy to adapt to the overall game environment.
[0111] Specifically, in joint training, the goal of the agent is not only to maximize its own utility function, but also to consider the reactions of other participants. For example, when the public's travel choice strategy changes, the government's traffic control strategy also needs to be adjusted synchronously. This dynamic interaction is achieved through iterative updates during the training process until the strategies of all agents converge to the Nash equilibrium of the dynamic game.
[0112] In one possible implementation, to meet the real-time requirements of urban operations, this module adopts an online learning strategy. Unlike offline training, online learning updates the agent's policy network based on feedback data collected in real time. This approach can quickly respond to emergencies, such as traffic jams or abnormal fluctuations in energy demand.
[0113] In addition, this module also integrates a dynamic constraint processing mechanism. In some embodiments, in order to ensure that the optimization results meet the policy or physical constraints of urban planning, such as carbon emission limits or energy supply caps, the control variable u is i (t) Setting the constraint range By introducing the constrained Lagrange multiplier method, the optimization objective function is redefined as:
[0114] L i (x,u,λ)=J i -λ(h(u)-b)
[0115] Among them, L i (x,u,λ): Lagrangian function; h(u): constraint function, such as total carbon emissions; b: constraint threshold; λ: Lagrangian multiplier.
[0116] In general, the optimization process completes the strategy update while satisfying the constraints to ensure that the optimization results are practical. The output of this module is the optimal strategy of each participant. These strategies will serve as inputs to subsequent feedback modules to guide actual city operation adjustments. Through multi-agent reinforcement learning methods, this module can effectively handle high-dimensional strategy optimization problems in complex game environments, achieve dynamic balance of multi-objective conflicts and real-time generation of global optimal solutions.
[0117] The system feedback module is used to apply the optimization strategy to urban operations, monitor the optimization results in real time, and dynamically update the game model and optimization strategy based on the monitoring data.
[0118] The optimal strategy generated by the optimization decision module needs to be verified and implemented in actual urban operations. The role of the system feedback module is to implement the application of the strategy, monitor its execution effect, and continuously update the game model and optimization strategy through feedback data. Through this closed-loop process, the system can dynamically adapt to emergencies and long-term changes in urban operations and provide support for continuous optimization. This module is closely connected with the aforementioned optimization module, taking the optimization results as input, and further improving the accuracy of the model and the real-time performance of the system through feedback data.
[0119] In this embodiment, the system feedback module first optimizes the strategy Directly applied to specific scenarios of urban operation. As an option, the optimization strategy can be issued to relevant executive agencies through the command interface. For example, the cycle adjustment strategy of traffic lights is updated in real time through the central control module of the traffic control system. The energy distribution strategy is directly distributed to each electricity user or region through the smart grid. In addition, pollutant emission restriction measures can be issued by issuing specific parameters to the emission control unit of the enterprise's production equipment.
[0120] In general, the execution of optimization strategies will directly lead to changes in the city state. In order to monitor such changes, the feedback module obtains the state variables after execution through sensor networks and other real-time data acquisition systems. For example, the improvement of traffic congestion can be evaluated by real-time monitoring of vehicle speed and road traffic. The effect of energy distribution optimization can be judged by the time series data of power consumption and load balancing indicators.
[0121] Specifically, the data of feedback monitoring include key state variables x(t) of urban operation, such as traffic flow x1(t), energy consumption x2(t), air pollution index x3(t), etc. These data are used to compare with the warning status to evaluate the effect of strategy execution. In one possible implementation, the feedback module calculates the deviation between the execution effect and the expected target, and defines the deviation function:
[0122] Δx(t)=x 实际 (t)-x 期望 (t),
[0123] Among them, x 实际 (t): the actual status monitored after the strategy is executed; x 期望 (t): expected state predicted by the optimization module; Δx(t): state deviation, used to quantify the degree of optimization effect.
[0124] As an option, if the state deviation exceeds the preset threshold, the feedback module will trigger the dynamic adjustment mechanism to update the game model and optimization strategy. In general, the model update includes parameter adjustment of the state dynamic evolution equation f(x,u) and the immediate benefit function g i The weight of (x,u) is recalibrated. For example, in traffic light control, if the actual traffic flow does not reach the expected improvement target, the system can increase the weight coefficient of the signal cycle adjustment to the state equation to strengthen the control effect.
[0125] In this embodiment, the feedback data can also be used to train the multi-agent policy network in the optimization module, thereby achieving online learning capabilities. In some embodiments, an incremental training method is used to update the policy network of the agent based only on the latest feedback data. This method not only retains historical optimization experience, but also can quickly adapt to new environmental changes.
[0126] In another possible implementation, the feedback module also includes an anomaly detection function to identify possible emergencies in urban operations. For example, if the fluctuation of traffic flow x1(t) in the feedback data significantly exceeds the historical range, the module will automatically determine it as a traffic accident or sudden congestion event. In this case, the system will prioritize re-optimizing related strategies to quickly respond to abnormal conditions.
[0127] Generally, the frequency of operation of the feedback module is determined by the specific application scenario. For example, for highly dynamic traffic control scenarios, the feedback data may be updated once a minute. For relatively stable scenarios such as energy distribution optimization, the feedback cycle can be set to every hour or every day.
[0128] In addition, this module also has a long-term trend analysis function. In some embodiments, the system will generate state trend forecasts based on the accumulated feedback data, such as long-term changes in air quality or seasonal fluctuations in energy demand. These forecast results can be directly integrated into the optimization module to provide additional reference for subsequent strategy generation.
[0129] Through the above process, the feedback module realizes the closed-loop verification and dynamic adjustment function of the strategy. The optimization results are applied in real time, and the feedback data continuously enriches the dynamic characteristics of the model. The system shows a high degree of adaptability in actual operation, and can continuously improve the optimization effect and adapt to the rapidly changing urban environment. The design of the feedback module makes the system operation more precise, while improving its robustness and practical operability.
[0130] Please refer to the attached Figure 2 The present invention also provides a data-driven urban sustainable development planning decision optimization method. The specific implementation methods of each step are described below in conjunction with the workflow of the optimization method of the present invention.
[0131] S1. Data collection and preprocessing:
[0132] Collect multi-source data on the city’s operating status, including traffic flow, energy consumption, air quality, and land use, through urban sensor networks and IoT devices;
[0133] The collected multi-source data are cleaned to remove outliers, dimension reduction is performed to simplify data complexity, and formatting is performed to generate structured urban state variables for modeling;
[0134] S2. Construction of dynamic game model:
[0135] Based on the city state variables and multi-party control variables, the city state dynamic evolution equation is established to describe the dynamic changes of the city state over time;
[0136] Define an immediate benefit function to reflect the benefits and costs of urban states and control strategies to multiple parties;
[0137] Construct a cumulative utility function to describe the optimization objectives of multiple parties during the entire planning period and establish a differential dynamic game model involving multiple parties;
[0138] S3. Optimization strategy solution:
[0139] The differential dynamic game model is transformed into a multi-agent reinforcement learning problem, and each participant is modeled as an agent;
[0140] The city state variable is used as the input of the intelligent agent, the control variable is used as the output of the intelligent agent, and the immediate benefit function is used as the reward of the intelligent agent;
[0141] Based on the deep reinforcement learning algorithm, the agent strategy is trained so that each agent's strategy gradually converges to the Nash equilibrium of the dynamic game;
[0142] S4. Optimize strategy execution and feedback:
[0143] Control the city's operating status according to optimization strategies, including adjusting traffic signal cycles, optimizing energy distribution and implementing pollution control measures;
[0144] Monitor changes in city status after the optimization strategy is implemented, and verify the optimization effect through feedback data;
[0145] The state evolution equation and profit function in the dynamic game model are adjusted according to the feedback data, and the optimization strategy is updated to adapt to the real-time changes in urban operations.
[0146] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A data-driven urban sustainable development planning decision optimization system, characterized by: The system includes: The data collection and preprocessing module is used to collect real-time multi-source urban operation data, including traffic flow, energy consumption, air quality, and land utilization rate, and to perform data cleaning, dimension reduction, and formatting to generate urban state variables; The differential dynamic game modeling module is used to construct the state dynamic evolution equation, instant benefit function and cumulative utility function based on the city state variables and multi-party control strategy variables, and establish a dynamic game model involving multiple parties; The decision-making optimization module is used to optimize the high-dimensional strategy space of the dynamic game model based on multi-agent reinforcement learning technology to solve the optimal dynamic strategy of each participant; The system feedback module is used to apply the optimization strategy to urban operations, monitor the optimization results in real time, and dynamically update the game model and optimization strategy based on the monitoring data.
2. The data-driven urban sustainable development planning decision optimization system according to claim 1 is characterized in that: The data acquisition and preprocessing module includes: Sensor networks and IoT devices to collect dynamic data such as traffic flow, energy consumption, and pollutant emissions; Data fusion unit, used to uniformly transmit multi-source data to the data center; The data preprocessing unit is used to perform data cleaning to remove outliers, data dimension reduction to simplify data complexity, and formatting to generate structured urban status data for modeling.
3. The data-driven urban sustainable development planning decision optimization system according to claim 1 is characterized in that: The differential dynamic game modeling module is used to establish a game model involving multiple parties, wherein: Urban state variables are used to describe the urban operation status of traffic flow, energy consumption, and air quality; Multiple parties include the government, enterprises and the public, who influence the city state by controlling variables; The state dynamics evolution equation is used to describe the changes of urban state variables over time, and its changes are jointly determined by the control variables and the internal dynamics of the system; The immediate benefit function is used to reflect the interest objectives of each participant in the current state, including the positive contribution of the state to utility and the negative impact of control variables.
4. The data-driven urban sustainable development planning decision optimization system according to claim 1 is characterized in that: The cumulative utility function is used to describe the target optimization results of the participants in the entire planning period. It is the integration result of the instant benefit function over time. The instant benefit function includes state utility, control variable cost and the joint utility of state and control variable.
5. The data-driven urban sustainable development planning decision optimization system according to claim 1 is characterized in that: The optimization decision module is based on multi-agent reinforcement learning technology and includes: Construct a reinforcement learning framework, model the game participants as intelligent agents, use the city state variables as the input of the intelligent agents, and use the control variables as the output of the intelligent agents; Define the immediate benefit as the agent's reward function, and use the participant's immediate benefit function as part of the reinforcement learning reward function; The strategies of intelligent agents are trained through deep reinforcement learning algorithms, so that the strategies of each intelligent agent gradually converge to a dynamic Nash equilibrium.
6. The data-driven urban sustainable development planning decision optimization system according to claim 1 is characterized in that: The optimization decision module uses a deep Q-learning algorithm or a policy gradient optimization method to optimize the agent strategy. Based on the joint training method of reinforcement learning, multiple agents update their strategies alternately while sharing state information, and finally converge to the Nash equilibrium solution of the dynamic game.
7. The data-driven urban sustainable development planning decision optimization system according to claim 1 is characterized in that: The system feedback module comprises: The optimization strategy execution unit is used to execute the control measures of urban operation according to the optimal strategy generated by the optimization decision module; Feedback monitoring unit, used to monitor the changes in urban status after the optimization strategy is executed in real time; The model updating unit is used to adjust the state evolution equation and profit function in the dynamic game model based on feedback data.
8. The data-driven urban sustainable development planning decision optimization system according to claim 1 is characterized in that: The model updating unit dynamically adjusts the game model based on the feedback monitoring data, and updates the state dynamic evolution equation parameters and the weight factor of the instant benefit function in the game model to adapt to the real-time changing urban operation status.
9. The data-driven urban sustainable development planning decision optimization system according to claim 1 is characterized in that: The optimization decision module can make trade-offs for multi-objective conflict issues, including the government's social equity goal, the enterprise's profit maximization goal, and the public's quality of life improvement goal, and optimize the balance between the goals through the cumulative utility function in the dynamic game model.
10. A data-driven urban sustainable development planning decision optimization method, applied to the system according to any one of claims 1 to 9, characterized in that: The following steps are involved: S1. Data collection and preprocessing: Collect multi-source data on the city’s operating status, including traffic flow, energy consumption, air quality, and land utilization, through urban sensor networks and IoT devices; The collected multi-source data are cleaned to remove outliers, dimension reduction is performed to simplify data complexity, and formatting is performed to generate structured urban state variables for modeling; S2. Construction of dynamic game model: Based on the city state variables and multi-party control variables, the city state dynamic evolution equation is established to describe the dynamic changes of the city state over time; Define an immediate benefit function to reflect the benefits and costs of urban states and control strategies to multiple parties; Construct a cumulative utility function to describe the optimization objectives of multiple parties during the entire planning period and establish a differential dynamic game model involving multiple parties; S3. Optimization strategy solution: The differential dynamic game model is transformed into a multi-agent reinforcement learning problem, and each participant is modeled as an agent; The city state variable is used as the input of the intelligent agent, the control variable is used as the output of the intelligent agent, and the immediate benefit function is used as the reward of the intelligent agent; Based on the deep reinforcement learning algorithm, the agent strategy is trained so that each agent's strategy gradually converges to the Nash equilibrium of the dynamic game; S4. Optimize strategy execution and feedback: Control the city's operating status according to optimization strategies, including adjusting traffic signal cycles, optimizing energy distribution and implementing pollution control measures; Monitor changes in city status after the optimization strategy is implemented, and verify the optimization effect through feedback data; The state evolution equation and profit function in the dynamic game model are adjusted according to the feedback data, and the optimization strategy is updated to adapt to the real-time changes in urban operations.
Citation Information
Cited By
Land reclamation collaborative planning system based on GIS and BIM fusion
CN120833040A
Membrane pollution control method
CN121148507A