An urban land use intelligent planning method for a production-city integration scene and a related device
Patent Information
- Application Number
- CN202610937332.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-18
AI Technical Summary
然而,现有深度强化学习框架多依赖于针对固定偏好权重优化的静态策略网络,导致规划偏好一旦调整,需耗费大量计算资源对整个模型进行重新训练,效率低下
本申请提供了一种面向产城融合场景的城市土地利用智能规划方法及相关装置,通过获取城市空间数据与规划偏好向量并划分地块单元数据,解决了规划数据杂乱、单元边界不清晰的问题,实现了国土空间、土地现状多源数据标准化分块处理;通过将地块单元数据作为全局规划环境状态并构建地块单元状态图,解决了城市地块空间关联关系难以建模的问题,实现了用地单元空间拓扑关系完整表征;通过条件图神经编码器融合地块状态图与规划偏好向量提取多维度状态特征,解决了全局用地信息、地块局部上下文、规划偏好无法同步融合编码的问题,实现了兼顾全局、局部地块、规划诉求、规划阶段的精细化特征提取;通过约束感知动作掩码机制搭配条件策略网络输出地块规划动作,解决了土地规划合规约束难以嵌入决策过程、易生成违规用地方案的问题,实现了符合国土管控规则的地块规划动作精准输出;通过条件价值网络输出偏好导向下的环境价值估计,解决了单一评价维度无法匹配差异化规划诉求的问题,实现了不同规划偏好下全局用地方案价值实时预判;通过执行规划动作迭代更新环境状态、循环完成全部地块规划并计算即时奖励,解决了地块分步规划交互过程缺少正向激励引导的问题,实现了逐地块规划行为动态优化;通过耦合三生空间复合奖励模型生成多维度奖励向量并结合规划偏好聚合总奖励,搭配价值估计更新策略与价值网络参数,解决了生产、生活、生态空间效益割裂、难以协同优化的问题,实现了三生空间功能均衡协同的模型迭代训练;通过训练完成的网络输出多偏好最优规划方案,解决了传统规划方案单一、难以适配多元发展诉求的问题,实现了多导向、均衡化、多样化的城市土地智能规划方案生成。
Smart Images

Figure CN122779486A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of land and space planning technology, and in particular to an intelligent urban land use planning method and related devices for the integration of industry and city. Background Technology
[0002] Optimizing and coordinating the functions of production, living, and ecological spaces through urban land use is a crucial path to achieving sustainable urban development. Maximizing the synergistic effects of these three spaces, and balancing economic development efficiency, social equity, and ecological environmental safety, has become a core objective in the field of land spatial planning. In the process of rapid urbanization, the integrated development model of industry and city is becoming increasingly prominent. This model, guided by the deep integration of industrial functions and urban infrastructure, signifies a shift in urban governance paradigms from incremental expansion to stock renewal and optimized allocation of existing resources. However, this comprehensive integration process requires precise control over the compatibility of production, living, and ecological spaces. Therefore, ensuring the dynamic equilibrium of these three spaces has become a significant technical challenge for contemporary urban governance.
[0003] Compared to traditional urban functional zoning models, the integrated industry-city model advocates mixed land use. However, achieving a coordinated balance between economic production, public life, and ecological security under this model is significantly more challenging due to the inherent complexity of its spatial system. Urban planning is essentially a dynamic, sequential decision-making process, rather than a static, one-off allocation of spatial resources. Its implementation alters existing regional environmental conditions, thereby constraining the future adaptability and optimization capabilities of urban space. Planners need to continuously and dynamically adjust the trade-off between "economic development priority" and "ecological protection priority" to adapt to constantly evolving policy directions. This time-dependent characteristic cannot be effectively captured by traditional static suitability assessment methods. Furthermore, the "NIMBY effect" further exacerbates planning complexity. For example, the adjacent distribution of high-intensity industrial land and residential land easily creates spatial conflict hotspots, reducing residents' quality of life and damaging regional ecological security. To address these challenges, planning tools with the ability to analyze complex topological relationships are urgently needed to proactively resolve spatial conflicts, rather than simply optimizing quantitative indicators.
[0004] Therefore, urban land use planning should be redefined as a complex combinatorial optimization problem. While traditional operations research methods and heuristic algorithms are widely used in this field, they have significant limitations in handling the inherent nonlinear constraints of urban planning, easily getting trapped in local optima and failing to meet the needs of global planning optimization. In recent years, deep reinforcement learning has emerged as a promising alternative technology. It extracts spatial features through deep neural networks and can adaptively formulate planning decision-making strategies. However, existing deep reinforcement learning frameworks mostly rely on static policy networks optimized for fixed preference weights. This leads to inefficiencies, requiring significant computational resources to retrain the entire model once planning preferences are adjusted. Furthermore, existing research has failed to adequately adapt to the multiple policy constraints in land use planning, especially in resolving topological "NIMBY" conflicts at the boundaries of industrial and residential areas, easily leading to irrational decisions. Simultaneously, the reward functions used in previous studies often focus on a single dimension or objective, failing to fully consider the core needs of multi-dimensional balanced development and making it difficult to support comprehensive planning optimization in the context of industry-city integration.
[0005] In summary, how to balance dynamic adjustments to planning preferences, resolution of NIMBY (Not In My Backyard) conflicts, and the integration of industry and city, and build an efficient and practical intelligent urban land planning optimization system, is an urgent problem to be solved. Summary of the Invention
[0006] The purpose of this application is to provide an intelligent urban land use planning method and related devices for the integration of industry and city, which can realize intelligent planning of urban land use in the context of industry-city integration.
[0007] To achieve the above objectives, this application provides the following solution: Firstly, this application provides an intelligent urban land use planning method for industry-city integration scenarios, including: Obtain urban spatial data and planning preference vectors; divide the urban spatial data into plot unit data; the urban spatial data includes at least land spatial planning data and land use status data.
[0008] The land parcel unit data is used as the global planning environment state at the current time step.
[0009] Construct a plot unit state diagram for the current time step based on the global planning environment state at the current time step.
[0010] Based on the state diagram of the land parcel unit and the planning preference vector, the state features of the current decision land parcel unit are obtained using a conditional graph neural encoder. The state features of the current decision land parcel unit include: embedding encoding of global features, graph-level summary feature representation of nodes and edges, current decision land parcel context vector after attention enhancement, stage information, and planning preference vector.
[0011] Based on the aforementioned state characteristics, the planning action for the current decision-making plot is obtained through a conditional policy network using a constraint-aware action masking mechanism.
[0012] Based on the aforementioned state characteristics, the value estimate of the global planning environment state under planning preferences at the current time step is obtained through the conditional value network.
[0013] Execute the planning action, update the global planning environment state, obtain the global planning environment state of the next time step, and calculate the immediate reward; use the global planning environment state of the next time step as the global planning environment state of the current time step and return "Construct the land parcel unit state diagram of the current time step based on the global environment planning state of the current time step" until all land parcel units have completed planning; the immediate reward is used to encourage effective interaction.
[0014] After a round of planning is completed, a reward vector is calculated based on the reward model of the coupled three-life space, and aggregated into a total reward scalar according to the planning preference vector; based on the value estimate at the current time step and the total reward scalar, the network parameters of the conditional policy network and the network parameters of the conditional value network are updated; the reward model of the coupled three-life space is a composite reward model that includes production space rewards, living space rewards, ecological space rewards and functional synergy rewards of the three-life space.
[0015] The updated conditional policy network outputs optimal planning and decision schemes with different preference orientations.
[0016] Secondly, this application provides an intelligent urban land use planning device for the integration of industry and city, including: a planning status perception module, a planning intelligent agent module, and an urban environment simulation module; The planning status perception module is used to acquire urban spatial data and planning preference vectors; divide the urban spatial data into plot unit data; the urban spatial data environment includes at least land spatial planning data and land use status data; use the plot unit data as the global planning environment status at the current time step; and construct the plot unit status map at the current time step based on the global planning environment status at the current time step.
[0017] The planning agent module is used to obtain the state features of the current decision-making plot unit based on the plot unit state graph and the planning preference vector using a conditional graph neural encoder. The state features of the current decision-making plot unit include: embedded encoding of global features, graph-level summary feature representation of nodes and edges, attention-enhanced current decision-making plot context vector, stage information, and planning preference vector. Based on the state features, the planning action of the current decision-making plot is obtained through a conditional policy network using a constraint-aware action masking mechanism. Based on the state features, the value estimate of the global planning environment state under the planning preference at the current time step is obtained through a conditional value network.
[0018] The urban environment simulation module is used to execute the planning actions, update the global planning environment state, obtain the global planning environment state of the next time step, and calculate the immediate reward; it uses the global planning environment state of the next time step as the global planning environment state of the current time step and returns "constructing the land parcel unit state diagram of the current time step based on the global environment planning state of the current time step" until all land parcel units have completed planning; the immediate reward is used to encourage effective interaction; after a round of planning, a reward vector is calculated based on the reward model of the coupled three-life space, and aggregated into a total reward scalar according to the planning preference vector; based on the value estimate of the current time step and the total reward scalar, the network parameters of the conditional policy network and the network parameters of the conditional value network are updated; the reward model of the coupled three-life space is a composite reward model that includes production space rewards, living space rewards, ecological space rewards and three-life space functional synergy rewards; based on the updated conditional policy network, the optimal planning decision scheme with different preference orientations is output.
[0019] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described intelligent urban land use planning method for the integration of industry and city.
[0020] Fourthly, this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the aforementioned intelligent urban land use planning method for the integration of industry and city.
[0021] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned intelligent urban land use planning method for the integration of industry and city.
[0022] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides an intelligent urban land use planning method and related devices for industry-city integration scenarios. By acquiring urban spatial data and planning preference vectors and dividing the data into plot units, it solves the problems of messy planning data and unclear unit boundaries, achieving standardized block processing of multi-source data on land space and current land status. By treating plot unit data as the global planning environment state and constructing a plot unit state diagram, it solves the problem of difficulty in modeling spatial relationships between urban plots, achieving a complete representation of the spatial topological relationships of land use units. Through a conditional graph neural encoder, it fuses the plot state diagram and planning preference vectors to extract multi-dimensional state features, solving the problem of the inability to synchronously fuse and encode global land use information, local plot context, and planning preferences, achieving refined feature extraction that considers the global picture, local plots, planning demands, and planning stages. By using a constraint-aware action mask mechanism combined with a conditional policy network to output plot planning actions, it solves the problem of difficulty in embedding land planning compliance constraints into the decision-making process and the easy generation of illegal land use schemes, achieving compliance... The system accurately outputs land planning actions that comply with land management regulations; it solves the problem that a single evaluation dimension cannot match differentiated planning demands by outputting preference-oriented environmental value estimates through a conditional value network, and achieves real-time prediction of the value of global land use schemes under different planning preferences; it solves the problem of lack of positive incentive guidance in the step-by-step planning interaction process of land parcels by iteratively updating the environmental status through the execution of planning actions, cyclically completing the planning of all land parcels and calculating immediate rewards, and achieves dynamic optimization of planning behavior for each land parcel; it solves the problem of fragmented benefits and difficulty in coordinating optimization of production, living and ecological spaces by generating multi-dimensional reward vectors through coupling the three-life space composite reward model and aggregating the total reward in combination with planning preferences, and combining it with value estimation update strategies and value network parameters, and achieves iterative training of the model to achieve balanced and coordinated functions of the three-life spaces; it solves the problem of traditional planning schemes being singular and difficult to adapt to diverse development demands by outputting multi-preference optimal planning schemes through the trained network, and achieves the generation of multi-oriented, balanced and diversified intelligent urban land planning schemes. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A flowchart illustrating an intelligent urban land use planning method for an integrated industry-city development scenario, provided as an embodiment of this application; Figure 2 A schematic diagram of a module for an intelligent urban land use planning device for an integrated urban-industrial development scenario, provided in an embodiment of this application; Figure 3 A schematic diagram illustrating the framework of an intelligent urban land use planning method for an integrated urban-industrial development scenario, provided as another embodiment of this application; Figure 4 A schematic diagram of a conditional graph neural network encoder provided in an embodiment of this application; Figure 5 A schematic diagram of Pareto optimal land use layout generated in an experimental area using the intelligent urban land use planning method provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0026] To address the aforementioned issues, this technology proposes a constraint-aware deep reinforcement learning framework specifically adapted to integrated urban-industrial development scenarios. Within this framework, a conditional graph neural network encoder explicitly integrates Pareto preference weight vectors into a graph-based land use state representation, enabling the planning agent to generate diverse trade-off optimization strategies through a single training iteration, thus adapting to planning needs across different scenarios. To enhance the framework's perception module, this technology introduces a constraint-aware occlusion mechanism. This mechanism directly embeds land use regulations, constraints, and restrictions into the decision-making process, performing feasibility verification before the planning agent makes action choices. This effectively reduces the decision search space, thereby avoiding "NIMBY" conflicts and ensuring the rationality of planning decisions. Finally, the framework's optimization process is guided by a coupled production-living-ecological space reward model. This model can quantitatively evaluate the functional synergy of the three spaces—production, living, and ecological—ensuring overall balanced optimization of urban land use space.
[0027] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0028] In one exemplary embodiment, such as Figure 1 As shown, an intelligent urban land use planning method for industry-city integration scenarios is provided, including the following steps 101 to 109: Step 101: Obtain urban spatial data and planning preference vector; divide the urban spatial data into plot unit data; the urban spatial data includes at least land spatial planning data and land use status data.
[0029] Step 102: Use the plot unit data as the global planning environment state at the current time step.
[0030] Step 103: Construct a plot unit state diagram for the current time step based on the global planning environment state at the current time step.
[0031] Step 104: Based on the state diagram of the land parcel unit and the planning preference vector, the state features of the current decision land parcel unit are obtained using a conditional graph neural encoder; the state features of the current decision land parcel unit include: embedding encoding of global features, graph-level summary feature representation of nodes and edges, current decision land parcel context vector after attention enhancement, stage information, and planning preference vector.
[0032] Step 105: Based on the state characteristics, the planning action of the current decision plot is obtained through the conditional policy network using the constraint-aware action masking mechanism.
[0033] Step 106: Based on the state characteristics, obtain the value estimate of the global planning environment state under planning preferences at the current time step through the conditional value network.
[0034] Step 107: Execute the planning action, update the global planning environment state, obtain the global planning environment state of the next time step, and calculate the immediate reward; use the global planning environment state of the next time step as the global planning environment state of the current time step and return "Construct the land parcel unit state diagram of the current time step based on the global environment planning state of the current time step" until all land parcel units have completed planning; the immediate reward is used to encourage effective interaction.
[0035] Step 108: After a round of planning, calculate the reward vector based on the reward model of the coupled three-life space, and aggregate it into a total reward scalar according to the planning preference vector; based on the value estimate at the current time step and the total reward scalar, update the network parameters of the conditional policy network and the network parameters of the conditional value network; the reward model of the coupled three-life space is a composite reward model that includes production space rewards, living space rewards, ecological space rewards and functional synergy rewards of the three-life space.
[0036] Step 109: Output the optimal planning decision schemes with different preference orientations based on the updated conditional policy network.
[0037] By implementing steps 101 to 109 above, this application relies on graph neural networks to model spatial relationships of land parcels, use constraint masks to standardize planning actions, and implement multi-dimensional benefit synergistic optimization through composite rewards for the three-life space. It can automatically generate diverse and balanced land planning schemes while taking into account differentiated planning preferences. Overall, it is significantly superior to traditional planning benchmark models in terms of planning diversity, decision-making compliance and rationality, and land use layout balance. The model is adaptable to various integrated urban-industrial development scenarios, has strong transferability, and can provide planning staff with multiple sets of scientific planning references that take into account the coordination of industry, residence, and ecology. It makes up for the problems of lack of diversity of schemes, lack of rationality of decision-making, and insufficient balance of planning decisions in the process of intelligent planning of urban land use in existing technologies.
[0038] In another exemplary embodiment of this application, in order to accurately obtain the state characteristics of the current decision-making plot unit using a conditional graph neural encoder based on the plot unit state diagram and the planning preference vector, the above step 104 is replaced by the following steps 1041 to 10410: Step 1041: Obtain the global feature vector based on the state diagram of the land parcel unit.
[0039] Step 1042: Based on the global feature vector, the embedding encoding of the global feature is obtained after processing by a multilayer perceptron.
[0040] Step 1043: Based on all node features and all edge features in the state diagram of the land parcel unit, obtain the graph-level summary feature representation of the nodes and edges.
[0041] Step 1044: Obtain stage information based on the current time step or decision stage.
[0042] Step 1045: Construct an initial embedding space based on the embedding encoding of the global features.
[0043] Step 1046: Project the feature vector of each node in the state diagram of the land parcel unit at the current time step to the initial embedding space through linear transformation to obtain the initial node representation.
[0044] Step 1047: Project the preference vector into a latent planning preference embedding through a multilayer perceptron, scale the planning preference embedding by a scaling factor, broadcast it, and add it to the initial node representation to obtain the preference-enhanced initial node features.
[0045] Step 1048: Based on the spatial adjacency relationship depicted by the edge set in the state graph of the land parcel unit and the initial node features of the preference enhancement, the neighborhood information is aggregated through a stacked conditional graph neural network to obtain the feature vector set of all land parcel units.
[0046] Step 1049: Based on the set of feature vectors, extract the node features corresponding to the current decision plot; at the same time, using the node features of the current decision plot as the query and the set of feature vectors of all plot units as the key and value, calculate the attention-enhanced context vector of the current decision plot through a global attention mechanism.
[0047] Step 10410: Concatenate the global feature embedding encoding, the graph-level summary feature representation of nodes and edges, the attention-enhanced current decision plot context vector, stage information, and planning preference vector to obtain the state features of the current decision plot unit.
[0048] In another exemplary embodiment of this application, in order to accurately utilize the constraint-aware action masking mechanism to perform a masking operation on the logarithmic probability of the action and obtain the masked logarithmic probability of the action, the above step 105 is replaced by the following steps 1051-1052: Step 1051: Calculate the Boolean action mask vector based on the three rules in the constraint-aware action masking mechanism; the three rules are the hygiene and safety isolation restriction rule, the functional heterogeneity and distribution equilibrium rule, and the NIMBY (Not In My Backyard) avoidance rule; the Boolean action mask vector is used to mark the legality of candidate actions in the action space.
[0049] Step 1052: Perform an action masking operation on the action log probability through the Boolean action mask vector to obtain the masked action log probability.
[0050] In one exemplary embodiment, such as Figure 2 As shown, an intelligent urban land use planning device for the integration of industry and city is provided, including: a planning status perception module, a planning intelligent agent module, and an urban environment simulation module.
[0051] The planning status perception module is used to acquire urban spatial data and planning preference vectors; divide the urban spatial data into plot unit data; the urban spatial data environment includes at least land spatial planning data and land use status data; use the plot unit data as the global planning environment status at the current time step; and construct the plot unit status map at the current time step based on the global planning environment status at the current time step.
[0052] The planning agent module is used to obtain the state features of the current decision-making plot unit based on the plot unit state graph and the planning preference vector using a conditional graph neural encoder. The state features of the current decision-making plot unit include: embedded encoding of global features, graph-level summary feature representation of nodes and edges, attention-enhanced current decision-making plot context vector, stage information, and planning preference vector. Based on the state features, the planning action of the current decision-making plot is obtained through a conditional policy network using a constraint-aware action masking mechanism. Based on the state features, the value estimate of the global planning environment state under the planning preference at the current time step is obtained through a conditional value network.
[0053] The urban environment simulation module is used to execute the planning actions, update the global planning environment state, obtain the global planning environment state of the next time step, and calculate the immediate reward; it uses the global planning environment state of the next time step as the global planning environment state of the current time step and returns "constructing the land parcel unit state diagram of the current time step based on the global environment planning state of the current time step" until all land parcel units have completed planning; the immediate reward is used to encourage effective interaction; after a round of planning, a reward vector is calculated based on the reward model of the coupled three-life space, and aggregated into a total reward scalar according to the planning preference vector; based on the value estimate of the current time step and the total reward scalar, the network parameters of the conditional policy network and the network parameters of the conditional value network are updated; the reward model of the coupled three-life space is a composite reward model that includes production space rewards, living space rewards, ecological space rewards and three-life space functional synergy rewards; based on the updated conditional policy network, the optimal planning decision scheme with different preference orientations is output.
[0054] This application provides an intelligent urban land use planning method for industry-city integration scenarios. This application models the urban planning problem as a constrained Markov decision process, whose features are defined by tuples. Description. Details are as follows: State space (S): a dynamic topological graph Represented in the form of . Where, Indicates time step The following is a status map of urban land parcel units. Represents the set of plot unit nodes; This represents the set of edges formed by the spatial adjacency relationships of land parcels. (Node set) This represents the set of land parcels containing static geometric attributes (such as area and shape) and dynamic functional states, while the edge set... This represents the set of spatial adjacency relationships. The topology map is dynamically updated as the state of the land parcels changes.
[0055] Action Space ( Defined as At time step An action Involves the set of unallocated land parcels Select a specific feasible plot of land. and designate the target land use category for it. (e.g., residential, industrial, commercial) A set representing the types of land parcel units.
[0056] Dynamic conversion ( ): Execute action This process updates the functionality of the land parcel. If the area exceeds the threshold for that land use category, a geometric segmentation algorithm (polygon partitioning) is triggered, generating new polygons while preserving the remaining portion. Therefore, this process... Involving Attribute updates and topology changes.
[0057] Constraint set ( This ensures spatial feasibility and regulatory compliance. It includes topological constraints. (Restricting operations to feasible blocks) and regulatory constraints (Enforcing partitioning rules, such as security buffers). In form, A validity mask space is defined, and invalid actions are eliminated before execution through an action mask mechanism.
[0058] Reward function ( ): by a time step Scalar step-by-step rewards (Encourage effective interaction) and a time step High-dimensional final reward vector (Calculated at the end of each planning decision round, its components correspond to production space rewards, living space rewards, ecological space rewards, and synergistic rewards for the functions of the three spaces, respectively.) During training, this is achieved through the planning preference vector. right Scalar processing is performed to estimate the advantage.
[0059] Discount factor ( ): A scalar This value is used to determine the time span of the agent. When this value approaches 1, it indicates that it focuses more on the results of long-term planning.
[0060] Preference space ( ): Preference weight vector Used to control multi-objective trade-offs ( When used as a conditional input, The guiding strategy generates non-dominated schemes that are close to the Pareto front.
[0061] The overall framework for implementing the intelligent urban land use planning method for the integrated development of industry and city scenarios proposed in this application is as follows: Figure 3 As shown, the framework mainly consists of three parts: a planning state perception module, a planning agent module, and an urban environment simulation module. The planning agent module includes a conditional graph neural network encoder, a conditional policy network, and a conditional value network. The urban environment simulation module includes a constraint-aware action masking mechanism and a reward function model coupled with the three-dimensional space. The core of this application is mainly reflected in three aspects: the conditional graph neural network encoder in the planning agent module, the constraint-aware action masking mechanism in the urban environment simulation module, and the reward function model coupled with the three-dimensional space.
[0062] Planning Status Awareness Module: Acquires and preprocesses urban spatial data, which includes at least national land spatial planning data and land use status data. It divides the data into land parcel units according to the national land spatial planning uses, and integrates the spatial structure and land use status of the land parcel units into a graph structure data to represent the status of the land use units.
[0063] The planning agent module comprises a conditional graph neural network encoder, a conditional policy network, and a conditional value network. Taking graph-structured data and planning preferences as input, the module iteratively learns the optimal planning strategy through dynamic changes in the current planning state and feedback from the urban environment simulation module, outputting optimal planning decision schemes oriented towards different preferences.
[0064] The urban environment simulation module includes a constraint-aware action masking mechanism and a reward function model coupled with the three-dimensional space (life, ecology, and environment). The constraint-aware action masking mechanism generates action masks based on urban planning constraints and feeds them back to the conditional policy network; the reward function model coupled with the three-dimensional space generates action masks based on the actions of the planning agent. Compared with the current planning environment and update the environment status to Calculate instant rewards This feedback is sent to the planning status perception module, forming a closed-loop iteration.
[0065] In the process of graph neural network encoder, firstly, the graph structure data is input into the conditional graph neural network encoder, which encodes it into a land use unit state embedding encoding representation that can be perceived by the intelligent agent planning module (including the feature information of the land parcel unit and the spatial relationship information of the land parcel unit); then, the conditional graph neural network encoder maps the planning preference vector into a conditional embedding vector; finally, the conditional graph neural network encoder fuses the state embedding encoding of the land parcel unit with the conditional embedding vector, and outputs the state feature representation of the fused planning preference of the land parcel unit.
[0066] In the conditional policy network processing, firstly, the conditional policy network maps the state feature representation of the fused planning preference output by the conditional graph neural network encoder into the log probability of actions; then, the conditional policy network receives the constraint-aware action mask from the urban environment simulation module, performs a masking operation on the log probability of actions, and filters out invalid actions that do not comply with urban planning regulations and constraints; finally, the conditional policy network normalizes the probability distribution of the masked log probability of actions and samples the planning actions of the planning agent.
[0067] In the process of Conditional Value Network processing, the Conditional Value Network maps the state features output by the Conditional Graph Neural Network encoder into value estimates, and evaluates the long-term expected returns of taking subsequent actions under the current state of the land parcel unit and planning preferences.
[0068] Urban environment simulation module: includes a constraint-aware action masking mechanism and a reward function model coupled with three-dimensional space.
[0069] The constraint-aware action masking mechanism generates action masks based on urban planning constraints and rules, and feeds them back to the conditional policy network in the agent planning module.
[0070] The reward function model of the coupled three-life space calculates the instant reward based on the action output by the planning module of the planning agent, the state of the fusion planning preference of the plot unit, and the current planning environment state, and updates the planning environment state, which is then fed back to the planning state perception module to form a closed loop iteration.
[0071] The overall process is as follows: Step 201: Obtain the state characteristics of the current decision-making plot unit (e.g., Figure 4 (As shown).
[0072] In the processing of the conditional graph neural network encoder, the encoder projects urban planning state features and planning preferences into a unified latent feature space, generating embedding vectors that can capture complex spatial details and policy changes. First, the conditional graph neural network encoder in the planning state perception module encodes the state of land parcel units into a state representation perceptible to the planning agent (including feature information of land parcel units and spatial relationship information between them). Then, the multilayer perceptron in the planning agent module maps the planning preference vectors into conditional embedding vectors. Finally, the conditional graph neural network fuses the state encoding of land parcel units with the conditional embedding vectors, outputting a state feature representation of the land parcel units that incorporates planning preferences.
[0073] Specifically as follows: Step 2011: Perform graph-based representation of land parcel unit states. (At time step...) At that time, the planning status of urban land parcel units was modeled as a topological graph: ; in, Indicates time step The following is a status diagram of urban land parcel units; Represents the set of plot unit nodes; This represents the set of edges derived from the spatial adjacency relationships of the current land parcel. Node Represents a single land parcel unit, edge This represents the spatial adjacency relationship between land parcel units. Each node carries a feature vector. It includes categorical semantics (such as land use type) and geometric attributes (such as area and shape). This is the dimension index of the vector.
[0074] A global feature vector is obtained based on the state diagram of the land parcel unit. Describes the overall planning environment at a macro level. As the number and shape of plots dynamically change during the planning process, the state of each node and edge changes synchronously to ensure consistency with the current graph structure.
[0075] Step 2012: Perform feature initialization. At time step... global feature vector The global feature embedding encoding is obtained after processing through a multilayer perceptron. The expression is: ; in, Indicates the activation function; , This represents a two-layer linear mapping used to encode global statistical features. Simultaneously, to capture the local features of individual plot units, a linear transformation is used to transform each node of the plot unit state diagram. At time step eigenvectors Projected onto the initial embedding space: ; in, For nodes The initial hidden state (represented by the initial node) is used for subsequent message passing; and These are the learnable weight matrix and the bias vector, respectively.
[0076] Step 2013: Inject preference-aware features. To enable the planning agent to adapt to the ever-evolving planning orientation requirements, the planning preference vector is... Embedding latent planning preferences through multilayer perceptron projection It can capture the nonlinear trade-offs between preference vectors under different development orientation requirements: ; in, This represents the planning preference weight vector; This represents the latent preference embedding obtained by mapping the planning preference weight vector; and This represents a linear mapping in a multilayer perceptron; This represents a non-linear activation function. The planning preference embedding is then broadcast and added to the initial node representation, placed before the first layer of the conditional graph neural network, and scaled using a scaling factor. Controlling the injection intensity: ; in, Represents a node At time step Initial node embedding; This indicates the embedding of conditional nodes after integrating planning preferences; This indicates the embedding of planning preferences; This indicates a preference for embedding scaling factors.
[0077] Step 2014 involves edge-based feature updating. Each layer of the conditional graph neural network directly uses a stacked graph neural network with residual connections for message passing. In the... Layer, through belt The edge multilayer perceptron of the activation function calculates the plot unit. With adjacent land parcel units Interaction features between : ; in, and They represent the first In-layer plot unit and The hidden feature vectors, Indicates the first A multilayer perceptron is used to calculate the interaction features between adjacent land parcel nodes.
[0078] A stacked graph neural network with residual connections is used for message passing, and the node features are updated as follows: ; in, It is a plot of land The neighborhood set, No. Layered plot unit and The edge; Indicates the first -1st floor middle plot unit The hidden feature vector.
[0079] Step 2015, Global Attention Mechanism and Output. The feature vector of the current decision-making parcel unit is used as the query object. The set of feature vectors of all land parcel units then serves as the key. ,value .
[0080] ; in, Indicates the current decision-making plot unit at time step Graph neural network embeddings (where (Embedding dimension for nodes). Representing the status diagram of urban land parcel units Embedding matrix of each plot unit, This is a validity mask for nodes, used to exclude invalid nodes from affecting the attention mechanism; This represents the multi-head attention calculation function; This represents the current decision plot context vector after attention enhancement. Finally, the conditional graph neural network encoder outputs the state features of the current decision plot unit; the state features of the current decision plot unit include: embedded encoding of global features, graph-level summary feature representation of nodes and edges, the current decision plot context vector after attention enhancement, stage information, and planning preference vector.
[0081] Step 202: Input the state features of the current decision plot unit output by the conditional graph neural network encoder into the policy network and the value network.
[0082] First, the conditional policy network generates an action logits probability distribution for candidate actions in action space A based on the state features output by the conditional graph neural network encoder. Then, it receives a constraint-aware action mask from the urban environment simulation module, performs a masking operation on the action log probabilities, and sets the logits corresponding to actions that do not meet planning constraints to a minimum penalty value, thereby filtering out invalid actions that do not comply with urban planning regulations and constraints. Finally, it normalizes the probability distribution of the masked action log probabilities and the masked logits using a normalized exponential function, obtaining a probability distribution of reasonable actions, and samples the planning actions of the planning agent module. .
[0083] Specifically, the constraint-aware action masking mechanism receives the proposed planning actions from the agent's planning module, generates action masks in real-time based on urban planning constraints, assesses the compliance of the proposed planning actions, and feeds this information back to the conditional policy network. It introduces constraint-aware action masks into the planning decision-making loop, intervening after the policy network outputs the original logarithmic probability value but before probability sampling, effectively narrowing the action space and ensuring compliance with the planning constraint set. This ensures that the random exploration of the planning agent remains consistent with mandatory planning compliance requirements, specifically including: The constraint-aware action masking mechanism functions in two ways: in the topological feasibility constraint set ( In terms of ensuring that the actions of the planning intelligent agent only target "unallocated" and "feasible" land parcels, maintaining the stability of the existing urban structure; and in terms of the control buffer constraint set ( In terms of ensuring that the actions of the planning intelligence agent comply with spatial rules for urban safety and environmental health, this mechanism employs a dynamic algorithm to verify the rationality of planning actions based on three rules: Health and safety isolation restrictions: To reduce the risk of pollution exposure, strict buffer zone controls are implemented around industrial and logistics sites. For example, residential areas and sensitive public facilities (such as schools and hospitals) are prohibited within the safety radius of industrial sites to ensure environmental safety and public health.
[0084] Functional Heterogeneity and Balanced Distribution Rules: To avoid functional redundancy and promote fairness in service coverage, minimum spacing constraints are set for similar public facilities. Specifically, if a newly built primary school is located within the service radius of an existing school, that plot of land will not be allocated, thereby forcing the planning agency to distribute educational resources in a more balanced manner, rather than concentrating them in high-density areas.
[0085] The rule to avoid the NIMBY effect: In order to reduce social conflict, noise-sensitive facilities (such as hospitals and schools) should be located away from noise sources such as entertainment venues or busy logistics hubs.
[0086] set up Indicates the policy network in state and preference vector The original output probability generated below. Represents the action space, Indicates the number of candidate actions. (At time step) Construct Boolean action mask vector ,in Indicates action satisfy All of the aforementioned conditions, and This indicates a violation. The masking operation sets the logical value of the invalid action as the penalty value. Corrected output probability Defined as: ; The urban environment simulation module first calculates a Boolean action mask based on these spatial rules. The policy network applies this mask to its output logic value before maximizing activation. Invalid actions are assigned extremely low valid logic values, with a low selection probability. Approaching zero: ; Ensure that the exploration space of the planning agent is strictly limited to the boundary of feasible solutions, where Representing the action space Any candidate action in the list.
[0087] The conditional value network takes the state features output by the conditional graph neural network encoder as input and uses a multilayer perceptron to map these state features into a value estimate. This value estimate On the one hand, it is used to construct value loss to update the conditional value network; on the other hand, it serves as a value baseline in proximal policy optimization to assist in updating the conditional policy network.
[0088] Step 203: At the end of each complete process of the planning agent's interaction with the simulation environment, a four-dimensional reward vector is calculated based on the reward model of the coupled three-dimensional space, and these vectors are aggregated into a total reward scalar using the planning preference vector. This total reward scalar, together with the step-by-step rewards, is used for proximal policy optimization updates to update the conditional policy network and the conditional value network.
[0089] The formula for calculating the total reward scalar is as follows: ; in, The total reward scalar; This represents the production space reward value; This indicates the weight of production space incentives under planning preferences; This represents the reward value for living space; This indicates the weight of living space incentives under planning preferences; Indicates the reward value for ecological space; This indicates the weight of ecological space reward items under planning preferences; This represents the reward value for the synergistic function of the three-life space; This indicates the weight of the reward item for the synergistic function of the three-life space under the planning preference.
[0090] After completing a full planning decision through multiple time steps, the conditional value network outputs the current global planning environment state. In planning preferences Value estimation below This value estimate, along with the scalarized reward, is used to approximate the long-term expected return of this state. The value estimate and the scalarized reward are used together to calculate the dominance function and the reward objective, where the dominance function is used to update the conditional policy network, and the reward objective is used to update the conditional value network. The reward model for the coupled three-life space is a composite reward model that includes rewards for production space, living space, ecological space, and synergistic rewards for the functions of the three-life space.
[0091] Specifically, the reward function model for the coupled three-space system assesses the balance of production, living, and ecological spaces from the perspective of industry-city integration. The reward function divides land use into functional groups to evaluate spatial balance in the urban planning process. The three-space system includes: production space (industrial and warehousing land), living space (residential land including public facilities), and ecological space (green space). To assess the balance of the three-space system in the urban layout, a composite reward model is adopted, incorporating rewards for production space, living space, ecological space, and synergistic functional rewards for the three-space system. This is based on the actions of the planning intelligent agent. With environmental conditions Update environment status to Calculate instant rewards The data is then fed back to the spatial data input module to form a closed-loop iteration. The final reward value is a weighted combination of production space rewards, living space rewards, ecological space rewards, and functional synergy rewards for the three spaces, reflecting their spatial interaction and synergistic effects.
[0092] The reward function model of the coupled three-dimensional space specifically includes: Production Space Reward Function: This reward function is based on industrial agglomeration theory and uses the global Moran index I to quantify the structural compactness of production space. Unlike primary indicators that treat all industrial land equally, this framework implements a hierarchical optimization strategy, achieving a balance between macro-level agglomeration effects and micro-level functional differentiation. Core Indicator: Global Moran Index The definition is as follows: ; in, This indicates the total number of feasible land parcel units within the planning area; Representing land parcel units The binary functional attribute (if the land parcel unit belongs to the target production category, then it is) Otherwise ); It is the global mean of the functional attribute ( ), representing the overall proportion of the target land use; Represents the spatial weight matrix The elements of this matrix are constructed using row-normalized adjacency to capture topological adjacency relationships. To ensure metric compatibility within the reward system, the original index... Through linear transformation Mapping to normalized interval The reward calculation function operates on two levels: at the macro level, it optimizes the aggregation effect of all production functions (including manufacturing and warehousing) to achieve the target value. At the micro level, it distinguishes between subtypes with different objectives: manufacturing functions have more stringent target values for achieving industrial-scale efficiency. The warehousing function sets appropriate target values for providing logistical support. .
[0093] Final production space reward ( ) is the overall aggregate score ( ) and subtype weighted specificity score ( The average value of ) ; in, , and These represent the normalized Moran's index values for the entire production space, industrial manufacturing land, and warehousing land, respectively. This composite objective function drives the planning agent to explore layout schemes that are both spatially compact and structurally rational, aligning with the "intensive and efficient" land use policy requirements.
[0094] Living space reward function ( Based on the concept of a "15-minute community life circle," the reward function is determined by coverage ( ) and diversity ( Service accessibility is evaluated using two dimensions. The reward value is calculated using the following formula: ; ; in, This indicates the coverage rate of public services for residential land. Represents a collection of residential land spaces; It represents a collection of public service facilities. Indicates the service buffer; Indicates the area of a spatial region; This represents the radius of the service buffer zone constructed centered on the collection of public service facilities; here, the value is taken as 500 meters. The Shannon entropy value represents the reachability of service types within the buffer zone, ensuring a balanced allocation of education and healthcare resources. This weighting scheme (coverage 0.6, diversity 0.4) reflects the planning principle of "supply first, optimization second," meaning that basic service accessibility is prioritized before improving the diversity of service functions.
[0095] Ecological space reward function ( In response to the national strategy for spatial equity in ecosystem services and the "Park City" initiative (which emphasizes spatial equity rather than sheer quantity), this function assesses green space coverage by calculating the proportion of residential areas accessible to green space buffer zones, ensuring that high-quality green public facilities can equitably benefit all residents. ; in, It refers to a collection of ecological green spaces or green areas; This represents the radius of the service buffer zone centered on a collection of ecological green spaces or green areas; here, the value is 300 meters.
[0096] Three-Life Space Functional Synergistic Reward Function ( This reward function employs compatibility-weighted buffer analysis to evaluate the spatial logical relationships between residential, industrial, and service areas. The method uses a policy-driven piecewise linear mapping to transform the original spatial coverage rate into a standardized reward signal. : ; in, This indicates an averaging operation. Each sub-indicator The calculation process is as follows: First, the original coverage rate is calculated based on spatial intersection analysis. Then, it is mapped to a score through a function. The key is that, to ensure the robustness of the algorithm, when a specific land use category is missing (e.g., a residential area has not yet been designated), the function assigns a neutral score of 0.5 by default. This mechanism prevents noise or invalid gradients from entering the value network, thereby stabilizing the value evaluation process during the initial training phase.
[0097] in, The measurement focuses on the proportion of residential areas that offer convenient access to public services while avoiding proximity to industrial zones. ; in, This represents the ratio of residential area located within the 500-meter service buffer zone but strictly outside the 300-meter industrial zone buffer zone to the total residential area.
[0098] This measures the proportion of residential areas that are green and ecologically sound and far from industrial zones. ; in, This represents the ratio of residential area covered by the 300-meter service buffer zone of ecological green space (excluding green space within 200 meters of industrial areas) to the total residential area.
[0099] It measures the proportion of residential areas that simultaneously possess the advantages of both service facilities and green space coverage, promoting the formation of a high-quality "15-minute living circle": ; in, This indicates the proportion of residential area that is simultaneously covered by a 500-meter public service buffer zone and a 300-meter green space buffer zone to the total residential area.
[0100] The above is an example illustrating the intelligent urban land use planning method for the industry-city integration scenario in this embodiment. Its key feature is the use of a constraint-aware deep learning method for intelligent planning decisions of land parcel units. As a method employing a deep learning model, the intelligent urban land use planning method requires pre-setting hyperparameters and undergoing training.
[0101] For example, in the intelligent planning method for urban land use, the network result is set as follows: the hidden layer dimension of the graph neural network is 16, the node embedding encoding dimension is 16, the hidden unit dimension of the policy network head is [32,1], and the hidden unit dimension of the value network head is [32,32,1]. The optimizer adopts the adaptive moment estimation optimizer, with a learning rate of 4.0e-4, a small constant of 1.0e-5, a gradient pruning of 0.05, and a minimum batch size of 128. In the near-end policy optimization algorithm, the discount factor is 0.99, the generalized advantage estimation parameter is 0.95, the pruning parameter is 0.2, the entropy coefficient is 0.01, and the value loss coefficient is 0.5. The training principle is that the agent interacts with the environment for a maximum of 500 rounds, the maximum number of steps is 30, the update cycle is 4, the total maximum number of decisions is 15 million, and the training is iterated 1000 times.
[0102] In summary, this application constructs a constraint-aware deep reinforcement learning framework for complex industry-city integration scenarios. This framework includes a conditional graph neural network encoder, a constraint-aware masking mechanism, and a reward model coupling the three spaces (life, ecology, and environment). The conditional graph neural network encoder incorporates Pareto weight vectors into the land use state representation, enabling the planning agent to generate multiple sets of trade-off optimization strategies at once. The constraint-aware masking mechanism verifies the feasibility of decision-making actions based on land use constraints and restrictions, ensuring the rationality of planning decisions. The reward model coupling the three spaces guides the agent's decision-making behavior, promoting balanced regional spatial development. This framework outperforms existing benchmarks in terms of scheme diversity, decision rationality, and planning balance, and possesses strong transferability, providing planners with diverse balanced planning solutions.
[0103] This application addresses the challenge of coordinated development planning for the "production-living-ecology" three-dimensional spaces within the context of industry-city integration. Utilizing territorial spatial planning data and current land use data, it constructs an intelligent urban land use planning model. This model reconstructs urban land use planning into an optimization-driven sequential decision-making process, capable of both perceiving spatial topological logic and satisfying planning constraints. It overcomes many shortcomings of traditional planning methods in dealing with dynamic policy adjustments and strict planning constraints. The core ideas of this application are mainly reflected in three aspects. First, the constructed conditional graph neural network encoder successfully integrates spatial geometry and planning preferences. This idea enables the model to understand the "spatial grammar" of the urban environment and generate Pareto optimal solutions adapted to different planning needs without retraining, providing multiple planning schemes. Second, the designed constraint-aware action mask mechanism effectively bridges the gap between algorithm exploration and planning constraints. By directly embedding planning constraints into the decision-making closed loop, it ensures that the generated planning scheme is not only numerically optimal but also feasible in practical planning. Third, the proposed reward function model for coupled industrial, ecological, and environmental spaces quantifies the complex synergistic effects among these spaces, guiding the model to autonomously discover solutions that resolve neighborhood conflicts while balancing spatial elements. Based on these three aspects, in the context of industry-city integration, compared with traditional baselines and heuristic algorithms, the planning decision-making scheme presented in this application maintains high rewards for industrial spaces while achieving coordinated development of living, ecological, and other spatial functions. The method proposed in this application can serve as an interactive decision-making support tool to help planners address the challenges of complex future urbanization planning needs.
[0104] To evaluate the effectiveness of the proposed method, this embodiment uses traditional planning methods as benchmarks for comparison. These benchmarks include: centralized global optimization algorithms, decentralized local heuristic algorithms, greedy space coverage algorithms, and expert planning schemes. This embodiment analyzes the comparative experiments between these methods and four representative planning benchmarks. All methods operate under a unified reward definition and weight configuration, and their planning decision performance is quantified through four dimensions: production space reward assessment, residential space reward assessment, ecological space reward assessment, and synergistic reward assessment of the three functions (production, living, and ecological functions). (See Table 1 for details.)
[0105] Table 1 Performance Comparison of Various Models
[0106] The proposed method demonstrates significant advantages over all algorithmic benchmarks. It achieves excellent results with production space rewards of 0.954, living space rewards of 0.419, ecological space rewards of 0.603, and a synergy reward of 0.931. Compared to centralized global optimization algorithms and decentralized local heuristic algorithms, this method achieves a qualitative leap in ecological space and synergy, with the ecological space reward reaching nearly 3.7 times that of the centralized benchmark (from 0.165 to 0.603), and the synergy reward also doubling. Even against rule-based greedy space covering algorithms, this method maintains a robust leading advantage, improving ecological space rewards and synergy rewards by 22.8% and 20.6%, respectively. This demonstrates that data-driven strategies can learn more complex and refined spatial configuration patterns than static rules.
[0107] In experiments comparing the proposed solution with that of human experts, although the expert solution, as a reference, outperformed the human expert solution in terms of production space reward (0.937) and living space reward (0.460), the proposed method achieved comprehensive superiority in production space (+1.8%), ecological space (doubling from 0.300 to 0.603), and synergy (+22.7%), only making a strategic, moderate concession in quality of life (-8.9%). This demonstrates the proposed method's superior ability to discover optimal solutions.
[0108] Furthermore, a significant advantage of the method in this application is its ability to generate a series of Pareto optimal land use strategies from a single training strategy. To demonstrate this, this embodiment selects four representative solutions from the final output Pareto set of the study area units, showing how the method responds to different preference weight settings within the three-dimensional space.
[0109] Scenario A: Production-oriented configuration. For example... Figure 5As shown in (a), under a production-first weighting, the model adopts a highly aggressive infill strategy. Utilizing existing orange industrial land in the southeast as its core, the model extensively expands industrial functions into the blank areas in the central and western regions, forming a continuous industrial block occupying half of the jurisdiction. This high-density, contiguous layout maximizes the industrial agglomeration effect. However, this expansion squeezes out potential living space—the newly added industrial zone directly fills the south side of the northern residential area, lacking the necessary buffer transition, resulting in low spatial synergy.
[0110] Scenario B: Service-oriented configuration. When preferences shift towards quality of life ( Figure 5 In (b) of the model, the logic for filling vacant plots underwent a qualitative change. The previously undeveloped central area was not filled with single-function plots, but rather occupied in a decentralized manner by red commercial land and blue public service facilities. Particularly in the northeast and central regions, the penetration of service facilities significantly enhanced coverage of existing residential areas in the north. This hybrid infill strategy effectively eliminated the original service blind spots, reflecting a strategic shift from "production scale" to "service accessibility."
[0111] Scenario C: Eco-oriented configuration. In an ecologically priority scenario ( Figure 5 In (c) of the map, green ecological land has become the dominant force in spatial restructuring. Compared with the initial map, the existing green patches in the south have not only been preserved but have also extended significantly northward, like a wedge inserted into the undeveloped central area. At the same time, a continuous green buffer zone has also emerged on the western boundary. This layout strategy actively cuts off the path of disorderly westward expansion of the southeastern industrial zone and uses the expanded green space to divide the construction land. Although it limits the increase in industrial land, it greatly enhances environmental safety and ecological resilience by constructing a "green skeleton".
[0112] Scenario D: Collaborative (Balanced) Configuration. When functional synergy is emphasized ( Figure 5 (d) The model demonstrates an optimal ability to coordinate the existing and incremental aspects. Utilizing the western blank area, the model consolidates and expands the yellow residential land, creating a coherent residential community with the existing northern area. The orange industrial land is strictly confined to the existing southeastern base, maintaining compactness while curbing its westward expansion. Crucially, the model constructs a composite buffer zone of green ecological patches and blue / red service facilities in the central blank area between the "new western residential area" and the "existing southeastern industrial area." This "sandwich" spatial structure (residential-buffer-industrial) adheres to the initial spatial base while effectively addressing potential pollution disturbances, representing the optimal comprehensive compromise at the Pareto front "knee."
[0113] These experimental results demonstrate that the method proposed in this application achieves coordinated development of living, ecological, and other spatial functions while maintaining high industrial space incentives. It provides an interactive decision-making tool for intelligent space planning, helping planners to meet the challenges of complex planning needs in future urbanization.
[0114] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 6 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores data used in the process of implementing an intelligent urban land use planning method for an integrated industry-city development scenario. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements an intelligent urban land use planning method for an integrated industry-city development scenario.
[0115] Figure 6 The structures shown are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0116] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0117] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0118] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with relevant regulations and be authorized by the owner of the corresponding device.
[0119] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0120] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0121] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0122] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A smart urban land use planning method for industry-city integration scenarios, characterized in that, The method includes: Acquire urban spatial data and planning preference vectors; divide the urban spatial data into land parcel unit data; the urban spatial data includes at least land spatial planning data and current land use data; The land parcel unit data is used as the global planning environment state at the current time step; Construct a plot unit state diagram for the current time step based on the global planning environment state at the current time step. Based on the state diagram of the land parcel unit and the planning preference vector, the state features of the current decision land parcel unit are obtained using a conditional graph neural encoder. The state features of the current decision land parcel unit include: embedding encoding of global features, graph-level summary feature representation of nodes and edges, current decision land parcel context vector after attention enhancement, stage information, and planning preference vector. Based on the aforementioned state characteristics, the planning action of the current decision-making plot is obtained through a conditional policy network using a constraint-aware action masking mechanism. Based on the aforementioned state characteristics, the value estimate of the global planning environment state under planning preferences at the current time step is obtained through the conditional value network. Execute the planning action, update the global planning environment state, obtain the global planning environment state of the next time step, and calculate the immediate reward; use the global planning environment state of the next time step as the global planning environment state of the current time step and return "Construct the plot unit state diagram of the current time step based on the global environment planning state of the current time step" until all plot units have completed planning; the immediate reward is used to encourage effective interaction; After a round of planning is completed, a reward vector is calculated based on the reward model of the coupled three-life space, and aggregated into a total reward scalar according to the planning preference vector; based on the value estimate at the current time step and the total reward scalar, the network parameters of the conditional policy network and the network parameters of the conditional value network are updated; the reward model of the coupled three-life space is a composite reward model that includes production space rewards, living space rewards, ecological space rewards and functional synergy rewards of the three-life space. The updated conditional policy network outputs optimal planning and decision schemes with different preference orientations.
2. The intelligent urban land use planning method for industry-city integration scenarios according to claim 1, characterized in that, Based on the land parcel unit state diagram and the planning preference vector, the state features of the current decision land parcel unit are obtained using a conditional graph neural encoder, specifically including: A global feature vector is obtained based on the state diagram of the land parcel unit. The global feature vector is processed by a multilayer perceptron to obtain the embedding encoding of the global features; Based on all node features and all edge features in the state diagram of the land parcel unit, a graph-level summary feature representation of nodes and edges is obtained; Based on the current time step or stage of the decision-making process, obtain stage information; An initial embedding space is constructed based on the embedding encoding of the global features; The feature vector of each node in the state diagram of the land parcel unit at the current time step is projected to the initial embedding space by linear transformation to obtain the initial node representation; The preference vector is projected into a latent planning preference embedding through a multilayer perceptron. The planning preference embedding is then scaled by a scaling factor, broadcast, and added to the initial node representation to obtain the preference-enhanced initial node features. Based on the spatial adjacency relationship depicted by the edge set in the state graph of the land parcel unit and the initial node features of the preference enhancement, neighborhood information is aggregated through a stacked conditional graph neural network to obtain the feature vector set of all land parcel units. Based on the set of feature vectors, the node features corresponding to the current decision plot are extracted; at the same time, the node features of the current decision plot are used as the query, and the set of feature vectors of all plot units are used as the key and value, and the attention-enhanced context vector of the current decision plot is calculated through a global attention mechanism. The state features of the current decision plot unit are obtained by concatenating the global feature embedding encoding, the graph-level summary feature representation of nodes and edges, the attention-enhanced current decision plot context vector, stage information, and planning preference vector.
3. The intelligent urban land use planning method for industry-city integration scenarios according to claim 1, characterized in that, Based on the aforementioned state characteristics, a constraint-aware action masking mechanism is used to obtain the planning action for the current decision-making plot through a conditional policy network, specifically including: The state features are input into a multilayer perceptron in a conditional policy network and mapped to action log probabilities. Using a constraint-aware action masking mechanism, the logarithmic probability of the action is masked to obtain the masked logarithmic probability of the action. By using a normalized exponential function and a sampling function, the probability distribution of the logarithmic probability of the masked action is normalized, and the planning action of the current decision plot is obtained by sampling.
4. The intelligent urban land use planning method for industry-city integration scenarios according to claim 3, characterized in that, Using a constraint-aware action masking mechanism, the logarithmic probability of the action is masked to obtain the masked logarithmic probability of the action, specifically including: Boolean action mask vectors are calculated based on three rules in the constraint-aware action masking mechanism: the hygiene and safety isolation restriction rule, the functional heterogeneity and distribution equilibrium rule, and the NIMBY (Not In My Backyard) avoidance rule. The Boolean action mask vectors are used to mark the legality of candidate actions in the action space. The action log probability is masked by the Boolean action mask vector to obtain the masked action log probability.
5. The intelligent urban land use planning method for industry-city integration scenarios according to claim 4, characterized in that, The expression for the action mask operation is: ; in, This indicates the corrected probability of the action output; Indicates the original action; Indicates a time step; In time step Construct a Boolean action mask vector; Indicates the number of candidate actions; Indicates at time step primitive action Boolean action mask vector; This indicates that a violation has occurred; This represents the probability of the original action output.
6. The intelligent urban land use planning method for industry-city integration scenarios according to claim 1, characterized in that, The expression for the production space reward function is: ; in, This represents the production space reward value; , and These represent the normalized Moran index values for the entire production space, industrial manufacturing land, and warehousing land, respectively. The target value for production space; The target value for industrial manufacturing land; The target value for warehousing land; Indicates the overall aggregate score; This represents the subtype-weighted specificity score; The expression for the living space reward function is: ; ; in, This indicates the coverage rate of public services for residential land. Indicates the area of a spatial region; Represents a collection of residential land spaces; Indicates the service buffer; It represents a collection of public service facilities. This indicates the radius of the service buffer zone built around the public service facility; This represents the reward value for living space; Indicates radius as The Shannon entropy value of the service type that can be reached within the buffer; The expression for the ecological space reward function is: ; in, Indicates the reward value for ecological space; It refers to a collection of ecological green spaces or green areas; The radius of the service buffer zone centered on the collection of ecological green spaces or green areas; The expression for the collaborative reward function of the three-life space is as follows: ; in, This represents the reward value for the synergistic function of the three-life space; This indicates an averaging operation; It measures the proportion of residential areas that offer convenient access to public services while avoiding proximity to industrial zones; It measures the proportion of residential areas that serve the green ecology of residents and are far away from industrial areas; It measures the proportion of residential areas that have the dual advantages of both service facilities and green space coverage; The formula for calculating the total reward scalar is as follows: ; in, The total reward scalar; This represents the production space reward value; This indicates the weight of the production space incentive items under the planning preference; This represents the reward value for living space; This indicates the weight of living space incentives under planning preferences; Indicates the reward value for ecological space; This indicates the weight of ecological space reward items under planning preferences; This represents the reward value for the synergistic function of the three-life space; This indicates the weight of the reward item for the synergistic function of the three-life space under the planning preference.
7. An intelligent urban land use planning device for industry-city integration scenarios, characterized in that, The device is used to implement the intelligent urban land use planning method for the integration of industry and city as described in any one of claims 1-6. The device includes: a planning status perception module, a planning intelligent agent module, and an urban environment simulation module. The planning status perception module is used to acquire urban spatial data and planning preference vectors; divide the urban spatial data into plot unit data; the urban spatial data environment includes at least land spatial planning data and land use status data; use the plot unit data as the global planning environment status at the current time step; and construct the plot unit status map at the current time step based on the global planning environment status at the current time step. The planning agent module is used to obtain the state features of the current decision-making plot unit based on the plot unit state diagram and the planning preference vector using a conditional graph neural encoder. The state features of the current decision-making plot unit include: embedding encoding of global features, graph-level summary feature representation of nodes and edges, attention-enhanced current decision-making plot context vector, stage information, and planning preference vector. Based on the state features, the planning action of the current decision-making plot is obtained through a conditional policy network using a constraint-aware action masking mechanism. Based on the state features, the value estimate of the global planning environment state under the planning preference at the current time step is obtained through a conditional value network. The urban environment simulation module is used to execute the planning actions, update the global planning environment state, obtain the global planning environment state of the next time step, and calculate the immediate reward; it uses the global planning environment state of the next time step as the global planning environment state of the current time step and returns "constructing the land parcel unit state diagram of the current time step based on the global environment planning state of the current time step" until all land parcel units have completed planning; the immediate reward is used to encourage effective interaction; after a round of planning, a reward vector is calculated based on the reward model of the coupled three-life space, and aggregated into a total reward scalar according to the planning preference vector; based on the value estimate of the current time step and the total reward scalar, the network parameters of the conditional policy network and the network parameters of the conditional value network are updated; the reward model of the coupled three-life space is a composite reward model that includes production space rewards, living space rewards, ecological space rewards and three-life space functional synergy rewards; based on the updated conditional policy network, the optimal planning decision scheme with different preference orientations is output.
8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that the processor executes the computer program to implement the intelligent urban land use planning method for the integrated urban-industrial development scenario as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the intelligent urban land use planning method for the integrated development of industry and city scenario as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the intelligent urban land use planning method for the integrated development of industry and city scenario as described in any one of claims 1-6.