Territorial space intelligent collaborative management system
By constructing an intelligent collaborative management system for national land space, and utilizing a digital twin environment, a strategy emergence engine, and a human-machine collaborative cockpit, the systemic fragmentation and computing resource bottlenecks have been resolved, enabling efficient multi-objective collaborative optimization and sustainable development decision-making.
Patent Information
- Application Number
- CN202511056996.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-04
AI Technical Summary
The existing territorial spatial planning system suffers from systemic fragmentation, limited decision-making space, static bottlenecks in computing resources, and a disconnect from knowledge application, making it difficult to achieve efficient, multi-objective collaborative optimization and sustainable development decision-making.
Construct an intelligent collaborative management system for national land space, including a digital twin environment, a strategy emergence engine, and a human-machine collaborative cockpit. Through multi-objective interactive learning between intelligent agents and the environment, generate national land space collaborative development strategies with Pareto optimality and spatiotemporal consistency.
It has significantly improved the level of intelligence in land and space management, solved the problems of systemic fragmentation, limited decision-making space, and bottlenecks in computing resources, and achieved efficient multi-objective collaborative optimization and sustainable development decision-making.
Smart Images

Figure CN120893984A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of land space planning and artificial intelligence technology, in particular to a land space intelligent collaborative management system. BACKGROUND
[0002] Currently, the planning and sustainable development of land space is a complex system engineering involving multiple dimensions such as economy, society, ecology, and resources. The technical means supporting such decisions are mainly based on the integration of geographic information systems, various professional field models (such as hydrology, land use, macroeconomic models), and expert knowledge bases. These technologies lay a solid foundation in data management, status analysis, and single-field deduction. In recent years, artificial intelligence, especially data-driven machine learning methods, have been introduced into this field, significantly improving the prediction accuracy in specific tasks such as land use change simulation, urban expansion prediction, and natural disaster risk assessment.
[0003] However, the existing technology has the following problems: systematic fragmentation: various field models are like independent computing modules, although they can run separately, but the depth coupling and real-time feedback mechanism between them is weak; limitations of decision space; lack of ability to actively "create" and "conceive" new planning paradigms; static bottleneck of computing resources: fine land system dynamics models often have high computational cost, a complete and long-period simulation may take several days, and the disconnection of knowledge application and other problems.
[0004] Therefore, there is an urgent need for a more intelligent management system to support the current efficient land planning and management needs. SUMMARY
[0005] The land space intelligent collaborative management system provided by the present application aims to build a computable and self-evolving land space management system, which can emerge and generate a set of land space collaborative development strategy combination with Pareto optimality, spatiotemporal consistency and robustness to uncertainty in a digital twin environment deeply integrated with "geoscience-ecology-society" multidimensional knowledge through long-term, multi-objective interaction learning between agents and environment.
[0006] The system includes: a digital twin environment for providing a simulation environment for decision-making agents to learn and deduce, and building a highly realistic and efficiently computable land system "mirror"; a strategy emergence engine for generating, evaluating and optimizing land space planning strategies, and realizing the autonomous "emergence" of excellent strategies; a human-machine collaborative cockpit for human-machine interaction and realizing efficient and intuitive unbiased human-machine collaborative decision-making.
[0007] The digital twin environment includes: a high-fidelity and dynamically evolving national knowledge system, which provides an upper intelligent agent with a simulation environment as close to the real world as possible and reflecting the inherent laws of the system; and an efficient and computable national system dynamics module, which improves the computational efficiency of simulation by several orders of magnitude while ensuring key dynamic characteristics and scientific mechanisms.
[0008] The high-fidelity and dynamically evolving national knowledge system includes: a deep quality control unit for identifying and quantifying the three core problems of knowledge redundancy, concept imbalance, and semantic ambiguity in the national knowledge system; a knowledge unit with feedback and self-learning capabilities for dynamically improving the self-learning capabilities of the national knowledge system; and an "online concept balancing" loss function for dynamically training and correcting the models in the national knowledge system to overcome the bias problem caused by the long-tail distribution of data.
[0009] The formula of the loss function is: The total loss of the knowledge system is E data ~ real distribution [concept imbalance response distance (predicted state, planning concept) distance (real state, predicted state)].
[0010] The efficient and computable national system dynamics module includes: a hierarchical agent model architecture based on the "teacher-student" paradigm, which is used to build one or more lightweight and extremely fast computing proxy models (student models) from a large, detailed, and slow computing complete national system dynamics model (teacher model) through advanced model compression and knowledge transfer techniques; A sliding window iterative refinement unit is used to maintain spatiotemporal continuity and logical self-consistency in long-term simulation periods, suppress error accumulation and divergence, and ensure the fast computation and stable simulation of the proxy model while maintaining scientific robustness. A "principle-guided-knowledge-distillation" review training unit is used to train a proxy dynamics model with high speed and ensure its accuracy, scientificity, and long-term stability.
[0011] The strategy emergence engine includes: a structured strategy generation module for simulating the thinking process of human experts during planning; and a collaborative multi-objective learning module for learning and generating a strategy set covering the entire Pareto optimal frontier and introducing advanced semantic understanding capabilities to handle complex objectives that are difficult to quantify.
[0012] The structured strategy generation module includes: a decoupled-autoregressive strategy generation pipeline for decoupling the generation process of the territorial space strategy into an autoregressive pipeline containing three stages. The autoregressive means that the decision of the next stage is conditioned on the output of the previous stage, thereby ensuring the internal logical coherence of the entire strategy; a cross-scale hierarchical strategy learning framework for simulating the multi-scale governance structure of territorial space planning; and a structured strategy "hierarchical-autoregressive" generation probability model.
[0013] The calculation formula of the probability model is: Probability (complete strategy | initial state) = P (strategic intent | initial state) Π_{k=1 to K} P (spatial layout unit_k | strategic intent, layout unit_{<k}) Π_{j=1 to J} P (policy configuration_j | complete spatial layout, policy configuration_{<j}); Where P(...) represents conditional probability, and Π represents the multiplication operation; Π_{k=1 to K} P (spatial layout unit_k | strategic intent, layout unit_{<k}) represents an autoregressive process; and Π_{j=1 to J} P (policy configuration_j | complete spatial layout, policy configuration_{<j}) represents a second regression process.
[0014] The collaborative multi-objective learning module includes: a Pareto frontier exploration unit based on unified supervised contrastive learning for converting discrete, multi-dimensional target evaluation into a unified, internally collaborative learning framework, thereby driving the agent to efficiently explore and approximate the entire Pareto optimal solution; a complex target quantification unit based on "agent as judge" for using the powerful semantic understanding and logical reasoning capabilities of large language models (LLM) to convert complex, qualitative targets into calculable reward signals; and a multi-objective collaborative "contrast-judgment" reward function for collaborative optimization of multiple conflicting objectives and incorporating evaluation of complex qualitative targets.
[0015] The reward function is: Collaborative reward (generated strategy) = log[(Σ_{positive samples} exp(-D_ embedding / τ)) / (Σ_{positive samples} exp(-D_ embedding / τ) + Σ_{negative samples} exp(-D_ embedding / τ))]; Where log[...] and exp(...) are the logarithm and exponential functions, respectively; positive samples and negative samples represent "good" and "bad" benchmark strategy sets, respectively; τ is a temperature hyperparameter for adjusting the sharpness of contrast; and D_ embedding is the weighted distance of two strategies in the target embedding space.
[0016] The human-machine collaborative cockpit comprises: an adaptive computing and instructive exploration module, which is used for intelligently managing and allocating computing resources, so that the system can respond to macro exploration requests of a user in near real time and perform high-fidelity fine simulation on specific and in-depth problems concerned by the user, thereby achieving the best balance between "interaction efficiency" and "analysis depth"; and an unbiased objective decision support module, which is used for taking "eliminating bias" and "ensuring objectivity" as the first principle of system design and building them into the core architecture.
[0017] The adaptive computing and instructive exploration module comprises: a cascaded adaptive simulation unit based on "instructive timing and spatial positioning", which is used for intelligent on-demand allocation of computing resources, greatly improving the efficiency and feasibility of deep interaction analysis; an agent active proof unit based on uncertainty perception, which is used for giving the system self-reflection and information completion ability in the self-learning process through agent active proof; and an adaptive computing decision rule function of human-machine interaction.
[0018] The rule function is: Call model (current state, user instruction) = {complete model, if (focus area (user instruction) ≠ ∅) or (U_ uncertainty (current state) > δ_ threshold); agent model, otherwise} The complete model and the agent model represent high-precision and high-speed dynamic models respectively; the focus area (user instruction) is an analytical function; the U_ uncertainty (current state) is an evaluation function used for calculating the uncertainty of the agent in the "current state" using the agent model for prediction; and the δ_ threshold is a preset uncertainty threshold.
[0019] The unbiased objective decision support module comprises: a reference system bias elimination unit based on replacement isomorphism, which is used for building a completely replacement isomorphic model to ensure that the system does not have any reference system bias at the architecture level; and a knowledge source bias identification and mitigation unit based on benchmark reliability review, which is used for actively identifying and mitigating objective bias caused by unbalanced data acquisition, processing and expression, to ensure the reliability and comprehensiveness of the model decision basis.
[0020] The beneficial effects of the present application are: different from the prior art, the land space intelligent collaborative management system provided by the present application comprises: a digital twin environment, which is used for providing a simulation environment for learning and deduction of a decision intelligent agent, and constructing a highly real and computable land system "mirror"; a strategy emergence engine, which is used for generating, evaluating and optimizing land space planning strategies, and realizing autonomous "emergence" of excellent strategies; a man-machine collaborative cockpit, which is used for realizing interaction of a decision maker with the strategy emergence engine module, and realizing efficient and intuitive unbiased man-machine collaborative decision. Through the above intelligent collaborative management system, the problems of systematic fragmentation, limited decision space, static bottleneck of computing resources, disconnection of knowledge application and the like in current land space management can be solved, and the intelligent management level of land space multi-objective collaborative optimization and sustainable development decision is significantly improved. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Among them: Figure 1 is a structural schematic diagram of an embodiment of the land space intelligent collaborative management system provided by the present application; Figure 2 is a structural schematic diagram of an embodiment of the digital twin environment provided by the present application; Figure 3 is a structural schematic diagram of an embodiment of the strategy emergence engine provided by the present application; Figure 4 is a structural schematic diagram of an embodiment of the man-machine collaborative cockpit provided by the present application. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings of the embodiments of the present application. It can be understood that the specific embodiments described here are only used to explain the present application, not to limit the present application. In addition, it should be noted that only the parts related to the present application are shown in the drawings, not all the structures. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0023] Reference to“an embodiment” herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment. The appearances of the phrase“in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily all referring to a common set of embodiments, of the other embodiments. It will be explicitly understood that the embodiments described herein can be combined with other embodiments in various ways.
[0024] In today's era of popularization of intelligent systems, the planning and management of territorial space has also entered the era of artificial intelligence management, but the existing technology has the following four technical problems: 1. Systematic fragmentation: Various domain models are like independent computing modules, although they can run separately, but the depth coupling and real-time feedback mechanism between them are weak. Simply superimposing the results of each model not only makes it difficult to obtain the global optimal solution of the system, but also may cause "synthetic fallacy" due to ignoring the nonlinear correlation across domains, that is, the combination of local optimal solutions damages the overall system.
[0025] 2. Limitation of decision space: The traditional decision-making mode is to compare and select among a few planning schemes preset by experts, which is like examining only a few known small islands in the vast ocean. The vast unknown decision space that may exist better solutions cannot be effectively explored. The application of artificial intelligence is also limited to learning and fitting historical patterns, lacking the ability to actively "create" and "imagine" new planning paradigms.
[0026] 3. Static bottleneck of computing resources: Fine-grained territorial system dynamics models often have high computational costs, and a complete, long-period simulation may take several days. This makes it impractical to conduct large-scale, multi-round "if-then" scenario simulations, thereby severely restricting the application of modern artificial intelligence methods that require massive "trial and error" to learn.
[0027] 4. Disconnection of knowledge application: Although knowledge graphs and other technologies can structurally store vast amounts of territorial space knowledge, including entities, relationships, and rules, this knowledge often remains at the level of querying, retrieving, and static reasoning, making it difficult to dynamically and real-time guide and constrain an ongoing simulation optimization process, resulting in the separation of "knowledge" and "action".
[0028] Therefore, the present application aims to solve the above-mentioned core technical problems, i.e. how to build a computable and self-evolutionary intelligent collaborative management system for national space, so that it can emerge and generate a set of national space collaborative development strategy combination with Pareto optimality, spatiotemporal consistency and robustness to uncertainty in a digital twin environment deeply integrated with "geoscience-ecology-society" multidimensional knowledge through long-term and multi-objective interaction learning between intelligent agents and the environment. Therefore, an intelligent system capable of autonomous learning, discovering and generating sustainable development paths needs to be created. For details, see the following embodiments.
[0029] Referring to Figure 1 , Figure 1 is a structural schematic diagram of an embodiment of the national space intelligent collaborative management system provided by the present application. The system 100 includes three very core subsystems: Core subsystem one: digital twin environment 10, which is a "sand table world" for decision-making agents to learn and deduce. Its goal is to build a highly realistic and efficiently computable "mirror" of the national system.
[0030] Referring to Figure 2 , Figure 2 is a structural schematic diagram of an embodiment of the digital twin environment provided by the present application.
[0031] In some embodiments, the digital twin environment 10 includes a high-fidelity and dynamically evolving national knowledge system 11, and an efficiently computable national system dynamics module 12.
[0032] Among them, the high-fidelity and dynamically evolving national knowledge system 11 is mainly realized through three levels of innovative technology: Level one: deep quality control unit, used to identify and quantify the three core problems of knowledge redundancy, concept imbalance and semantic ambiguity in the national knowledge system; traditional knowledge system construction often stops at the aggregation and integration of multi-source data, which easily leads to the problem of "garbage in, garbage out", i.e. errors, biases and inconsistencies hidden in the data will directly pollute subsequent analysis and decision-making. To achieve high fidelity, we have established a systematic, life-cycle deep quality control mechanism throughout the knowledge system.
[0033] Among them, the quality control mechanism mainly includes a systematic benchmark quality audit and meta-knowledge annotation mechanism, which is realized through an automatic quality audit process, including: 1. Knowledge redundancy detection, for example, there may be slight differences in the classification of the same plot of land in land use data released by different government departments, but they refer to the same thing; or in time series data, the statistical indicators change slightly year after year. The system automatically detects such overlapping, approximate or low-information knowledge units and fuses or marks them, avoiding wasting computing resources on redundant information.
[0034] 2. Concept imbalance judgment: Land space data naturally has an extreme long-tail distribution. For example, data on economic activities in megacities is extremely rich (head concept), while data on rare species habitats in remote areas or specific types of geological disasters is very sparse (tail concept). Our audit process systematically quantifies this imbalance in data distribution and records it as an inherent property of the knowledge system.
[0035] 3. Semantic ambiguity evaluation: A large amount of land space knowledge comes from natural language texts such as policy documents and planning outlines, which are full of fuzzy concepts such as "high-quality development", "ecological civilization construction", and "regional coordination" that are difficult to understand directly at the machine level. The system maps these fuzzy expressions to a set of quantifiable and computable index systems through large language models and domain ontologies, and evaluates the uncertainty of this mapping.
[0036] After completing the audit through the above strict quality audit process, the system will attach a meta-knowledge layer to each entity, relationship or rule in the knowledge graph. This layer contains key meta-information such as confidence, timeliness, data density and applicable boundaries. This allows the knowledge system to evolve from a simple "collection of fact statements" into a "knowledge network with self-awareness and uncertainty description", providing a foundation for the robustness of upper-layer models and the reliability of decision-making.
[0037] Level two: knowledge units with feedback and self-learning ability, used to dynamically improve the self-learning ability of the land knowledge system; "high fidelity" ensures the accuracy of the knowledge system at a certain moment, while "dynamic evolution" gives it the ability to self-improve over time and through external interaction. This marks the transition of the knowledge system from a passive data "repository" to an active "learning system".
[0038] Among them, the knowledge unit with feedback and self-learning ability contains two important mechanisms: 1. Paradigm shift from offline data management to online concept balancing. Traditionally, the approach to data imbalance is to conduct complex offline data resampling or manually adjust weights. This way not only costs a lot, but also is one-time and cannot adapt to dynamic environment. We introduce an online concept balancing idea, which deeply binds the "evolution" of the knowledge system with its "use" process.
[0039] 2. Concept response strength driven dynamic weight adjustment mechanism. The core workflow of this mechanism is as follows: A. When the upper-level decision engine (e.g., reinforcement learning agent) uses this knowledge system for training or reasoning, the knowledge system will continuously monitor its performance; B. The system will pay special attention to the performance of tasks related to "tail concepts" (i.e., concepts with sparse data but strategic significance, such as the "rare species habitat" mentioned above); C. If the system finds that the model continues to perform poorly in decisions involving these tail concepts (e.g., low prediction accuracy, poor strategy effect), it will calculate a "concept imbalance response distance" - the difference between the model's performance on this concept and its performance on the head concepts with abundant data; D. This "distance" will be dynamically converted into a loss weight in real time, which will be fed back to the training process. Specifically, the system will automatically increase the importance of knowledge units related to the tail concept (including its relationships, rules, and constraints) in subsequent model training.
[0040] For example: A planning model simulates the impact of a large-scale project on a small, data-sparse nature reserve. Due to the low weight of related knowledge, its prediction results are greatly biased. The dynamic evolution mechanism of this knowledge system will capture this large error signal related to the "nature reserve" concept. In the next round of training iteration, the system will automatically increase the weight of all geological, ecological, and hydrological knowledge related to the nature reserve in the knowledge graph. This forces the model to "pay more attention" and "respect" these sparse but critical knowledge during the learning process, gradually correcting its behavior and generating strategies that better meet the requirements of ecological protection. Through this online, closed-loop feedback mechanism, the knowledge system achieves dynamic evolution: instead of passively waiting for data updates, it actively discovers and remedies its "cognitive shortfalls" and "biases" through continuous interaction with the upper-level intelligent application, making its description of the national space more balanced and more in line with the internal requirements of sustainable development.
[0041] Level three: "online concept balancing" loss function, used to dynamically train and correct the model in the national knowledge system to overcome the bias problem caused by the long-tail distribution of data, the formula of the function is: Total loss of knowledge system = E_{data ~ real distribution} [concept imbalance response distance (predicted state, planning concept) distance (real state, predicted state)].
[0042] E_{data ~ real distribution} [... ] means to calculate the expectation of all real world data samples (in practice, it is approximated by the mean of small batch samples).
[0043] Real state and predicted state represent the real multi-dimensional attribute vector and the model predicted attribute vector of a certain unit (such as a plot or a region) in territorial space, respectively.
[0044] Distance (...) is a metric function, such as Euclidean distance, used to calculate the error between prediction and real value.
[0045] Concept imbalance response distance (...) is the core creative point of this formula, which is a dynamic weight calculated online, defined as: Concept imbalance response distance (state, concept) = ||gradient_{input} L(model(input|concept)) - gradient_{input} L(model(input|empty))||; Gradient_{input} L(...) represents the gradient of the loss function L with respect to the input data, which reflects how sensitive the model is to small changes in input, i.e. the "attention" of the model; Model(input|concept) represents the model's prediction of the input given a specific "planning concept" (such as "ecological protection red line"); Model(input|empty) represents the model's prediction without any concept constraints in the "unconditional" case; ||...|| represents the norm operation, which is used to calculate the difference between two gradient vectors.
[0046] Further, the explanation of the above formula is as follows: The essence of this formula is that it no longer uses fixed, historical data-based weights, but dynamically adjusts in real time according to the "performance" of the model itself. For a "tail concept" with sparse data (such as "geological disaster hidden points"), the difference in "attention" of the model's prediction under conditional and unconditional circumstances can be very large, resulting in a larger "response distance". This larger distance value acts as a weight, amplifying the contribution of this sample in the total loss, forcing the model to pay more attention to such sparse but critical concepts in subsequent training. This achieves the adaptive and dynamic evolution of the knowledge system in the interaction with the model. This formula is the core mathematical expression for the dynamic evolution of the knowledge system and the solution to the data imbalance problem. It is not directly applied to the construction of the knowledge graph, but rather to all downstream models that utilize this knowledge system for training (e.g., dynamic agent models, policy networks or value networks of reinforcement learning). At each step of these model training, the loss function calculates the weights for different "concepts" in the current data samples (i.e., concept imbalance response distance) online and dynamically, forcing the model to pay more attention to "tail concepts" that are sparse but crucial. This process also reflects the dynamic nature of the knowledge system - it continuously adjusts and optimizes the "influence" of its knowledge in model training through interaction with the upper-layer model, achieving a functional and task-oriented dynamic evolution.
[0047] Among them, the efficient and computable national system dynamics module 12 is mainly composed of the following four technical pillars: Pillar One: Hierarchical Agent Model Architecture Based on "Teacher-Student" Paradigm, used to build one or more lightweight, extremely fast computing agent models (i.e., "student models") through advanced model compression and knowledge transfer technology from the large, detailed, and slow computing complete national system dynamics model (i.e., "teacher model").
[0048] Specific implementation method: 1. Knowledge Distillation (KD): The core task of the student model is not to learn the complex laws of the national system from scratch, but to "imitate" the behavior of the teacher model. We will run the teacher model in an extremely wide input parameter space to generate a large number of "input-output" pairs. The student model (usually a deep optimized neural network) learns to fit these data generated by the teacher model, which contains complex system dynamics knowledge. The goal is to have the student model produce highly similar outputs to the teacher model for the same input in most cases.
[0049] 2. Structured Pruning & Low-Rank Approximation: To compress the size and computational load of the student model to the extreme, we employ systematic model compression techniques. Rather than simply reducing the number of network layers or neurons, we identify and "prune" connections, channels, and even entire structural blocks in the model that contribute less to the system's dynamic evolution through meticulous sensitivity analysis. Meanwhile, for high-dimensional parameter matrices in the model, we use techniques such as low-rank decomposition for approximation, significantly reducing storage and computational requirements without significantly compromising key information.
[0050] Through these series of operations, we create a proxy model that can complete a simulation in milliseconds. It becomes the main "training ground" for reinforcement learning agents to conduct high-frequency and large-scale exploration and learning.
[0051] Pillar Two: Sliding Window Iterative Refinement Unit, used to maintain spatiotemporal continuity and logical self-consistency in long-term simulation periods, suppressing error accumulation and divergence. The sliding window iterative refinement unit ensures the spatiotemporal consistency of the simulation process through the following mechanisms: 1. Spatiotemporal context awareness: When predicting the state of the national space at a future time (e.g., T+1), the model not only inputs the current state (T), but also inputs the historical state (e.g., T-1). This forms a "sliding window" with temporal context.
[0052] 2. Iterative correction: More importantly, this process is iterative. The model performs multiple internal iterations within a sliding window. It first generates a preliminary prediction of T+1, then uses this preliminary prediction as future information to correct and refine the state representation within the current window, and then makes a new prediction based on the refined state. This process is like an internal "self-negotiation" mechanism, ensuring that the state transition from T to T+1 is smooth, gradual, and logical, rather than a one-time, isolated "jump".
[0053] This mechanism ensures that even in a simulation period of several decades, the national space evolution sequence generated by the proxy model can maintain high spatiotemporal continuity and logical self-consistency, effectively suppressing error accumulation and divergence, making the results of long-term strategy simulation realistic and reliable.
[0054] Pillar Three: Self-supervised learning unit based on scientific first principles, used to ensure fast calculation and stable simulation of the proxy model, while maintaining scientific robustness. The self-supervised learning unit mainly includes the following two learning paradigms: 1. Formalize scientific principles into computable constraints: In the land space system, there are a large number of widely recognized rules based on physical, ecological and economic first principles. For example: the conservation of mass constraint: the total water quantity (precipitation, runoff, evaporation, infiltration) must be balanced within a closed watershed. Energy balance constraint: the total energy consumption of a region cannot exceed the total of its energy production and input without reason. Economic input-output rules: the input-output relationship between different industrial sectors can be described and constrained by the classic Leontief input-output matrix.
[0055] 2. Dual loss function driven training: When training the agent model, we designed a dual loss function: Mimicking Loss: the difference between the output of the agent model and the output of the teacher model, which ensures the basic simulation accuracy. Principle Violation Loss: the extent to which the output of the agent model violates the pre-set scientific principle constraints.
[0056] By minimizing the weighted sum of the two loss functions, we force the agent model to learn to mimic the teacher model while its behavior must strictly follow the basic scientific laws. This self-supervised mechanism brings two great benefits: First, it provides a strong "scientific guardrail" for the model's learning process, preventing it from learning illogical pseudo-relationships and greatly improving the model's generalization ability and reliability in unknown situations. Second, it reduces the dependence on massive "teacher model" simulation data to some extent, as a large amount of training signals can be directly obtained from these scientific principle constraints.
[0057] Pillar four: Principle-guided-knowledge-distillation review training unit, used to train proxy dynamic models with high speed and high accuracy, to ensure their accuracy, scientificity and long-term stability.
[0058] Agent model total loss = λ_mimicking L_mimicking + λ_principle L_principle + λ_timing L_timing Where: 1. L_mimicking = E_scene [|| teacher_model (scene) - agent_model (scene) || 2 ] This is the knowledge distillation loss. It requires the output of the agent model (student) to approach the output of the high-precision but slow teacher model as much as possible under various "input scenarios". This ensures the basic simulation accuracy of the agent model.
[0059] 2. L_principle = E_scene [Σ_all principles | scientific_principle_constraint_function_i (agent_model (scene)) |] This is the principle violation loss. The scientific principle constraint function _i is a function that formalizes the first principles of physics, ecology, or economics (such as the conservation of matter and energy balance), and its output should ideally be zero. This loss term penalizes any behavior of the proxy model that violates these basic scientific laws, constituting principle-guided self-supervision and ensuring the scientific validity and generalization ability of the model.
[0060] 3. L_time series = E_{scenario,t}[||surrogate model(state_t)-iterative refinement(surrogate model(state_{t-1}),surrogate model(state_t))|| 2 ] This is the time-consistency loss. Iterative refinement(...) is an internal iterative function that uses historical and current predictions to smooth out... Slippage and correction of state transitions. This loss term requires that the result of a single-step prediction must maintain consistency with the result refined through multiple iterations. This ensures the spatiotemporal continuity and stability of long-cycle simulations, avoiding error accumulation.
[0061] 4. λ_imitation, λ_principle, and λ_timing are hyperparameters used to balance the importance of the three loss terms.
[0062] Explanation: This composite training objective is a highly innovative fusion. It simultaneously shapes the surrogate model from three dimensions: mimicking the teacher to ensure its "formal resemblance," adhering to principles to ensure its "essence," and temporal consistency to ensure its "stable and sustainable development." This multi-task, multi-constraint learning framework enables us to construct efficient and computable dynamic models that are fast, accurate, and robust. This formula is the theoretical cornerstone for building the entire dynamic model system; it precisely defines how to train that fast, accurate, and robust "surrogate model" (i.e., the student model).
[0063] L_imitation (knowledge distillation loss) corresponds to the process by which the agent model learns the core dynamic features of the system by imitating the "teacher model".
[0064] The L-principle (principle violation loss) corresponds to the "self-supervised learning paradigm based on the first principles of science" that we introduced, ensuring that the surrogate model conforms to basic scientific laws.
[0065] L_time (time consistency loss) corresponds to the "sliding window iterative refinement mechanism" designed to ensure long-term simulation stability. This composite loss function perfectly unifies multiple core technical requirements in the surrogate model training process under a single optimization objective.
[0066] Core Subsystem 2: Strategy Emergence Engine 20, used to generate, evaluate and optimize territorial spatial planning strategies, and to enable the autonomous "emergence" of excellent strategies.
[0067] Referring to Figure 3 , Figure 3 is a structural schematic diagram of an embodiment of the strategy emergence engine provided in the present application.
[0068] Among them, the strategy emergence engine 20 includes a structured strategy generation module 21 for simulating the thinking process of human experts when planning; a synergistic multi-objective learning module 22 for learning and generating a strategy set covering the entire Pareto optimal frontier, and introducing advanced semantic understanding ability to process complex objectives that are difficult to quantify.
[0069] Regarding the structured strategy generation module 21, it mainly includes three mechanisms to support the operation of the module: Mechanism 1: Decoupled-autoregressive strategy generation pipeline; Inspired by the fact that humans usually follow the process of “first composition, then sketch, and finally color” when creating complex works (such as painting and design), we decouple the generation process of land space strategy into an autoregressive pipeline consisting of three stages. Here, “autoregressive” means that the decision of the next stage is conditioned on the output of the previous stage, thereby ensuring the internal logical coherence of the entire strategy. Mechanism 1 includes the following execution stages: Stage 1: Strategic intent generation; Objective: First, the agent needs to determine the core strategic direction of this planning.
[0070] Implementation: In this stage, the task of the agent is not to perform specific spatial operations, but to generate one or a set of high-level semantic strategic intentions from a predefined strategy set. These intentions are similar to natural language instructions, such as “prioritize the development of metropolitan areas centered on the digital economy”, “strengthen the ecological function of the upper reaches of the xx river basin”, and “promote urban-rural integration, focusing on improving rural infrastructure levels”. This step establishes the “soul” and top-level goal of the entire planning scheme.
[0071] Stage 2: Spatial layout autoregressive planning; Objective: Based on the determined strategic intent, the agent needs to “draw” the macro layout of the functional zoning on the digital land space.
[0072] Implementation: This is the most creative part of the entire framework. We borrow the idea of three-dimensional model part-by-part autoregressive generation and transform the spatial planning process into a sequential, region-by-region generation task. The specific process is as follows: 1. Sowing initial core: The agent first “places” an initial core functional area in the land space, for example, according to the strategic intent of “developing the digital economy”, it determines an initial core of a “digital industrial park” in a certain location.
[0073] 2. Component-wise conditioning: Then, like building with blocks, the agent starts to generate recursively. In the next step, it generates a next related functional zone, such as a “high-end talent community”, based on the existing “digital industrial park”. In the next step, it may generate an “ecological green corridor” between the community and the park.
[0074] 3. Contextual awareness: In each step of generating decisions, the agent will fully consider all existing spatial layouts. For example, when deciding the location and type of the next functional zone, it will not only see the “digital industrial park” and “high-end talent community”, but also perceive their adjacency, distance, and the overall spatial form that has been formed.
[0075] 4. Termination signal: This recursive process will continue until the agent generates a special “termination” signal, indicating that the macro spatial layout planning is complete.
[0076] Through this component-wise, recursive way, the spatial layout generated by the agent is no longer a chaotic patchwork of color blocks, but an organically grown, structurally reasonable spatial blueprint. It can automatically learn planning common sense such as “industrial areas and residential areas should be kept at an appropriate distance and separated by green spaces”, because this layout pattern can achieve higher long-term returns in the training data.
[0077] Phase Three: Specific Policy Recursive Configuration; Objective: After the spatial layout is determined, configure specific, quantitative policy tools for each planned spatial unit.
[0078] Implementation: This stage also adopts a recursive mode. The agent will “traverse” each spatial unit generated in the second stage and configure a series of policy parameters for it. For example, for the “digital industrial park”, it may generate “maximum land use ratio: 3.5”, “R&D investment subsidy ratio: 15%”, and “high-tech enterprise tax reduction policy: Class A”. For the “ecological green corridor”, it may generate “development intensity limit: 0.1” and “species diversity protection target: Class B”. This process is also context-aware, meaning that the policy configuration for an area will take into account the existing policies of its surrounding areas to ensure policy coordination.
[0079] Through this “strategic intent → spatial layout → policy configuration” decoupling-recursive pipeline, we break down an extremely complex, high-dimensional strategy generation problem into three relatively simple but logically closely connected sequential decision-making problems, greatly reducing the learning difficulty and structurally ensuring the completeness and reasonableness of the generated strategy.
[0080] Mechanism Two: Hierarchical Strategy Learning Architecture across Scales Territorial spatial planning has a natural cross-scale hierarchical feature from the national to the regional and then to the urban level. Macro-level decisions, such as the delineation of strategic regions at the national level, directly affect and constrain micro-level planning. To simulate this multi-scale governance structure in our agent framework, a hierarchical reinforcement learning architecture is introduced in some embodiments.
[0081] 1. Multi-level Agent Structure: We construct a multi-level agent system, such as a one-level structure: High-level Agent: Responsible for decision-making at the macro scale. Its action space may be limited to deciding the overall development orientation of major regions (such as the Yangtze River Delta, Pearl River Delta, and Beijing-Tianjin-Hebei), the cross-regional layout of major infrastructure, and the macro allocation of key resources (such as annual construction land indicators and fiscal transfer payment amounts) among regions.
[0082] Low-level Agent: Operates at the regional scale. It receives decisions from the high-level agent as its goals and constraints. For example, after receiving the instruction "this region obtains 1000 square kilometers of construction land indicators and develops as a national advanced manufacturing base", the low-level agent will execute the "decoupling-autoregressive strategy generation pipeline" within its jurisdiction to conduct specific spatial layout planning and policy configuration.
[0083] 2. Cross-level goal transmission and reward allocation: The high-level agent has a longer decision-making cycle, and it receives rewards and updates its strategy based on long-term, macro-system evolution results (such as future 20-year national GDP growth, total carbon emissions, etc.).
[0084] The low-level agent receives rewards based on its performance in achieving high-level goals within its region in the short and medium term.
[0085] This design mimics the real administrative management system, breaking down a large, difficult-to-solve national optimization problem into "one macro resource allocation problem" and "multiple regional planning problems under constraints".
[0086] Through this hierarchical architecture, we greatly simplify the complexity of the learning task. The high-level agent only needs to focus on strategic "big games", without getting bogged down in micro details; the low-level agent can efficiently learn specific planning implementation schemes under clear goals and constraints. This not only improves learning efficiency, but also ensures that the final generated strategy set maintains high consistency and coordination at both macro and micro levels.
[0087] Therefore, the "structured strategy generation framework" decouples the autoregressive strategy generation pipeline, decomposing the complex planning scheme generation process into an ordered, logically coherent sequence of decision-making tasks; at the same time, through the cross-scale hierarchical strategy learning architecture, the grand national planning problem is decomposed into multi-level, manageable, and collaborative sub-problems. The organic combination of these two mechanisms enables our agent to break free from the shackles of traditional action spaces and learn to generate and create truly structured, logically consistent, and multi-level coordinated land space sustainable development strategies.
[0088] Mechanism three: "hierarchical-autoregressive" generation probability model of structured strategy; This formula describes how the agent generates a structured, logically consistent land space planning strategy.
[0089] Probability (complete strategy | initial state) = P (strategic intent | initial state) Π_{k=1 to K} P (space layout unit_k | strategic intent, layout unit_{<k}) Π_{j=1 to J} P (policy configuration_j | complete space layout, policy configuration_{<j}) Where: 1. P(...) represents conditional probability.
[0090] 2. Π represents the product operation.
[0091] 3. The first term: P (strategic intent | initial state) The agent first generates a macro "strategic intent" based on the "initial state" of the land. This is the starting point of the entire strategy generation.
[0092] 4. The second term: Π_{k=1 to K} P (space layout unit_k | strategic intent, layout unit_{<k}) This is an autoregressive process. When generating the kth "space layout unit" (such as an industrial park), the agent's decision is conditioned on the "strategic intent" and all k-1 layout units (layout unit_{<k}) that have been generated before it. This ensures the organic growth and internal coordination of the space layout.
[0093] 5. The third term: Π_{j=1 to J} P (policy configuration_j | complete space layout, policy configuration_{<j}) After the "complete space layout" is generated, the agent again configures specific "policy parameters" for each space unit in an autoregressive manner. The configuration of the jth policy considers all j-1 policies that have been configured.
[0094] Explanation: This formula ingeniously decomposes an extremely complex, parallel strategy design problem into an ordered, sequential conditional probability generation process. It strictly follows the logical chain of "strategy first, then space, and finally policy", and through a self-recurrent mechanism, fully utilizes the information generated in history at each decision-making step. This not only greatly reduces the learning difficulty, but also mathematically guarantees the completeness, hierarchy, and internal logical coherence of the final generated strategy.
[0095] This formula is a precise mathematical description of the "decoupling-self-recurrent strategy generation pipeline". It decomposes the task of generating a complete strategy, which is complex and unstructured, into three ordered, sequential generation steps through the chain rule of conditional probability: 1. Generate strategic intent.
[0096] 2. Recursively generate spatial layout units under the condition of strategic intent and generated historical layout.
[0097] 3. Recursively generate policy configuration under the condition of complete spatial layout and configured historical policy.
[0098] This formula is the core of the agent's action space design, which defines the structured process of how the agent "thinks" and "acts".
[0099] Regarding the collaborative multi-objective learning module 22: The core idea is to abandon the pursuit of a single optimal solution and instead learn and generate a strategy set covering the entire Pareto optimal frontier, and introduce advanced semantic understanding capabilities to handle complex objectives that are difficult to quantify. In some embodiments, the collaborative multi-objective learning module 22 includes the following three mechanism units: Mechanism one: Pareto frontier exploration unit based on unified supervised contrast learning, used to convert discrete, multi-dimensional target evaluation into a unified, internally collaborative learning framework, thereby driving the agent to efficiently explore and approximate the entire Pareto optimal solution; The Pareto frontier exploration unit based on unified supervised contrast learning mainly explores the optimal solution set from three aspects: 1. From scalar reward to embedding space contrast paradigm shift, we no longer calculate a single reward score for each strategy generated by the agent, but "embed" it into a high-dimensional target space. In this space, each dimension corresponds to a planning objective (such as GDP growth rate, carbon sink volume, etc.). The merits of a strategy are no longer determined by its absolute score, but by its relative position relationship with other "benchmark strategies" in this high-dimensional space.
[0100] 2. Workflow of unified supervised contrast reward: A. Constructing a dynamic benchmark set: In each round of learning, the system maintains a dynamic set of "benchmark strategies". This set contains: Positive samples: historical planning cases that have been verified as successful, and other excellent strategies that have been found to be on the Pareto frontier.
[0101] Negative samples: historical planning cases that have failed, or inferior strategies that are "dominated" by other strategies on all objectives (i.e. all metrics are worse).
[0102] B. Calculating the contrast reward signal: For a new strategy generated by the agent, the system calculates its distance in the objective embedding space: From all positive samples; From all negative samples.
[0103] C. Construction of the reward function: The reward signal is constructed as a contrast loss function. The agent's goal is to minimize the distance between the strategy it generates and the positive samples, while maximizing the distance from the negative samples. This process is similar to an "pull good, push bad" optimization process.
[0104] 3. Technical advantages and internal synergy, including: A. Avoiding gradient conflicts: This contrast learning framework naturally solves the problem of gradient conflicts between multiple objectives. Because it no longer requires the model to optimize in multiple directions at the same time (such as maximizing GDP and maximizing carbon sinks), it integrates all objectives into a unified "distance" metric. The model's optimization direction is only one: moving to the area where "good" strategies are located in the embedding space.
[0105] B. Driving exploration of the Pareto frontier: By taking multiple different strategy points on the Pareto frontier as positive samples, this mechanism encourages the agent to generate diverse strategies. For example, the agent will find that a strategy that focuses too much on economic development and a strategy that focuses too much on ecological protection can both be "good" (because they can both be on the Pareto frontier), and learn to generate various strategies with different trade-off preferences between the two.
[0106] C. Unified loss function design: The design of this reward mechanism is essentially inspired by the idea of unified loss function in supervised contrast learning. It elegantly combines supervised signals (which strategy is better) and self-supervised signals (similarity between strategies) under the same mathematical framework, making the multi-objective learning process stable, efficient and goal synergistic.
[0107] Mechanism Two: Complex Goal Quantification Unit Based on "Agent as Judge", which utilizes the powerful semantic understanding and logical reasoning capabilities of Large Language Models (LLMs) to transform complex, qualitative goals into computable reward signals. One: Workflow 1. Generating Comprehensive Evaluation Report: For each strategy generated by the agent, our decision engine drives a "highly efficient and computable national system dynamics model" to conduct a complete simulation. After the simulation, the system automatically aggregates multi-source, multi-modal simulation results (such as economic growth data tables, environmental monitoring reports, population migration spatial distribution maps, and simulation-generated social media public opinion texts) and generates a structured comprehensive evaluation report.
[0108] 2. Instruction-based Qualitative Evaluation: This comprehensive report, along with pre-set high-level evaluation criteria (such as "regional coordinated development strategy points" defined in natural language), is input into a large language model that has been specifically instructed (Instruction-Tuning).
[0109] 3. Output Quantitative Reward Score: This large language model as "judge" is not designed for numerical calculation, but for advanced semantic and logical judgment. It will "read and understand" the comprehensive report according to the evaluation criteria and finally output one or more qualitative evaluation scores (for example, 85 points in the "social fairness" dimension and 92 points in the "regional coordination" dimension).
[0110] Two: Technical Advantages and Capability Expansion 1. Bridging the Quantitative Gap: This mechanism successfully translates complex, value-laden goals in human society that are difficult to formalize into quantifiable reward signals that machines can understand and optimize through the "translation" of large language models, greatly expanding the range of goals that the decision engine can handle.
[0111] 2. Dynamic and Flexibility: Decision-makers can adjust the top-level goals of the plan at any time by modifying or adding evaluation criteria in natural language, without the need to rewrite complex reward function code. For example, "promoting innovation" or "ensuring people's livelihoods" can be added as new evaluation criteria in different development stages to guide the agent to generate strategies that meet the new strategic direction.
[0112] 3. Enhancing Decision Interpretability: Since the evaluation process of the "judge" is based on language and logic, it can output the reasons and explanations for its scores. This provides clear and readable explanations for why the AI generates a particular strategy and how well it performs in various complex goals, enhancing the transparency of the entire decision-making process.
[0113] Therefore, the synergistic multi-objective learning module 22 transforms the multi-objective optimization problem from solving a point to exploring a "face" (Pareto frontier) through unified supervised contrast learning, fundamentally solving the conflict between objectives and achieving internal synergy. At the same time, through the "agent as evaluator", the semantic understanding ability of large language models is used to seamlessly connect human complex and qualitative strategic objectives into the learning loop of machines. The synergistic effect of these two major technical pillars enables our decision engine not only to achieve synergistic optimization in quantifiable hard indicators, but also to achieve intelligent alignment in unquantifiable soft targets, thereby generating truly sustainable development-oriented, comprehensive and balanced planning strategies.
[0114] Mechanism three: Multi-objective synergistic "contrast-evaluation" reward function, used for synergistic optimization of multiple conflicting objectives and incorporating evaluation of complex qualitative objectives; the function formula is: Synergistic reward (generated strategy) = log[(Σ_{positive samples}exp(-D_ embedding / τ)) / (Σ_{positive samples}exp(-D_ embedding / τ)+Σ_{negative samples}exp(-D_ embedding / τ))]; log[...] and exp(...) are logarithmic and exponential functions, respectively; Positive samples and negative samples represent "good" and "bad" benchmark strategy sets, respectively; τ is a temperature hyperparameter used to adjust the sharpness of contrast; D_ embedding is the weighted distance between two strategies in the target embedding space.
[0115] Specifically, D_ embedding(strategy 1, strategy 2) = J_ big model(report 1, report 2) ||E_ target(strategy 1)-E_ target(strategy 2)| E_ target(...) is an encoder that maps the performance of a strategy on all quantifiable hard indicators (GDP, carbon sink, etc.) into a high-dimensional vector.
[0116] ||...|| calculates the embedding vector distance between two strategies in hard indicators.
[0117] J_ big model(...) is the "agent as evaluator" function. It receives the "comprehensive evaluation report" generated after the deduction of two strategies, and outputs a scalar between 0 and 1 representing the similarity or superiority of the two strategies in all complex and unquantifiable soft targets (such as social fairness, regional coordination).
[0118] Explanation of the formula: It converts the multi-objective optimization problem into a problem of "choosing sides" in the embedding space: the agent's goal is to make the strategy it generates, after synthesizing the hard indicators (through E_target) and soft targets (through J_large_model), as close as possible to the "good" strategy group, while being as far away as possible from the "bad" strategy group. The evaluation results of the large language model, as a kind of dynamic, semantic-aware weight, adjust the importance of the hard indicator distance, achieving seamless collaboration between quantitative and qualitative.
[0119] This formula is the core calculation method of the reward signal (RewardSignal) when the reinforcement learning agent performs multi-objective optimization.
[0120] The overall contrast loss structure of the formula corresponds to the "Pareto frontier exploration based on unified supervised contrast learning" mechanism, which drives the agent to learn to generate diverse strategies on the Pareto frontier.
[0121] The J_large_model term in the formula is a mathematical formalization of the "complex target quantification based on 'agent as evaluator'" mechanism. It converts the semantic judgment of large language models on complex, qualitative targets into a calculable weight that can adjust the hard indicator distance, achieving seamless collaboration between qualitative and quantitative targets. This reward function is the "value compass" of the entire strategy emergence engine, guiding the agent to evolve towards the optimal direction of multi-objective collaboration.
[0122] Core subsystem three: human-machine collaborative cockpit 30, for human-computer interaction and realizing efficient and intuitive and unbiased human-computer collaborative decision-making. Refer to Figure 4 , Figure 4 is a structural schematic diagram of an embodiment of the human-machine collaborative cockpit provided by the present application.
[0123] Among them, the human-machine collaborative cockpit 30 includes: an adaptive computing and instructive exploration module 31 for intelligently managing and allocating computing resources, so that the system can respond to users' macro exploration requests in near real time, and can also perform high-fidelity fine simulation on specific and in-depth problems of user interest, thereby achieving the best balance between "interaction efficiency" and "analysis depth"; an unbiased objective decision support module 32 for building "eliminating bias" and "ensuring objectivity" as the first principle of system design, and building them into the core architecture.
[0124] Among them, the adaptive computing and instructive exploration module 31 mainly includes the following three main support units: Pillar one: cascaded adaptive simulation unit based on "instructive time sequence and spatial positioning", for intelligent on-demand allocation of computing resources, greatly improving the efficiency and feasibility of deep interactive analysis; Among them, the adaptive simulation unit mainly carries out rigorous scenario deduction through the following three aspects of technology: 1. Hierarchical calculation model. In some embodiments, the system background simultaneously deploys two types of dynamic models: High-speed agent model: that is, the aforementioned "student model", which is deeply optimized and compressed, and has extremely fast calculation speed, and can complete a long-period macro scenario deduction in seconds or even milliseconds.
[0125] High-precision complete model: that is, the "teacher model", which contains complete scientific mechanism and fine parameters, and has high calculation cost, but can provide the highest fidelity simulation results.
[0126] 2. Default macro exploration and cascading triggering. In the default human-computer interaction state, when the decision maker conducts macro-level "What-if" analysis (for example, "If the average urbanization rate in the country is increased by 2% in the next ten years, what impact will it have on the total arable land area?"), the system only calls the high-speed agent model for deduction. This ensures that the decision maker can smoothly and without delay explore various macro policy options and quickly obtain trend and directional judgments about the future direction of the system. When the decision maker's exploration behavior focuses on a more specific and detailed issue, the system will automatically trigger cascading calls. This triggering process is based on instruction-based timing and spatial positioning technology.
[0127] 3. Instruction-based timing and spatial positioning, mainly including: A. Instruction analysis: the system can understand the exploration intention expressed by the decision maker through natural language, chart interaction (such as framing an area on a map), or parameter adjustment in real time. For example, when the decision maker proposes "Please simulate in detail the possible ecological impact of this water conservancy project on the rare fish protection zone 30 kilometers downstream in the next 50 years", the system will analyze this instruction through its built-in natural language understanding module into a set of precise spatiotemporal coordinates and objects of interest: Spatial positioning: the scope of the water conservancy project and its downstream 30-kilometer protection zone.
[0128] Time positioning: a time span of 50 years in the future.
[0129] Objects of interest / processes: hydrological changes, water quality evolution, fish population dynamics, etc.
[0130] B. Precise calling of high-precision model: once the instruction is analyzed, the system no longer simulates the entire national space, but only in the specific spatiotemporal region and related physical processes "positioned" by the instruction, automatically calls the high-precision complete model. For areas outside this range, the macro background field data provided by the high-speed agent model is still used as boundary conditions.
[0131] This mechanism is like providing a "computational magnifying glass" for the decision maker. At the macro scale, the system provides a "satellite map" (generated by the surrogate model) of fast response; when the decision maker needs to carefully observe a certain building on the map, the system can immediately seamlessly switch to a high-precision "microscope" mode (generated by the full model) without generating micro-level images for the entire city. This achieves intelligent on-demand allocation of computing resources, greatly improving the efficiency and feasibility of in-depth interactive analysis.
[0132] Pillar II: Agent-initiated verification unit based on uncertainty perception, used to give the system the ability of self-reflection and information completion in the self-learning process through agent-initiated verification; Among them, the agent-initiated verification unit mainly performs adaptive computing through the following three aspects of technology: 1. Uncertainty assessment in decision-making process. When the agent performs strategy deduction on the high-speed surrogate model, it not only predicts a single future result, but also assesses the confidence or uncertainty of this prediction simultaneously. This uncertainty may come from multiple aspects: Intrinsic uncertainty of the model: As an approximation of the full model, the surrogate model itself has inherent uncertainty in its prediction ability under certain complex scenarios.
[0133] Scenario sensitivity: At certain "critical points" or "bifurcation points", a small change in input can lead the system to a completely different future. In this high-sensitivity area, the prediction results of the surrogate model often show large variance.
[0134] 2. Active request for high-precision verification. The agent sets an uncertainty threshold internally. During its strategy exploration, if it finds that the uncertainty of a key decision exceeds this threshold, it will judge that "the information based on the surrogate model is not sufficient to make a reliable decision". At this time, the agent will actively trigger a call request for high-precision full model, like asking the system: "I need a more accurate verification for the specific decision consequence." After receiving this request, the system will run a high-precision simulation at this specific decision point as if responding to a user instruction, and feed back the results to the agent.
[0135] 3. Dynamic trade-off and learning efficiency improvement. After receiving the high-precision verification result, the agent updates its internal beliefs and strategy network with this "real" data. This enables it to make more accurate judgments when encountering similar scenarios in the future. More importantly, this mechanism enables the agent to balance between "extensive and rapid cheap exploration" (on the surrogate model) and "deep and accurate verification" (on the full model), so that it can learn more efficiently and effectively. between the "cheap and accurate proof" (call the complete model) to achieve dynamic, intelligent trade-offs. It only requests expensive computing resources when needed, thereby greatly improving the overall learning efficiency while ensuring the final policy quality.
[0136] Therefore, the "adaptive computing and instructive exploration" mechanism gives the system the ability to respond to human decision-makers "point to shoot" through cascading adaptive simulation; at the same time, through active proof of the agent, the system is given the ability of self-reflection and information completion in the process of self-learning. These two technical pillars work together to turn the original static and rigid computing process into a dynamic, intelligent, and human-machine collaborative exploration process, making it possible to conduct in-depth simulation analysis and intelligent decision optimization of the complex and huge system of territorial space in practice.
[0137] Pillar three: human-computer interaction "adaptive computing" decision rule function, the rule function is: Call model (current state, user instruction) = {complete model, if (focus area (user instruction) ≠ ∅) or (U_ uncertainty (current state) > δ_ threshold); agent model, otherwise} The complete model and the agent model represent high-precision and high-speed dynamic models, respectively; Focus area (user instruction) is an analytical function; U_ uncertainty (current state) is an evaluation function used to calculate the uncertainty of the agent in the "current state" using the agent model for prediction; δ_ threshold is a preset uncertainty threshold.
[0138] Explanation of the function formula: This is a simple and powerful gating rule. It clearly defines the two conditions that trigger expensive computing (call the complete model): either "external instructions" from human users, requiring in-depth drilling analysis, or "internal requests" from the agent itself, because it finds that its current cognition is in a highly uncertain state, and needs more accurate information to assist decision-making (active proof). In all other cases, the system defaults to using the efficient agent model. This rule realizes the extreme on-demand allocation of computing resources and the deep integration of human-machine intelligence.
[0139] This formula is a "gating switch" that realizes intelligent scheduling of computing resources. It precisely defines when the system should use the efficient "agent model" and when it must call the high-cost "complete model".
[0140] The condition that the attention region (user instruction) ≠ ∅ corresponds to the "instructional exploration" or "cascading adaptive simulation" in response to the deep analysis needs of human decision makers.
[0141] The condition that U_ uncertainty (current state) > δ_ threshold corresponds to the "agent's active proof based on uncertainty perception" in the self-learning process due to insufficient cognition.
[0142] This rule is the key to efficient human-computer interaction and improving the learning efficiency of the agent, which ensures that valuable computing resources are used "on the cutting edge".
[0143] Among them, the unbiased objective decision support module 32 mainly includes the following two main pillar units: Pillar one: reference system bias elimination unit based on permutation equivariance, which is used to ensure that the system does not have any reference system bias from the architecture level by constructing a completely permutation equivariant model; Traditional spatial analysis and reconstruction methods, whether classic "Structure-from-Motion" or modern deep learning models, usually need to first specify a "reference system" or "anchor point". For example, when conducting multi-city collaborative development planning, the model may implicitly or explicitly take a certain city (such as the provincial capital) as the center of analysis and the origin of the coordinate system, and the indicators and spatial relationships of other cities are measured relative to this center. The serious defect of this design is that the choice of the reference system is arbitrary, but it will deeply affect the final result. If a non-traditional city is chosen as the reference system, the internal calculation process of the entire model will change, and ultimately may lead to fundamental changes in the evaluation of the development potential of each city, resource allocation recommendations, and even the pros and cons of the planning scheme. This is like observing the same scenery from different mountain tops, and the conclusions drawn will be completely different. This inconsistency in conclusions due to the artificial setting of the analysis perspective is a major obstacle to objective decision-making.
[0144] Therefore, to completely solve this problem, we introduce "permutation equivariance" as the fundamental design principle of the core neural network architecture of our engine.
[0145] Core definition: "permutation equivariance" is a mathematical property of a lattice, which requires that for a sequence of input elements (for example For example, a list containing statistical data of all cities in the country), no matter how the elements in the sequence are arranged, the output sequence of the model will maintain the same arrangement as the input sequence. In other words, the arrangement of the input is equal to the arrangement of the output.
[0146] Implementation: In network architecture design, we abandoned all components that depended on the input order. For example, we did not use traditional recurrent neural networks (RNNs) or Transformer modules that relied on fixed-position encoding. Instead, we employed an attention mechanism that alternated between different perspectives and the global context for information processing. In this architecture, each input unit (e.g., data from each city) was treated equally, and its relationship with all other units was dynamically calculated through its intrinsic properties, regardless of its "position" in the input list.
[0147] Effects and Significance: By constructing a fully permuted and equivariant model, we ensured at the architectural level that the system is free from any reference system bias. Regardless of which city's data decision-makers use as the first element of the input list, and regardless of how they shuffle the data presentation order, the system's assessment of each city's development potential, judgment of inter-city synergies, and final planning strategy recommendations will be completely consistent and stable. This ensures the objectivity of decision support, making the conclusions rely solely on the inherent laws of the data itself, and completely eliminating human interference from the analytical perspective.
[0148] Pillar Two: A knowledge source bias identification and mitigation unit based on benchmark reliability review, used to proactively identify and mitigate objective biases caused by imbalances in data collection, processing, and representation, ensuring the reliability and comprehensiveness of the model's decision-making basis.
[0149] In the actual generation and collection of land spatial data and knowledge, various biases are inevitably present. For example: Data collection bias: Data from economically developed regions is usually more comprehensive, more detailed, and updated more frequently, while data from underdeveloped regions may be missing, coarse, or outdated.
[0150] Semantic bias: Policy documents often describe certain "star" areas in a positive and detailed manner, while other areas may be mentioned only briefly or in a negative way.
[0151] Sample representativeness bias: The cases used to train the model may be overly concentrated on certain common planning scenarios (such as new town development), while insufficiently covering other less common but equally important scenarios (such as the transformation of old industrial areas and the restoration of ecologically fragile areas).
[0152] To this end, in some embodiments, this application establishes a baseline reliability review process: A. Automated bias identification: The system will automatically and periodically scan the entire knowledge graph and training dataset, drawing on the idea of systematic problem analysis of benchmark datasets. It will actively seek and quantify various biases present in the data, such as data redundancy, imbalance in concept distribution (e.g., "head" and "tail" concepts as mentioned earlier), and sentiment bias and ambiguity in text descriptions.
[0153] B. Explicit bias: The results of the review will not be hidden, but will be explicitly labeled as metadata on the relevant knowledge units. For example, data from a certain region may be labeled as "data sparse, low confidence," and a policy text may be labeled as "positive sentiment bias."
[0154] C. Triggering of bias mitigation mechanisms: These identified and quantified biases will directly trigger various bias mitigation mechanisms within the system. For example, identified data imbalance will trigger the aforementioned "online concept balancing" mechanism, which dynamically adjusts weights during model training; identified data gaps may trigger the system to actively suggest data collection tasks.
[0155] Through this mechanism, our system has made the transition from "passively accepting data" to "actively reviewing and managing knowledge bias." It not only tells decision-makers "what is it," but also reveals "what we know 'what is it' may have biases," thereby establishing a more transparent and reliable knowledge base for decision-making.
[0156] The technical effects formed are as follows: Through the above-mentioned systematic technical means, the decision engine will form a logically rigorous closed loop and produce a revolutionary technical effect that surpasses traditional decision support systems: 1. From "strategy optimization" to "path emergence" paradigm revolution: The biggest non-obviousness of this solution is that it does not search for the optimal solution in a fixed decision space, but creates an evolutionary environment that allows excellent development paths to emerge autonomously. High-fidelity digital twins are the soil of evolution, structured strategy generation provides rich variation, and a unified multi-objective reward function plays the role of "natural selection." This "simulation-learning-generation" closed loop enables the engine to discover new patterns of land space organization and collaborative development paths that human experts based on intuition and experience may never have imagined, realizing a fundamental leap in decision support from "solving problems" to "discovering the future."
[0157] 2. From “data black box” to “computational causality”: Traditional predictive models are often “black boxes”. But after the training of our engine, the internal agent model and policy network essentially build a computable causal model of how the “geo-ecological-social” complex system interacts. We can not only use it to generate decisions, but also use it as a “computational experiment platform” to “dissect” the decision-making logic of this intelligent brain, to uncover hidden and nonlinear causal relationships between different subsystems, and to provide unprecedented tools for theoretical innovation in land space science.
[0158] 3. From “static planning” to “dynamic adaptive governance”: The traditional idea that “planning is a blueprint” has been difficult to cope with the high uncertainty of the future. The “continuous learning-dynamic deduction-human-machine collaboration” closed loop of this program endogenously supports adaptive governance. Decision makers can inject various future “black swan” events (such as extreme climate, technological mutation) into the engine, which can not only assess the vulnerability of the existing pattern, but also generate a set of dynamic and phased response strategies in real time. This makes land space planning no longer a one-time blueprint, but a “life process” of continuous evolution and dynamic adjustment, greatly improving the forward-looking and resilience of national governance.
[0159] The above is only an embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation using the content of the present application specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. A land space intelligent collaborative management system, characterized in that, include: Digital twin environments are used to provide a simulated environment for decision-making agents to learn and infer, and to build a highly realistic and computationally efficient "mirror" of the land system; The strategy emergence engine is used to generate, evaluate, and optimize territorial spatial planning strategies, and to enable the autonomous "emergence" of excellent strategies. The human-machine collaborative cockpit is used for human-machine interaction and to achieve efficient, intuitive and unbiased human-machine collaborative decision-making.
2. The intelligent collaborative management system for territorial space according to claim 1, characterized in that, The digital twin environment includes: A high-fidelity and dynamically evolving land knowledge system is used to provide upper-level intelligent agents with a simulation environment that is as close to the real world as possible and can reflect the inherent laws of the system. The highly efficient and computable land system dynamics module is used to improve the computational efficiency of simulation by several orders of magnitude while ensuring key dynamic characteristics and scientific mechanisms.
3. The intelligent collaborative management system for territorial space according to claim 2, characterized in that, The high-fidelity and dynamically evolving land knowledge system includes: The deep quality control unit is used to identify and quantify the three core problems in the land knowledge system: knowledge redundancy, conceptual imbalance, and semantic ambiguity. Knowledge units with feedback and self-learning capabilities are used to dynamically improve the self-learning capabilities of the aforementioned land knowledge system; The "online concept equilibrium" loss function is used to dynamically train and refine the model in the aforementioned land knowledge system to overcome the bias problem caused by the long-tail distribution of data. The formula for the function is as follows: Total knowledge system loss = E_{data ~ actual distribution}[concept imbalance response distance (predicted state, planned concept) distance (actual state, predicted state)].
4. The intelligent collaborative management system for territorial space according to claim 2, characterized in that, The efficient and computable land system dynamics module includes: The hierarchical proxy model architecture based on the "teacher-student" paradigm is used to transform the large, detailed, and computationally slow complete land system dynamics model, i.e. the "teacher model," into one or more lightweight and computationally fast proxy models, i.e. the "student models," through advanced model compression and knowledge transfer technologies. The sliding window iterative refining unit is used to maintain spatiotemporal continuity and logical consistency during long simulation cycles, and to suppress the accumulation and divergence of errors. A self-supervised learning unit based on the first principles of science is used to ensure the fast computation and stable simulation of the surrogate model, while also possessing scientific robustness. The "Principle-Guided - Knowledge Distillation" verification training unit is used to train the surrogate dynamics model of several speed blocks, ensuring its accuracy, scientific nature and long-term stability.
5. The intelligent collaborative management system for territorial space according to claim 1, characterized in that, The strategy emergence engine includes: A structured strategy generation module is used to mimic the thought process of human experts when making plans; The collaborative multi-objective learning module is used to learn and generate a set of policies covering the entire Pareto optimal frontier, and introduces advanced semantic understanding capabilities to handle complex objectives that are difficult to quantify.
6. The intelligent collaborative management system for territorial space according to claim 5, characterized in that, The structured strategy generation module includes: A decoupled-autoregressive strategy generation pipeline is used to decouple the generation process of territorial spatial strategies into a three-stage autoregressive pipeline. The autoregressive nature of the pipeline means that the decision in the next stage is conditional on the output of the previous stage, thus ensuring the internal logical coherence of the entire strategy. A hierarchical strategy learning framework across scales is used to simulate the multi-scale governance structure of territorial spatial planning. The "Hierarchical - Autoregressive" generative probability model of the structured strategy, and the calculation formula of this model is: Probability(complete strategy|initial state) = P(strategic intention|initial state) Π_{k = 1 to K} P(spatial layout unit_k|strategic intention, layout unit_{<k}) Π_{j = 1 to J} P(policy configuration_j|complete spatial layout, policy configuration_{<j}); P(...) represents the conditional probability, and Π represents the product operation; Π_{k = 1 to K} P(spatial layout unit_k|strategic intention, layout unit_{<k}) represents an autoregressive process; Π_{j = 1 to J} P(policy configuration_j|complete spatial layout, policy configuration_{<j}) represents a re - regression process.
7. The intelligent collaborative management system for territorial space according to claim 5, characterized in that, The collaborative multi - objective learning module described above includes: A Pareto - front exploration unit based on unified supervised contrast learning, which is used to transform the discrete and multi - dimensional objective evaluation into a unified and inherently collaborative learning framework, so as to drive the intelligent agent to efficiently explore and approximate the entire Pareto optimal solution; A complex objective quantification unit based on "agent as judge", which is used to utilize the powerful semantic understanding and logical reasoning ability of the large - language model (LLM) to transform complex and qualitative objectives into computable reward signals; A "contrast - judge" reward function for multi - objective collaboration, which is used to collaboratively optimize multiple conflicting objectives and incorporates the evaluation of complex qualitative objectives; the function is: Collaborative reward(generated strategy) = log[(Σ_{positive samples} exp(-D_embedding / τ)) / (Σ_{positive samples} exp(-D_embedding / τ) + Σ_{negative samples} exp(-D_embedding / τ))]; log[...] and exp(...) are the logarithmic and exponential functions respectively; Positive samples and negative samples respectively represent the sets of "good" and "bad" benchmark strategies; τ is a temperature hyperparameter used to adjust the sharpness of the contrast; D_embedding is the weighted distance between two strategies in the objective embedding space.
8. The intelligent collaborative management system for territorial space according to claim 1, characterized in that, The human - machine collaborative cockpit described above includes: An adaptive computing and imperative exploration module, which is used to intelligently manage and allocate computing resources, enabling the system to not only respond almost in real - time to the user's macroscopic exploration requests, but also perform high - fidelity refined simulations on specific and in - depth issues that the user is concerned about, so as to achieve the best balance between "interaction efficiency" and "analysis depth"; An unbiased objective decision - making support module, which takes "eliminating bias" and "ensuring objectivity" as the first principles of system design and builds them into the core architecture.
9. The intelligent collaborative management system for territorial space according to claim 8, characterized in that, The adaptive computing and imperative exploration module described above includes: A cascaded adaptive simulation unit based on "imperative timing and spatial positioning", which is used for intelligent on - demand allocation of computing resources, greatly improving the efficiency and feasibility of in - depth interactive analysis; An agent - active verification unit based on uncertainty perception, which is used to endow the system with the ability of self - reflection and information complementation in the self - learning process through agent - active verification; An "adaptive computing" decision - making rule function for human - machine interaction, and the rule function is: Calling model(current state, user instruction) = {complete model, if (region of interest(user instruction) ≠ ∅) or (U_uncertainty(current state) > δ_threshold); surrogate model, otherwise} The complete model and the surrogate model represent high-precision and high-speed dynamic models, respectively; The region of interest (user instruction) is a parsing function; U_uncertainty(current state) is an evaluation function used to calculate the uncertainty of the agent's prediction using the surrogate model in the "current state"; δ_threshold is a preset uncertainty threshold.
10. The intelligent collaborative management system for territorial space according to claim 8, characterized in that, The unbiased, objective decision support module includes: The reference frame bias elimination unit based on permutation equivariance is used to ensure, from the architectural level, that the system is free from any reference frame bias by constructing a fully permutation equivariant model. The knowledge source bias identification and mitigation unit based on benchmark reliability review is used to proactively identify and mitigate objective biases caused by imbalances in data collection, processing, and representation, ensuring the reliability and comprehensiveness of the model's decision-making basis.