A park operation decision analysis method and device
By using generative causal Bayesian networks and an uncertainty-guided active research loop, the decision-making challenges of counterfactual and hypothetical scenarios in park operations have been solved, enabling reliable prediction and scientific decision-making for unimplemented plans, and reducing the risks and costs of park operations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHERN ENG DESIGN & RES INST CO LTD
- Filing Date
- 2026-05-08
- Publication Date
- 2026-06-02
AI Technical Summary
Existing park operation decision analysis methods cannot make reliable predictions when faced with counterfactual or hypothetical scenarios, resulting in high investment costs and increased operational risks. In particular, in the absence of historical data, traditional methods are difficult to effectively assess the causal effects of unimplemented projects.
A candidate scenario library is constructed by combining Generative Causal Bayesian Network (CGBNN) with variational autoencoder. Predictions are made through causal inference and generative modeling. High-quality data is obtained through an active survey loop guided by uncertainty, and the model is iteratively updated to improve prediction accuracy.
It enables reliable causal prediction of counterfactual and hypothetical scenarios, outputs predicted effect values and uncertainty indicators, quantifies prediction credibility, reduces investment risks and decision-making costs in park operations, and improves the robustness and efficiency of decision-making.
Smart Images

Figure CN122134158A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart park operation and management technology, and in particular to a park operation decision analysis method and equipment. Background Technology
[0002] As the demand for refined operation of closed / semi-closed management areas such as smart parks, science and technology parks, cultural and sports venues, and university campuses continues to upgrade, park managers must conduct a preliminary assessment of the possible changes in personnel behavior that may be caused by the intervention before implementing operational intervention measures such as space renovation, addition of service facilities, adjustment of business layout, and optimization of pricing strategies, in order to reduce investment costs and operational risks.
[0003] Currently, the technical support for park operation decisions mainly falls into two categories: one is the correlation analysis method based on historical data. This method relies on historical operational data such as access control, WiFi, consumption records, and location to mine the statistical correlation between facility layout and personnel behavior. It can only be used to explain the existing operational status or extrapolate short-term trends. Its core flaw is that when the scheme to be evaluated belongs to a counterfactual / hypothetical scenario that has never been implemented in the past and has no corresponding observation data, such as adding a new business format in an area without similar services or structural changes in the service population, it can only learn statistical correlation models. When faced with external scenarios with distributional shifts, the prediction results will produce serious systematic biases and cannot support the pre-decision-making of schemes without historical data at all.
[0004] The second method is the demand assessment method based on static questionnaires. This type of method usually distributes a general questionnaire in the later stages of solution design to collect users' preferences for a limited number of alternative solutions. It can only be used for solution comparison and support assessment. Essentially, it is a post-voting decision support tool that cannot cover all scenarios of high-dimensional parameter combinations in operational solutions, nor can it predict the causal effects of unimplemented counterfactual solutions, nor can it supplement effective data for counterfactual scenarios.
[0005] Therefore, current decision analysis methods rely on subjective experience and are unable to make decisions in counterfactual / hypothetical scenarios. Summary of the Invention
[0006] This invention provides a method and equipment for park operation decision analysis, which solves the problem of inability to make decisions in counterfactual and hypothetical scenarios, and realizes decision analysis for counterfactual and hypothetical scenarios.
[0007] In a first aspect, the present invention provides a method for park operation decision analysis. This method includes: identifying the operational problem to be decided within a target park, and acquiring multidimensional features of the background environment and individual audiences within the target park, as well as a set of scheme variables consisting of different solutions and parameters corresponding to the operational problem; constructing a candidate scheme scenario library using the multidimensional features as covariates and the set of scheme variables as intervention variables; based on the candidate scheme scenario library, and combined with a pre-constructed generative causal Bayesian network, predicting counterfactual and hypothetical solutions without historical data through causal inference and generative modeling, outputting the prediction result for each scenario, including the predicted effect value and uncertainty index; based on the prediction result, selecting a high-value incremental sample set from the candidate scheme scenario library, iteratively updating the generative causal Bayesian network to obtain an updated generative causal Bayesian network; and based on the updated generative causal Bayesian network, predicting the final decision solution for the operational problem.
[0008] In a second aspect, embodiments of the present invention provide an electronic device including a memory and a processor. The memory stores a computer program, and the processor is configured to call and run the computer program stored in the memory to perform the steps of the method as described in the first aspect and any possible implementation thereof.
[0009] Thirdly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of the method as described in the first aspect and any possible implementation thereof.
[0010] This invention provides a method and device for park operation decision analysis. It constructs a full-spectrum candidate scenario library through structured definitions of covariates and intervention variables, covering various operational intervention scenarios to be evaluated. Relying on generative causal Bayesian networks that integrate causal inference and generative modeling capabilities, it achieves reliable causal prediction of counterfactual and hypothetical scenarios without historical observation data. It simultaneously outputs predicted effect values and uncertainty indicators to quantify prediction credibility and decision risk. Through uncertainty-guided model iteration updates, it achieves scientific decision analysis for counterfactual and hypothetical scenarios, effectively reducing investment risks and decision-making costs in park operations, and improving the robustness and efficiency of refined operational decisions. This invention addresses the problem of unreliable prediction and difficulty in proactive decision-making for counterfactual and hypothetical scenarios without historical data in park operations, achieving a closed-loop, end-to-end park operation decision analysis. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating a park operation decision analysis method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0014] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.
[0015] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the steps or modules listed, but may optionally include other steps or modules not listed, or may optionally include other steps or modules inherent to such process, method, product, or device.
[0016] To make the objectives, technical solutions, and advantages of the present invention clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.
[0017] As described in the background section, in the refined operation of closed or semi-closed management areas such as smart parks, science parks, and university campuses, managers usually need to conduct a preliminary assessment of the possible changes in personnel behavior caused by intervention measures before implementing space renovations or service facility adjustments (such as optimizing the layout of canteens, adding shared office areas, introducing business formats and transforming functions, etc.), including space access frequency, stay duration, path selection and time distribution, in order to reduce investment and operational risks.
[0018] Taking the post-Games Winter Olympics park as an example, the operator plans to add a 24-hour unmanned coffee bar on the east side of the first floor of Hall A, with a price of 15 yuan per cup and a maximum daily supply of 200 cups. Historically, this area has never had a beverage service point, and the post-Games visitor population has shifted from athletes / staff during the Games to a mixed group of local residents, tourists, and office workers. This results in a lack of observation of historical access control / WiFi / consumption / location data for this "location-business type-service method-price" combination, making it difficult to directly support a reliable prediction of "how many times the target group (e.g., park office workers) will use the coffee bar per week in the future." Essentially, this problem involves causal inference of the intervention effect: the intervention variable t is "the addition of an unmanned coffee bar on the east side of the first floor of Hall A and its parameter combination," the covariate x is user and environmental characteristics (such as identity category, department / company, gender, spending power, past dining behavior characteristics, accessibility, time constraints, etc.), and the outcome variable y is the user's expected weekly visit frequency to the coffee bar.
[0019] Since the intervention never occurred, existing methods based on historical correlation analysis or general questionnaires are insufficient to make credible inferences about the scenario of this unimplemented solution. This makes it difficult to make reliable predictions about operational optimization solutions that have not yet been implemented. If only subjective experience is relied upon, it can easily lead to idle facilities or wasted investment, resulting in high time, financial and operational risk costs.
[0020] Currently, decision support for park operations mainly relies on two technical approaches: one is the correlation analysis method based on historical data. This method utilizes historical trajectory data such as access control, WiFi, consumption, and location to mine statistical correlations between facility layout and personnel behavior, which are used to explain the existing state or extrapolate short-term trends. Its limitation lies in the fact that when the proposed intervention is an intervention that has not appeared in historical data or deviates significantly from the existing distribution, the model lacks the ability to model counterfactual scenarios. For example, in the post-Games transformation phase of the Winter Olympics park, if a plan is made to add a 24-hour unmanned coffee bar in a completely new location on the east side of the first floor of Hall A (historically, there were no beverage service facilities), and the target audience changes from athletes / staff during the Games to the public, tourists, and park office workers outside the Games, there are no user access records for this location in the existing historical data, nor is there any behavioral feedback on a similar service model. In this case, a model relying solely on historical correlations cannot reliably answer the question, "How many times a week will the target users use this coffee bar in the future?" The prediction results are prone to systematic bias due to distribution shifts.
[0021] The second method is static demand assessment based on questionnaires. This type of method is typically used when the operational plan is designed or nearing completion. It involves distributing general or semi-customized questionnaires to collect user preferences and subjective attitudes towards a limited number of pre-defined alternatives, used for support assessment or alternative comparison. However, this method is essentially a post-event evaluation or voting-based decision support, rather than a causal prediction tool for intervention effects. Taking the coffee bar in Building A as an example, if the manager only provides coarse-grained options such as "Do you support setting up a coffee bar in Building A?" or "Choose one between Building A and Building B", the questionnaire cannot cover the potential optimal configuration under high-dimensional parameter combinations such as spatial location (e.g., east side vs. west side on the first floor of Building A, east side vs. west side on the third floor), service hours (24 hours vs. 8:00–20:00), and price (15 yuan vs. 10 yuan / cup). More importantly, it cannot predict the effects of solutions that have not yet been conceived (such as "setting up a self-service coffee machine at 10 yuan per cup on the east side of Hall A"), nor does it support dynamic adjustment and re-evaluation of new solutions based on user feedback to achieve joint optimization of high-dimensional parameter combinations (location, service time, price, capacity, service method, etc.), which severely limits its applicability in the early stages of generating and optimizing operational solutions.
[0022] Therefore, current decision analysis methods lack reliable predictive ability for unimplemented counterfactual / hypothetical scenarios: existing methods struggle to answer forward-looking questions that exceed historical observations, such as, "If a 24-hour unmanned coffee bar (priced at 15 yuan per cup) is added on the east side of the first floor of Hall A in the Winter Olympics park, how many times a week will office workers in the park use it?" Since this location historically lacked beverage service, and the user group has shifted from athletes during the Games to a mixed public afterward, traditional correlation-based models cannot establish a causal mechanism between intervention and behavior, leading to unreliable predictions. This invention constructs a Generative Causal Bayesian Neural Network (CGBNN), which explicitly models "intervention-outcome" causal inference based on a latent outcome framework and integrates a variational autoencoder to enhance the representation ability of unobserved intervention scenarios, thus enabling structured inference of counterfactual scenarios even in the absence of historical data.
[0023] The prediction results lack a usable expression of uncertainty. Existing technologies typically output only a single predicted value (e.g., "expected to be used 2.3 times per week"), but do not provide the credibility or risk level of the prediction. Here, "expression of uncertainty" specifically refers to epistemic uncertainty, i.e., the degree of model knowledge gaps caused by insufficient coverage of training data in a specific intervention area (e.g., the east side of Hall A + 15-yuan combination), usually quantified in the form of the standard deviation, entropy, or BALD (Bayesian Active Learning by Disagreement) score of the prediction distribution. Without such an expression, managers cannot determine the credibility of the model's prediction results, making risk-sensitive decision-making difficult. This invention introduces a Bayesian neural network mechanism (implemented using MC Dropout) into CGBNN, enabling the model to simultaneously output a corresponding uncertainty metric (e.g., BALD value) while outputting behavioral response predictions, to identify high-risk scenarios or situations requiring further investigation.
[0024] The data acquisition strategy is mismatched with the needs of "external scenario modeling": Data collection using traditional surveys and intelligent sensing methods such as the Internet of Things typically aims to characterize the current situation, employing general questionnaires or uniform sampling, which cannot specifically obtain effective feedback on hypothetical solutions such as "setting up a 15-yuan unmanned coffee bar on the east side of Hall A." To improve the model's predictive ability in this high-dimensional intervention space, it is precisely necessary to actively collect high-quality labels in the areas where the model is most uncertain (such as specific location-price combinations). Currently, there is a lack of a survey organization and data collection strategy that matches the modeling needs for predicting the effects of "counterfactual / hypothetical" interventions in park operations; continuing with traditional paradigms may result in significant waste of time and economic costs. This invention addresses this issue through an "uncertainty-guided proactive research loop": the system calculates the predictive uncertainty of various hypothetical scenarios (such as different locations and price combinations) based on CGBNN, automatically filters high-uncertainty scenarios (such as the east side of Hall A + 15 yuan), dynamically generates customized questions (such as "If you often work in Hall A, assuming there is a 24-hour coffee bar here priced at 15 yuan per cup, how many times a week do you expect to use it?"), and pushes them to relevant subsets of people (such as office workers in Hall A), thus realizing a data acquisition paradigm of "on-demand collection and precise modeling".
[0025] Based on the above-mentioned technical problems, this invention provides a proactive survey and causal modeling method for optimizing park operations. This method is used to conduct proactive surveys guided by uncertainty before intervention implementation, and at a lower survey cost, to efficiently predict the intervention effects of operational intervention plans that have not been implemented or lack observational data support (such as the "location-time-price-capacity" combination of adding unmanned coffee bars).
[0026] like Figure 1 As shown, the present invention provides a method for park operation decision analysis. The method includes steps S101-S104.
[0027] S101. Identify the operational problems to be decided within the target park, and obtain the multidimensional characteristics of the background environment and individual audiences within the target park, as well as the set of solution variables consisting of different solutions and parameters corresponding to the operational problems.
[0028] In some embodiments, the multidimensional features include the spatial environment features of the target park, the individual identity attributes of the audience, the statistical features of audience behavior, and the spatiotemporal constraint features; wherein, the spatial environment features include the distribution of facility locations and spatial accessibility parameters; the individual identity attributes of the audience include user type, affiliated organization, and spending power tiers; the statistical features of audience behavior include historical visit frequency, consumption behavior habits, and work and rest patterns; and the spatiotemporal constraint features include weekday / weekend time division and spatial activity range.
[0029] In some embodiments, the set of scheme variables includes the spatial location of the facilities to be deployed, service type, service time rules, pricing strategy, supply capacity, and service method; wherein, spatial location includes specific locations of different buildings and floors within the target park; service type includes unmanned retail, catering services, shared office space, and public service facilities; service time rules include 24-hour operation and fixed-time operation; pricing strategy includes service pricing and fee rules at different tiers; supply capacity includes maximum daily service volume and peak carrying capacity; and service method includes unmanned self-service and manual service.
[0030] In some embodiments, operational issues include optimizing the layout and pricing strategies of service facilities in smart parks and technology parks; adjusting the configuration of canteens, study spaces, and campus transportation services on university campuses; reconstructing the business formats and service points for post-event operations in sports venues and cultural and sports parks; and optimizing the layout of public spaces and traffic-generating facilities in industrial parks and commercial complexes.
[0031] For example, optimizing the service facility layout and pricing strategies in smart parks and technology parks includes adding or adjusting service facilities (unmanned retail, coffee bars, shared meeting rooms), adjusting business format combinations, and optimizing service hours and pricing strategies. Adjusting the configuration of canteens, study spaces, and campus transportation services on university campuses includes: adjusting the location and scale of canteens / convenience stores, setting up shared study spaces, and configuring campus shuttle and micro-transportation services. Reconstructing the business formats and service points for the post-event operation of sports venues and cultural and sports parks includes: reconstructing supporting services during the post-event comprehensive utilization phase and optimizing the location and capacity configuration of public service points. Optimizing the layout of public spaces and traffic-generating facilities in industrial parks and commercial complexes includes: deploying traffic-generating facilities, renovating public spaces, and optimizing leasing formats.
[0032] The present invention can be deployed as an operational decision support system. The system accesses historical operational data, constructs candidate intervention scenarios, and outputs predicted effects and uncertainties; it automatically generates targeted survey tasks for highly uncertain solutions, and collects supplementary samples through online questionnaires / mini-programs / APP pushes, offline QR codes, etc.; after iteratively updating the model, it provides operators with closed-loop support of "predicted effects - risks - suggested surveys".
[0033] For example, the system can be deployed on the park's operation platform server or in the cloud. The data access layer accesses data such as access control, WiFi probes, consumption records, visit statistics, spatial POI information, organizational structure, and user profiles via API / database / file methods; survey and advertising can be integrated with channels such as SMS / WeChat Work / Official Accounts / Mini Programs / Apps.
[0034] S102. Construct a candidate scenario library using multidimensional features as covariates and the set of scenario variables as intervention variables.
[0035] As one possible implementation, step S102 can be specifically implemented as steps S1021-S1023.
[0036] S1021. Based on the operation configuration rules or preset scheme rule library of the target park, enumerate the candidate values of each type of parameter in the scheme variable set to generate a set of candidate operation schemes composed of multiple different parameter combinations.
[0037] S1022. For the target audience and environmental conditions of the target park, enumerate the candidate values of multidimensional features to generate a set of covariate combinations composed of multiple different feature combinations.
[0038] S1023. Perform a Cartesian product combination of the covariate combination set and the candidate operation scheme set to form a structured candidate scheme scenario library.
[0039] In some embodiments, each scenario in the candidate scenario library consists of a one-to-one combination of covariates and intervention variables, and each scenario corresponds to a park operation intervention scenario to be evaluated.
[0040] S103. Based on the candidate scenario library, combined with a pre-built generative causal Bayesian network, predict counterfactual and hypothetical scenarios without historical data through causal inference and generative modeling, and output the prediction results for each scenario.
[0041] In this embodiment of the application, the prediction results include the predicted effect value and the uncertainty index.
[0042] In some embodiments, the present invention can, based on the candidate scenario library and combined with a pre-built generative causal Bayesian network, predict counterfactual and hypothetical scenarios awaiting evaluation by causal inference and generative fusion modeling after initial training based on historical observation samples, and output the effect prediction and uncertainty assessment of each scenario.
[0043] As one possible implementation, step S103 can be specifically implemented as steps S1031-S1037.
[0044] S1031. Traverse all scenarios to be evaluated in the candidate scenario library and extract the covariate combinations and intervention variable combinations corresponding to each scenario.
[0045] S1032. Based on the one-to-one combination of covariates and intervention variables corresponding to each scenario, causal inference and generative modeling are completed through generative causal Bayesian networks. Structured causal inference is performed on counterfactual and hypothetical schemes without historical data, that is, causal effect estimation and response prediction are performed.
[0046] In this process, the covariate combination is mapped to the latent causal representation space via network encoding, and the intervention variable combination serves as a conditional input to participate in the causal relationship modeling of behavioral response.
[0047] S1033. During the inference phase, keep the Dropout layer of the generative causal Bayesian network continuously open, perform multiple random forward propagations on the same combination of covariates and intervention variables to obtain multiple sets of prediction results corresponding to the scenario.
[0048] For example, in the embodiment of the present invention, while keeping the parameters of the generative causal Bayesian network fixed during the inference phase, a random deactivation mechanism can be enabled to perform multiple random forward propagations on the same input to obtain multiple sets of prediction results corresponding to the scenario.
[0049] S1034. Perform statistical analysis on multiple sets of prediction results to obtain statistical results.
[0050] In some embodiments, the statistical results include the mean and the statistical distribution.
[0051] S1035. Based on the mean in the statistical results, determine the predicted value of the effect of the scenario.
[0052] S1036. Based on the statistical distribution in the statistical results, calculate the uncertainty index of the scenario.
[0053] In some embodiments, uncertainty metrics include information entropy, prediction standard deviation, and BALD metric based on information gain.
[0054] For example, prediction uncertainty can be represented by one or more of variance, standard deviation, and entropy; when selecting high-value scenarios based on uncertainty, the BALD index can be further used for ranking.
[0055] S1037. Based on the predicted effect value and uncertainty index of each scenario, generate the prediction result for each scenario.
[0056] S104. Based on the prediction results, select a high-value incremental sample set from the candidate solution scenario library, and iteratively update the generative causal Bayesian network to obtain the updated generative causal Bayesian network.
[0057] As one possible implementation, step S104 can be specifically implemented as steps one through nine.
[0058] Step 1: Based on the uncertainty index in the prediction results of each scenario in the candidate solution scenario library, sort the scenarios in descending order to obtain the sorting results.
[0059] Step 2: Based on the ranking results, select multiple scenarios with uncertainty greater than a preset threshold to form a set of proactive research tasks corresponding to the high-value incremental sample set.
[0060] Step 3: For each scenario in the set of proactive research tasks, generate customized research questions based on the combination of covariates and intervention variables for each scenario.
[0061] Step 4: Based on the combination of covariates for each scenario, filter to obtain the target audience.
[0062] Step 5: Based on customized research questions and target audience, conduct targeted research and collect effective behavioral feedback results.
[0063] For example, step five can be implemented as steps A1-A4.
[0064] A1. For each scenario in the set of proactive research tasks, based on the combination of covariates for that scenario, match the corresponding target audience subset from the audience database of the target park.
[0065] A2. Standardize the customized research questions corresponding to this scenario into a structured questionnaire, which includes quantifiable answer options.
[0066] A3. Through online channels or offline locations suitable for the target park, the structured questionnaire is pushed to a subset of the target audience for them to answer.
[0067] A4. Collect the responses to the structured questionnaires, and filter the responses to obtain valid behavioral feedback results.
[0068] Step 6: Based on the covariate and intervention variable combinations for each scenario in the active survey task set, as well as the effective behavioral feedback results, generate a high-value incremental sample set.
[0069] Step 7: Based on the high-value incremental sample set, use the joint loss function to retrain the generative causal Bayesian network in the current iteration process to complete the single-round model update.
[0070] Step 8: Calculate the overall prediction uncertainty of the generative causal Bayesian network for all scenarios in the candidate solution scenario library after a single round of model update, and the prediction uncertainty of the key target scenario corresponding to the operational problem to be decided.
[0071] Step 9: If the overall prediction uncertainty or the prediction uncertainty of the key target scenario is less than the preset convergence threshold, or if the maximum number of iterations is reached, stop the iterative update and use the generative causal Bayesian network updated in the current single round as the updated generative causal Bayesian network; otherwise, use the generative causal Bayesian network updated in the current single round as the generative causal Bayesian network in the next iteration process, and repeat steps 1 to 9 until the iterative update stops.
[0072] S105. Based on the updated generative causal Bayesian network, the final decision-making scheme for the operational problem is predicted.
[0073] In some embodiments, the present invention can output the effect evaluation results and risk assessment results of candidate solutions based on the updated generative causal Bayesian network, so as to support solution comparison, screening or decision-making, and finally obtain the decision solution for the operational problem.
[0074] As one possible implementation, step S105 can be specifically implemented as steps S1051-S1054.
[0075] S1051. Based on the updated generative causal Bayesian network, predict each scenario in the candidate solution scenario library to obtain the effect prediction value, uncertainty index and aggregated statistical results of the target audience's behavior response for each scenario.
[0076] S1052. Based on the predicted effect values of each scheme for the corresponding scenario, prioritize the operational revenue to obtain the priority ranking result.
[0077] S1053. Based on the uncertainty indicators of the corresponding scenarios of each scheme, the decision risk level is classified to obtain the risk level classification result.
[0078] S1054. Based on the priority ranking results and risk level classification results, the final decision-making schemes for the operational problems to be decided are obtained.
[0079] For example, after completing steps S201-S206 and initializing the model training based on historical observation data, this invention enters the uncertainty-guided proactive survey cycle phase. The goal is to proactively select the most informative survey subjects and questions for hypothetical operational intervention schemes that have not been implemented or lack observation data, gradually acquiring key data and improving the model's predictive accuracy and reliability across the entire scenario space. This cycle includes the following steps: constructing a hypothetical intervention scenario library, i.e., a candidate scheme scenario library; assessing scenario uncertainty and screening high-value samples; dynamically generating targeted survey questions and implementing data collection; integrating new samples to update the model; and repeating the cycle until the model uncertainty converges.
[0080] Construction of a Hypothetical Intervention Scenario Library: This involves systematically enumerating all combinations of hypothetical operational intervention schemes that are yet to be evaluated and have not been implemented, forming a large scenario library. This library is a structured set of intervention scenarios formed by systematically combining all operational intervention schemes to be evaluated but not yet implemented or lacking observational data support, along with their target groups. Each intervention scenario can be formally represented as: ;in: Characteristics of covariates representing the subjects of intervention and their background environment; This represents the vectorized representation of the intervention plan. This represents the j-th scene.
[0081] The scenario library is used to systematically represent all potential operational intervention scenarios, including counterfactual scenarios that have not occurred in combination and hypothetical intervention scenarios that have not been implemented at all. The scenario library provides a set of candidate research subjects for subsequent proactive learning. Combined with park operation application scenarios, the scenario library is constructed through the following steps: Step 1: Define the intervention variable space.
[0082] Intervention variables may include: facility location (e.g., east side of the first floor of Building A, north area of the third floor of Building B); service type (e.g., unmanned coffee bar, vending machine, shared office space); service time parameters (e.g., 24-hour operation, 8:00–20:00); pricing strategy (e.g., 15 yuan per cup, 20 yuan per cup); and capacity parameters (e.g., maximum daily supply of 200 cups, unlimited supply). For example, taking "adding an unmanned coffee bar" as an example, the intervention variable t can be broken down into several adjustable dimensions: location, service time, price, capacity, service method, etc. The system reads the candidate value set for each dimension from the operator's configuration or rule base and generates a candidate intervention set through enumeration. Forming a space for combinations of intervention variables: . Let m represent the m-th intervention variable.
[0083] Example: Location = {East side of the first floor of Hall A, West side of the first floor of Hall A, North area of the third floor of Hall B}; Service hours = {24h, 8-20}; Price = {10, 15, 20}; Capacity = {100, 200, 300}.
[0084] Step 2: Define the covariate space.
[0085] The covariate space represents the target user group and environmental conditions, including: user type (office workers, residents, tourists); spatial location (meeting space in Area A, office space in Area C, park restaurant in Area B); and time period (weekdays / weekends). This forms the covariate space. . Let n represent the nth covariate.
[0086] This invention constructs a set of target covariate types X based on organizational structure, permanent location, and behavioral profiles, such as "employees working in Hall A" and "weekend visitors", and can limit the time period and environmental conditions.
[0087] Step 3: Combine to form a complete scene library Will and This is combined to form a scenario library Q*. Each scenario corresponds to a complete description of "who (target group / profile) + when and where (covariates) + what intervention method is used (intervention parameters)," which is used for subsequent evaluation, screening, and survey question generation. Through combination: ; Obtain a complete library of hypothetical intervention scenarios: ; This represents the Kth candidate scenario.
[0088] Each scenario corresponds to a potential intervention survey question. The scenario consists of an intervention plan (a 24-hour unmanned coffee bar on the east side of the first floor of Building A, priced at 15 yuan per cup, with a daily limit of 200 cups) and covariates (office workers, the park restaurant in Zone B, and weekdays). The corresponding intervention survey questions are: Hello, "Office worker," if the park plans to open a "24-hour unmanned coffee bar" on the "east side of the first floor of Building A," priced at "15 yuan per cup" and "limited to 200 cups per day," would you visit this coffee bar before or after dining at the "park restaurant in Area B" on "weekdays"? How often? Scenario uncertainty assessment and high-value sample selection: Each hypothetical scenario in the "scenario library" is decomposed into covariates and intervention variables, which are then input into the pre-initialized and trained CGBNN model to calculate the corresponding prediction uncertainty (preferably using the BALD index, as it can accurately locate the region with the highest cognitive uncertainty). The system automatically selects several scenarios with the highest uncertainty.
[0089] Specifically, for each scene in the scene library: .
[0090] Input a trained CGBNN model and perform MC Dropout inference: . Let represent the prediction function during the k-th random forward propagation; This represents the random mask or random state corresponding to the k-th random forward propagation.
[0091] Calculate the uncertainty index and use it as a basis to screen high-value scenarios. Select: ; to obtain the set of scenarios with the most research value. This represents a set of high-value scenarios. This represents the uncertainty index corresponding to the j-th scenario.
[0092] For example, this invention can input each scenario in the candidate scenario library Q* into a CGBNN to obtain the predicted value and uncertainty U(q). The system sorts by U or BALD and selects the Top-K scenarios as the set of tasks for this round of research.
[0093] Dynamically generate and collect targeted survey questions: For selected high-uncertainty scenarios, generate customized questions consistent with intervention parameters and push these questions to relevant subsets of people (e.g., employees / students in specific buildings, users during specific time periods) to obtain direct feedback data on the "hypothetical intervention effect" (e.g., expected usage frequency, acceptable price range, alternative choices). For example, based on a set of high-value scenarios... The following question can be automatically generated: If the park plans to open a 24-hour unmanned coffee bar on the east side of the first floor of Hall A, priced at 15 yuan per cup and limited to 200 cups per day, would you visit this coffee bar before or after dining at the park restaurant in Area B on weekdays? How often? And based on... Covariate matching of survey subjects: for example, office staff in Hall A.
[0094] For example, this invention can automatically fill in a questionnaire template based on t and x for each survey scenario, generating customized questions that include scenario descriptions, price / location / time parameters, and setting the answer format to quantifiable labels (such as "expected 0 / 1 / 2 / 3 / 4 times or more per week" or continuous numerical input) so that it can be directly used as y for training. Example question: "Are you an office worker in Building A of the park? If so, assuming a new 24-hour unmanned coffee bar is added on the east side of the first floor of Building A, priced at 15 yuan / cup, how many times do you expect to use it per week?" This invention selects respondents from the user pool according to the target audience screening rules of the scenario (e.g., "permanent office workers in Building A"), and distributes them online or offline via QR codes; after collecting the answers, a new sample (x, t, y) is formed.
[0095] Obtaining corresponding behavioral response results through surveys and with This forms a new triplet of covariates, intervention variables, and behavioral effects.
[0096] Model updates are achieved by incorporating new samples: Survey feedback is used as a label for scene effects, and together with corresponding covariates and intervention variables, it constitutes new training samples. These samples are incrementally added to the training set, and the model is retrained or updated online. For example, this invention adds new samples to the training set and performs incremental training or periodic retraining of CGBNN to obtain an updated model. The system monitors whether the validation set error and overall uncertainty (e.g., average uncertainty on Q* or uncertainty in key scenes) converge. It stops when the uncertainty decreases below a threshold or reaches the budget / round limit; the final prediction and recommendations are output.
[0097] New data: ; This represents the covariates in the newly added samples. This represents the intervention variable in the newly added sample. This represents the outcome variable in the newly added sample.
[0098] Add to training set: D represents the original training set. This represents the training set after adding new samples.
[0099] Update the model: ; Represents the loss function. This means that the updated model parameters are obtained by minimizing the loss function.
[0100] Model uncertainty is reduced. The cycle of evaluation, screening, investigation, and updating is repeated to gradually improve the model's predictive accuracy and uncertainty characterization across the entire scenario library (including factual, counterfactual, and hypothetical scenarios).
[0101] For example, the decision support output, after several iterations, can output the following for any operational plan to be evaluated within the scenario library: predicted behavioral response results (which can be a single indicator or a vector of multiple indicators); and predicted confidence level / uncertainty level (used for risk alerts and ranking). Managers can then compare, screen, and make robust decisions based on this information.
[0102] For example, regarding the target solution "to add a 24-hour unmanned coffee bar on the east side of the first floor of Hall A, priced at 15 yuan per cup, with a daily production capacity of 200 cups," the system outputs: 1) Individual-level prediction of the target population and group aggregation statistics (mean, quantiles, etc.); 2) Predict the uncertainty U and risk level (e.g., low / medium / high); 3) A list of recommended high-value research scenarios for the next round, used to further reduce uncertainty or compare alternative solutions.
[0103] This invention provides a method for park operation decision analysis. It constructs a full-spectrum candidate scenario library through structured definitions of covariates and intervention variables, covering various operational intervention scenarios to be evaluated. Relying on generative causal Bayesian networks that integrate causal inference and generative modeling capabilities, it achieves reliable causal prediction of counterfactual and hypothetical scenarios without historical observation data. It simultaneously outputs predicted effect values and uncertainty indicators to quantify prediction credibility and decision risk. Through uncertainty-guided model iterative updates, it achieves scientific decision analysis for counterfactual and hypothetical scenarios, effectively reducing investment risks and decision-making costs in park operations, and improving the robustness and efficiency of refined operational decisions. This invention addresses the problem of unreliable prediction and difficulty in proactive decision-making for counterfactual and hypothetical scenarios without historical data in park operations, achieving a closed-loop, end-to-end park operation decision analysis.
[0104] Optionally, the park operation decision analysis method provided in this embodiment of the invention includes steps S201-S206 before step S103, that is, before the initial training is completed based on the candidate scheme scenario library, combined with the pre-built generative causal Bayesian network, and through causal inference and generative fusion modeling to predict counterfactual and hypothetical scheme waiting for evaluation scenarios without historical data, and before outputting the effect prediction and uncertainty assessment of each scenario.
[0105] S201. Obtain existing survey data on various operational issues in the target park during historical periods.
[0106] In some embodiments, the existing survey data includes combinations of covariates, combinations of intervention variables, and actual behavioral response results for each operational issue.
[0107] In some embodiments, existing survey data, i.e., historical sample data, including survey data and / or behavioral observation data, is available. S202. Using the covariate combination and intervention variable combination corresponding to each operational problem as input, and the actual behavioral response results corresponding to each operational problem as output, construct training samples.
[0108] In some embodiments, each training sample is represented as a triple (x, t, y). Covariate x: Characterizes the "person-time-space" contextual conditions, such as: user category (office / resident / tourist), organizational department, gender, spending power tier, past frequency of restaurant visits, walking distance from the usual location to the candidate location, accessibility indicators, weekday / weekend, etc. Intervention variable t: Characterizes operational plan parameters, such as: facility type (unmanned coffee bar / self-service coffee machine), spatial location code (east side of the first floor of Hall A, etc.), service hours (24 hours, etc.), price (15 yuan / cup, etc.), capacity (200 cups / day, etc.), payment method, etc. Outcome variable y: Behavioral response indicators, such as: weekly visits to the target facility, or the probability / number of visits within a given time period.
[0109] S203. Construct the initial network with encoder-multiple decoder as the main network structure.
[0110] In some embodiments, the network main structure includes a covariate encoding module, an intervention reconstruction decoding module, a behavior prediction decoding module, and a covariate reconstruction decoding module.
[0111] The covariate encoding module maps the input covariate combination to the latent causal representation space. The intervention reconstruction decoding module, behavior prediction decoding module, and covariate reconstruction decoding module are used to complete the intervention feature reconstruction, behavior response prediction, and covariate feature reconstruction, respectively, based on the features of the latent causal representation space.
[0112] For example, the covariate encoding module maps the input covariate combination to the latent causal representation space, including: inputting the input covariate combination into the covariate encoding module, obtaining the distribution parameters of the latent variables through a multi-layer feedforward neural network, wherein the distribution parameters include the mean and standard deviation, or the mean and log-variance parameters; obtaining the latent variables based on the mean and variance parameters or the log-variance parameters through a reparameterization sampling operation; inputting the latent variables into each decoding module; the latent variables are used to capture key structural information related to the intervention-behavior causal relationship in the covariate combination.
[0113] In some embodiments of the present invention, the generative causal Bayesian neural network (CGBNN) model provides the following inputs: Covariates: describing the multidimensional characteristics of the park's background environment and individual audiences, including identity attributes, department / college, work and rest habits, past behavioral statistical characteristics, spatial accessibility, and time constraints; Treatment: describing different sets of parameters for different plans to add unmanned coffee bars in the office area during the operation period, including but not limited to combinations of facility type, different spatial locations, service time, capacity / supply, price / charging rules, and service methods.
[0114] In some embodiments, the CGBNN model is based on the Rubin Casual Model framework in the field of causal inference, deeply integrating the generative capabilities of a variational autoencoder with the uncertainty quantification capabilities of a Bayesian neural network. By fusing causal inference and generative models, the model enhances its extrapolation and generalization capabilities in counterfactual intervention scenarios. By employing Monte Carlo Dropout (MCDropout) technology to transform the model into a Bayesian neural network model, it can output not only predictions of human behavioral responses (such as the frequency of visits to a facility) but also the uncertainty of those predictions.
[0115] This invention provides a deep learning-based latent outcome model fusion generative model to enhance the model's predictive ability in counterfactual inference. The model is rewritten as a Bayesian Neural Network, enabling the model to output the predicted value while providing the uncertainty of the prediction. The proposed "uncertainty-guided active survey loop" can calculate the prediction uncertainty of each hypothetical scenario, automatically filter high uncertainty scenarios, dynamically generate customized questions, and push them to relevant subsets of people, realizing a data acquisition paradigm of "on-demand collection and precise modeling".
[0116] In some embodiments, the CGBNN model uses a parameterized representation vector t of the intervention plan, along with parameterized representation vectors x of the park's background environment and the individual audience member, as input to predict the individual audience member's response y. Specifically, the model employs a deep neural network structure based on generative causal modeling, consisting of an encoder module and multiple decoder modules. The overall structure includes: an encoder module; an intervention reconstruction decoder module; a behavior prediction decoder module; and a covariate reconstruction decoder module.
[0117] For example, the encoder module uses VAE-style random encoding. It inputs the covariance x into the encoder E and outputs latent distribution parameters: the logarithmic form of the latent distribution mean parameter μ and the latent distribution variance parameter logσ². The latent variable z is obtained through reparameterized sampling. Random encoding improves robustness to unseen combinations and noise perturbations. The encoder module is used to process the input individual features... Mapping to the latent representation space yields the distribution parameters of the latent variables: Latent distribution mean parameter ; Latent distribution standard deviation ; And latent variables are obtained through sampling operations: ; The mean is The standard deviation is It follows a normal distribution.
[0118] in: This represents the latent representation vector, used to capture key structural information in individual features. express A real space of dimension , representing latent variables It is A 3D real-valued vector. This latent variable is used for subsequent decoding tasks.
[0119] For example, the intervention reconstruction decoder module is used to reconstruct intervention features based on latent variables. Intervention Reconstruction Decoder D t Based on z, the intervention representation t' is reconstructed to guide the encoder to learn representations related to the intervention (which can be understood as representations of propensity score / intervention selection mechanism), thereby mitigating the influence of confounding factors.
[0120] ; This represents the intervention variable after reconstruction. Indicates intervention reconstruction decoder D t The output value of this module is used to constrain latent variables to effectively express the association structure between individuals and interventions, thereby improving the model's causal expressiveness and generalization ability.
[0121] For example, the behavior prediction decoder module is used to predict behavioral response outcomes based on latent variables and intervention variables. Behavior Prediction Decoder D r Predict y' using (z,t) as conditional input. To avoid the attenuation of intervention information in deep networks, t can be repeatedly injected into multiple layers (e.g., splicing or AdaIN / dynamic fully connected layers) to enhance the expression of different combinations of intervention parameters.
[0122] ; This represents the predicted behavioral response value. The behavior prediction decoder uses latent variables. Intervention variables The input condition is the behavior response prediction result.
[0123] This module is a condition generation module, where intervention variables serve as condition inputs and directly participate in the prediction process, enabling the model to clearly define the impact of interventions on behavioral outcomes. This module can be constructed using a multi-layer feedforward neural network.
[0124] For example, the covariate reconstruction decoder module is used to reconstruct input covariates from latent variables. Covariate Reconstruction Decoder D c Reconstructing x' based on z to constrain the potential spatial structure reduces overfitting and improves generalization.
[0125] This module is used to constrain the structural rationality of the potential representation space, thereby improving the generalization ability and stability of the model. Represents the reconstructed covariates. Describing the covariate reconstruction decoder D x The output is the calculation result of the reconstructed covariates, with the latent variable z as input.
[0126] This model, based on the causal generative network described above, introduces the concept of a Bayesian neural network and employs Monte Carlo Dropout (MC Dropout) to approximate the point estimation of network weights as sampling of the posterior distribution. This expands the model output from a single-point prediction to a "prediction distribution + uncertainty measure." The key to this modification lies in placing Dropout layers in all hidden layers of the model and setting a deactivation probability. This makes each forward propagation equivalent to sampling a prediction from a "random subnetwork". Structurally, this can be summarized as: the same set of model parameters... +Dropout random mask They jointly decide on a forward propagation, and the output is .in Randomly generated by Dropout at each level.
[0127] Model training: Model training is based on a set of data samples. The model parameters are trained by jointly optimizing multiple objective functions, including: Behavioral prediction loss: ; Intervention and reconstruction losses: ; Covariate reconstruction loss: ; Latent variable distribution constraint loss: ; The total loss function is: ,in: Here are the weight parameters. Training uses mini-batch stochastic gradient descent (e.g., the Adam optimizer) to iteratively update the parameters until convergence. E represents the encoder. This indicates intervention to reconstruct the decoder. This indicates a behavior prediction decoder. This indicates the covariate reconstruction decoder.
[0128] Specifically, during the training phase, Dropout is enabled in all hidden layers. By randomly deactivating some neurons, the model avoids relying on specific neuron combinations, thereby improving generalization ability; simultaneously, this "random sub-network training" also provides the foundation for the subsequent Bayesian approximation of MC Dropout. The training samples are... The model is optimized using a multi-task joint loss: minimizing intervention reconstruction, response prediction, covariate reconstruction, and KL divergence (to balance prediction accuracy, propensity score / intervention representation, and covariate representation). The training process involves mini-batch sampling, forward propagation to obtain the encoded representation and each decoded output, calculating the loss, and updating the parameters using Adam until convergence.
[0129] Unlike conventional deep networks that disable Dropout during inference, this model keeps Dropout enabled during the inference phase and its parameters remain fixed without being updated, applying the same input scenario. A random forward propagation is used to obtain the prediction sample set. The supplementary material explicitly states that inference will only include unlabeled data containing the input features. (This can be understood as a candidate intervention scenario) (Combination of inputs) into the trained model, execute Subsequent random forward propagation, and based on The uncertainty of the prediction results is calculated.
[0130] Specifically, let's assume the same input scenario conduct Random forward propagation (with a different Dropout mask each time): The prediction set {y(1)…y(T)} is obtained. The mean is used as the predicted value; the variance or entropy is used as the uncertainty index; when necessary, the BALD index can be further calculated to measure the uncertainty of the model parameters and as a sampling criterion for active surveys.
[0131] in, These are the fixed model parameters after training. For the first The mask is randomly generated by Dropout during the next inference. This represents the predicted value of the same input during the k-th random forward propagation.
[0132] The model's final prediction value is taken The mean of the outputs: ; Uncertainty can be estimated by calculating the entropy of multiple predictions. This is when the model output is a class probability vector. (Or, when discretizing continuous outputs into probability distributions), one can first... Calculate the mean of the probability outputs: ; Then use entropy to measure the total uncertainty: ; in, For discrete result state / range index, The mean probability is at the th The components on the class. The greater the entropy, the "flatter" the predicted distribution, and the more uncertain the model. Let represent the probability vector obtained in the k-th forward propagation. T represents the number of forward propagations.
[0133] S204. Set the joint loss function in the initial network.
[0134] In some embodiments, the joint loss function is composed of behavior prediction loss, intervention reconstruction loss, covariate reconstruction loss, and a weighted sum of KL divergence terms.
[0135] For example, the joint loss function is composed of a weighted average of the intervention prediction loss, response prediction loss, covariate reconstruction loss, and KL divergence loss.
[0136] S205. Perform Bayesian transformation on the initial network after setting the joint loss function by setting Dropout layers in all hidden layers of the network to obtain the Bayesian transformed network.
[0137] In some embodiments, the Dropout layer is used to enable the network’s Bayesian inference capability through the Monte Carlo Dropout mechanism, that is, to quantify the uncertainty of approximate Bayesian inference, so that the effect prediction value and uncertainty index can be output synchronously through multiple random forward propagations during the inference phase.
[0138] S206. Based on the training samples and combined with the joint loss function, the Bayesian modified network is trained to obtain a pre-constructed generative causal Bayesian network.
[0139] For example, embodiments of the present invention can perform model initialization training based on historical observation data. Historical observation data under the existing operational conditions of the park is acquired. This historical observation data includes behavioral records and their extractable feature information collected during actual operation, such as access control records, consumption records, facility visit records, and other operational data reflecting behavioral responses. Combined with existing operational settings, covariates, intervention variables, and the effects of observed behaviors are vectorized to construct an initial training sample set. The causal generative Bayesian neural network (CGBNN) is then preliminarily trained using these training samples to obtain the model's fitting ability and basic generalization ability to the observed scenarios.
[0140] In the process of initializing and training the model using historical observation data, the covariate, intervention variable, and behavioral response effect triplet constructed based on historical data reflects the "background conditions of the behavior," the intervention variable is not an artificially designed experimental variable, but corresponds to the state of facilities or strategies already implemented under real operational conditions, and the behavioral response effect is the observed behavioral outcome under real operational settings. Historical observation data provides behavioral response data under real operational conditions, enabling the model to: fit observed scenarios; learn the latent representation space; and establish preliminary causal structure expression capabilities. Through joint training on historical samples, the model acquires: the ability to fit observed scenarios; the ability to interpolate adjacent scenarios; and the preliminary extrapolation ability for some unseen combinations. This basic generalization ability is the starting point for subsequent active survey learning and iterative optimization.
[0141] This invention offers the following technical advantages: It enhances the predictive ability for counterfactual and hypothetical solutions. By combining a causal inference framework with generative modeling, it can still perform structured inferences about uninterrupted solutions even when historical observations are insufficient or solutions deviate from existing distributions. This alleviates the problem of traditional correlation models failing in out-of-domain scenarios, allowing operators to conduct comparable effect assessments of innovative solutions before implementation. It enables the quantitative expression and indication of predicted risks: While outputting effect predictions, this invention also outputs uncertainty metrics, enabling managers to identify high-risk solutions / high-risk groups / high-risk spatial conditions and decide whether to "directly adopt the predicted results" or "further research is needed before making a decision," thereby reducing the risk of blind decision-making. It reduces research costs with higher information gain: Through proactive research guided by uncertainty, this invention concentrates research resources on the key scenarios where the model lacks the most knowledge and which have the greatest impact on decision-making, avoiding the inefficient collection of general questionnaires. This achieves higher modeling gain with fewer samples, improving the return on investment in research and shortening the decision-making cycle. It forms a closed-loop intelligent decision-making system of "prediction-evaluation-learning-optimization": This invention is not only a predictive tool but also a continuously evolving decision support system. It can improve itself with each proactive survey, providing solid technical support for the long-term, dynamic, and refined operation of the park.
[0142] For example, to verify the technical effectiveness of the Causal Generative Bayesian Neural Network (CGBNN) model and its uncertainty-guided active learning framework in spatial planning decision-making, this invention systematically validated its solution through model performance experiments and framework efficiency experiments. The experiments were conducted on both synthetic and real-world datasets to verify the model's predictive capabilities and the effectiveness of the active learning mechanism in reducing data collection costs. The following section describes the technical implementation process and verification results of this invention in the application scenario of predicting the impact of optimizing the layout of public service facilities in urban functional parks on visitor behavior.
[0143] (1) Application Scenarios and Technical Issues. In the process of optimizing the operation of cities or industrial parks, managers usually need to assess the potential impact of different spatial intervention schemes on people's behavior patterns before implementing operational adjustment plans. For example, when adding catering facilities, adjusting the layout of public service facilities, or changing the spatial accessibility structure in the park, it is necessary to predict whether these interventions will change the frequency of visits or activity paths of people. However, the above planning schemes often belong to planning scenarios that have not yet been implemented, and therefore lack historical observation data to support them.
[0144] Existing deep learning-based behavior prediction methods primarily rely on historical observation data to establish correlations between variables. When the prediction scenario exceeds the original data distribution, these methods struggle to reliably extrapolate and cannot provide an assessment of the prediction's credibility. Furthermore, planning practices often require supplementary data collection through questionnaires or public participation, but traditional survey methods lack specificity and struggle to identify which scenarios' data are most critical for model training, resulting in high survey costs and low data utilization efficiency. Therefore, a technical method is needed that can reliably predict unprecedented planning intervention schemes under limited observation data conditions, quantify prediction uncertainty, and guide subsequent data collection processes.
[0145] (2) Example of Implementation. In this embodiment, the applicant constructed an experimental dataset based on behavioral survey data from a certain urban park. The data sample included covariate information of the target audience, such as gender, age range, occupation type, activity area, and commuting mode, and also recorded the frequency of visits or behavioral feedback under different facility layout conditions. Spatial intervention factors were represented by quantifiable spatial indicators, such as facility walking distance, service radius, or facility density. Through questionnaire surveys and spatial data processing, a total of 3,874 valid sample records were obtained.
[0146] During the experiment, the data samples were first divided into training dataset, validation dataset, and test dataset. The training dataset was used for model parameter learning, the validation dataset was used for model parameter tuning, and the test dataset remained independent throughout the experiment and was used to evaluate the model's predictive performance.
[0147] During the model training phase, the covariate features and intervention variables from the training data are input into the CGBNN model proposed in this invention. The model first learns representations of the covariate features through an encoding network, mapping them to a latent feature space, and constructs a causal relationship model between the intervention variables and behavioral responses within this latent space. A variational autoencoder (VAE) mechanism is further introduced into this encoder-decoder structure, enabling the model to learn the latent probability distribution of the covariates, thereby enhancing the model's generalization ability under unobserved intervention conditions.
[0148] The Monte Carlo Dropout mechanism is introduced into the model structure, enabling the neural network to obtain the probability distribution of the prediction results through multiple random samplings during the prediction phase, and to calculate the prediction uncertainty accordingly. In this way, the model can not only provide predicted behavioral responses but also provide prediction reliability information for identifying high-risk prediction scenarios.
[0149] After the model completes its initial training, this invention further constructs an uncertainty-guided active learning mechanism. Specifically, the currently trained CGBNN model is first used to predict candidate intervention scenarios, and the prediction uncertainty index for each candidate scenario is calculated. Intervention combinations with higher prediction uncertainty are selected based on the magnitude of uncertainty, and targeted survey questions are generated accordingly, such as asking specific groups about their preferred frequency of visits under a new facility layout.
[0150] In the experimental implementation, a data pool containing potential survey questions was set up to simulate the questionnaire question set in planning practice. In each iteration, the model selected several samples from the data pool as new survey samples based on the uncertainty index, and added their corresponding real responses to the training dataset. Subsequently, the model was retrained on the expanded training dataset, and the model's predictive performance was evaluated using the unchanged test dataset. By repeatedly executing the process of "high uncertainty sample selection - data supplementation - model retraining - performance evaluation," the training data was gradually expanded, and the model's predictive ability was continuously improved.
[0151] (3) Verification of Technical Effectiveness. The applicant first conducted a model performance experiment, comparing the proposed CGBNN model with existing deep learning causal inference models, including DragonNet, DRNet, and VCNet. The experiments were conducted on both synthetic and real-world campus datasets, and the model prediction performance was evaluated using the mean squared error (AMSE). The experimental results show that on the real-world campus dataset, the proposed CGBNN model has an average mean squared error of approximately 0.0395 on the test set, while the errors of DragonNet, DRNet, and VCNet are approximately 0.0440, 0.0446, and 0.0528, respectively. The model of this invention exhibits higher prediction accuracy and more stable prediction performance under different data environments.
[0152] In the Framework Experiment, the applicant further verified the impact of the uncertainty-guided data collection mechanism on model training efficiency. In this experiment, the model was first trained using a small initial training sample pool. Then, training samples were gradually supplemented from the candidate data pool, and the model was retrained after each round of data supplementation. The predictive performance was then evaluated using the same test dataset. Experimental results show that when using a sample selection strategy based on the uncertainty index (BALD), the method of this invention can achieve predictive performance comparable to or even higher than that of a random sampling strategy with a smaller number of training samples. For example, in real-world data experiments, only about 960 supplementary samples were needed to achieve optimal predictive performance, while the random sampling strategy required about 1160 supplementary samples to achieve similar predictive accuracy, reducing the required training samples by about 21%. In experiments with partially synthetic data, the method of this invention reduced the required training samples by up to about 74% compared to the random sampling strategy.
[0153] As can be seen from the above embodiments, the causal generative Bayesian neural network model proposed in this invention can effectively predict the behavioral effects of spatial intervention measures under limited observation data conditions, and quantifies the uncertainty of prediction through Bayesian inference mechanisms. Simultaneously, through an uncertainty-guided active learning strategy, priority can be given to collecting data samples most critical to improving model performance, thereby reducing data collection costs while ensuring predictive performance. Therefore, this invention not only improves the accuracy and stability of spatial intervention response prediction, but also provides predictive credibility information for planning decisions, and significantly improves data utilization efficiency and reduces planning survey costs through active learning mechanisms, thus providing reliable technical support for spatial intervention decisions in urban planning and park operation management.
[0154] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0155] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 300 includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, it implements the steps in the above-described method embodiments. Alternatively, when the processor 301 executes the computer program 303, it implements the functions of each module / unit in the above-described device embodiments.
[0156] For example, the computer program 303 may be divided into one or more modules / units, which are stored in the memory 302 and executed by the processor 301 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program 303 in the electronic device 300.
[0157] The processor 301 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0158] The memory 302 can be an internal storage unit of the electronic device 300, such as a hard disk or memory of the electronic device 300. The memory 302 can also be an external storage device of the electronic device 300, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device 300. Furthermore, the memory 302 can include both internal and external storage units of the electronic device 300. The memory 302 is used to store the computer program and other programs and data required by the terminal. The memory 302 can also be used to temporarily store data that has been output or will be output.
[0159] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for analyzing park operation decisions, characterized in that, include: Identify the operational problems to be decided within the target park, and obtain the multidimensional characteristics of the background environment and individual audiences within the target park, as well as the set of solution variables consisting of different solutions and parameters corresponding to the operational problems; A candidate scenario library is constructed using multidimensional features as covariates and a set of scenario variables as intervention variables. Based on the candidate scenario library, combined with a pre-built generative causal Bayesian network, the system predicts counterfactual and hypothetical scenarios without historical data through causal inference and generative modeling, and outputs the prediction results for each scenario, which include the predicted effect value and uncertainty index. Based on the prediction results, a high-value incremental sample set is selected from the candidate solution scenario library, and the generative causal Bayes network is iteratively updated to obtain the updated generative causal Bayes network. Based on the updated generative causal Bayesian network, the final decision-making solution for the operational problem is predicted.
2. The park operation decision analysis method according to claim 1, characterized in that, The multidimensional features include the spatial environment features of the target park, individual audience identity attributes, audience behavioral statistical features, and spatiotemporal constraint features; wherein, the spatial environment features include facility location distribution and spatial accessibility parameters; the individual audience identity attributes include user type, affiliated organization, and spending power tiers; the audience behavioral statistical features include historical visit frequency, consumption habits, and daily routines; and the spatiotemporal constraint features include weekday / weekend time division and spatial activity range; The set of variables for the proposed solution includes the spatial location of the facilities to be deployed, service type, service time rules, pricing strategy, supply capacity, and service method. Specifically, the spatial location includes the specific locations of different buildings and floors within the target park; the service type includes unmanned retail, catering services, shared office space, and public service facilities; the service time rules include 24-hour operation and fixed-time operation; the pricing strategy includes different service pricing and fee rules; the supply capacity includes the maximum daily service volume and peak carrying capacity; and the service method includes unmanned self-service and manual service. The process of constructing a candidate scenario library using multidimensional features as covariates and a set of scenario variables as intervention variables includes: Based on the target park's operation configuration rules or preset scheme rule library, enumerate the candidate values of each type of parameter in the scheme variable set to generate a set of candidate operation schemes composed of multiple different parameter combinations. For the target audience and environmental conditions of the target park, candidate values for multidimensional features are enumerated to generate a set of covariate combinations consisting of multiple different feature combinations; The set of covariate combinations and the set of candidate operation schemes are combined by Cartesian product to form a structured candidate scheme scenario library. Each scenario in the candidate scheme scenario library consists of a one-to-one covariate combination and an intervention variable combination, and each scenario corresponds to a park operation intervention scenario to be evaluated.
3. The park operation decision analysis method according to claim 1, characterized in that, Based on the candidate scenario library, and combined with a pre-built generative causal Bayesian network, the system predicts counterfactual and hypothetical scenarios without historical data through causal inference and generative modeling, outputting the prediction results for each scenario, including: Traverse all scenarios to be evaluated in the candidate scenario library and extract the covariate combination and intervention variable combination corresponding to each scenario; Based on the one-to-one combination of covariates and intervention variables for each scenario, causal inference and generative modeling are completed through the generative causal Bayesian network. Structured causal inference is performed on counterfactual and hypothetical schemes without historical data. The covariate combination is mapped to the latent causal representation space through network encoding, and the intervention variable combination participates in the causal relationship modeling of behavioral response as a conditional input. During the inference phase, the Dropout layer of the generative causal Bayesian network is kept on, and multiple random forward propagations are performed on the same combination of covariates and intervention variables to obtain multiple sets of prediction results corresponding to the scenario. Statistical analysis is performed on the multiple sets of prediction results to obtain statistical results, which include the mean and statistical distribution. Based on the mean of the statistical results, the predicted effect value for this scenario is determined; Based on the statistical distribution in the statistical results, the uncertainty index of the scenario is calculated; the uncertainty index includes information entropy, prediction standard deviation, and BALD index based on information gain; Based on the predicted effect value and uncertainty index for each scenario, a prediction result is generated for each scenario.
4. The park operation decision analysis method according to claim 1, characterized in that, Before the step of predicting counterfactual and hypothetical scenarios without historical data based on the candidate scenario library, combined with a pre-built generative causal Bayesian network, and using causal inference and generative modeling to output the prediction result for each scenario, the following steps are also included: Obtain existing survey data on various operational issues in the target park during historical periods. The existing survey data includes the covariate combination, intervention variable combination, and actual behavioral response results for each operational issue. Training samples are constructed by taking the covariate combination and intervention variable combination corresponding to each operational problem as inputs and the actual behavioral response results corresponding to each operational problem as outputs. An initial network is constructed using an encoder-multi-decoder architecture. The main network structure includes a covariate encoding module, an intervention reconstruction decoding module, a behavior prediction decoding module, and a covariate reconstruction decoding module. The covariate encoding module maps the input covariate combinations to a latent causal representation space. The intervention reconstruction decoding module, the behavior prediction decoding module, and the covariate reconstruction decoding module are used to perform intervention feature reconstruction, behavior response prediction, and covariate feature reconstruction, respectively, based on the features of the latent causal representation space. In the initial network, a joint loss function is set, which is composed of behavior prediction loss, intervention reconstruction loss, covariate reconstruction loss and KL divergence term weighting. The initial network after setting the joint loss function is modified by Bayesianization. Dropout layers are set in all hidden layers of the network to obtain the modified network. The dropout layers are used to realize the network's Bayesian inference capability through the Monte Carlo Dropout mechanism, so that the effect prediction value and uncertainty index can be output synchronously through multiple random forward propagation during the inference stage. Based on the training samples and combined with the joint loss function, the Bayesian modified network is trained to obtain the pre-constructed generative causal Bayesian network.
5. The park operation decision analysis method according to claim 3 or 4, characterized in that, The covariate encoding module maps the input covariate combinations to a latent causal representation space, including: The input covariates are combined and fed into the covariate encoding module, and the distribution parameters mean and standard deviation of the latent variables are obtained through mapping by a multi-layer feedforward neural network. By performing a reparameterized sampling operation, latent variables are obtained from a normal distribution that conforms to the distribution parameters mean and standard deviation. The latent variables are input into each decoding module; the latent variables are used to capture key structural information related to the intervention-behavior causal relationship in the covariate combination.
6. The park operation decision analysis method according to claim 1, characterized in that, The step of selecting a high-value incremental sample set from the candidate solution scenario library based on the prediction results and iteratively updating the generative causal Bayesian network to obtain the updated generative causal Bayesian network includes: Step 1: Based on the uncertainty index in the prediction results of each scenario in the candidate solution scenario library, sort the scenarios in descending order to obtain the sorting results; Step 2: Based on the ranking results, select multiple scenarios with uncertainty greater than a preset threshold to form a set of proactive research tasks corresponding to the high-value incremental sample set; Step 3: For each scenario in the set of proactive research tasks, generate customized research questions based on the combination of covariates and intervention variables for each scenario; Step 4: Based on the combination of covariates for each scenario, filter to obtain the target audience; Step 5: Based on the customized survey questions and target audience, conduct targeted surveys and collect valid behavioral feedback results; Step 6: Based on the covariate and intervention variable combinations for each scenario in the active survey task set, as well as the effective behavioral feedback results, generate a high-value incremental sample set; Step 7: Based on the high-value incremental sample set, the generative causal Bayesian network in the current iteration process is retrained using the joint loss function to complete a single round of model update; Step 8: Calculate the overall prediction uncertainty of the generative causal Bayesian network for all scenarios in the candidate solution scenario library after a single round of model update, and the prediction uncertainty of the key target scenario corresponding to the operational problem to be decided. Step 9: If the overall prediction uncertainty or the prediction uncertainty of the key target scenario is less than the preset convergence threshold, or if the maximum number of iterations is reached, stop the iterative update and use the generative causal Bayesian network updated in the current single round as the updated generative causal Bayesian network; otherwise, use the generative causal Bayesian network updated in the current single round as the generative causal Bayesian network in the next iteration process, and repeat steps 1 to 9 until the iterative update stops.
7. The park operation decision analysis method according to claim 6, characterized in that, The targeted survey, based on the customized survey questions and target audience, and the collection of valid behavioral feedback results, include: For each scenario in the set of proactive survey tasks, based on the combination of covariates for that scenario, a corresponding subset of the target audience is matched from the audience database of the target park; The customized research questions corresponding to this scenario are standardized into a structured questionnaire, which includes quantifiable answer options. The structured questionnaire will be pushed to the target audience subset through online channels or offline locations that are suitable for the target park, so that the audience in the target audience subset can answer the questions. The responses to the structured questionnaire are collected and filtered to obtain valid behavioral feedback results.
8. The park operation decision analysis method according to claim 1, characterized in that, The final decision-making scheme for the operational problem, based on the updated generative causal Bayesian network, includes: Based on the updated generative causal Bayesian network, predictions are made for each scenario in the candidate scenario library, and the predicted effect value, uncertainty index and aggregated statistical results of the behavioral response of the target audience are obtained for each scenario. Based on the predicted effects of each solution for the corresponding scenario, the operational revenue is prioritized and the priority ranking results are obtained. The decision risk level is classified based on the uncertainty indicators of the corresponding scenarios for each scheme, and the risk level classification results are obtained. Based on the priority ranking results and risk level classification results, the final decision-making solutions for the operational problems to be decided are obtained.
9. The park operation decision analysis method according to any one of claims 1 to 8, characterized in that, The operational issues mentioned include optimizing the layout and pricing strategies of service facilities in smart parks and technology parks; adjusting the configuration of canteens, study spaces, and campus transportation services on university campuses; reconstructing the business formats and service points for post-event operations in sports venues and cultural and sports parks; and optimizing the layout of public spaces and traffic-generating facilities in industrial parks and commercial complexes.
10. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor being configured to invoke and run the computer program stored in the memory to perform the method as described in any one of claims 1 to 9.