Park resource scheduling method, equipment and medium

By constructing a cluster of intelligent agents for park resources and a multi-objective optimization model, the dynamic and real-time problems of traditional park resource scheduling systems have been solved, achieving efficient resource scheduling and improved user experience.

CN120806468APending Publication Date: 2025-10-17山东浪潮智慧建筑科技有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510894786.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-17

Smart Images

  • Figure CN120806468A_ABST
    Figure CN120806468A_ABST
Patent Text Reader

Abstract

The invention relates to the field of resource management, and discloses a park resource scheduling method and device and a medium, and the method comprises the steps: constructing a park resource agent cluster based on park resource information; constructing a resource space-time state diagram based on a preset rule of the park resource agent cluster and the park resource information; based on a pre-constructed strategy cooperation rule and the resource space-time state diagram, determining an initial scheduling strategy of the resource scheduling event when the plurality of resource scheduling events are triggered; and optimizing the initial scheduling strategy through a multi-target optimization model to obtain a target scheduling strategy of the resource scheduling event. By constructing the agent cluster, when the resource scheduling event is triggered, the initial scheduling strategy can be quickly generated through the resource space-time state diagram of the park resource agents, the initial scheduling strategy is optimized when the resource scheduling conflict occurs, and the resource scheduling efficiency is improved while the current resource scheduling event is met. And optimizing resource scheduling strategies of the current resource scheduling event and other resource scheduling events.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of resource scheduling, in particular to a park resource scheduling method, device and medium. BACKGROUND

[0002] In recent years, under the background of smart city construction, the park as an important carrier of economic activities, its resource scheduling efficiency directly affects the operating cost and user experience. The current park resource management has transitioned from manual mode to information stage, but the defects of traditional technical system in dynamic, real-time and multi-target coordination are increasingly prominent.

[0003] The mainstream solution of park resource scheduling relies on static rule engine (such as first come first served, fixed priority) to realize resource allocation, although relying on relational database and state machine logic to realize automation, but there are still problems of significant information update delay caused by rough environment perception granularity, lack of dynamic adjustment ability of strategy rigidity, and high proportion of manual intervention caused by inefficient conflict resolution. SUMMARY

[0004] In order to solve the above problems, the present application provides a park resource scheduling method, device and medium, wherein the method comprises: Based on the park resource information, a park resource agent cluster is constructed, the park resource includes charging pile resource, elevator resource, conference room resource and parking space resource; based on the preset rules of the park resource agent cluster and the park resource information, a resource space-time state diagram is constructed, the resource space-time state diagram contains the resource type, interaction weight and historical coordination success rate of the park resource agent cluster; based on the pre-constructed strategy coordination rule and the resource space-time state diagram, the initial scheduling strategy of the resource scheduling event is determined when multiple resource scheduling events are triggered; through a multi-objective optimization model, the initial scheduling strategy is optimized to obtain the target scheduling strategy of the resource scheduling event.

[0005] In one example, the event type of the resource scheduling event includes at least one of a strong trigger event, a prediction event, and a conflict event; and based on the pre-constructed policy coordination rule and the resource space-time state diagram, an initial scheduling policy of the resource scheduling event is determined when multiple resource scheduling events are triggered, specifically including: when the resource scheduling event is a strong trigger event, through a local model, a policy confidence of each pre-stored policy in a pre-stored policy pool is determined; a pre-stored policy with a policy confidence not lower than a preset threshold is taken as the initial scheduling policy; when all policy confidences are lower than the preset threshold, a latest global policy is requested from the cloud based on federated learning and taken as the initial scheduling policy; after outputting a scheduling instruction, a pre-stored policy pool is updated using a decision trajectory corresponding to the scheduling instruction; when the resource scheduling event is a prediction event, a resource demand prediction diagram within a future preset time period is generated through a prediction model; based on the resource demand prediction diagram, an initial scheduling policy is generated; when the resource scheduling event is a conflict event, a bidding strategy of each park resource agent is obtained, the bidding strategy including a bidding budget, a bidding time window demand, and a bidding resource specification; an initial equilibrium solution corresponding to the bidding strategy is learned through Nash-Q, and taken as the initial scheduling policy.

[0006] In one example, the initial scheduling policy is optimized through a multi-objective optimization model to obtain a target scheduling policy of the resource scheduling event, specifically including: based on a resource real-time state, a resource space-time state diagram, and the initial scheduling policy, a state space is defined; an action type in the initial scheduling policy is determined, and an action space is constructed based on the action type, the action type including discrete actions and continuous actions; based on a preset multi-objective reward function, the action space, and the state space, the initial scheduling policy is optimized to obtain the target scheduling policy of the resource scheduling event.

[0007] In one example, the initial scheduling policy is optimized based on a preset multi-objective reward function, an action space, and a state space, specifically including: the preset multi-objective reward function is Chebyshev scalar quantization decomposed to define a scalar quantization reward function; based on the action space, a double-channel policy network is constructed, the double-channel policy network including a discrete action branch and a continuous action branch; based on the scalar quantization reward function and the double-channel policy network, a multi-objective advantage function is determined; and based on the multi-objective advantage function, the initial scheduling policy is optimized.

[0008] In one example, after the resource space-time state diagram is constructed based on the preset rules of the park resource agent cluster and the park resource information, the method further includes: current state information of the park resource agent is obtained through a preset protocol at intervals of a preset time; and based on the current state information, the resource space-time state diagram is updated.

[0009] In an example, before determining the initial scheduling strategy of the resource scheduling event based on the pre-constructed policy coordination rule and the resource space-time state graph when the multiple resource scheduling event triggers, the method further comprises: receiving state information and demand information from a user agent; determining a condition that meets the multiple resource scheduling event triggers based on the state information, the demand information of the user agent and the resource space-time state graph.

[0010] In an example, after optimizing the initial scheduling strategy by the multi-objective optimization model to obtain the target scheduling strategy of the resource scheduling event, the method further comprises: monitoring that a deviation degree of a current user state of a user agent from the target scheduling strategy is higher than a preset deviation threshold; starting a breach weight calculation, and regenerating the initial scheduling strategy based on the current user state.

[0011] In an example, after optimizing the initial scheduling strategy by the multi-objective optimization model to obtain the target scheduling strategy of the resource scheduling event, the method further comprises: obtaining a topological association change of each park resource agent cluster in the park through a pre-deployed time sequence graph convolution network; and correcting model parameters in the multi-objective optimization model based on the topological association change.

[0012] The application also provides a park resource scheduling device, comprising: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform: constructing a park resource agent cluster based on park resource information, the park resource information comprising charging pile resource information, elevator resource information, conference room resource information and parking space resource information; constructing a resource space-time state graph based on preset rules of the park resource agent cluster and park resource information, the resource space-time state graph comprising resource types, interaction weights and historical coordination success rates of the park resource agent cluster; determining an initial scheduling strategy of a resource scheduling event based on a pre-constructed policy coordination rule and the resource space-time state graph when multiple resource scheduling event triggers; and optimizing the initial scheduling strategy by a multi-objective optimization model to obtain a target scheduling strategy of the resource scheduling event.

[0013] The application further provides a non-volatile computer storage medium, which stores computer executable instructions configured to: construct a park resource intelligent agent cluster based on park resource information, the park resource information including charging pile resource information, elevator resource information, conference room resource information and parking space resource information; construct a resource space-time state diagram based on preset rules of the park resource intelligent agent cluster and the park resource information, the resource space-time state diagram including resource types, interaction weights and historical coordination success rates of the park resource intelligent agent cluster; determine an initial scheduling strategy of a resource scheduling event when the resource scheduling event is triggered based on a pre-constructed strategy coordination rule and the resource space-time state diagram; and optimize the initial scheduling strategy through a multi-objective optimization model to obtain a target scheduling strategy of the resource scheduling event.

[0014] The method provided by the application can bring the following beneficial effects: by constructing the park resource intelligent agent cluster, the initial scheduling strategy can be quickly generated through the resource space-time state diagram of the park resource intelligent agent when the resource scheduling event is triggered, and the initial scheduling strategy can be optimized through the multi-objective optimization model when the resource scheduling conflict occurs, so that the resource scheduling strategy of the current resource scheduling event and other resource scheduling events can be optimized to the greatest extent while meeting the current resource scheduling event. BRIEF DESCRIPTION OF DRAWINGS

[0015] The accompanying drawings, which are included to provide a further understanding of the application, constitute a part of this application and help to explain the application together with the specification. The illustrative embodiments of the application and their description serve to explain the application. In the drawings: Figure 1 FIG. 1 is a flowchart of a park resource scheduling method according to an embodiment of the application; Figure 2 FIG. 2 is a structural diagram of a park resource scheduling device according to an embodiment of the application. DETAILED DESCRIPTION

[0016] To make the objectives, technical solutions and advantages of the application clearer, the following will describe the technical solutions of the application with reference to the embodiments of the application and the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the application, but not all the embodiments of the application. Based on the embodiments of the application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the application.

[0017] Under the background of smart city construction, as an important carrier of economic activities, the resource scheduling efficiency of the park directly affects the operating cost and user experience. The current park resource management has transitioned from manual mode to information stage, but the defects of traditional technical system in dynamic, real-time and multi-target coordination are increasingly prominent. The mainstream scheme relies on static rule engine (such as first come first served, fixed priority) to realize resource allocation, although it relies on relational database and state machine logic to realize automation, but it faces three core bottlenecks: First, the environmental perception granularity is rough, relying on discrete event triggering update (such as parking sensor signal), which cannot model continuous state (charging pile power fluctuation, human flow density gradient change), resulting in significant information update delay (such as the probability of parking space being occupied when the user arrives is 35% under the average 5-minute refresh cycle of parking space state).

[0018] Second, the strategy is rigid and lacks dynamic adjustment ability (such as pricing elasticity, dynamic priority weighting), for example, the conference room management, based on the calendar reservation system, cannot detect the actual use (such as no-show or early departure), and the resource idle rate is more than 40%.

[0019] Third, the conflict resolution is inefficient, and the proportion of manual intervention is more than 30%, with a median time of 15 minutes. To break through the limitations of rule system, some schemes introduce operations research optimization methods (such as integer programming, genetic algorithm), but are limited by real-time deficiency (genetic algorithm convergence time is more than 10 seconds in thousands of node scenarios), dimension disaster (state space is more than 10^6 when multiple resource coordination optimization), and missing physical constraints (simplified spatial relationship as Euclidean distance, ignoring elevator waiting time, fire passage no-parking area and other real heterogeneity), which is difficult to scale.

[0020] The technical solutions provided by the embodiments of the present application are described in detail below with reference to the drawings.

[0021] Figure 1 A flowchart of a park resource scheduling method provided by one or more embodiments of the present specification. The method can be applied to different park resource regulation in the park, such as parking space resources, charging pile resources related to parking space, conference room resources, and elevator resources, etc. The flow can be executed by a computing device in the corresponding field (such as an edge computing server deployed in the cloud or a park resource scheduling server arranged in the park, etc.), and some input parameters or intermediate results in the flow allow manual intervention to adjust to help improve accuracy.

[0022] The implementation of the analysis method related to the embodiments of the present application can be a terminal device or a server, and the present application does not make special limitations. For the convenience of understanding and description, the following embodiments are described in detail taking the server as an example.

[0023] It should be noted that the server can be a single device or a system composed of multiple devices, that is, a distributed server, and this application does not make any specific restrictions on this.

[0024] like Figure 1 As shown, the embodiment of the present application provides a campus resource scheduling method, including: S101: Construct a park resource intelligent body cluster based on park resource information.

[0025] First, the server obtains the current park resource information and builds a lightweight intelligent entity with autonomous decision-making capabilities for various types of resources in the park. The park resources here include charging pile resources, elevator resources, conference room resources, parking space resources and other controllable resources.

[0026] Take charging piles, elevators, and conference rooms as examples to illustrate: When building a park resource intelligent cluster, a state vector can be constructed for the charging piles. .in, is the charging pile state vector, For electricity, The appointment timestamp. is the real-time electricity price, is the coordinate of the charging pile.

[0027] Similarly, when constructing an elevator agent, a state vector can be constructed for the elevator ,in, The floor where the elevator currently stops. is the waiting queue length, is the historical average response time.

[0028] When building a conference room agent, you can build a state vector for the conference room ,in, The duration of time the conference room has been used. For the number of participants, Department authority level.

[0029] The park resource intelligent agent cluster refers to the cluster composed of all park resource intelligent agents in the park, for example, it includes 200 elevator intelligent agents, 100 charging pile intelligent agents, and 300 conference room intelligent agents in the park.

[0030] It should be noted that the above-mentioned intelligent agent construction method is only an example. When constructing an intelligent agent, the parameters in the state vector can be replaced. For example, the usage time of the conference room in the conference room intelligent agent can be replaced with the current usage status of the conference room. This application does not limit this.

[0031] S102: Construct a resource space-time state graph based on preset rules of the park resource agent cluster and park resource information.

[0032] After the park resource agent is constructed, a resource space-time state graph can be constructed based on preset rules of the park resource agent cluster and park resource information. The resource space-time state graph contains the resource type, interaction weight, and historical coordination success rate of the park resource agent cluster, and can be expressed as: wherein the node is a resource, the edge is an interaction weight, and is a quantized historical coordination success rate. The preset rules can include a space-time proximity rule combining Euclidean distance and vertical attenuation factor, an event causality rule quantifying causal strength based on transfer entropy, and a business logic constraint rule injected by artificial strategy. The above rules can be fused by linear regression to generate a dynamic adjacency matrix. Further, an incremental update (local weight recalculation) and periodic learning (causal model retraining) mechanism can be designed to realize low-cost real-time evolution of the rule base.

[0033] The above preset rules can be pre-stored in the storage device of the computer device. When it is necessary to construct the resource space-time state graph, the computer device can select the preset rules from the storage device. Of course, the computer device can also obtain the preset rules from other external devices. For example, the preset rules are stored in the cloud, and when it is necessary to construct the resource space-time state graph, the computer device can obtain the preset rules from the cloud. The embodiment does not limit the obtaining method of the preset rules.

[0034] S103: Determine an initial scheduling strategy of a resource scheduling event when multiple resource scheduling events are triggered based on the pre-constructed strategy coordination rule and the resource space-time state graph.

[0035] After the space-time state graph of the park resource agent is determined, when multiple resource scheduling events are triggered, an initial scheduling strategy can be generated based on the pre-constructed strategy coordination rule and the resource space-time state graph. The types of resource scheduling events include strong triggering events, prediction events, and conflict events.

[0036] For different types of resource scheduling events, the boundary and triggering threshold need to be clearly defined. For example, for device state mutation type request (strong triggering event), the triggering condition is sensor data jump (such as zero of charging gun current), user App operation signal (such as end of charging confirmation), and response delay requirement ≤200ms (such as charging pile plug-in, elevator emergency stop button triggering).

[0037] For resource conflicts that can be predicted based on historical rules (predicted events), the trigger condition is that the timing prediction confidence is greater than 75% (such as using prediction 15 minutes before the end of the meeting room) or the space congestion index breaks through the threshold (such as the number of queued vehicles at the entrance of the parking lot is greater than 5 for 3 minutes).

[0038] For multiple users initiating competitive requests for the same resource (conflict events), the trigger condition is that the number of resource lock failures is greater than or equal to 2 (such as consecutive appointment conflicts for a conference room) or the bid strategy difference is greater than 40% (such as a large budget gap between different departments).

[0039] When determining the type of resource scheduling event based on the obtained park data, the park data can be obtained through the fusion of Internet of Things device real-time flow (such as through the MQTT protocol), business system event log (such as through the Kafka message queue), and user behavior embedding (such as through App click flow) data input layer, and through the use of a lightweight XGBoost model (features include event source device type, timing burst index, and spatial density gradient), millisecond-level event type labeling can be completed on the edge computing node. For resource scheduling events that arrive at the same time, resource scheduling events can be weighted and sorted to determine the resource scheduling event that needs to be processed first.

[0040] At the same time, different types of resource scheduling events determine the initial scheduling strategy in different ways. Specifically, when the resource scheduling event is a strong trigger event, through the local model, the strategy confidence of each pre-stored strategy in the pre-stored strategy pool is determined, and the pre-stored strategy whose strategy confidence is not lower than the preset threshold is used as the initial scheduling strategy. When all strategy confidences are lower than the preset threshold, the latest global strategy is requested from the cloud based on federated learning and used as the initial scheduling strategy. After outputting the scheduling instruction, the pre-stored strategy pool is updated using the decision trajectory corresponding to the scheduling instruction.

[0041] For example, when the resource scheduling event is a charging pile insertion and extraction request, the server triggers the DDPG algorithm based on the improved priority experience replay (PER) at this time, thereby generating an initial scheduling instruction within a short time (such as 100 ms).

[0042] When the resource scheduling event is a prediction event, a resource demand prediction graph within a preset future time period is generated by a prediction model; and an initial scheduling strategy is generated based on the resource demand prediction graph. For example, when the trigger condition of the prediction event is met 15 minutes before the end of the meeting, the LSTM-Attention prediction model is started to generate a resource demand heat map (such as meeting room usage intensity and elevator peak floor) within 15-30 minutes in the future; and then pre-scheduling is performed based on the heat map to make the elevator empty to the target floor in advance and the parking lot guide sign display the reserved parking space number. When the prediction accuracy is less than 60% for three times in succession, online fine-tuning of the model is triggered, and the LSTM weight is updated through incremental learning, so as to improve the prediction accuracy of the model.

[0043] When the resource scheduling event is a conflict event, the bidding strategies of the resource agents of each park are obtained, which include bidding budget, bidding time window demand, and bidding resource specification, and finally the initial equilibrium solution corresponding to the bidding strategy is learned through Nash-Q to serve as the initial scheduling strategy. For example, each agent submits a bidding strategy (including budget, time window demand, and resource specification), and the validity is verified through a smart contract. Then, the initial equilibrium solution is calculated through the Nash-Q learning protocol, and if the user satisfaction score is less than 70 points, the NSGA-II multi-objective optimizer is enabled, and the final scheme is pushed through WeChat Enterprise / sms to provide alternative schemes and compensation points (such as “allocating a B conference room on the same floor, and giving 20 credit points”).

[0044] S104: The initial scheduling strategy is optimized through a multi-objective optimization model to obtain a target scheduling strategy of the resource scheduling event.

[0045] After obtaining the initial scheduling strategy, if a conflict event is involved, the initial scheduling strategy is optimized through a multi-objective optimization model to obtain a target scheduling strategy, so that the park resources are scheduled through the target scheduling strategy.

[0046] In one embodiment, when the initial scheduling strategy is optimized, the state space needs to be defined based on the real-time state of the resource, the resource space-time state graph, and the initial scheduling strategy, and then the action types in the initial scheduling strategy are determined, and the action space is constructed based on the action types, where the action types include discrete actions and continuous actions; the initial scheduling strategy is optimized based on a preset multi-objective reward function, the action space, and the state space to obtain a target scheduling strategy of the resource scheduling event.

[0047] In defining the state space, the resource space-time state graph of each park resource agent and the initial scheduling strategy need to be integrated. Specifically, the resource space-time state graph and the initial scheduling strategy are constructed into a hybrid state vector:

[0048] wherein, is a state space, is a real-time state of resources (e.g. charging pile occupancy, elevator waiting queue length), is a set of actionable actions output by the game layer (initial scheduling strategy), is an embedding feature of a GAT encoded spatio-temporal graph node.

[0049] The action space here is a discrete-continuous hybrid action space: discrete actions correspond to resource allocation mode selection (e.g. "VIP users are given priority service in charging piles" "time zoning and partitioning scheduling of elevators"); continuous actions correspond to parameter fine-tuning (e.g. charging power gradient adjustment, elevator stop time interval).

[0050] The multi-objective reward function is:

[0051] wherein is a reward function value, , , is a preset weight, is an efficiency reward function, which can be calculated based on resource utilization (e.g. number of vehicles served per hour by charging piles / theoretical maximum value); is a tool fairness reward function, which can be obtained by quantifying the difference in user waiting time using the Gini coefficient; is a cost reward function, which can be calculated by combining energy consumption (e.g. charging pile power integration) and equipment wear and tear (e.g. start-stop number penalty).

[0052] By calculating a variety of initial scheduling strategies through the state space, action space and multi-objective reward function, the optimal initial scheduling strategy among the variety of initial scheduling strategies can be determined, so as to take the initial scheduling strategy as the target scheduling strategy.

[0053] In one embodiment, when determining the target scheduling strategy, the preset multi-objective reward function can also be subjected to Chebyshev scalarization decomposition to define a scalarized reward function, and then based on the action space, a double-channel policy network is constructed, which includes a discrete action branch and a continuous action branch. At the same time, a multi-objective advantage function can be determined based on the scalarized reward function and the double-channel policy network; and the initial scheduling strategy is optimized based on the multi-objective advantage function.

[0054] Specifically, by a multi-objective proximal policy optimization algorithm, the multi-objective reward is projected to a preference space to define a scalarized reward:

[0055] wherein, is a scalarized reward value, actual performance indicators (e.g. elevator energy consumption, charging pile utilization rate) of the ideal Pareto frontier reference points (theoretical limits or expert experience settings), dynamic preference weights for reflecting the preference of different objectives in the optimization process (adjusted in real time through user feedback).

[0056] In defining the multi-objective advantage function, the multi-objective advantage function can be defined as:

[0057] wherein, is the comprehensive advantage function, measuring the net benefit of action in state relative to the global strategy ; is the dynamic weight coefficient for target priority control, which can be dynamically adjusted through game protocols or artificial rules (e.g. energy consumption weight 0.6 vs user waiting time weight 0.4); long-term benefit prediction of the th optimization target (e.g. energy saving benefit, user satisfaction), used to predict the cumulative benefit of the target after performing a certain action (e.g. the number of kilowatt-hours of total energy consumption reduced by charging pile power allocation); is the average baseline benefit of the th optimization target in state , used to establish a benefit benchmark: as a comparison benchmark (e.g. average waiting time under the current elevator dispatching strategy); is the single-target advantage difference, which can be used to reflect the improvement of the action on a single target (e.g. if the action improves fairness but is not beneficial to energy consumption, the difference is negative), is the component Q function of each target.

[0058] The above scalar reward and multi-objective advantage function can be used to solve the quantitative trade-off problem of multi-objective conflict. The following is a specific implementation scenario of charging pile power allocation, which has the following three objectives: Objective 1: Maximize overall charging efficiency (total power utilization rate); Objective 2: Ensure the charging experience of high-priority users (VIPs) (minimum power guarantee); Objective 3: Avoid overloading the power distribution network (current threshold constraint).

[0059] ​The key constraint at this time is the maximum deviation of target 2 (insufficient VIP current), which can be solved by preferentially adjusting the VIP charging pile power to 48A, while reducing the power of ordinary users to maintain the safety of the total load. During the calculation process, the Pareto optimal guidance is used to maximize the minimum satisfaction (Minimax criterion) to avoid system imbalance caused by excessive optimization of a single target. At the same time, the weight can be automatically adjusted according to the time period (such as increasing w2 during the evening peak and increasing w1 during the holiday). The system bottleneck target can be intuitively located by the Chebyshev radius value, and the interpretability is stronger.

[0060] In one embodiment, after constructing the resource space-time state diagram, it is necessary to update the resource space-time state diagram in real time. For example, the current state information of the park resource agent can be obtained through a preset protocol at a preset interval, so as to update the resource space-time state diagram based on the current state information. Wherein, each park resource agent can broadcast the local state to the edge gateway at a period of 50ms through the lightweight MQTT protocol.

[0061] In one embodiment, based on the pre-constructed policy coordination rule and the resource space-time state diagram, before determining the initial scheduling strategy of the resource scheduling event when a plurality of resource scheduling events are triggered, it is necessary to determine whether the trigger condition of the resource scheduling event is met. At this time, the state information and demand information from the user agent need to be received, and based on the state information, demand information of the user agent and the resource space-time state diagram, it is determined that the trigger condition of the plurality of resource scheduling events is met.

[0062] In one embodiment, when the state of the user agent does not respond to the target scheduling strategy as pre-agreed, the corresponding target scheduling strategy needs to be regenerated based on the state of the user agent. Specifically, when it is monitored that the current user state of the user agent deviates from the target scheduling strategy by more than a preset deviation threshold, the breach weight calculation is started, and the initial scheduling strategy is regenerated based on the current user state.

[0063] In one embodiment, after generating the target scheduling strategy, the topological association relationship between each agent can be monitored, and the topological association change of each park resource agent cluster in the park can be obtained through the pre-deployed time series graph convolution network. Then, based on the topological association change, the model parameters in the multi-objective optimization model are corrected.

[0064] The strategy of the park resource agent is described below by abstracting the above process into a game coordination layer, a reinforcement learning optimization layer, and a space-time correction layer: In the process of dynamic scheduling of charging piles and user path coordination, the game coordination layer generates strategies based on the dynamic game between the charging pile Agent cluster and the user Agent based on asymmetric information. The pile end Agent publishes real-time state (remaining power, queuing time, price gradient), and the user Agent uploads mobile trajectory (LBS positioning) and demand preference (fast charging / slow charging). Both sides generate an initial matching strategy through iterative Q-learning, such as triggering the breach weight calculation when the user deviates from the reservation path, and the pile end Agent starts the secondary auction mechanism. The reinforcement learning optimization layer builds a multi-objective deep deterministic policy gradient (MADDPG) model, taking grid load balancing (variance minimization), user waiting time (FIFO correction weighting), and operator revenue (dynamic pricing elasticity) as joint optimization objectives. Through the experience replay pool, historical load curve data of the power grid is injected to train the strategy network, and the Pareto optimal power distribution scheme is output. The space-time correction layer will deploy a time series graph convolution network (T-GCN) to dynamically perceive the topological association changes of charging piles-parking lots-elevators in the region (such as heavy rain causing basement congestion), real-time correct the availability weight coefficient of the pile position, and prioritize emergency vehicle charging rights, and push navigation information of nearby idle piles to users who deviate from the path (V2X vehicle-road cooperation data interface triggers).

[0065] In the process of elevator group control and cross-floor resource linkage, the game coordination layer generates strategies based on the distributed incomplete information game model constructed by the elevator Agent cluster. Each elevator independently decides the stop floor and start-stop timing, competes for high-priority tasks (such as VIP user direct, medical material emergency transportation) through a virtual bidding mechanism, and allocates revenue weights based on Shapley value to balance load differences. The reinforcement learning optimization layer will use the multi-agent proximal policy optimization (MAPPO) algorithm to build a three-dimensional reward function with energy consumption (motor start-stop times x load power consumption), response time (user waiting time standard deviation), and equipment wear (bearing stress accumulation) as optimization objectives. Curriculum learning strategy is used to gradually expand from single elevator scheduling to multi-elevator coordination, avoiding the shock effect caused by "order grabbing conflicts". The space-time correction layer will use a dynamic graph attention network (DGAT) to model the elevator shaft topology (distance between adjacent elevators, floor flow density heat map) in real time. When a certain elevator fails or a sudden peak of people is detected, the space-time constraints are automatically reconstructed (such as limiting non-emergency task cross-zone scheduling), triggering the reinforcement learning model to fine-tune the strategy network parameters online, ensuring system disaster recovery capability.

[0066] In the intelligent allocation of conference rooms and dynamic priority adjustment process, the game coordination layer generates strategies through multiple rounds of sealed bid games between conference room agents and department agents. The conference room publishes available time periods, device configurations (projector / teleconference), and historical usage rates. Department agents submit the number of attendees, security levels, and cross-department collaboration needs. Based on the VCG mechanism, an incentive-compatible initial allocation is achieved. The reinforcement learning optimization layer designs a hierarchical reinforcement learning (HRL) framework. The high-level strategy network decides the conference room allocation priority (calculates the weight of the emergency meeting queue), and the bottom-level strategy network optimizes the device pre-start timing (the air conditioner starts early to save energy). Through adversarial training, it generates adversarial samples (simulates malicious occupation scenarios) to improve the robustness of the strategy. The space-time correction layer constructs a space-time propagation model (such as the message passing neural network MPNN). When a conference room is temporarily requisitioned (such as an emergency meeting for senior management), it automatically calculates the migration cost of affected meetings (the degree of overlap of participant movement paths and the time-consuming device reset), generates a minimum disturbance adjustment scheme, and synchronously pushes the change reminder and alternative conference room navigation path through the WeChat API.

[0067] As shown in Figure 2 The embodiments of the present application also provide a park resource scheduling device, which comprises at least one processor and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: construct a park resource agent cluster based on park resource information, the park resource information including charging pile resources, elevator resources, conference room resources, and parking space resources; construct a resource space-time state diagram based on preset rules of the park resource agent cluster and the park resource information, the resource space-time state diagram containing resource types, interaction weights, and historical collaboration success rates of the park resource agent cluster; determine an initial scheduling strategy of a resource scheduling event when a plurality of resource scheduling events are triggered based on a pre-constructed strategy collaboration rule and the resource space-time state diagram; and optimize the initial scheduling strategy through a multi-objective optimization model to obtain a target scheduling strategy of the resource scheduling event.

[0068] The embodiments of the present application also provide a non-volatile computer storage medium, which stores computer executable instructions, and the computer executable instructions are configured to: Construct a park resource intelligent agent cluster based on park resource information, the park resource including charging pile resource, elevator resource, conference room resource and parking space resource; construct a resource space-time state diagram based on preset rules of the park resource intelligent agent cluster and park resource information, the resource space-time state diagram containing resource type, interaction weight and historical coordination success rate of the park resource intelligent agent cluster; determine an initial scheduling strategy of a resource scheduling event when the resource scheduling event is triggered based on a pre-constructed strategy coordination rule and the resource space-time state diagram; optimize the initial scheduling strategy through a multi-objective optimization model to obtain a target scheduling strategy of the resource scheduling event.

[0069] Each of the embodiments in the present application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment mainly describes the difference from other embodiments. Especially, the device and medium embodiments are described simply because they are basically similar to the method embodiments, and the related parts can be referred to the part of the method embodiments.

[0070] The device and medium provided by the embodiments of the present application are one-to-one corresponding to the method, and therefore, the device and medium also have the similar beneficial technical effects as the method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device and medium will not be described here.

[0071] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0072] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices produce a device implemented in the flowcharts and / or block diagrams. Figure 1 The device that implements the function specified in one flow or multiple flows and / or blocks Figure 1 The device that implements the function specified in one flow or multiple flows and / or blocks

[0073] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.

[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.

[0075] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0076] The memory can include non-persistent memory and / or volatile memory, such as a random access memory (RAM) including a cache area for the temporary storage of data. The memory can also include non-volatile memory, such as a read only memory (ROM), EPROM, EEPROM, or flash memory. The memory can be another form of computer-readable media.

[0077] Computer-readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.

[0078] It should also be noted that the terms "comprising," "including," or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0079] The above description is merely illustrative of the application, and not restrictive. Various modifications and changes can become apparent to those skilled in the art. Incorporating any modification, equivalent substitution, improvement, etc. within the spirit and principle of the application, shall be included in the scope of the claims of the application.

Claims

1. A park resource scheduling method, characterized in that: include: Building a cluster of park resource agents based on park resource information, where the park resources include charging pile resources, elevator resources, conference room resources, and parking space resources; Based on the preset rules of the park resource agent cluster and the park resource information, a resource spatiotemporal state diagram is constructed, wherein the resource spatiotemporal state diagram includes the resource type, interaction weight and historical collaboration success rate of the park resource agent cluster; Based on the pre-built policy coordination rules and the resource spatiotemporal state diagram, determining the initial scheduling policy of the resource scheduling event when multiple resource scheduling events are triggered; The initial scheduling strategy is optimized through a multi-objective optimization model to obtain a target scheduling strategy for the resource scheduling event.

2. The method according to claim 1, characterized in that The event type of the resource scheduling event includes at least one of a strong trigger event, a prediction event, and a conflict event; The method of determining the initial scheduling strategy for a resource scheduling event based on the pre-built strategy coordination rules and the resource spatiotemporal state diagram when multiple resource scheduling events are triggered specifically includes: When the resource scheduling event is a strong trigger event, the policy confidence of each pre-stored policy in the pre-stored policy pool is determined through the local model; The pre-stored strategy whose strategy confidence is not lower than the preset threshold is used as the initial scheduling strategy; When the confidence of all strategies is lower than the preset threshold, the latest global strategy is synchronously requested from the cloud based on federated learning as the initial scheduling strategy; After outputting the scheduling instruction, the decision trajectory corresponding to the scheduling instruction is used to update the pre-stored strategy pool; When the resource scheduling event is a forecast event, a forecast graph of resource demand within a preset time period in the future is generated through a forecast model; generating an initial scheduling strategy based on the resource demand forecast graph; When the resource scheduling event is a conflict event, obtaining the bidding strategy of each park resource agent, wherein the bidding strategy includes a bidding budget, a bidding time window requirement, and a bidding resource specification; An initial equilibrium solution corresponding to the bidding strategy is learned through Nash-Q and used as the initial scheduling strategy.

3. The method according to claim 1, characterized in that The initial scheduling strategy is optimized by a multi-objective optimization model to obtain a target scheduling strategy for the resource scheduling event, specifically including: Defining a state space based on the real-time state of the resources, the spatiotemporal state diagram of the resources, and the initial scheduling strategy; Determining action types in the initial scheduling strategy and constructing an action space based on the action types, wherein the action types include discrete actions and continuous actions; Based on a preset multi-objective reward function, the action space, and the state space, the initial scheduling strategy is optimized to obtain a target scheduling strategy for the resource scheduling event.

4. The method according to claim 3, characterized in that The optimizing the initial scheduling strategy based on the preset multi-objective reward function, the action space, and the state space specifically includes: Performing Chebyshev scalar decomposition on the preset multi-objective reward function to define a scalar reward function; Based on the action space, constructing a dual-channel strategy network, wherein the dual-channel strategy network includes a discrete action branch and a continuous action branch; Determining a multi-objective advantage function based on the scalar reward function and the dual-channel policy network; The initial scheduling strategy is optimized based on the multi-objective advantage function.

5. The method according to claim 1, wherein After constructing the resource spatiotemporal state diagram based on the preset rules of the park resource agent cluster and the park resource information, the method further includes: Obtaining the current status information of the park resource agent through a preset protocol at preset intervals; Based on the current status information, the resource spatiotemporal status diagram is updated.

6. The method according to claim 1, characterized in that Before determining the initial scheduling strategy for a resource scheduling event when multiple resource scheduling events are triggered based on the pre-built strategy coordination rules and the resource spatiotemporal state diagram, the method further includes: Receive status information and demand information from user agents; Based on the state information of the user agent, the demand information and the resource spatiotemporal state diagram, it is determined that the trigger conditions of the multiple resource scheduling events are met.

7. The method according to claim 1, characterized in that After optimizing the initial scheduling strategy through the multi-objective optimization model to obtain the target scheduling strategy for the resource scheduling event, the method further includes: It is detected that the deviation between the current user state of the user agent and the target scheduling strategy is higher than a preset deviation threshold; Default weight calculation is initiated, and the initial scheduling policy is regenerated based on the current user state.

8. The method according to claim 1, characterized in that After optimizing the initial scheduling strategy through the multi-objective optimization model to obtain the target scheduling strategy for the resource scheduling event, the method further includes: The topological correlation changes of each cluster of resource agents in the park are obtained through the pre-deployed temporal graph convolutional network; Based on the topological association changes, the model parameters in the multi-objective optimization model are modified.

9. A park resource scheduling device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, wherein the instructions are executed by the at least one processor to enable the at least one processor to perform: Building a cluster of park resource agents based on park resource information, where the park resources include charging pile resources, elevator resources, conference room resources, and parking space resources; Based on the preset rules of the park resource agent cluster and the park resource information, a resource spatiotemporal state diagram is constructed, wherein the resource spatiotemporal state diagram includes the resource type, interaction weight and historical collaboration success rate of the park resource agent cluster; Based on the pre-built policy coordination rules and the resource spatiotemporal state diagram, determining the initial scheduling policy of the resource scheduling event when multiple resource scheduling events are triggered; The initial scheduling strategy is optimized through a multi-objective optimization model to obtain a target scheduling strategy for the resource scheduling event.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Building a cluster of park resource agents based on park resource information, where the park resources include charging pile resources, elevator resources, conference room resources, and parking space resources; Based on the preset rules of the park resource agent cluster and the park resource information, a resource spatiotemporal state diagram is constructed, wherein the resource spatiotemporal state diagram includes the resource type, interaction weight and historical collaboration success rate of the park resource agent cluster; Based on the pre-built policy coordination rules and the resource spatiotemporal state diagram, determining the initial scheduling policy of the resource scheduling event when multiple resource scheduling events are triggered; The initial scheduling strategy is optimized through a multi-objective optimization model to obtain a target scheduling strategy for the resource scheduling event.

Citation Information

Cited By

  • Office dispatching system based on behavior recognition

    CN121303757A

  • An office scheduling system based on behavior recognition

    CN121303757B