Elevator group control dynamic scheduling method and system based on reinforcement learning, electronic equipment and storage medium
By using a reinforcement learning-based dynamic scheduling method for elevator group control, and leveraging the learning of spatiotemporal patterns by intelligent agents to automatically generate scheduling strategies, the spatiotemporal adaptability and multi-objective optimization problems of elevator group control in existing technologies are solved, thereby achieving efficient utilization of elevator resources and reduced energy consumption.
Patent Information
- Application Number
- CN202511145106.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-18
AI Technical Summary
Existing elevator group control technology lacks spatiotemporal adaptability, has high manual parameter adjustment costs, a disconnect between prediction and response capabilities, and an imbalance in multi-objective optimization, making it difficult to meet the needs of efficient, flexible, and low-cost scheduling for complex passenger flow patterns in smart buildings.
An elevator group control dynamic scheduling method based on reinforcement learning is adopted. An intelligent closed loop is constructed through data collection, spatiotemporal modeling and dynamic strategy generation. The intelligent agent learns the spatiotemporal patterns of elevator ride data, automatically generates and dynamically updates the scheduling strategy, and iteratively optimizes the scheduling strategy by combining the reward function.
It achieves efficient utilization of elevator resources and improves usage efficiency, reduces system operation and maintenance difficulty and cost, ensures optimal system performance in the long term, and reduces energy consumption and waiting time.
Smart Images

Figure CN120964536A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of artificial intelligence, in particular to an elevator group control dynamic scheduling method and system based on reinforcement learning, an electronic device and a storage medium. BACKGROUND
[0002] The elevator resource scheduling process is to optimize the response strategy of multiple elevators through real-time algorithms to minimize passenger waiting time and energy consumption. The core steps include: demand perception: capturing passenger requests through floor buttons (up / down) and in-car destination signals; dynamic allocation: based on the current position, running direction, load status and predicted traffic flow of the elevator, such as morning and evening peak, using scheduling algorithms such as SCAN, LOOK or reinforcement learning-based strategies to allocate new requests to the optimal elevator; path planning: adjusting the elevator stop sequence, merging requests in the same direction, and avoiding invalid stops such as "piggyback" strategy; real-time adjustment: recalculating priorities in the event of sudden demand, such as emergency calls or traffic flow changes, to ensure global efficiency. This process relies on multi-objective optimization (speed, energy consumption, fairness) and intelligent scheduling through centralized controllers or distributed collaboration.
[0003] The elevator group control system is the core transportation hub of modern high-rise buildings, and its scheduling efficiency directly affects passenger experience and building energy consumption. Existing elevator group control technology mainly relies on preset rules, such as shortest waiting time priority or simple time period division strategies, such as peak / flat peak mode, which has the following inherent defects: static rules lack spatial and temporal adaptability: fixed algorithms cannot perceive the dynamic changes in building passenger flow space-time rules, resulting in mismatch between scheduling strategies and real demand; high cost of manual parameter adjustment: engineers need to manually adjust parameters based on experience, which not only lags in response, but also is difficult to cover complex building scenarios; prediction and response capabilities are fragmented: traditional methods only respond to real-time requests and cannot predict future passenger flow trends based on historical data, missing the pre-scheduling optimization window; multi-objective optimization imbalance: rule engines are difficult to optimize conflicting objectives such as waiting time, energy consumption, and elevator balanced wear, easily falling into local optimum, such as shortening waiting time but increasing energy consumption by 30%. Especially with the development of smart buildings, the diversification of building functions leads to increasingly complex passenger flow patterns, such as cross-floor travel, holiday peak fluctuations, and existing technologies cannot meet the scheduling needs of efficiency, flexibility and low consumption. SUMMARY
[0004] The purpose of the present application is to provide an elevator group control dynamic scheduling method and system based on reinforcement learning, an electronic device and a storage medium, which realizes the autonomous evolution of dispatching logic with building passenger flow rules through the construction of an intelligent closed loop of "data collection → space-time modeling → dynamic strategy generation → continuous evolution".
[0005] An elevator group control dynamic scheduling method based on reinforcement learning, comprising:
[0006] S100: obtaining real-time elevator demand data;
[0007] S200: obtaining a time-space feature of a lift-taking behavior according to the real-time elevator demand data;
[0008] S300: generating an elevator resource dynamic scheduling strategy according to the time-space feature of the lift-taking behavior, comprising:
[0009] S310A: when the lift-taking demand meets a periodic feature, deploying a prepared scheduling strategy for elevator resource planning;
[0010] S310B: when the lift-taking demand does not meet the periodic feature, an elevator resource scheduling agent calculates a dynamic scheduling strategy in real time for elevator resource planning.
[0011] Preferably, before the real-time elevator demand data is obtained, the method further comprises:
[0012] S010: obtaining prior lift-taking data;
[0013] S020: calculating a prepared elevator resource scheduling strategy according to the prior lift-taking data.
[0014] Preferably, the real-time elevator demand data is obtained by:
[0015] S110: obtaining a timestamp, a lift request source floor and a lift request target floor of a real-time elevator request;
[0016] S120: dividing the obtained real-time data into time data and space data.
[0017] Preferably, the time-space feature of the lift-taking behavior is obtained according to the real-time elevator demand data by:
[0018] S210: obtaining a passenger flow direction matrix between floors in different time periods according to the time data and the space data;
[0019] S220: identifying a passenger flow path of a peak period and a valley period according to the passenger flow direction matrix in the different time periods;
[0020] S230: extracting a time-space feature of a lift-taking behavior according to the passenger flow path of the peak period and the valley period.
[0021] Preferably, the passenger flow direction matrix between floors in different time periods is obtained according to the time data and the space data by:
[0022] defining a time slice, dividing a day into T time periods;
[0023] generating an N×N matrix M for a time period t t, N is the number of floors, B is the starting floor, and L is the destination floor:
[0024]
[0025] Preferably, the calculation of the deployed elevator resource scheduling strategy based on prior elevator data includes:
[0026] S021: inputting the prior elevator data as a training set into a reinforcement learning-based agent model;
[0027] S022: training and optimizing the reinforcement learning-based agent model;
[0028] S023: the trained reinforcement learning-based agent model outputs a deployed elevator resource scheduling strategy.
[0029] Preferably, after generating the elevator resource dynamic scheduling strategy based on the spatiotemporal characteristics of the elevator behavior, the method further includes iteratively optimizing the elevator resource dynamic scheduling strategy through a reward function feedback mechanism:
[0030] The reward function is set as follows:
[0031] R = ω1R wait + ω2R energy + ω3R balance + ω4R emergency
[0032] where ω1, ω2, ω3, and ω4 are adjustable weights, R wait is the waiting time penalty, R energy is the energy consumption penalty, R balance is the load balancing reward, R emergency is the emergency demand reward.
[0033] A reinforcement learning-based elevator group control dynamic scheduling system, comprising:
[0034] a data acquisition module for acquiring real-time elevator demand data;
[0035] a feature extraction module for obtaining spatiotemporal characteristics of elevator behavior based on the real-time elevator demand data;
[0036] a strategy deployment module for generating an elevator resource dynamic scheduling strategy based on the spatiotemporal characteristics of elevator behavior, including:
[0037] when the elevator demand meets the periodic characteristics, a deployed scheduling strategy is used for elevator resource planning;
[0038] when the elevator demand does not meet the periodic characteristics, an elevator resource scheduling agent calculates a dynamic scheduling strategy in real time for elevator resource planning.
[0039] An electronic device, comprising: a chip, a processor and a memory, the memory being used to store computer program code, the computer program code comprising computer instructions, under the execution of the computer instructions by the chip, the electronic device executes a dynamic scheduling method of elevator group control based on reinforcement learning.
[0040] A computer readable storage medium, the computer readable storage medium storing a computer program, the computer program comprising program instructions, under the execution of the program instructions by a processor of an electronic device, the processor executes a dynamic scheduling method of elevator group control based on reinforcement learning.
[0041] The beneficial effects of the present application are that: the scheduling agent of the present application automatically generates and dynamically updates the scheduling strategy by autonomously learning the space-time law of elevator data, such as uplink traffic in the morning peak and dispersed travel across floors during the lunch break; without defining complex rules manually, the problem of relying on artificial experience is completely solved, and the difficulty and cost of system operation and maintenance are greatly reduced; the scheduling agent of the present application identifies subtle space-time laws by analyzing high-dimensional data such as timestamps and source / destination floors, generates space-time variable strategies, accurately models and dynamically responds to high-dimensional space-time passenger flow laws in buildings, and significantly improves scheduling efficiency; the system of the present application continuously iteratively optimizes the scheduling strategy through real-time data closed loop (collection-> training-> deployment), the agent automatically detects changes in passenger flow patterns, and actively triggers model retraining to ensure that the system maintains optimal performance for a long time and prolongs the technical life cycle; the multi-objective collaborative optimization mechanism of the present application can reduce the comprehensive operation cost of the system while ensuring the elevator experience, and reasonable scheduling arrangement can reduce the empty running of the elevator and reduce energy consumption. BRIEF DESCRIPTION OF DRAWINGS
[0042] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present application and, together with the specification, serve to explain the principles of the application.
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows, and obviously, other drawings can also be obtained by those skilled in the art without creative labor.
[0044] Figure 1 A flow chart of a dynamic scheduling method of elevator group control based on reinforcement learning of the present application;
[0045] Figure 2 A system architecture diagram of a dynamic scheduling system of elevator group control based on reinforcement learning of the present application;
[0046] Figure 3A hardware structure schematic diagram of an electronic device according to the present application. DETAILED DESCRIPTION
[0047] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort fall within the protection scope of the present application.
[0048] It should be noted that all directionality indications (such as up, down, left, right, front, back, and the like) in the embodiments of the present application are only used to explain the relative position relationship, movement condition, and the like between components in a certain specific posture (as shown in the drawings), and if the specific posture changes, the directionality indications also change accordingly.
[0049] In addition, the descriptions of “first”, “second”, and the like in the present application are only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features defined with “first”, “second” can explicitly or implicitly include at least one of the features. In addition, the technical solutions of each embodiment can be combined with each other, but it must be based on the fact that a person of ordinary skill in the art can realize it, and when the combination of technical solutions contradicts each other or cannot be realized, it should be considered that the combination of technical solutions does not exist and is not within the protection scope required by the present application.
[0050] The existing elevator group control technology mainly relies on preset rules, such as shortest waiting time priority or simple time period division strategy, such as peak / flat peak mode, and has the following inherent defects: Static rules lack space-time adaptability: Fixed algorithms cannot perceive the dynamic changing passenger flow space-time law in the building, resulting in mismatch between scheduling strategy and real demand; Artificial parameter adjustment is costly: Engineers need to manually adjust parameters according to experience, which not only has a lagging response, but also is difficult to cover complex building scenarios; Prediction and response ability is split: Traditional methods only respond to real-time requests and cannot predict future passenger flow trends based on historical data, missing the pre-scheduling optimization window; Multi-objective optimization is unbalanced: Rule engines are difficult to cooperatively optimize conflicting goals such as waiting time, energy consumption, and elevator balanced loss, and are prone to local optimization, such as shortening the waiting time but increasing the energy consumption by 30%. Especially with the development of smart buildings, the functional diversification of buildings leads to increasingly complex passenger flow patterns, such as cross-floor joint travel and holiday peak fluctuations, and existing technologies cannot meet the scheduling needs of high efficiency, flexibility, and low consumption.
[0051] To solve the above problems, the application provides an elevator group control adaptive scheduling system based on agent learning, which is used to execute an elevator group control dynamic scheduling method based on reinforcement learning.
[0052] Please refer to Figure 1 and Figure 2 The elevator group control adaptive scheduling system based on agent learning provided by the application includes a data acquisition module, a feature extraction module and a strategy deployment module, which is used to execute the conversion of elevator data into space-time features, and to construct an elevator resource scheduling agent to generate an elevator resource dynamic scheduling strategy according to the space-time features of the elevator behavior. The elevator resource dynamic scheduling strategy can be customized according to the different space-time features of each building.
[0053] In specific implementation, before acquiring elevator demand data, it further includes:
[0054] S010: Acquire prior elevator data.
[0055] S020: Calculate the deployed elevator resource scheduling strategy according to the prior elevator data.
[0056] S021: Input the prior elevator data into the elevator resource scheduling agent as a training set;
[0057] Train the prior data: pre-train the agent using historical elevator data, accelerate the convergence of online iterative optimization: continuously update the strategy through new data to realize adaptive scheduling.
[0058] S022: Train and optimize the agent model based on reinforcement learning.
[0059] Iterative optimization of the elevator resource scheduling agent parameters through a reward function feedback mechanism includes:
[0060] Set the reward function, specifically:
[0061] R = ω1R wait + ω2R energy + ω3R balance + ω4R emergency
[0062] Where ω1, ω2, ω3 and ω4 are adjustable weights, R wait is the waiting time penalty, R energy is the energy consumption penalty, R balance is the load balancing reward, R emergency is the emergency demand reward.
[0063] The waiting time penalty is a linear penalty within a certain waiting time, and the penalty sharply rises after the timeout to avoid extremely long waiting.
[0064]
[0065] wherein N: the number of pending requests in the current period, Tk: the cumulative waiting time of the kth request, a: the basic waiting penalty coefficient, β: the long waiting penalty intensity, τ: the tolerance threshold.
[0066] The energy consumption penalty is a penalty for both running distance (positively correlated with energy consumption) and frequent start-stop (mechanical loss).
[0067]
[0068] wherein E: the total number of elevators, De: the running distance of elevator e, We: the load weight of elevator e, Se: the start-stop times of elevator e, γ, C: the energy consumption coefficient.
[0069] Load balancing reward, the more balanced the load, the higher the reward, to prevent some elevators from being overloaded or idle.
[0070]
[0071] wherein Le: the current load rate of elevator e (0-1), σ: the standard deviation of the load rate, μ: the mean of the load rate, ε: the smoothing factor, δ: the reward intensity.
[0072] Emergency demand reward, such as peak, emergency, and other special scenarios.
[0073]
[0074] wherein Kemerg: the set of emergency / priority demands, τ emerg : the maximum tolerance time for emergency demand, [X] + = max(x, 0): ensure non-negative, η: priority coefficient.
[0075] In the embodiments of the present application, the reinforcement learning-based agent model can also be trained and optimized multiple times in use. Before being deployed to the elevator group control system, the reinforcement learning-based agent model has been trained according to the current building environment and has a certain elevator resource scheduling capability. In the subsequent use process, the reinforcement learning-based agent model can be trained through the real-time data of the elevator obtained continuously, so as to better adapt to the changes in the use environment. The training time can be selected at leisure. If it is an office building, the training can be performed at weekends. If it is a mall, the training can be performed on weekdays or closed days. The reinforcement learning-based agent model can continuously update the deployed elevator resource dynamic scheduling strategy to adapt to the changing use environment.
[0076] S023: The trained reinforcement learning-based intelligent agent model outputs a deployed elevator resource scheduling strategy.
[0077] In specific implementation, the elevator resource scheduling intelligent agent can be used as the trained reinforcement learning-based intelligent agent model to formulate an elevator resource scheduling strategy that can meet most of the elevator demand according to the prior elevator demand of the current building. With the deployed elevator resource scheduling strategy, the system can respond very quickly in most cases, improving the efficiency of elevator scheduling.
[0078] In specific implementation, scenario one: a high-end office building, the scene characteristics are as follows: 8:00-10:00 on weekdays is the uplink peak, 12:00-14:00 is the scattered cross-layer flow, and 18:00-19:00 is the downlink peak.
[0079] Deploy the elevator resource scheduling intelligent agent in the office building group control system.
[0080] Collect and record 4 weeks of data continuously, including:
[0081] Early peak: 1-5 floors to 20-30 floors concentrated uplink request (70%).
[0082] Lunch: Random requests between floors (restaurants / meeting rooms).
[0083] Based on the collected and processed elevator demand data, extract the spatio-temporal feature data and generate a "period-floor flow heat map", and extract the "midday cross-layer request dispersion" index.
[0084] Using the DDQN algorithm, the average waiting time and energy saving rate are used as the reward function performance index design, and the elevator resource deployment scheme is obtained after training:
[0085] Early peak: Assign an empty elevator to stand in the low layer in advance, skip the middle layer response and directly reach the target high area.
[0086] Lunch: Enable "regional circulation" mode, single elevator serving adjacent 5 floors.
[0087] The elevator group control system applies the dynamic scheduling strategy generated by the scheduling intelligent agent to control the elevator main control system to execute elevator scheduling, such as Monday morning 8:30, a new request "1→25F" triggers the strategy, the system assigns A elevator that is waiting on 1 floor to directly reach (the traditional system needs to wait for B elevator to downlink from 15 floor).
[0088] Automatic retraining every weekend.
[0089] Scenario two:
[0090] A certain general hospital, the scene characteristics are as follows: outpatient floor (2-4 layers) all day random peak, inpatient department (8-12 layers) timing meal delivery / inspection requirements, emergency elevator.
[0091] The elevator resource scheduling intelligent agent accesses the hospital elevator group control system.
[0092] Continuous collection records emergency elevator data (sudden, high priority), statistics meal delivery period (11:00 / 17:00 inpatient department down peak).
[0093] Based on the collected and processed elevator demand data, the spatio-temporal feature data is extracted, the "emergency response delay" and the elevator position correlation are identified, and the "meal delivery period inpatient department down demand prediction model" is constructed.
[0094] Using PPO reinforcement learning, taking emergency response time and ordinary passenger satisfaction as reward function performance index design, the elevator resource deployment scheme is obtained after training:
[0095] Normal: reserve 1 elevator on 1 floor, cover emergency.
[0096] 10 minutes before meal delivery: schedule elevator to inpatient department high floor standby in advance.
[0097] Scenario three:
[0098] A certain shopping center, the scene characteristics are as follows: weekend family passenger flow surge, high floor catering / cinema concentrated down demand, sparse passenger flow in the afternoon on weekdays.
[0099] Deploy elevator scheduling intelligent agent in shopping mall elevator group control system.
[0100] Compare weekday / weekend data: weekend cinema dispersal period (21:00-22:00) 7-9F→1 layer request quantity increased by 300%.
[0101] Based on the collected and processed elevator demand data, the spatio-temporal feature data is extracted, the "down request aggregation index" is calculated, and the "cross-layer joint demand" is extracted.
[0102] Using Actor-Critic framework, taking long waiting rate and joint matching degree as reward function performance index design, the elevator resource deployment scheme is obtained after training:
[0103] Weekend down peak: enable "funnel scheduling" - high floor elevator only down, middle zone elevator relay transfer.
[0104] Detect joint request: if the passenger presses "4F→7F", assign the same elevator service, reduce transfer waiting.
[0105] After the elevator resource scheduling intelligent agent is constructed and trained, the trained elevator resource scheduling intelligent agent is deployed in the elevator control system in the application scene, and the trained elevator resource scheduling intelligent agent outputs an elevator resource scheduling strategy according to the received elevator demand data.
[0106] In specific implementation, the data acquisition module is configured to perform step S100 and acquire real-time elevator demand data.
[0107] The data acquisition module includes a data extraction unit and a data classification unit.
[0108] The data extraction unit is configured to perform step S110 and acquire the timestamp of the real-time elevator request, the source floor of the elevator request, and the target floor of the elevator request.
[0109] In specific implementation, the timestamp of the elevator request, the source floor of the elevator request, and the target floor of the elevator request can be acquired through a camera, an elevator button, or the like. The elevator demand data that can be acquired can also include one or more of the following: a call direction, a call type (car call or landing call), passenger quantity information, and current state information of the elevator car (including position, running direction, load, and door state).
[0110] The data classification unit is configured to perform step S120 and classify the acquired real-time data into time data and space data.
[0111] The time data includes the timestamp of the elevator request, and the space data includes the source floor of the elevator request and the target floor of the elevator request.
[0112] The feature extraction module is configured to perform step S200 and acquire the time-space features of the elevator behavior according to the real-time elevator demand data.
[0113] The data processing module includes a flow direction matrix construction unit, a passenger flow path construction unit, and a time-space feature extraction unit.
[0114] The matrix construction unit is configured to perform step S210 and acquire the passenger flow direction matrix between floors in different time periods according to the time data and the space data.
[0115] In specific implementation, the elevator scheduling intelligent agent acquires the passenger flow direction matrix between floors in different time periods according to the elevator demand data.
[0116] Specific implementation steps include: defining a time slice: dividing a day into T time periods, such as one time period every 15 minutes.
[0117] Constructing a passenger flow matrix: for time period t, generating an N×N matrix M t (where N is the number of floors, the starting floor of the behavior is the row, and the destination floor is the column):
[0118]
[0119] After the passenger flow matrix is constructed, normalization processing can be performed.
[0120] The core path is identified through the passenger flow matrix. In the peak period, the high-frequency floor pairs are located, and in the trough period, the low-frequency requests are filtered. The passenger flow rule is predicted, the periodic characteristic is extracted, the passenger flow matrix of the same period on different days is compared, such as the morning peak on weekdays, the abnormality is detected, and the deviation of the current matrix from the historical average matrix exceeds the threshold value → the real-time scheduling is triggered. The resource allocation is optimized, and the directional dispatching is performed. For the high-frequency path, such as 1 floor → 20 floor, the exclusive elevator dynamic path planning is pre-allocated. The elevator stop priority is calculated by combining the matrix weight, and the high-weight path is preferentially responded.
[0121] The passenger flow path construction unit is configured to perform step S220 and identify the passenger flow paths in the peak period and the trough period according to the passenger flow direction matrix in the different time periods.
[0122] In specific implementation, the elevator dispatching intelligent agent is adopted to identify the passenger flow paths in the peak period and the trough period according to the passenger flow direction matrix in the different time periods.
[0123] The space-time feature extraction unit is configured to perform step S230 and extract the space-time features of the elevator taking behavior according to the passenger flow paths in the peak period and the trough period.
[0124] In specific implementation, the elevator dispatching intelligent agent is adopted to extract the space-time features of the elevator taking behavior according to the passenger flow paths in the peak period and the trough period.
[0125] The strategy deployment module is configured to perform step S300 and generate the elevator resource dynamic scheduling strategy according to the space-time features of the elevator taking behavior.
[0126] In specific implementation, the elevator dispatching intelligent agent is adopted to generate the elevator resource dynamic scheduling strategy according to the space-time features of the elevator taking behavior.
[0127] The strategy deployment module includes a decision unit configured to perform step S310A and adopt the deployed scheduling strategy to plan the elevator resource when the elevator taking demand meets the periodic characteristic.
[0128] In specific implementation, if it is judged that the current elevator taking demand meets the prior periodic characteristic, the above-mentioned elevator resource scheduling scheme can be directly adopted to schedule the elevator resource of each application scenario.
[0129] S310B, when the elevator taking demand does not meet the periodic characteristic, the elevator resource dispatching intelligent agent calculates the dynamic scheduling strategy in real time to plan the elevator resource.
[0130] In specific implementation, if it is judged that the current elevator demand does not conform to the prior periodic characteristics, the elevator resource scheduling intelligent agent calculates a dynamic scheduling strategy in real time to plan the elevator resources, and different weight coefficients of the reward function can be adjusted and set according to different scene demands, for example, in an emergency scene, the emergency demand coefficient can be adjusted to 0.8 and the like.
[0131] Scenario one:
[0132] The elevator scheduling intelligent agent is applied to the office building group control system, if a current emergency meeting or activity and the like is requested, the elevator scheduling intelligent agent can extract the time and space characteristics of the emergency condition according to the meeting or activity information entered by the property, and perform emergency scheduling on the current elevator resources according to the time and space characteristics of the emergency condition.
[0133] For example, a company needs to hold an activity, then when the activity is about to start, the office building group control system actively arranges the idle elevators to the first floor to receive the guests, and when the activity is over, the office building group control system actively arranges the idle elevators to the activity floor.
[0134] Scenario two:
[0135] The elevator scheduling intelligent agent deployed elevator resource dynamic scheduling strategy is applied to the elevator group control system to control the elevator main control system to perform elevator scheduling: for example, at 15:08 in the afternoon, a “3F→1F emergency request” is suddenly triggered, the elevator scheduling intelligent agent immediately assigns a reserved standby elevator to respond, and the traditional system needs to be recalled from the 12th floor, which delays for 22 seconds.
[0136] In a hospital scene, the request priority can be directly set, when an emergency request is triggered, the elevator group control system will actively ignore other elevator requests, and directly go to the request floor, and after receiving the emergency patient, directly send to the target floor without stopping on the way.
[0137] When a new inpatient building is put into use, causing the traffic mode to change, the intelligent agent automatically triggers incremental training after detecting data deviation, and adapts to the new passenger flow distribution within 2 weeks.
[0138] Scenario three:
[0139] The elevator group control system applies the dynamic scheduling strategy generated by the scheduling intelligent agent to control the elevator main control system to perform elevator scheduling: for example, when a cinema is added on the weekend, the elevator scheduling intelligent agent acquires the opening information and closing information published by the cinema to perform real-time emergency scheduling of the elevator resources on the weekend.
[0140] When the movie is about to start, the elevator group control system actively schedules idle elevators to the first floor standby.
[0141] When the movie is over, the elevator group control system actively schedules idle elevators from the first floor to the 8th floor standby, and the traditional system needs to respond after the passenger calls the elevator.
[0142] Pre-train the model in combination with holiday historical data.
[0143] Based on the same inventive concept, the application also provides an electronic device, comprising a chip, a processor and a memory, the memory being used to store computer program codes, the computer program codes comprising computer instructions, under the condition that the chip executes the computer instructions, the electronic device executes a dynamic scheduling method of elevator group control based on reinforcement learning.
[0144] Reference Figure 3 The electronic device 2 comprises a processor 21, a memory 22, an input device 23 and an output device 24. The processor 21, the memory 22, the input device 23 and the output device 24 are coupled through a connector, which comprises various interfaces, transmission lines or buses, etc., and the embodiments of the application are not limited thereto. It should be understood that in various embodiments of the application, coupling refers to mutual connection in a specific manner, including direct connection or indirect connection through other devices, for example, various interfaces, transmission lines, buses, etc.
[0145] The processor 21 can be one or more graphics processing units (GPUs), and in the case that the processor 21 is a GPU, the GPU can be a single-core GPU or a multi-core GPU. Alternatively, the processor 21 can be a processor group composed of multiple GPUs, and the multiple processors are coupled to each other through one or more buses. Alternatively, the processor can also be other types of processors, etc., and the embodiments of the application are not limited thereto.
[0146] The memory 22 can be used to store computer program instructions, and various types of computer program codes for executing the schemes of the application. Alternatively, the memory includes but is not limited to random access memory (RAM), read-only memory (ROM), erasable programmable read only memory (EPROM), or compact disc read-only memory (CD-ROM), which is used for related instructions and data.
[0147] The input device 23 is used to input data and / or signals, and the output device 24 is used to output data and / or signals. The output device 24 and the input device 23 can be independent devices, or can be an integral device.
[0148] Based on the same inventive concept, the application also provides a computer readable storage medium, corresponding to the aforementioned embodiment of the elevator group control dynamic scheduling method based on reinforcement learning, the computer readable storage medium has a computer program stored thereon, and the program is executed by a processor to implement the steps of the elevator group control dynamic scheduling method based on reinforcement learning described in any of the aforementioned embodiments.
[0149] The application can adopt the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing program code. The computer usable storage medium includes permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0150] The scheduling agent of the application automatically generates and dynamically updates the scheduling strategy by autonomously learning the space-time rules of the elevator data, such as the uplink traffic in the morning peak and the dispersed travel across floors during the lunch break; without defining complex rules by humans, the problem of relying on human experience is completely solved, and the difficulty and cost of system operation and maintenance are greatly reduced; the scheduling agent of the application identifies subtle space-time rules by analyzing high-dimensional data such as time stamp + source / destination floor, generates space-time variable strategies, accurately models and dynamically responds to high-dimensional space-time passenger flow rules in buildings, and significantly improves scheduling efficiency; the system of the application realizes real-time data closed loop (collection→training→deployment), continuously iterates and optimizes the scheduling strategy, automatically detects changes in passenger flow patterns, actively triggers model retraining, ensures that the system maintains optimal performance for a long time, and prolongs the technical life cycle; the multi-objective collaborative optimization mechanism of the application can reduce the overall operation cost of the system while ensuring the elevator experience, and reasonable scheduling arrangement can reduce the empty running of the elevator and reduce energy consumption.
[0151] The foregoing is considered as illustrative only of the principles of the application. Numerous modifications and changes will readily occur to those skilled in the art, and it is intended to embrace all such modifications and changes that fall within the scope of the application. Accordingly, the application is not to be restricted in scope to the specific embodiments disclosed herein but is to be accorded the full scope that the principles and novel features request appropriately granted.
Claims
1. A dynamic scheduling method for elevator group control based on reinforcement learning, characterized in that, include: S100: Obtain real-time elevator demand data; S200: Obtain the spatiotemporal characteristics of elevator riding behavior based on the real-time elevator demand data; S3 00: Generate a dynamic elevator resource scheduling strategy based on the spatiotemporal characteristics of the elevator riding behavior, including: S310A: When elevator demand exhibits cyclical characteristics, elevator resource planning is performed using a pre-deployed scheduling strategy; S310B: When elevator demand does not conform to periodic characteristics, the elevator resource scheduling intelligent agent calculates dynamic scheduling strategies in real time to plan elevator resources.
2. The elevator group control dynamic scheduling method based on reinforcement learning according to claim 1, characterized in that, Before obtaining real-time elevator demand data, the process also includes: S010: Obtain prior elevator data; S020: Calculate the deployed elevator resource scheduling strategy based on prior elevator riding data.
3. The elevator group control dynamic scheduling method based on reinforcement learning according to claim 1, characterized in that, The acquisition of real-time elevator demand data includes: S110: Obtain the timestamp, source floor, and target floor data of the real-time elevator request; S120: The acquired real-time data is divided into temporal data and spatial data.
4. The elevator group control dynamic scheduling method based on reinforcement learning according to claim 1, characterized in that, The step of obtaining the spatiotemporal characteristics of elevator riding behavior based on the real-time elevator demand data includes: S210: Obtain the passenger flow direction matrix between floors within different time periods based on time and spatial data; S220: Identify passenger flow paths during peak and off-peak hours based on the passenger flow direction matrix for different time periods; S230: Extract the spatiotemporal features of elevator riding behavior based on the passenger flow paths during peak and off-peak hours.
5. The elevator group control dynamic scheduling method based on reinforcement learning according to claim 1, characterized in that, The process of obtaining the passenger flow direction matrix between floors within different time periods based on time and spatial data includes: Define time slices to divide a day into T time slots; For time period t, generate an N×N matrix M. t N represents the floor number, rows represent the starting floor, and columns represent the destination floor.
6. The elevator group control dynamic scheduling method based on reinforcement learning according to claim 1, characterized in that, The elevator resource scheduling strategy calculated and deployed based on prior elevator riding data includes: S021: Use the prior elevator data as a training set to input into the reinforcement learning-based agent model; S022: Train and optimize the reinforcement learning-based agent model; S023: The trained reinforcement learning-based agent model outputs the deployed elevator resource scheduling strategy.
7. The elevator group control dynamic scheduling method based on reinforcement learning according to claim 1, characterized in that, After generating the elevator resource dynamic scheduling strategy based on the spatiotemporal characteristics of the elevator riding behavior, the method further includes iteratively optimizing the elevator resource dynamic scheduling strategy through a reward function feedback mechanism. Set the reward function as follows: R=ω1R wait +ω2R energy +ω3R balance +ω4R emergency Where ω1, ω2, ω3, and ω4 are adjustable weights, and R wait As a waiting time penalty, R energy As a penalty for energy consumption, R balance As a load balancing reward, R emergency Rewards are given for urgent needs.
8. A dynamic scheduling system for elevator group control based on reinforcement learning, characterized in that, include: The data acquisition module is used to acquire real-time elevator demand data; The feature extraction module is used to obtain the spatiotemporal features of elevator riding behavior based on the real-time elevator demand data; The strategy deployment module is used to generate a dynamic elevator resource scheduling strategy based on the spatiotemporal characteristics of the elevator riding behavior, including: When elevator demand exhibits cyclical characteristics, elevator resource planning is performed using pre-deployed scheduling strategies. When elevator demand does not conform to periodic characteristics, the elevator resource scheduling intelligent agent calculates dynamic scheduling strategies in real time to plan elevator resources.
9. An electronic device, characterized in that, include: The electronic device comprises a chip, a processor, and a memory, the memory being used to store computer program code, the computer program code including computer instructions, wherein, when the chip executes the computer instructions, the electronic device performs a reinforcement learning-based dynamic scheduling method for elevator group control as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which includes program instructions that, when executed by a processor of an electronic device, cause the processor to perform a dynamic scheduling method for elevator group control based on reinforcement learning, as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Intelligent control system of elevator group
CN110155827A
Traffic mode recognition method based on genetic algorithm and fuzzy neural network
CN114707587A
Communication system based on elevator Internet of Things and data sending method
CN118479315A
Elevator group control method and device, computer equipment and storage medium
CN118619024A
Elevator group control method and device and storage medium
CN120328279A
Cited By
Elevator group control energy-saving dispatching method and system based on load prediction
CN122276550A