Elevator group control method and device and storage medium
Through the elevator group control system of the perception layer, decision-making layer and execution layer, and using reinforcement learning to optimize elevator scheduling, the problems of long elevator time, high energy consumption and low comfort in traditional elevator scheduling methods are solved, and more efficient and energy-saving elevator operation is achieved.
Patent Information
- Application Number
- CN202510483326.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional elevator group scheduling methods rely on fixed rules or experience, resulting in long-term escalators waiting for the elevator, high energy consumption, low comfort and poor adaptability.
The elevator group control system of the perception layer, decision-making layer and execution layer is adopted to sense the status information of the elevator through sensors, and the near-end strategy optimization of reinforcement learning is used to select decision-making strategies in the hybrid action space, and control elevator execution to compress the time and energy consumption of the elevator and improve comfort.
The elevator operation efficiency is optimized, the elevator rider time is reduced, the elevator energy consumption is reduced, and the elevator rider experience is improved.
Smart Images

Figure CN120328279A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of elevator dispatching, and in particular, to an elevator group control method, device, and storage medium. Background Art
[0002] With the continuous advancement of the urbanization process, the scale of buildings has gradually increased, especially the construction of high-rise buildings has become more and more common. In order to ensure the traffic fluency inside the building, as a vertical transportation vehicle, the operation efficiency and service quality of the elevator affect people's travel experience.
[0003] Traditional elevator group dispatching methods mainly rely on fixed rules or experience-based heuristic algorithms, such as first come first served, round-robin scheduling, etc. These methods determine the elevator dispatching strategy through preset rules or conditions, and these dispatching strategies are usually fixed, resulting in poor adjustment ability of the elevator dispatching strategy; traditional elevator dispatching methods usually formulate dispatching strategies based on real-time collected data, which may lead to poor adaptability of the dispatching strategy to the building, and may result in long waiting times for elevator passengers, high elevator energy consumption, and low comfort of elevator passengers. Summary of the Invention
[0004] The present invention provides an elevator group control method, device, and storage medium, which are used to optimize the long waiting time for elevator passengers, high elevator energy consumption, and low comfort of elevator passengers as a whole.
[0005] In a first aspect, an embodiment of the present invention provides an elevator group control method, which is applied to a group control system of multiple elevators in a building. The group control system includes a sensing layer, a decision-making layer, and an execution layer; the method includes:
[0006] In the sensing layer, call sensors in each of the elevators to sense status information related to the waiting time of elevator passengers, the energy consumption of multiple elevators, and the comfort of elevator passengers in the historical, present, and future tenses;
[0007] In the decision-making layer, perform proximal policy optimization under reinforcement learning based on the status information to select one action in the mixed action space as the decision-making strategy; the mixed action space includes discrete actions and continuous actions related to call control of each elevator;
[0008] In the execution layer, control multiple elevators to execute the decision-making strategy to compress the waiting time of elevator passengers, the energy consumption of multiple elevators, and improve the comfort of elevator passengers.
[0009] In a second aspect, an embodiment of the present invention further provides an elevator group control device, which is applied to a group control system of multiple elevators in a building. The group control system includes a sensing layer, a decision-making layer, and an execution layer; the device includes:
[0010] A status information perception module, configured to call sensors in each of the elevators in the perception layer to perceive status information related to the waiting time of passengers taking the elevator, the energy consumption of multiple elevators, and the comfort of passengers taking the elevator in the historical tense, present tense, and future tense;
[0011] A decision strategy selection module, configured to execute proximal policy optimization under reinforcement learning in the decision layer according to the status information, so as to select one action from a mixed action space as a decision strategy; the mixed action space includes discrete actions and continuous actions related to elevator call control for each of the elevators;
[0012] A decision strategy execution module, configured to control multiple elevators to execute the decision strategy in the execution layer, so as to compress the waiting time of passengers taking the elevator, the energy consumption of multiple elevators, and improve the comfort of passengers taking the elevator.
[0013] In a third aspect, an embodiment of the present invention further provides a computer device, where the computer device includes:
[0014] One or more processors;
[0015] A storage device, configured to store one or more programs;
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the elevator group control method provided in the first aspect of the present invention.
[0017] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the elevator group control method provided in the first aspect of the present invention is implemented.
[0018] In a fifth aspect, an embodiment of the present invention further provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, the elevator group control method provided in the first aspect of the present invention is implemented.
[0019] In an embodiment of the present invention, there is provided a group control system for multiple elevators in a building. The group control system includes a sensing layer, a decision-making layer, and an execution layer. In the sensing layer, sensors in each elevator are called to sense status information related to the waiting time of passengers, the energy consumption of multiple elevators, and the comfort of passengers in historical, current, and future time tenses. In the decision-making layer, proximal policy optimization under reinforcement learning is performed based on the status information to select one action from a mixed action space as the decision-making strategy. The mixed action space includes discrete and continuous actions related to each elevator and call control. In the execution layer, multiple elevators are controlled to execute the decision-making strategy to compress the waiting time of passengers, the energy consumption of multiple elevators, and improve the comfort of passengers. By collecting the status information of the elevator in multiple time tenses through the sensing layer, the operating conditions of the elevator and the needs of passengers are comprehensively understood, providing a data basis for the decision-making of the decision-making layer. The decision-making layer uses the proximal policy optimization algorithm in reinforcement learning to select the decision-making strategy from the mixed action space based on the status information provided by the sensing layer. Introducing the reinforcement learning technology enables the group control system to select the optimal decision-making strategy, improving the efficiency of elevator dispatching. The mixed action space provides flexible decision-making strategies, enabling the group control system to be more accurate and flexible when facing different types of control tasks. The execution layer controls each elevator to execute the corresponding decision-making strategy according to the output of the decision-making layer. The execution layer converts the theoretical decision-making strategy into actual actions, overall compressing the waiting time of passengers, the energy consumption of multiple elevators, and improving the comfort of passengers. Through multi-level collaboration, the group control system not only optimizes the elevator operation efficiency but also improves the user experience and reduces energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a flowchart of an elevator group control method provided in Embodiment 1 of the present invention;
[0021] Figure 2 It is an architecture diagram of an elevator group control system provided in Embodiment 1 of the present invention;
[0022] Figure 3 It is a schematic structural diagram of a basic state matrix provided in Embodiment 1 of the present invention;
[0023] Figure 4 It is a flowchart of an elevator group control method provided in Embodiment 2 of the present invention;
[0024] Figure 5 It is a schematic diagram of the time interval division structure provided in Embodiment 2 of the present invention;
[0025] Figure 6 It is a structural block diagram of an elevator group control device provided in Embodiment 3 of the present invention;
[0026] Figure 7 It is a schematic structural diagram of a computer device provided in Embodiment 4 of the present invention. Detailed implementation mode
[0027] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can cover the sequential implementations other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0029] Embodiment 1
[0030] See Figure 1 , which shows a flowchart of an elevator group control method provided in Embodiment 1 of the present invention. This embodiment can be adapted to the situation of elevator group control based on proximal policy optimization under reinforcement learning to optimize the waiting time of elevator passengers, high elevator energy consumption, and passenger comfort as a whole. This method can be executed by an elevator group control device, which can be implemented in the form of hardware and / or software, and the elevator group control device can be configured in a computer device. As Figure 1 shown, the method includes:[[]]
[0031] Step 101: In the perception layer, call the sensors in each elevator to perceive the state information related to the waiting time of elevator passengers, the energy consumption of multiple elevators, and the comfort of elevator passengers in the historical, present, and future tenses.
[0032] This embodiment is applied to a group control system of multiple elevators in a building. The group control system includes a perception layer, a decision layer, and an execution layer.
[0033] The perception layer is responsible for collecting and monitoring the operating status, environmental data, and passenger demands of the elevators. Its main task is to obtain and transmit all real-time information about the current status of the elevators. The perception layer includes various sensors and data collection devices.
[0034] The decision-making layer is responsible for analyzing and processing the real-time data transmitted by the perception layer and making scheduling decisions. The core task of the decision-making layer is to formulate the optimal elevator decision-making strategy based on the input of the perception layer and the target requirements to ensure the efficient operation of the group control system.
[0035] The execution layer is responsible for actual operations according to the elevator decision-making strategy of the decision-making layer. The execution layer controls the elevator's drive system and various control devices to ensure that the elevator accurately and timely completes tasks according to the scheduling instructions of the decision-making layer.
[0036] Exemplarily, such as Figure 2 shown is the architecture diagram of the elevator group control system. Figure 2 It includes a perception layer, a decision-making layer, and an execution layer; the perception layer is used for data collection. The perception layer is equipped with elevator sensors, floor call panels, and environmental perception devices (such as cameras or infrared sensors). The perception layer collects real-time status data, and the real-time status data includes the elevator's position, elevator speed, elevator load, up / down requests, etc.
[0037] The collected real-time status data is transmitted to the decision-making layer. The decision-making layer acts as a reinforcement learning agent, and the decision-making layer includes a state space module, a PPO algorithm module, and a dynamic reward calculation module. The state space module is used to expand the real-time status data, including the basic state and the extended state; the PPO algorithm module analyzes the basic state and the extended state based on the policy network and the value network to generate action instructions. The action instructions include dispatching a certain elevator to execute the call request, picking up and dropping off passengers to the destination floor, and adjusting the speed of the dispatched elevator.
[0038] The generated action instructions are transmitted to the execution layer. The execution layer is used for elevator control and feedback. The execution layer includes an elevator control unit and an elevator actuator. The elevator control unit includes motor control and door machine control. The elevator actuator includes a car and a drive system. The elevator control unit receives the action instructions generated by the decision-making layer and issues control signals according to the action instructions. The elevator actuator receives the control signals and picks up and drops off passengers to the destination floor. The elevator actuator feeds back the execution result to the dynamic reward calculation module of the decision-making layer. The dynamic reward calculation module includes a multi-objective reward function to calculate the reward value generated by this execution result, and updates the PPO algorithm module based on the reward value. The elevator actuator also feeds back the execution result to the perception layer, and the perception layer receives the status data of the elevator group at this time.
[0039] In this embodiment, the perception layer collects multi-dimensional state information related to the waiting time of elevator passengers, elevator energy consumption, and passenger comfort by invoking sensors in the elevator, understands the operating conditions of the group control system in the historical, current, and future tenses, can accurately capture the dynamic changes of the group control system, and thus schedule and optimize the group control system. The waiting time can reflect the waiting situation of elevator passengers, the elevator energy consumption helps evaluate the resource consumption of the elevator system, and the comfort reflects the smoothness during the elevator ride and the passenger experience.
[0040] Exemplarily, the state information includes a basic state matrix and extended state variables.
[0041] As Figure 3 shown is the structural schematic diagram of the basic state matrix. Figure 3 The rows represent each floor in the building, with a total of m rows (1, 2.... m), where m represents a positive integer. The columns include each floor in the building, with m columns for each floor (1, 2.... m). The columns also include the waiting time when calling the elevator upward (B11, B21......Bm1), the waiting time when calling the elevator downward (B12, B22......Bm2), the internal call labels of each elevator (C11....Cm1), the labels of the floors not stopped at when moving upward (C12....Cm2), the labels of the floors not stopped at when moving downward (C13.....Cm3), and the labels of the floors where the elevator stops (C14....Cm4); each elevator contains these 4 columns of information: internal call label, label of floors not stopped at when moving upward, label of floors not stopped at when moving downward, and label of floors where the elevator stops. There are a total of N elevators arranged in this building, where N is a positive integer; the floors in the building (A11...Amm) form a binary matrix, represented by 0 and 1 in binary for easy computer calculation and storage; the waiting time when calling the elevator upward and the waiting time when calling the elevator downward (B11...Bm2) form a real number matrix. The waiting time of elevator passengers is a dynamic and continuously changing quantity, so using real numbers can accurately capture the changing characteristics of the waiting time of elevator passengers. For example, in an actual scenario, the waiting time of elevator passengers may be a specific real value (such as 5.3 minutes), which is important for the system to optimize the waiting time and energy consumption. By using real numbers, the response time of each elevator and the passenger experience can be accurately calculated, compared, and optimized, thereby enabling the group control system to more effectively reduce the waiting time of elevator passengers and improve the efficiency of elevator scheduling; the internal calls, non-stop upward, non-stop downward, stops, and each floor in the building of all elevators form a binary matrix, and the number of elevators is N, where N is represented as a positive integer.
[0042] Before the basic status matrix collects information, each element in the basic status matrix is assigned a value of 0, and the collected information is filled into the corresponding element to facilitate efficient observation of which part of the information is collected. The basic status matrix is used to collect the current status information of multiple elevators related to the waiting time of passengers, the energy consumption of multiple elevators, and the comfort of passengers.
[0043] When the element with the floor in the behavior building and the column as the floor in the building is 0, the element is invalid. When the element with the floor in the behavior building and the column as the floor in the building is 1, the element indicates that the floor in the building belonging to the row is the departure floor and the floor in the building belonging to the column is the destination floor; for example, if the element in the i-th row and j-th column is 1, it means that there is a call signal with the departure floor as the i-th floor and the destination floor as the j-th floor. The call signals include external call signals and internal call signals. The internal call signal means that the passengers inside the elevator press the floor button, requesting the elevator to go to the floor where the passengers want to arrive. The external call signal means that the passengers on the floor press the elevator button, requesting the elevator to go up or down. Some external call signals only contain the departure floor, and some external call signals contain both the departure floor and the destination floor, but all internal call signals contain the destination floor.
[0044] When the element with the floor in the behavior building and the column as the waiting time for upward call is the first value, the first value indicates that the floor in the building belonging to the row is the departure floor, and the waiting time for upward call belonging to the column is the waiting time for the passenger to call the elevator upward at the departure floor, which is convenient for counting the time spent by the passenger calling the elevator this time.
[0045] When the element with the floor in the behavior building and the column as the waiting time for downward call is the second value, the second value indicates that the floor in the building belonging to the row is the departure floor, and the waiting time for downward call belonging to the column is the waiting time for the passenger to call the elevator downward at the departure floor.
[0046] When the element with the floor in the behavior building and the column as the label of the internal call of a certain elevator is 0, the element is invalid. When the element with the floor in the behavior building and the column as the label of the internal call of a certain elevator is 1, the element indicates that the internal call signal triggered by the elevator belonging to the column points to the floor in the building belonging to the row.
[0047] When the element with the floor in the behavior building and the column as the label of the floors that an elevator does not stop at when moving upward is 0, the element is invalid. When the element with the floor in the behavior building and the column as the label of the floors that an elevator does not stop at when moving upward is 1, the element indicates that the elevator belonging to the column does not stop at the floors in the building belonging to the row when moving upward.
[0048] When the element of the label indicating the floors that an elevator does not stop at during its downward movement in a building is 0, the element is invalid. When the element of the label indicating the floors that an elevator does not stop at during its downward movement in a building is 1, the element represents the floors in the building belonging to the row that the elevator belonging to the column does not stop at during its downward movement.
[0049] When the element of the label indicating the floors where an elevator stops in a building is 0, the element is invalid. When the element of the label indicating the floors where an elevator stops in a building is 1, the element represents the floors in the building belonging to the row that the elevator belonging to the column stops at during its movement.
[0050] Exemplarily, the extended state variables include a heat map representing the demand for elevators in space and time, the cumulative energy consumption of multiple elevators, and the comfort index of passengers. The space-time demand heat map can display the distribution of elevator demand on each floor in real time, helping the group control system to accurately dispatch elevators, avoid overcrowding or empty running of elevators on certain floors, and thus improve the response efficiency of the group control system. The cumulative energy consumption of multiple elevators helps the system identify high-energy-consuming elevator operation modes, optimize the decision-making strategy for elevator dispatching, reduce energy waste, and extend the service life of elevator equipment. At the same time, combined with the comfort index of passengers, the system can balance the response speed of the elevator with the waiting time and comfort of passengers, enabling passengers to have a better experience. Through the comprehensive analysis of these extended state variables, the elevator group control system can effectively save energy, improve the operation efficiency, and enhance the satisfaction of passengers on the premise of ensuring service quality. The extended state variables are used to collect state information related to the waiting time of passengers, the energy consumption of multiple elevators, and the comfort of passengers in the present tense, historical tense, and future tense.
[0051] Specifically, historical passenger flow information is collected; collecting historical passenger flow information is to collect and accumulate relevant data, which can reflect the passenger flow changes on each floor over a past period of time.
[0052] The historical passenger flow information is input into a pre-set long short-term memory network (LSTM). LSTM is particularly good at processing time series data. The temporal pattern of historical passenger flow can help predict the distribution information of the floors in the building as the destination floors in a future period of time, serving as a heat map representing the demand for elevators in space and time. The heat map is used as the state information collected in the future tense.
[0053] Define the time interval between the previous execution decision strategy and the current execution decision strategy as the time period; add up the energy consumption of multiple elevators during the time period to obtain the cumulative energy consumption of multiple elevators; the cumulative energy consumption is an indicator to measure the operation efficiency of the elevator. Through this cumulative amount, the elevator dispatching strategy can be further optimized to reduce energy consumption and improve energy utilization efficiency. The cumulative energy consumption is used as the state information collected in the present tense.
[0054] Calculate the root mean square (RMS) of the accelerations of multiple elevators when they stop at the floors of a building multiple times in the past as an indicator of the comfort level of elevator passengers. The RMS value of acceleration is used as an indicator of passenger comfort because excessive acceleration may cause discomfort to passengers. By calculating the acceleration of the elevator when it stops at different floors, the smoothness of the elevator operation can be evaluated, and thus the elevator group control system can be further improved. The RMS is the state information collected in the historical tense.
[0055] Step 102: In the decision-making layer, based on the state information, perform proximal policy optimization under reinforcement learning to select one action in the mixed action space as the decision-making policy.
[0056] In this embodiment, in the decision-making layer of the elevator group control system, based on the state information obtained by the perception layer, proximal policy optimization (PPO) in reinforcement learning is used to select an action. This is to select the optimal decision-making policy in the mixed action space to achieve more efficient resource utilization, reduce the waiting time of passengers, reduce elevator energy consumption, and improve the comfort level of passengers. The mixed action space includes discrete actions and continuous actions, and the mixed action space includes discrete actions and continuous actions related to each elevator and call control.
[0057] Proximal Policy Optimization (PPO) is a policy optimization method in reinforcement learning. Its core idea is to control the step size of the decision-making policy update to avoid instability caused by excessive updates. PPO ensures that each update is near the current decision-making policy through "proximal" constraints, which can reduce the risk of performance degradation or learning instability caused by over-updating. In practical applications, PPO can converge to a good decision-making policy in a relatively short time and has good robustness. This makes PPO an ideal choice in the elevator group control system, which can quickly and stably adapt to changing requirements and ensure the efficiency and stability of the elevator decision-making policy.
[0058] Reinforcement learning is a machine learning method that learns the optimal behavior policy by the interaction between an agent and an environment. At each time step, the agent observes the state of the environment, selects an action, then executes the action, which causes a change in the state of the environment, and obtains a reward according to the executed action. The goal of the agent is to maximize the long-term cumulative reward through continuous trial and error and optimization, so as to learn to select the most appropriate action in different states and finally achieve the optimal decision-making policy.
[0059] An Agent is the decision maker in reinforcement learning. It learns the optimal policy by interacting with the environment. The Agent selects actions and learns and optimizes based on the environmental feedback. In the present invention, the Agent is the elevator group control system. The Agent receives the state information from the sensing layer and makes decisions according to the reinforcement learning algorithm (such as Proximal Policy Optimization), so as to control the operation of the elevator (such as dispatching the elevator, adjusting the elevator moving speed, etc.).
[0060] The Environment refers to the external environment where the Agent is located. The environment gives rewards and new states according to the actions of the Agent. In the present invention, the environment can be regarded as the physical and operating environment of the elevator group control system in a building.
[0061] The State is an observation of the environment by the Agent. The state represents the situation of the Agent at a certain moment and determines the behavior selection of the Agent. The state corresponds to the state information collected by the sensing layer.
[0062] The Action is the behavior that the Agent can execute. The action directly affects the environment and has an impact on the future state of the Agent. In the present invention, the action corresponds to selecting an action from the mixed action space as the decision-making strategy in the elevator group control system.
[0063] The Reward is the feedback signal given by the environment after the Agent executes a certain decision-making strategy, indicating the quality of the decision-making strategy. A high reward value means that the decision-making strategy of the Agent is more in line with the goal, while a low reward value means that the decision-making strategy of the Agent does not meet the expectation. The reward in the present invention corresponds to the comprehensive evaluation index of elevator behavior. The group control system calculates multiple evaluation indexes (such as waiting time, elevator energy consumption, elevator comfort, etc.) to measure the effect of the elevator group control strategy.
[0064] Exemplarily, the mixed action space includes discrete actions and continuous actions. The discrete action can efficiently handle the basic control of the elevator, such as responding to the call signal and moving to the specified floor, and can simply correspond to the specific decision-making strategy in elevator dispatching. The discrete action provides clear decision-making choices, enabling the group control system to quickly select the optimal solution from multiple possible dispatching schemes. The continuous action, on the other hand, introduces fine-tuning of the elevator moving speed and acceleration, providing more fine-grained control and making the elevator dispatching more refined and flexible.
[0065] The discrete action is a one-dimensional matrix. The columns of the one-dimensional matrix represent each elevator. When the value of a column is 1, it means that when the elevator call signal is received, the dispatched elevator responds to the elevator call signal and moves to the floor in the building indicated by the elevator call signal. To keep the dimension of the discrete action space consistent, each decision-making strategy only faces one elevator call signal, and the elevator allocation adopts One-hot encoding. One-hot encoding is a method of converting categorical variables into numerical forms, which is widely used in machine learning and data processing. It represents each category as a binary vector of length N, where N is the total number of categories. In this vector, only the position representing the current category is 1, and the rest are 0. For example, in a case with three categories, category 1 can be represented as [1,0,0], category 2 as [0,1,0], and category 3 as [0,0,1]. This encoding method can avoid the order relationship between categories and ensure that each category is independently processed in the model.
[0066] The continuous action includes adding a speed adjustment coefficient to the moving speed of the dispatched elevator within a preset first adjustment range, and adding an acceleration coefficient to the moving acceleration of the dispatched elevator within a preset second adjustment range. The acceleration adjustment coefficient V ratio ∈[0.7,1.1] (70%-110% of the rated speed), and the acceleration coefficient a ∈ [0.8,1.2] m / s 2 .
[0067] Step 103: In the execution layer, control multiple elevators to execute the decision-making strategy to compress the waiting time of passengers taking the elevator, the energy consumption of multiple elevators, and improve the comfort of passengers taking the elevator.
[0068] In this embodiment, in the execution layer, the group control system controls multiple elevators to execute the decision-making strategy formulated by the decision-making layer, which can compress the waiting time of passengers taking the elevator, the energy consumption of multiple elevators to the greatest extent, improve the comfort of passengers taking the elevator, and improve the elevator group control efficiency.
[0069] In the embodiments of the present invention, a group control system applied to multiple elevators in a building includes a sensing layer, a decision-making layer, and an execution layer. In the sensing layer, sensors in each elevator are called to sense status information related to the waiting time of passengers, the energy consumption of multiple elevators, and the comfort of passengers in historical, present, and future tenses. In the decision-making layer, proximal policy optimization under reinforcement learning is performed based on the status information to select one action from a mixed action space as the decision-making strategy. The mixed action space contains discrete and continuous actions related to each elevator and call control. In the execution layer, multiple elevators are controlled to execute the decision-making strategy to compress the waiting time of passengers, the energy consumption of multiple elevators, and improve the comfort of passengers. By collecting the status information of the elevator in multiple tenses through the sensing layer, the operating conditions of the elevator and the needs of passengers are comprehensively understood, providing a data basis for the decision-making of the decision-making layer. The decision-making layer uses the proximal policy optimization algorithm in reinforcement learning to select the decision-making strategy from the mixed action space according to the status information provided by the sensing layer. Introducing reinforcement learning technology enables the group control system to select the optimal decision-making strategy, improving the efficiency of elevator dispatching. The mixed action space provides flexible decision-making strategies, enabling the group control system to be more accurate and flexible when facing different types of control tasks. The execution layer controls each elevator to execute the corresponding decision-making strategy according to the output of the decision-making layer. The execution layer transforms the theoretical decision-making strategy into actual actions, overall compressing the waiting time of passengers, the energy consumption of multiple elevators, and improving the comfort of passengers. Through multi-level collaboration, the group control system not only optimizes the operating efficiency of the elevator but also improves the user experience and reduces energy consumption.
[0070] Embodiment 2
[0071] Figure 4 It is a flowchart of an elevator group dispatching method provided by Embodiment 2 of the present invention. Based on the foregoing Embodiment 1, the process of offline update of reinforcement learning is described in detail, as Figure 4 shown, the method includes:
[0072] Step 401, in the sensing layer, sensors in each elevator are called to sense status information related to the waiting time of passengers, the energy consumption of multiple elevators, and the comfort of passengers in historical, present, and future tenses.
[0073] The present invention is applied to a group control system of multiple elevators in a building. The group control system includes a sensing layer, a decision-making layer, and an execution layer.
[0074] Step 402, in the decision-making layer, proximal policy optimization under reinforcement learning is performed based on the status information to select one action from a mixed action space as the decision-making strategy.
[0075] The mixed action space contains discrete and continuous actions related to each elevator and call control.
[0076] Step 403: In the execution layer, control multiple elevators to execute the decision-making strategy to compress the waiting time of passengers taking the elevator, the energy consumption of multiple elevators, and improve the comfort of passengers taking the elevator.
[0077] Step 404: In the execution layer, query the time points when non-decision-making strategies occur between two adjacent executions of the decision-making strategy.
[0078] In this embodiment, the decision-making process in the elevator group control system usually includes multiple steps, and each step involves different decisions and adjustments, such as dispatching elevators, controlling the speed and acceleration of elevators, or changing the operation sequence of elevators, etc. Querying the time points of non-decision-making strategies can help the system identify whether there are situations where the adjustment timing is inaccurate or the response is not timely due to changes in the state of the group control system (such as passenger demand, floor change, energy consumption control, etc.) between two decisions. Understand and monitor the performance of each elevator at different time points, so as to evaluate and feedback each time period.
[0079] Exemplarily, as Figure 5 shown in the schematic diagram of the time interval division structure, where both t and t' in the figure are the time points of executing the decision-making strategy, t represents the moment of the previous execution of the decision-making strategy, t' represents the moment of the current execution of the decision-making strategy, and both t1 and t2 are the time points of non-decision-making strategies. Non-decision-making strategies usually refer to the situation where no call signal is generated during the operation of the elevator, such as passengers arriving at the destination floor and leaving the elevator, elevator failure, etc. According to the time points of non-decision-making strategies occurring between two adjacent executions of the decision-making strategy, it is divided into multiple time periods (such as A, B, C).
[0080] The elevator group control problem can be regarded as a sequential decision-making process, and the state points of the sequence are mainly divided into discrete event points and decision event points. Discrete event points are when passengers get on or off the elevator, and decision event points vary according to the setting of the action space. However, the time interval between decision events is not fixed. That is, the decision-making problem MDP occurring in continuous time with a fixed discrete time step is not fully applicable to the destination floor scheduling problem. Therefore, it is necessary to introduce the continuous time extension of MDP - the Semi-Markov Decision Process (SMDP).
[0081] The Semi-Markov Decision Process (SMDP) is an extended Markov Decision Process (MDP) used to handle situations where the state transition times are not necessarily fixed and may be random. In a standard Markov Decision Process, state transitions occur based on fixed time steps, while in the Semi-Markov Decision Process, the time intervals between state transitions are random and may depend on the current state and the actions taken. Therefore, SMDP not only considers the transitions between states but also incorporates the time factor into the decision-making model, enabling it to better simulate and solve some decision-making problems involving time delays or persistence, such as the problem of elevator group scheduling. By introducing the time horizon, SMDP can more accurately capture dynamic changes and long-term effects in multi-stage decision-making.
[0082] Step 405: In the decision-making layer, jointly calculate the first sub-reward value representing the waiting time of elevator passengers at each time point in the dimension of instantaneity and prospect.
[0083] In this embodiment, by introducing the dimensions of instantaneity and prospect, the decision-making strategy for elevator scheduling can not only meet the current elevator demand but also anticipate and optimize future possible demands. This way of combining immediate and long-term demands quantifies the experience of elevator passengers during the waiting process, thereby optimizing the decision-making strategy for elevator scheduling, reducing the waiting time of passengers, and improving the elevator riding efficiency.
[0084] Specifically, at each time point, count the waiting time of each elevator passenger; if the waiting time is greater than the preset waiting threshold, calculate the difference between the waiting time and the waiting threshold as the first excess time; square the ratio between the first excess time and the preset penalty threshold to obtain the second excess time; take the negative sum of all the second excess times to obtain the waiting time penalty term.
[0085] At each time point, according to the heat map of the elevator in the state information, query multiple destination floors that the elevator passengers will go to in the future; for each elevator, take the absolute value of the difference between the position of the elevator floor and the positions of each destination floor to obtain multiple absolute distances; for each absolute distance, take the reciprocal of the sum of the absolute distance and 1 to obtain the influence value of the elevator on the destination floor;
[0086] Add up all the influence values to obtain the total influence value; add up the total influence values corresponding to each elevator to obtain the future elevator scheduling reward value;
[0087] Add the product of the waiting time penalty term, the future elevator scheduling reward value, and the dynamic decay factor as the first sub-reward value representing the waiting time of elevator passengers.
[0088] Exemplarily, the first sub-reward value is expressed as:
[0089]
[0090] In the formula, R wait is the first sub - reward value, is the waiting - time penalty term, is the future elevator scheduling reward value, δ is the dynamic attenuation factor, t is the time point, T total is the duration of a day (i.e., 86400 seconds), P wait is the set of passengers waiting for the elevator, t i is the waiting time for the elevator, T threshold is the preset waiting threshold, T threshold = 30s, T max is the preset penalty threshold, T max = 180s, ε is the set of elevators, F pred is the set of destination floors, floor(e j ) is the location of the elevator on the floor, e j is the j - th elevator.
[0091] Step 406: In the decision - making layer, jointly calculate the second sub - reward value representing the energy consumption of multiple elevators at each time point in the dimensions of start - stop and operation for the decision - making strategy.
[0092] In this embodiment, calculating the second sub - reward value representing the energy consumption of multiple elevators in the dimensions of start - stop and operation can accurately measure and optimize the overall energy efficiency of the elevator group control system. Since the energy consumption of the elevator affects the building operation cost and energy consumption, reasonable energy consumption assessment helps to reduce unnecessary energy waste in different scheduling decisions.
[0093] Specifically, at the time point, add the square value of the acceleration of the elevator going up and the square value of the acceleration of the elevator going down to obtain the elevator acceleration value; add the elevator acceleration values corresponding to each elevator to obtain the total acceleration value; take the opposite of the product of the total acceleration value and the elevator start - stop coefficient to obtain the elevator start - stop energy consumption; at each time point, square the difference between the running speed of the elevator and the preset elevator running speed to obtain the elevator running speed value; add the elevator running speed values corresponding to each elevator to obtain the total running speed value; take the opposite of the product of the total running speed value and the deviation - from - running - speed penalty coefficient to obtain the elevator running energy consumption; take the sum of the elevator start - stop energy consumption and the elevator running energy consumption as the second sub - reward value representing the energy consumption of multiple elevators;
[0094] Exemplarily, the second sub - reward value is expressed as:
[0095]
[0096] wherein, Renergy is the second sub - reward value, is the energy consumption of elevator start - stop, is the energy consumption of elevator operation, k1 is the elevator start - stop coefficient, k1 = 0.35, ε is the set of elevators, is the acceleration of the elevator going up, is the acceleration of the elevator going down, k2 is the deviation from the operating speed penalty coefficient, k2 = 0.2, is the operating speed of the i - th elevator, is the preset elevator operating speed of the i - th elevator. According to the motor efficiency MAP diagram, the elevator is encouraged to operate within the preset elevator operating speed range.
[0097] Step 407: In the decision - making layer, jointly calculate the third sub - reward value representing the comfort level of the passengers at each time point on the dimensions of the physiology and psychology of the passengers and the vibration of the elevator for the decision - making strategy.
[0098] In this embodiment, the physiological and psychological states of the passengers and the vibration characteristics of the elevator are combined to calculate the third sub - reward value of the comfort level of the passengers, reflecting the influencing factors of comfort during the elevator ride. This not only involves the physiological reaction of the passengers to the elevator acceleration but also includes the impact of elevator vibration on the psychological feeling. By evaluating these influencing factors, the discomfort of the passengers can be reduced, avoiding discomfort caused by elevator overloading or excessive vibration, and further optimizing the decision - making strategy of elevator dispatching.
[0099] Specifically, at each time point, integrate the square of the elevator acceleration from 0 to the time point to obtain the first acceleration integral; perform a power operation on the ratio between the first acceleration integral and the time point to obtain the second acceleration integral; add up the second acceleration integrals corresponding to each elevator to obtain the root - mean - square acceleration.
[0100] At each time point, integrate the fourth power of the elevator acceleration from 0 to the time point to obtain the third acceleration integral; perform a power operation on the third acceleration integral to obtain the fourth acceleration integral; add up the fourth acceleration integrals corresponding to each elevator to obtain the first vibration dose value; add the product of the root - mean - square acceleration, the first vibration dose value, and the preset vibration coefficient as the elevator smoothness value; take the opposite of the product of the elevator smoothness value and the preset comfort penalty coefficient to obtain the elevator smoothness penalty term.
[0101] At each time point, count the number of passengers in the carriages of multiple elevators; take the product of the rated passenger capacity of the elevator and a preset rated coefficient as the elevator overweight standard value; if the number of passengers in the carriage is greater than the elevator overweight standard value, calculate the difference between the number of passengers in the carriage and the elevator overweight standard value as the first elevator overweight value; add up all the first elevator overweight values to obtain the second elevator overweight value; take the opposite of the product of the second elevator overweight value and the congestion coefficient to obtain the carriage congestion penalty term.
[0102] Take the sum of the elevator smoothness penalty term and the carriage congestion penalty term as the third sub-reward value representing the comfort of passengers taking the elevator.
[0103] Exemplarily, the third sub-reward value is expressed as:
[0104]
[0105] In the formula, R comfort is the third sub-reward value, is the elevator smoothness penalty term, is the carriage congestion penalty term, RMS a is the root mean square of acceleration, ε is the set of elevators, T is the end point, a ω (t) is the acceleration of the elevator, VDV is the first vibration dose value, k3 is the preset comfort penalty coefficient, k3 = 1.2, 0.4 is the preset vibration coefficient, k4 is the carriage congestion penalty term, k4 = 0.5, N max is the rated passenger capacity of the elevator, N pass is the number of passengers in the carriage.
[0106] Step 408: In the decision layer, calculate the total reward value for the decision strategy based on the first sub-reward value, the second sub-reward value, and the third sub-reward value.
[0107] In this embodiment, by combining each sub-reward value (measuring waiting time, energy consumption, and comfort respectively) with preset weights, a comprehensive total reward value is obtained. This total reward value can effectively guide the decision layer to optimize the decision strategy of elevator dispatching, ensuring that the elevator group control system achieves the best balance between the experience of passengers taking the elevator and energy use under different circumstances.
[0108] Specifically, in the decision layer, add the product of the first sub-reward value and the preset first weight, the product of the second sub-reward value and the preset second weight, and the product of the third sub-reward value and the preset third weight to obtain a comprehensive reward index; integrate the product of the comprehensive reward index and the preset discount factor between two adjacent executions of the decision strategy to obtain the total reward value.
[0109] Exemplarily, the total reward value is expressed as:
[0110]
[0111] r τ = ω1R wait + ω2Renergy + ω3R comfort
[0112] In the formula, R(s,a) is the total reward value, t last is the moment of the previous execution decision strategy, t current is the moment of the current execution decision strategy, (t, t1), (t1, t2),...., (t n, t') are all time periods divided by the time points when non-decision strategies occur, e -β(τ-t) is the preset discount factor, r τ is the comprehensive reward index, ω1 is the first weight, ω2 is the second weight, ω3 is the third weight, R wait is the first sub-reward value, Renergy is the second sub-reward value, R comfort is the third sub-reward value.
[0113] In an embodiment of the present invention, in the perception layer, passenger flow data of multiple elevators is collected through sensor data in the elevator (such as passenger counters, speed and position sensors), floor button records, wireless network and video surveillance data, control system logs, external environment information, and smartphone APP data.
[0114] In the decision layer, the K-means algorithm is used to cluster the passenger flow data to obtain the elevator passenger flow pattern; the elevator passenger flow pattern includes morning upward peak, lunch upward peak, lunch downward peak, evening downward peak, normal inter-floor, and idle; in the decision layer, the first weight, the second weight, and the third weight are all adjusted to values adapted to the passenger flow pattern.
[0115]
[0116] ω2 = 1 - ω1, ω3 = 0.2.
[0117] Among them, ω1 is the first weight, ω2 is the second weight, ω3 is the third weight, and the peak period refers to the morning upward peak, lunch upward peak, lunch downward peak, and evening downward peak.
[0118] If it is detected that several indicators such as the waiting time of passengers waiting for the elevator, the energy consumption of multiple elevators, and the comfort of passengers do not reach the expected optimization effect, their weight values are adjusted. For example, when the elevator waiting time of passengers exceeds the threshold, execute ω1 → ω1 + 0.1.
[0119] Step 409, update the proximal policy optimization under reinforcement learning in the decision layer according to the total reward value.
[0120] In this embodiment, the core of reinforcement learning lies in gradually improving the effectiveness of decision-making strategies through a trial-and-error and feedback mechanism. The total reward value, as the feedback signal in reinforcement learning, represents the overall performance of the group control system within a certain time interval. By using Proximal Policy Optimization (PPO) within the reinforcement learning framework, the group control system can gradually achieve coordination among multiple elevators by continuously adjusting weights and decision-making strategy selections, thereby reducing passengers' waiting time, lowering energy consumption, and enhancing passengers' comfort.
[0121] Exemplarily, before applying the proximal policy optimization under reinforcement learning to actual elevator group control, the proximal policy optimization under reinforcement learning can be trained first to ensure that it can reach a certain performance level in a simulated environment and has sufficient capabilities to handle challenges in actual applications.
[0122] Specifically, different passenger flow scenarios are simulated for multiple elevators through a preset digital twin simulation environment. In different passenger flow scenarios, curriculum learning is used to select influencing factors related to scheduling multiple elevators for the proximal policy optimization under reinforcement learning; the influencing factors include passengers' waiting time, energy consumption of multiple elevators, and passengers' comfort.
[0123] The PPO algorithm includes an Actor network (policy network) and a Critic network (value network)
[0124] The Actor network (policy network) contains 3 fully connected layers (256 - 128 - 64), and outputs the joint actions of elevator group control, that is, the discrete action probabilities of the target floor assignment strategy and the continuous action parameters of acceleration.
[0125] The Critic network (value network) contains 2 layers of Long Short-Term Memory network layers (LSTM) (128 units) and a fully connected layer, which evaluates the long-term value of the state information - decision strategy pair and guides the update of the decision strategy.
[0126] Introduce curriculum learning, and gradually transition from a simple scenario (single elevator) to a complex scenario and combine it with the proximal policy optimization under reinforcement learning for training. Curriculum learning is divided into 4 stages, which are represented by the following table:
[0127]
[0128] Stage 1 Basic Scheduling: ω1 = 1, only optimize passengers' waiting time, fix the elevator speed at the rated value, and let the agent learn the basic elevator dispatching logic;
[0129] Stage 2 Energy Consumption Introduction: ω1 = 0.7, ω2 = 0.3, add the energy consumption index of the elevator, allow speed adjustment, and explore the speed - energy consumption balance strategy;
[0130] Stage 3 full objective optimization: ω3 = 0.2, enabling the comfort index of passengers taking the elevator, and dynamically adjusting the weights corresponding to influencing factors;
[0131] Stage 4 anti-interference training: Randomly shield one elevator to simulate a fault and improve the system robustness.
[0132] The stage transition conditions are as follows:
[0133] 1. Success rate threshold: The total reward value of the current stage reaches the preset value;
[0134] 2. Convergence detection: The total reward value fluctuates < 5% for 100 consecutive iterations;
[0135] 3. Forced transition: If the maximum number of iterations is exceeded (for example, at most 5000 times in stage 1), forcefully enter the next stage.
[0136] In an embodiment of the present invention, there is provided a group control system applicable to multiple elevators in a building. The group control system includes a sensing layer, a decision-making layer, and an execution layer. In the sensing layer, sensors in each elevator are called to sense status information related to the waiting time of passengers, the energy consumption of multiple elevators, and the comfort of passengers in historical, current, and future tenses. In the decision-making layer, proximal policy optimization under reinforcement learning is performed based on the status information to select one action from a mixed action space as the decision-making strategy. The mixed action space includes discrete and continuous actions related to elevator call control for each elevator. In the execution layer, multiple elevators are controlled to execute the decision-making strategy to compress the waiting time of passengers, the energy consumption of multiple elevators, and improve the comfort of passengers. In the execution layer, the time points when non-decision-making strategies occur between two adjacent executions of the decision-making strategy are queried. In the decision-making layer, the first sub-reward value representing the waiting time of passengers at each time point is calculated for the decision-making strategy jointly in the dimensions of instantaneous and long-term perspectives. In the decision-making layer, the second sub-reward value representing the energy consumption of multiple elevators at each time point is calculated for the decision-making strategy jointly in the dimensions of start-stop and operation. In the decision-making layer, the third sub-reward value representing the comfort of passengers at each time point is calculated for the decision-making strategy jointly in the dimensions of the physiology and psychology of passengers and the vibration of the elevator. In the decision-making layer, the total reward value is calculated for the decision-making strategy based on the first sub-reward value, the second sub-reward value, and the third sub-reward value. The proximal policy optimization under reinforcement learning in the decision-making layer is updated according to the total reward value. By querying the time points, the decision-making strategy of elevator dispatching can be accurately evaluated and optimized within different time windows. The first sub-reward value reflects the waiting experience of passengers, the second sub-reward value aims to improve energy efficiency and reduce unnecessary energy waste, while the third sub-reward value focuses on the riding experience of passengers in the elevator. Integrating these three sub-reward values into the total reward value can achieve a balance among multiple objectives, ensuring both improved user satisfaction and reduced energy consumption. The proximal policy optimization method is used to update the decision-making strategy of elevator dispatching efficiently and smoothly. The dispatching strategy is continuously adjusted through reinforcement learning technology to achieve the adaptive optimization of the decision-making strategy of the group control system, ultimately improving the overall efficiency of elevator dispatching and the passenger experience.
[0137] Embodiment III
[0138] Figure 6 FIG. is a schematic structural diagram of an elevator group control device provided in Embodiment III of the present invention, which is applicable to a group control system of multiple elevators in a building. The group control system includes a sensing layer, a decision-making layer, and an execution layer, as Figure 6 shown. The device includes:
[0139] The status information perception module 601 is used to call sensors in each elevator in the perception layer to perceive status information related to the waiting time of elevator passengers, the energy consumption of multiple elevators, and the comfort of elevator passengers in historical, present, and future tenses;
[0140] The decision strategy selection module 602 is used to perform proximal policy optimization under reinforcement learning in the decision layer based on the status information to select one action from the mixed action space as the decision strategy; the mixed action space includes discrete and continuous actions related to elevator call control for each elevator;
[0141] The decision strategy execution module 603 is used to control multiple elevators to execute the decision strategy in the execution layer to compress the waiting time of elevator passengers, the energy consumption of multiple elevators, and improve the comfort of elevator passengers.
[0142] In an embodiment of the present invention, the status information includes a basic status matrix and an extended status quantity;
[0143] The rows of the basic status matrix are the floors in the building, and the columns include the floors in the building, the waiting time when calling the elevator upward, the waiting time when calling the elevator downward, the in-call labels of each elevator, the labels of the floors that are not stopped when moving upward, the labels of the floors that are not stopped when moving downward, and the labels of the stopped floors;
[0144] When the element with the row being the floor in the building and the column being the floor in the building is 0, the element is invalid. When the element with the row being the floor in the building and the column being the floor in the building is 1, the element indicates that the floor in the building belonging to the row is the departure floor and the floor in the building belonging to the column is the destination floor;
[0145] When the element with the row being the floor in the building and the column being the waiting time when calling the elevator upward is the first value, the first value indicates that the floor in the building belonging to the row is the departure floor, and the waiting time when calling the elevator upward belonging to the column is the waiting time of the elevator passenger when calling the elevator upward at the departure floor;
[0146] When the element with the row being the floor in the building and the column being the waiting time when calling the elevator downward is the second value, the second value indicates that the floor in the building belonging to the row is the departure floor, and the waiting time when calling the elevator downward belonging to the column is the waiting time of the elevator passenger when calling the elevator downward at the departure floor;
[0147] When the element of the label for the in - call of a certain elevator in a floor of the building in question is 0, the element is invalid. When the element of the label for the in - call of a certain elevator in a floor of the building in question is 1, the element indicates that the in - call signal triggered by the elevator belonging to the column points to the floor in the building belonging to the row.
[0148] When the element of the label for the floors that an elevator belonging to a certain column in the building does not stop at during upward movement is 0, the element is invalid. When the element of the label for the floors that an elevator belonging to a certain column in the building does not stop at during upward movement is 1, the element indicates that the elevator belonging to the column does not stop at the floors in the building belonging to the row during upward movement.
[0149] When the element of the label for the floors that an elevator belonging to a certain column in the building does not stop at during downward movement is 0, the element is invalid. When the element of the label for the floors that an elevator belonging to a certain column in the building does not stop at during downward movement is 1, the element indicates that the elevator belonging to the column does not stop at the floors in the building belonging to the row during downward movement.
[0150] When the element of the label for the floors where an elevator belonging to a certain column in the building stops is 0, the element is invalid. When the element of the label for the floors where an elevator belonging to a certain column in the building stops is 1, the element indicates that the elevator belonging to the column stops at the floors in the building belonging to the row during movement.
[0151] The extended state quantity includes a heat map representing the demand for the elevator in space - time, the cumulative energy consumption of multiple elevators, and the comfort index of elevator riders.
[0152] In an embodiment of the present invention, the state information perception module 601 includes:
[0153] A historical passenger flow information collection module, used to collect historical passenger flow information;
[0154] A distribution information prediction module, used to input the historical passenger flow information into a preset long - short - term memory network to predict the distribution information of each floor of the building as the destination floor within a future period of time, as the heat map representing the demand for the elevator in space - time;
[0155] A time - period definition module, used to define the time interval between the previous execution of the decision - making strategy and the current execution of the decision - making strategy as the time period;
[0156] An energy - consumption cumulative - quantity acquisition module, used to add up the energy consumption of multiple elevators during the time period to obtain the cumulative energy consumption of multiple elevators;
[0157] The root mean square calculation module is used to calculate the root mean square of the accelerations of multiple elevators when they stop at the floors of the building multiple times in the past, as the comfort index of passengers taking the elevator.
[0158] In an embodiment of the present invention, the hybrid action space includes discrete actions and continuous actions;
[0159] The discrete action is a one-dimensional matrix, and the columns of the one-dimensional matrix represent each elevator. When the value of a column is 1, it means that when a call signal is received, the elevator is dispatched to respond to the call signal and move to the floor in the building indicated by the call signal.
[0160] The continuous action includes adding a speed adjustment coefficient to the moving speed of the dispatched elevator within a preset first adjustment range, and adding an acceleration coefficient to the moving acceleration of the dispatched elevator within a preset second adjustment range.
[0161] In an embodiment of the present invention, the device further includes:
[0162] The time point query module is used to query the time points of non-decision strategies that occur between two adjacent executions of the decision strategy in the execution layer;
[0163] The first sub-reward value calculation module is used to jointly calculate the first sub-reward value representing the waiting time of passengers taking the elevator at each of these time points for the decision strategy in the decision layer in the dimensions of instantaneous and long-term;
[0164] The second sub-reward value calculation module is used to jointly calculate the second sub-reward value representing the energy consumption of multiple elevators at each of these time points for the decision strategy in the decision layer in the dimensions of start-stop and operation;
[0165] The third sub-reward value calculation module is used to jointly calculate the third sub-reward value representing the comfort of passengers taking the elevator at each of these time points for the decision strategy in the decision layer in the dimensions of the physiology and psychology of passengers taking the elevator and the vibration of the elevator;
[0166] The total reward value calculation module is used to calculate the total reward value for the decision strategy in the decision layer based on the first sub-reward value, the second sub-reward value, and the third sub-reward value;
[0167] The reinforcement learning update module is used to update the proximal policy optimization under reinforcement learning in the decision layer according to the total reward value.
[0168] In an embodiment of the present invention, the first sub-reward value calculation module includes:
[0169] The waiting elevator time statistical module is used to statistically calculate the waiting elevator time of each rider at each of the time points;
[0170] The first excess time calculation module is used to calculate the difference between the waiting elevator time and the preset waiting threshold as the first excess time if the waiting elevator time is greater than the preset waiting threshold;
[0171] The second excess time calculation module is used to square the ratio between the first excess time and the preset penalty threshold to obtain the second excess time;
[0172] The waiting time penalty term acquisition module is used to take the opposite of the sum of all the second excess times to obtain the waiting time penalty term;
[0173] The destination floor query module is used to query multiple destination floors that the rider will go to within a future period of time according to the heat map of the elevator in the status information at each of the time points;
[0174] The absolute distance calculation module is used to take the absolute value of the difference between the position of the floor where the elevator is located and the positions of each of the destination floors for each elevator to obtain a plurality of absolute distances;
[0175] The influence value acquisition module is used to take the reciprocal of the sum of each absolute distance and 1 for each absolute distance to obtain the influence value of the elevator on the destination floor;
[0176] The total influence value calculation module is used to add up each of the influence values to obtain the total influence value;
[0177] The future elevator scheduling reward value calculation module is used to add up the total influence values corresponding to each elevator to obtain the future elevator scheduling reward value;
[0178] The first sub-reward value acquisition module is used to add the product of the waiting time penalty term, the future elevator scheduling reward value and the dynamic attenuation factor as the first sub-reward value representing the waiting time of the rider;
[0179] In an embodiment of the present invention, the second sub-reward value calculation module includes:
[0180] The elevator acceleration value calculation module is used to add the squared value of the acceleration of the elevator going up and the squared value of the acceleration of the elevator going down at each of the time points to obtain the elevator acceleration value;
[0181] The total acceleration value calculation module is used to add up the elevator acceleration values corresponding to each elevator to obtain the total acceleration value;
[0182] An elevator start - stop energy consumption calculation module, which is used to take the opposite of the product between the total acceleration value and the elevator start - stop coefficient to obtain the elevator start - stop energy consumption;
[0183] An elevator running speed value calculation module, which is used to square the difference between the running speed of the elevator and the preset elevator running speed at each of the time points to obtain the elevator running speed value;
[0184] A total running speed value calculation module, which is used to add up the elevator running speed values corresponding to each elevator to obtain the total running speed value;
[0185] An elevator running energy consumption calculation module, which is used to take the opposite of the product between the total running speed value and the deviation running speed penalty coefficient to obtain the elevator running energy consumption;
[0186] A second sub - reward value acquisition module, which is used to use the sum of the elevator start - stop energy consumption and the elevator running energy consumption as the second sub - reward value representing the energy consumption of multiple elevators;
[0187] In an embodiment of the present invention, the third sub - reward value calculation module includes:
[0188] A first acceleration integral acquisition module, which is used to integrate the square of the acceleration of the elevator from 0 to the time point at each of the time points to obtain the first acceleration integral;
[0189] A second acceleration integral acquisition module, which is used to perform a power operation on the ratio between the first acceleration integral and the time point to obtain the second acceleration integral;
[0190] An acceleration root - mean - square calculation module, which is used to add up the second acceleration integrals corresponding to each elevator to obtain the acceleration root - mean - square;
[0191] A third acceleration integral acquisition module, which is used to integrate the fourth power of the acceleration of the elevator from 0 to the time point at each of the time points to obtain the third acceleration integral;
[0192] A fourth acceleration integral acquisition module, which is used to perform a power operation on the third acceleration integral to obtain the fourth acceleration integral;
[0193] A first vibration dose value acquisition module, which is used to add up the fourth acceleration integrals corresponding to each elevator to obtain the first vibration dose value;
[0194] An elevator smoothness value calculation module, which is used to add up the product of the acceleration root - mean - square, the first vibration dose value and a preset vibration coefficient as the elevator smoothness value;
[0195] An elevator smoothness penalty term acquisition module, configured to take the opposite of the product of the elevator smoothness value and a preset comfort penalty coefficient to obtain an elevator smoothness penalty term;
[0196] A car passenger number statistics module, configured to count the number of passengers in the cars of multiple elevators at each of the time points;
[0197] An elevator overweight standard value setting module, configured to use the product of the elevator rated passenger capacity and a preset rated coefficient as the elevator overweight standard value;
[0198] A first elevator overweight value calculation module, configured to calculate the difference between the number of passengers in the car and the elevator overweight standard value as the first elevator overweight value if the number of passengers in the car is greater than the elevator overweight standard value;
[0199] A second elevator overweight value calculation module, configured to add up all the first elevator overweight values to obtain a second elevator overweight value;
[0200] A car congestion penalty term acquisition module, configured to take the opposite of the product of the second elevator overweight value and a congestion coefficient to obtain a car congestion penalty term;
[0201] A third sub - reward value acquisition module, configured to use the sum of the elevator smoothness penalty term and the car congestion penalty term as a third sub - reward value representing the comfort of elevator riders;
[0202] In an embodiment of the present invention, the total reward value calculation module includes:
[0203] A comprehensive reward index calculation module, configured to add up the product of the first sub - reward value and a preset first weight, the product of the second sub - reward value and a preset second weight, and the product of the third sub - reward value and a preset third weight in the decision layer to obtain a comprehensive reward index;
[0204] A total reward value acquisition module, configured to integrate the product of the comprehensive reward index and a preset discount factor between two adjacent executions of the decision strategy to obtain a total reward value.
[0205] In an embodiment of the present invention, the device further includes:
[0206] A passenger flow data acquisition module, configured to collect passenger flow data for multiple elevators in the perception layer;
[0207] A passenger flow pattern clustering module, configured to cluster the passenger flow data of each elevator in the decision layer to obtain passenger flow patterns; the passenger flow patterns include morning upward peak, lunch upward peak, lunch downward peak, evening downward peak, normal inter - floor, and idle;
[0208] A weight value adjustment module, configured to adjust the first weight, the second weight, and the third weight to values adapted to the passenger flow pattern in the decision-making layer.
[0209] In an embodiment of the present invention, the device further includes:
[0210] A passenger flow scenario simulation module, configured to simulate different passenger flow scenarios for multiple elevators through a preset digital twin simulation environment;
[0211] An influencing factor determination module, configured to use curriculum learning to optimize the proximal policy in reinforcement learning to select influencing factors related to the scheduling of multiple elevators in different passenger flow scenarios; the influencing factors include the waiting time of passengers taking the elevator, the energy consumption of multiple elevators, and the comfort of passengers taking the elevator.
[0212] The elevator group control device provided by the embodiment of the present invention can execute the elevator group control method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the elevator group control method.
[0213] Embodiment 4
[0214] Refer to Figure 7 , which shows a schematic structural diagram of a computer device provided by an embodiment of the present invention. The computer device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a blade server, a mainframe computer, and other suitable computers. The computer device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0215] As Figure 7 shown, the computer device 10 includes at least one processor 11, and a memory communicatively connected to at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by at least one processor, and the processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the computer device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0216] Multiple components in the computer device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the computer device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0217] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the elevator group control method.
[0218] In some embodiments, the elevator group control method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the computer device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the elevator group control method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the elevator group control method by any other suitable means (e.g., by means of firmware).
[0219] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, and the programmable processor can be a special or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0220] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0221] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0222] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, voice input, or tactile input).
[0223] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0224] A computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs that run on respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0225] Example Five
[0226] The embodiment of the present invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the elevator group control method provided in any embodiment of the present invention.
[0227] In the process of implementing the computer program product, the computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0228] It should be understood that various forms of the processes shown above can be used, with steps reordered, added or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0229] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An elevator group control method, characterized in that, A group control system applied to multiple elevators in a building, the group control system including a sensing layer, a decision-making layer and an execution layer; the method includes: In the sensing layer, call sensors in each of the elevators to sense status information related to the waiting time of elevator passengers, the energy consumption of multiple elevators, and the comfort of elevator passengers in the historical, present, and future tenses; In the decision-making layer, perform proximal policy optimization under reinforcement learning based on the status information to select one action in the mixed action space as the decision-making strategy; the mixed action space includes discrete and continuous actions related to elevator call control for each elevator; In the execution layer, control multiple elevators to execute the decision-making strategy to compress the waiting time of elevator passengers, the energy consumption of multiple elevators, and improve the comfort of elevator passengers.
2. The method according to claim 1, wherein The status information includes a basic status matrix and an extended status quantity; The rows of the basic status matrix are the floors in the building, and the columns include the floors in the building, the waiting time when calling for an upward elevator, the waiting time when calling for a downward elevator, the internal call labels of each elevator, the labels of floors that are not stopped when moving upward, the labels of floors that are not stopped when moving downward, and the labels of stopped floors; When the element with the row being the floor in the building and the column being the floor in the building is 0, the element is invalid; when the element with the row being the floor in the building and the column being the floor in the building is 1, the element indicates that the floor in the building belonging to the row is the departure floor and the floor in the building belonging to the column is the destination floor; When the element with the row being the floor in the building and the column being the waiting time when calling for an upward elevator is a first value, the first value indicates that the floor in the building belonging to the row is the departure floor, and the waiting time when calling for an upward elevator belonging to the column is the waiting time of the elevator passenger when calling for an upward elevator at the departure floor; When the element with the row being the floor in the building and the column being the waiting time when calling for a downward elevator is a second value, the second value indicates that the floor in the building belonging to the row is the departure floor, and the waiting time when calling for a downward elevator belonging to the column is the waiting time of the elevator passenger when calling for a downward elevator at the departure floor; When the element with the row being the floor in the building and the column being the internal call label of a certain elevator is 0, the element is invalid; when the element with the row being the floor in the building and the column being the internal call label of a certain elevator is 1, the element indicates that the internal call signal triggered by the elevator belonging to the column points to the floor in the building belonging to the row; When the element with the row being the floor in the building and the column being the label of floors that are not stopped when a certain elevator moves upward is 0, the element is invalid; when the element with the row being the floor in the building and the column being the label of floors that are not stopped when a certain elevator moves upward is 1, the element indicates that the elevator belonging to the column does not stop at the floor in the building belonging to the row when moving upward; When the element of the label of the floors in the building where the behavior is located and the floors that the certain elevator does not stop at during downward movement is 0, the element is invalid; when the element of the label of the floors in the building where the behavior is located and the floors that the certain elevator stops at during downward movement is 1, the element indicates that the elevator belonging to the column does not stop at the floors in the building belonging to the row during downward movement; When the element of the label of the floors in the building where the behavior is located and the floors that the certain elevator stops at is 0, the element is invalid; when the element of the label of the floors in the building where the behavior is located and the floors that the certain elevator stops at is 1, the element indicates that the elevator belonging to the column stops at the floors in the building belonging to the row during movement; The extended state quantity includes a heat map characterizing the demand for the elevator in space-time, the cumulative amount of energy consumption of multiple elevators, and the comfort index of elevator passengers.
3. The method according to claim 2, wherein In the perception layer, sensors in each elevator are called to perceive state information related to the waiting time of elevator passengers, the energy consumption of multiple elevators, and the comfort of elevator passengers in the historical, present, and future tenses, including: Collect historical passenger flow information; Input the historical passenger flow information into a pre-set long short-term memory network to predict the distribution information of the floors of the building as the destination floors in a future period of time, as a heat map characterizing the demand for the elevator in space-time; Define the time interval between the previous execution of the decision-making strategy and the current execution of the decision-making strategy as the time period; Add up the energy consumption of multiple elevators during the time period to obtain the cumulative amount of energy consumption of multiple elevators; Calculate the root mean square of the accelerations of multiple elevators when they stop at the floors of the building multiple times in the past as the comfort index of elevator passengers.
4. The method according to claim 1, wherein The hybrid action space includes discrete actions and continuous actions; The discrete action is a one-dimensional matrix, and the columns of the one-dimensional matrix represent each elevator. When the value of the column is 1, it means that when a call signal is received, the elevator is dispatched to respond to the call signal and move to the floor in the building indicated by the call signal; The continuous actions include adding a speed adjustment coefficient to the moving speed of the dispatched elevator within a preset first adjustment range and adding an acceleration coefficient to the moving acceleration of the dispatched elevator within a preset second adjustment range.
5. The method according to claim 1, characterized in that, It also includes: In the execution layer, query the time points of non-decision-making strategies that occur between two adjacent executions of the decision-making strategy; In the decision-making layer, jointly calculate the first sub-reward value characterizing the waiting time of elevator passengers at each of these time points for the decision-making strategy in the dimensions of instant and long-term perspective; In the decision-making layer, jointly calculate the second sub-reward value characterizing the energy consumption of multiple elevators at each of these time points for the decision-making strategy in the dimensions of start-stop and operation; In the decision-making layer, jointly calculate the third sub-reward value characterizing the comfort of elevator passengers at each of these time points for the decision-making strategy in the dimensions of the physiology and psychology of elevator passengers and the vibration of the elevator; In the decision-making layer, the total reward value of the decision-making strategy is calculated based on the first sub-reward value, the second sub-reward value, and the third sub-reward value. Based on the total reward value, the proximal policy optimization under reinforcement learning in the decision-making layer is updated.
6. The method according to claim 5, characterized in that, In the decision-making layer, the first sub-reward value representing the waiting time of the elevator riders at each time point is jointly calculated for the decision-making strategy in the dimensions of instantaneous and long-term perspectives, including: At each time point, the waiting time of each elevator rider is counted. If the waiting time of the elevator is greater than the preset waiting threshold, the difference between the waiting time of the elevator and the waiting threshold is calculated as the first excess time. The square of the ratio between the first excess time and the preset penalty threshold is taken to obtain the second excess time. The opposite value of the sum of all the second excess times is taken to obtain the waiting time penalty term. At each time point, according to the heat map of the elevator in the state information, multiple destination floors that the elevator riders will go to in a future period are queried. For each elevator, the absolute value of the difference between the position of the floor where the elevator is located and the positions of each destination floor is taken to obtain multiple absolute distances. For each absolute distance, the reciprocal of the sum of the absolute distance and 1 is taken to obtain the influence value of the elevator on the destination floor. All the influence values are added up to obtain the total influence value. The total influence values corresponding to each elevator are added up to obtain the future elevator scheduling reward value. The sum of the product of the waiting time penalty term, the future elevator scheduling reward value, and the dynamic decay factor is used as the first sub-reward value representing the waiting time of the elevator riders. In the decision-making layer, the second sub-reward value representing the energy consumption of multiple elevators at each time point is jointly calculated for the decision-making strategy in the dimensions of start-stop and operation, including: At each time point, the square value of the acceleration of the elevator going up and the square value of the acceleration of the elevator going down are added up to obtain the elevator acceleration value. The elevator acceleration values corresponding to each elevator are added up to obtain the total acceleration value. The opposite value of the product of the total acceleration value and the elevator start-stop coefficient is taken to obtain the elevator start-stop energy consumption. At each time point, the square of the difference between the running speed of the elevator and the preset elevator running speed is taken to obtain the elevator running speed value. The elevator running speed values corresponding to each elevator are added up to obtain the total running speed value. The opposite value of the product of the total running speed value and the deviation from the running speed penalty coefficient is taken to obtain the elevator running energy consumption. The sum of the elevator start-stop energy consumption and the elevator running energy consumption is used as the second sub-reward value representing the energy consumption of multiple elevators. In the decision-making layer, the third sub-reward value representing the comfort of the elevator riders at each time point is jointly calculated for the decision-making strategy in the dimensions of the physiology and psychology of the elevator riders and the vibration of the elevator, including: At each of the time points, integrate the squared acceleration of the elevator from 0 to the time point to obtain a first acceleration integral; Perform a power operation on the ratio between the first acceleration integral and the time point to obtain a second acceleration integral; Add the second acceleration integrals corresponding to each of the elevators to obtain a root mean square acceleration; At each of the time points, integrate the fourth power of the acceleration of the elevator from 0 to the time point to obtain a third acceleration integral; Perform a power operation on the third acceleration integral to obtain a fourth acceleration integral; Add the fourth acceleration integrals corresponding to each elevator to obtain a first vibration dose value; Add the product of the root mean square acceleration, the first vibration dose value, and a preset vibration coefficient as the elevator smoothness value; Take the opposite of the product of the elevator smoothness value and a preset comfort penalty coefficient to obtain an elevator smoothness penalty term; At each of the time points, count the number of passengers in the carriages of multiple elevators; Take the product of the rated passenger capacity of the elevator and a preset rated coefficient as the elevator overweight standard value; If the number of passengers in the carriage is greater than the elevator overweight standard value, calculate the difference between the number of passengers in the carriage and the elevator overweight standard value as a first elevator overweight value; Add all the first elevator overweight values to obtain a second elevator overweight value; Take the opposite of the product of the second elevator overweight value and a crowding coefficient to obtain a carriage crowding penalty term; Take the sum of the elevator smoothness penalty term and the carriage crowding penalty term as a third sub-reward value representing the comfort of passengers taking the elevator; In the decision layer, calculate the total reward value for the decision strategy based on the first sub-reward value, the second sub-reward value, and the third sub-reward value, including: In the decision layer, add the product of the first sub-reward value and a preset first weight, the product of the second sub-reward value and a preset second weight, and the product of the third sub-reward value and a preset third weight to obtain a comprehensive reward index; Integrate the product of the comprehensive reward index and a preset discount factor between two adjacent executions of the decision strategy to obtain the total reward value.
7. The method according to claim 6, wherein Further include: In the perception layer, collect passenger flow data for multiple elevators; In the decision layer, cluster the passenger flow data of each elevator to obtain a passenger flow pattern; the passenger flow pattern includes morning upward peak, lunch upward peak, lunch downward peak, evening downward peak, normal inter-floor, and idle; In the decision layer, adjust the first weight, the second weight, and the third weight to values adapted to the passenger flow pattern.
8. The method according to any one of claims 1 to 7, characterized in that, Further include: Simulate different passenger flow scenarios for multiple elevators through a preset digital twin simulation environment; In different passenger flow scenarios, use curriculum learning to optimize the proximal policy in reinforcement learning to select influencing factors related to the scheduling of multiple elevators; The influencing factors include the waiting time of passengers taking the elevator, the energy consumption of multiple elevators, and the comfort of passengers taking the elevator.
9. A computer device, characterized in that, The computer device includes: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the elevator group control method according to any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the elevator group control method according to any one of claims 1-8.
Citation Information
Cited By
Elevator control method and device and storage medium
CN120774297A
Elevator group control dynamic scheduling method and system based on reinforcement learning, electronic equipment and storage medium
CN120964536A
Multi-elevator intelligent scheduling method and device based on heterogeneous graph and PPO algorithm
CN122009927A