A smart construction scheduling method based on the Internet of Things

By modeling the construction lift scheduling problem as a Markov decision-making process and using stratified reinforcement learning and genetic algorithms for collaborative optimization, the problem of low intelligence in construction lift scheduling is solved, and an efficient and energy-saving scheduling strategy is achieved.

CN119476881BActive Publication Date: 2025-05-16AVIC CONSTR GRP CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510052226.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-16
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

The existing construction lift scheduling methods are low in intelligence and lack the ability to perceive and adapt to the real-time environmental state, making it difficult to meet the scheduling efficiency while taking into account energy saving and consumption reduction.

Method used

The construction lift scheduling problem is modeled as a Markov decision-making process, and a layered reinforcement learning algorithm is used, and the scheduling strategy is coordinated and optimized in combination with genetic algorithms to generate independent construction of scheduling strategies and dynamically optimized.

Benefits of technology

The intelligent level and comprehensive performance of construction lift scheduling have been improved, more efficient resource utilization and energy conservation have been achieved, and the construction needs have been adapted to dynamically changing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119476881B_ABST
    Figure CN119476881B_ABST
Patent Text Reader

Abstract

The present application discloses a smart construction scheduling method based on the Internet of Things, which relates to the field of smart construction, including: obtaining call extension data of each floor of a construction elevator, the call extension data including floor door status data and call data; converting the floor door status data using a binary processing algorithm to obtain floor door status parameters of each floor; counting the call data using a counter to obtain call status parameters of each floor; taking the floor door status parameters and the call status parameters as environmental states, modeling the construction elevator scheduling problem as a Markov decision process, taking the scheduling strategy as an action, taking the energy consumption of the construction elevator as a reward function, and performing reinforcement learning through a Q-learning algorithm to obtain a scheduling strategy; in view of the low degree of intelligence in the scheduling of construction elevators in the prior art, the present application uses a hierarchical reinforcement learning algorithm to obtain a global scheduling strategy and a local scheduling strategy for the scheduling problem of the construction elevator, and uses a genetic algorithm for collaborative optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of smart construction, and in particular to a smart construction scheduling method based on the Internet of Things. Background Art

[0002] As the scale of construction projects continues to expand, the demand for vertical transportation at construction sites is increasing. As an indispensable vertical transportation equipment in modern construction, construction elevators play a key role in ensuring construction progress and improving work efficiency. However, the scheduling of traditional construction elevators mainly relies on manual experience, and there are problems such as low scheduling efficiency and serious energy waste, which makes it difficult to meet the needs of modern construction.

[0003] In recent years, the development of Internet of Things technology has provided new ideas for solving the above problems. By deploying various sensors and communication equipment on construction elevators, data such as equipment operating status and floor call requirements can be collected in real time, laying a data foundation for realizing intelligent scheduling. However, how to use this data to make scientific scheduling decisions is still a problem that needs to be solved urgently.

[0004] Existing construction elevator scheduling methods mainly include rule-based methods and optimization-based methods. Rule-based methods schedule construction elevators according to pre-set scheduling rules, such as sequential scheduling, shortest search time priority, etc. This type of method is simple to implement, but lacks flexibility and adaptability, and it is difficult to cope with complex and changing construction environments. Optimization-based methods establish mathematical models and use heuristic algorithms or exact algorithms to solve the optimal scheduling solution. This type of method takes into account multiple optimization objectives, such as minimizing energy consumption and minimizing waiting time, but the solution process is complex and real-time performance is difficult to guarantee.

[0005] In general, the existing construction elevator dispatching methods are not highly intelligent and lack the ability to perceive and adapt to real-time environmental conditions, making it difficult to meet dispatching efficiency while taking into account energy conservation and consumption reduction. Therefore, a new type of intelligent dispatching method is urgently needed, which uses advanced artificial intelligence technology to analyze and learn massive IoT data, independently build dispatching strategies, and dynamically optimize according to environmental changes to improve the intelligence level and comprehensive performance of construction elevator dispatching. Summary of the invention

[0006] In response to the problem of low intelligence level in construction elevator scheduling in the prior art, the present application provides an intelligent construction scheduling method based on the Internet of Things, models the construction elevator scheduling problem as a Markov decision process, and uses a hierarchical reinforcement learning algorithm to obtain the global scheduling strategy and the local scheduling strategy, and uses a genetic algorithm to collaboratively optimize the scheduling strategies at the two levels, thereby improving the intelligence level of scheduling.

[0007] The purpose of this application is achieved through the following technical solutions.

[0008] The present application provides a smart construction scheduling method based on the Internet of Things, including: obtaining the call extension data of each floor of the construction elevator, the call extension data includes the floor door status data and the call data; the floor door status signal reflects the switch status of the floor door, and the call data reflects the call request of the floor; the floor door status data is converted by a binary processing algorithm to obtain the floor door status parameters of each floor; wherein the floor door status parameters include the floor door opening flag and the floor door closing flag, when the value of the floor door status data is greater than the threshold, the floor door opening flag is set to true, otherwise the floor door closing flag is set to true. The call data is counted by a counter to obtain the call status parameters of each floor, the call status parameters include the number of calls and the last call time; the floor door status parameters and the call status parameters are used as the environmental state, the construction elevator scheduling problem is modeled as a Markov decision process, the scheduling strategy is used as the action, the energy consumption of the construction elevator is used as the reward function, and the scheduling strategy is obtained by reinforcement learning through the Q-learning algorithm.

[0009] Further, the floor door status data is converted by a binary processing algorithm to obtain the floor door status parameters of each floor, including: digital processing of the acquired floor door status data, converting the floor door status data into a digital signal, and obtaining a floor door status digital quantity; setting a floor door opening threshold and a floor door closing threshold by a hysteresis comparator, and setting a floor door hysteresis interval according to the floor door opening threshold and the floor door closing threshold; judging whether the floor door status digital quantity is in the floor door hysteresis interval, if so, maintaining the current floor door state unchanged, if not, setting the floor door status flag: when the floor door status digital quantity is greater than the floor door opening threshold, setting the floor door opening flag of the corresponding floor to true, and setting the floor door closing flag to false; when the floor door status digital quantity is less than the floor door closing threshold, setting the floor door opening flag of the corresponding floor to false, and setting the floor door closing flag to true; generating the floor door status parameters of each floor according to the set floor door opening flag and floor door closing flag.

[0010] Furthermore, the call data is counted by a counter to obtain call status parameters of each floor. The call status parameters include the number of calls and the last call time, including: obtaining the number of calls by accumulating the call data; obtaining the last call time by obtaining the timestamp of the most recent call data.

[0011] Furthermore, the door state parameters and call state parameters are used as environmental states, the construction elevator scheduling problem is modeled as a Markov decision process, the scheduling strategy is used as the action, the energy consumption of the construction elevator is used as the reward function, and the Q-learning algorithm is used for reinforcement learning to obtain the scheduling strategy, including: constructing a state space S according to the door state parameters and the call state parameters; according to the state space S of the Markov decision process; and setting the action space A according to the user's scheduling needs; wherein the action space A includes responding to call requests, changing the running direction, stopping or skipping a certain floor; according to the state space S and the action space A , define the transition probability P between states; where the transition probability P represents the probability that the environment state is transferred to the next state after a certain scheduling action is taken in the current state; the energy consumption of the construction elevator is used as the reward function R; according to the state space S, action space A, transition probability P and reward function R, a Markov decision process MDP is constructed; based on the constructed Markov decision process MDP, a hierarchical reinforcement learning algorithm is used to perform reinforcement learning at the global scheduling layer and the local scheduling layer to obtain the global scheduling strategy and the local scheduling strategy; based on the global scheduling strategy and the local scheduling strategy, the final scheduling strategy is obtained through collaborative optimization.

[0012] Furthermore, a global scheduling strategy is obtained, including: using the state space S of the Markov decision process MDP as the input state feature of the DQN algorithm; wherein the state feature includes the layer gate state parameter and the call state parameter; constructing a multi-layer feedforward neural network as the Q network of the DQN algorithm, the Q network takes the state feature as input, and outputs the Q value of each scheduling action taken in the corresponding state; wherein the Q value represents the long-term cumulative return expectation of liking to take a certain scheduling action in the current state; in the Q network training process, according to the Bellman equation, the maximum Q value and immediate reward of the next state are used to calculate the target Q value of the current state-action pair; wherein the target Q value is used as a supervisory signal for Q network training to optimize the Q network parameters; using the mean square error loss function, calculating the difference between the estimated Q value output by the Q network and the target Q value, and updating the Q network parameters in the direction of the loss function gradient through the stochastic gradient descent algorithm; using the trained Q network, selecting the scheduling action with the largest Q value in each state as the global scheduling strategy in the corresponding state.

[0013] Furthermore, the mean square error loss function is expressed as follows: L ( θ ) = E [ ( Q ( s , a ; θ ) − y ) 2 ] ,in, represents the mean square error loss function; E [*] represents the expectation operator and represents the average of the training data. It represents the estimated Q value output by the Q network when taking action a in state s, which is a function of the network parameter θ. s represents the state variable, which represents a state in the Markov decision process MDP, and is composed of the layer gate state parameter and the call state parameter. a represents the action variable, which represents a scheduling action taken in state s. θ represents the parameters of the Q network, including the weights and biases of the neurons in each layer. y represents the target Q value, which is calculated according to the Bellman equation and is expressed as: ; Among them: r represents the immediate reward, which means the reward value obtained after taking action a in state s. γ represents the discount factor, which ranges from [0, 1] and represents the discount ratio of future rewards. Represents the next state variable, which indicates the new state to which the action a is transferred after taking action a in state s. Represents the next action variable, which represents the action taken in the next state s'. Represents the target parameters of the Q network, which is used to calculate the target Q value, and is usually a delayed updated version of the Q network parameters θ.

[0014] Furthermore, the local scheduling strategy includes: extracting a feature subset reflecting the state of a single construction elevator according to the state space S of the Markov decision process MDP; wherein the feature subset includes the current floor, running direction, speed and floor door state of the construction elevator; constructing a multi-layer feedforward neural network as the policy network of the PPO algorithm, the policy network takes the feature subset as input, and outputs the probability distribution of taking various scheduling actions under the current state; according to the current policy network parameters, a set of state-action-reward data is collected in the state space S and action space A of the MDP; during the data collection process, according to the probability distribution output by the policy network, the scheduling action is selected by random sampling; according to the collected state-action-reward data, the state-action-reward data is collected, and the state-action-reward data is collected. Action-reward data, calculate the importance weight between the current policy network and the data collection strategy; where the data collection strategy refers to the strategy of randomly sampling scheduling actions according to the action probability distribution; the importance weight is used to correct the deviation of data distribution; according to the importance weight, construct the objective function of the PPO algorithm; the objective function contains policy loss, value function loss and entropy regularization term; use the stochastic gradient ascent algorithm to update the parameters of the policy network in the gradient direction of the objective function; use the trained policy network to select a scheduling action from the probability distribution of the scheduling action output by the policy network through random sampling under each feature subset, and the scheduling action sequence selected by multiple samplings constitutes a local scheduling strategy.

[0015] Furthermore, the objective function of the PPO algorithm is expressed as follows: J ( θ ) = E [ r ( θ ) × A ( s , a ) − β × ( V ( s ; θ ) − V t arg et ) 2 + η × H ( π (*| s ; θ ))] ; The definitions of each parameter are as follows: Represents the objective function of the PPO algorithm and represents the target value under the policy network parameter θ. E [*] Represents the expectation operator, which represents the average of the collected data. Represents the importance weight, which represents the ratio between the current policy network and the data acquisition policy. The calculation formula is: ;in: Represents the probability that the current policy network takes action a in state s. It represents the probability that the data collection strategy takes action a in state s. Represents the advantage function, which represents the advantage value of taking action a in state s. The calculation formula is: ;in: represents the state-action value function, which represents the long-term cumulative reward of taking action a in state s. represents the state value function, which represents the long-term cumulative return of state s. Represents the weight coefficient of the value function loss, which is used to balance the ratio of policy loss and value function loss. Represents the output of the current value function network in state s, which is a function of the network parameters θ. represents the objective value function, usually an estimate of the Monte Carlo return or TD target. Represents the weight coefficient of the entropy regularization term, which is used to encourage the exploration of the strategy. It represents the entropy of the action probability distribution output by the policy network in state s, indicating the uncertainty of the policy.

[0016] The PPO algorithm maximizes the objective function , while optimizing the parameters of the policy network and the value function network. Among them, the policy loss term Encourage the policy network to generate actions with higher advantage values, and the value function loss term Make the prediction of the value function network close to the actual cumulative return, entropy regularization term The policy network is encouraged to output action probability distributions with higher entropy, thus improving the exploration ability of the policy.

[0017] Furthermore, according to the global scheduling strategy and the local scheduling strategy, the final scheduling strategy is obtained through collaborative optimization, including: encoding the global scheduling strategy into a chromosome, the global scheduling strategy chromosome is composed of scheduling actions under each state; taking multiple local scheduling strategies as a population, each local scheduling strategy represents an individual in the population, and the individual chromosome is composed of a scheduling action sequence under the feature subset of the corresponding local scheduling strategy; constructing a fitness function, the fitness function reflects the consistency of the local scheduling strategy and the global scheduling strategy, and taking the cumulative reward obtained by the local scheduling strategy running in the environment as the fitness value; using a genetic algorithm to select and cross the individuals in the local scheduling strategy population and mutation, and select a group of individuals from the current population as parents through the roulette wheel selection operator with the fitness value as the weight; the parent individuals are randomly combined in pairs, and the chromosomes are cross-recombined through the crossover operator to generate new offspring individuals; when the genetic algorithm meets the preset iteration termination condition, the individual with the highest fitness value is selected from the last generation of the population, and the selected individual chromosome is decoded into a local scheduling strategy, which is output as the local scheduling strategy that is most coordinated with the global scheduling strategy; the selected local scheduling strategy is combined with the obtained global scheduling strategy to obtain the final scheduling strategy; among them, the global scheduling strategy is responsible for selecting the target floor, and the local scheduling strategy is responsible for controlling the operation of a single construction elevator.

[0018] Furthermore, a fitness function is constructed, and the expression of the fitness function is as follows: ;in, Represents an individual The fitness function represents the local scheduling strategy Consistency with the global scheduling policy. Represents the i-th individual in the population and represents a local scheduling strategy. Represents the discount factor, with a value range of [0, 1], which represents the discount ratio of future rewards, and has the same definition as above. t represents the time step variable, which represents the local scheduling strategy The tth time step of running in the environment. Represents the local scheduling strategy The instant reward obtained at the tth time step. Fitness function Calculate local scheduling strategy The cumulative reward for running in the environment for T time steps, where is the discount factor, which discounts future rewards. The higher the cumulative reward, the better the local scheduling strategy. The better the consistency with the global scheduling strategy, the greater the fitness value.

[0019] In the optimization process of genetic algorithm, the fitness function As an individual The evaluation index is used for selection, crossover and mutation operations. Individuals with higher fitness values ​​have a greater probability of being selected as parents and generating new offspring individuals through crossover and mutation operations. After multiple iterations, the average fitness value of individuals in the population continues to increase, and eventually converges to the individual with the highest fitness value, that is, the local scheduling strategy that is most coordinated with the global scheduling strategy. Through the fitness function Establish the association between local scheduling strategies and global scheduling strategies, use the optimization capability of genetic algorithms to search for the best individuals from multiple local scheduling strategies, achieve coordinated optimization of local scheduling strategies and global scheduling strategies, and obtain the scheduling strategy combination with the best overall performance.

[0020] Compared with the prior art, the advantages of this application are:

[0021] This application models the construction elevator scheduling problem as a Markov decision process and uses a hierarchical reinforcement learning algorithm to solve it, so that scheduling decisions can be generated autonomously based on real-time status information, no longer relying on manual experience and fixed rules, and has stronger adaptability and flexibility.

[0022] This application adopts a layered architecture, using the DQN algorithm to optimize the selection of target floors at the global level, and the PPO algorithm to optimize the operation control of a single elevator at the local level. The genetic algorithm coordinates the strategies at the two levels, so that the scheduling decision can take into account both overall efficiency and individual energy consumption to achieve global optimization.

[0023] This application introduces energy consumption factors into the reinforcement learning objectives and minimizes the total energy consumption of the elevator by optimizing the scheduling strategy. Compared with the traditional method, the scheduling strategy generated by this method can effectively reduce the empty load rate and ineffective operation of the elevator, thereby greatly saving energy. By sensing the floor call demand in real time, the operation order and stop floors of the elevator are optimized and the waiting time of the user is minimized. At the same time, multiple elevators operate in coordination under the guidance of the global scheduling strategy, avoiding mutual interference and duplicate services, and improving the overall scheduling efficiency.

[0024] This application uses a hysteresis comparator to filter the sensor data to cope with the noise interference in the construction environment; the accumulation of time step rewards is used to evaluate the long-term effect of the scheduling strategy to cope with the uncertainty of the environment. At the same time, the proposed reinforcement learning algorithm has a low computational complexity and can update the scheduling strategy online in real time to adapt to dynamically changing construction needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The present application will be further described in the form of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same number represents the same structure, wherein:

[0026] Figure 1 is an exemplary flow chart of a smart construction scheduling method based on the Internet of Things according to some embodiments of the present application;

[0027] Figure 2 is an exemplary flow chart of obtaining the state parameters of the door of each floor according to some embodiments of the present application;

[0028] Figure 3 is an exemplary flow chart of obtaining a final scheduling strategy according to some embodiments of the present application;

[0029] Figure 4 is an exemplary flow chart of obtaining a global scheduling strategy according to some embodiments of the present application;

[0030] Figure 5 This is an exemplary flowchart for obtaining a local scheduling strategy according to some embodiments of the present application. DETAILED DESCRIPTION

[0031] The method and system provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0032] like Figure 1 As shown in the figure, the call extension data of each floor of the construction elevator is obtained, and the call extension data includes the floor door status data and the call data; the floor door status signal reflects the switch status of the floor door, and the call data reflects the call request of the floor; the floor door status data is converted by the binary processing algorithm to obtain the floor door status parameters of each floor; wherein, the floor door status parameters include the floor door opening flag and the floor door closing flag, and when the value of the floor door status data is greater than the threshold, the floor door opening flag is set to true, otherwise the floor door closing flag is set to true. The call data is counted by the counter to obtain the call status parameters of each floor, and the call status parameters include the number of calls and the last call time; the floor door status parameters and the call status parameters are used as the environmental state, and the construction elevator scheduling problem is modeled as a Markov decision process, with the scheduling strategy as the action and the energy consumption of the construction elevator as the reward function, and the scheduling strategy is obtained by reinforcement learning through the Q-learning algorithm.

[0033] The call extension data of each floor of the construction elevator is obtained, and a call extension is installed on each floor of the construction elevator. The call extension includes a floor door status sensor and a call button. The floor door status sensor is used to detect the switch status of the floor door. Common access control sensors such as magnetic switches and photoelectric switches can be used. The call button is used to receive the user's call request. Human-computer interaction devices such as contact switches or touch screens can be used. The call extension is connected to the data acquisition unit by wire or wireless means, and the collected floor door status data and call data are transmitted to the data acquisition unit in real time. Common communication methods include RS485 bus, CAN bus, Zigbee wireless transmission, etc. The received raw data is preprocessed, including data format conversion, timestamp alignment, deduplication and other operations to ensure data consistency and integrity. The preprocessed data is encapsulated in a fixed data structure to form a call extension data packet. The call extension data packet is transmitted to the monitoring server through the network. Encryption and verification mechanisms are used during the transmission process to ensure data security and reliability. Common transmission protocols include TCP / IP, MQTT, etc. After receiving the call extension data packet, the data is parsed and stored. The door status data and call data are extracted separately, indexed by floor and timestamp, and stored in the database. For the door status data, the monitoring server converts the original analog or digital signal into a unified logical value, with 0 representing the closed state and 1 representing the open state. For call data, the monitoring server records the timestamp and duration of each call and generates a unique call number. Continuous call requests on the same floor can be merged into one call event to reduce data redundancy.

[0034] like Figure 2 As shown in the figure, the floor door status data is converted using a binary processing algorithm to obtain the floor door status parameters of each floor. Specifically, the collected floor door status data is digitized. The floor door status data usually comes from access control sensors, which may be analog signals (such as voltage, current, etc.) or digital signals (such as high and low levels, pulse sequences, etc.). It is necessary to select a suitable digital-to-analog conversion (ADC) or digital signal processing (DSP) algorithm based on the type and characteristics of the sensor to convert the original signal into a unified digital quantity representation, that is, the digital quantity of the floor door status. The conversion process needs to consider parameters such as the signal range, resolution, and sampling frequency to ensure conversion accuracy and real-time performance.

[0035] The hysteresis comparator is used to set the door opening threshold and closing threshold, and the hysteresis interval is set according to the two thresholds. The hysteresis comparator is a comparator with hysteresis characteristics, which can effectively prevent frequent state switching when the signal jitters near the threshold. The door opening threshold means that when the digital value of the door state exceeds this value, the door is considered to be opened; the door closing threshold means that when the digital value of the door state is lower than this value, the door is considered to be closed. The area between the two thresholds is called the hysteresis interval, which can be adjusted according to the actual situation to balance sensitivity and stability.

[0036] Specifically, in this embodiment, the analog signal range of the floor door status sensor is 0-5V, and the corresponding digital range is 0-1023 (ADC resolution is 10 bits). Through observation and statistics of the actual floor door opening and closing process, the following data characteristics are obtained: when the floor door is fully closed, the analog output of the sensor is about 0.5V, and the corresponding digital value is about 102. When the floor door is fully opened, the analog output of the sensor is about 4.5V, and the corresponding digital value is about 921. During the floor door opening and closing process, the analog output of the sensor fluctuates between 1.0V and 4.0V, and the corresponding digital value varies between 205819. Due to factors such as mechanical vibration and electrical interference, the analog output of the sensor has a random noise of ±0.1V, and the corresponding digital noise is about ±20. According to the above data characteristics, the floor door opening threshold and closing threshold are set: considering the signal fluctuation and noise during the floor door opening process, the opening threshold is set to 80% of the digital value of the fully open state, that is, 921*80%≈737. This means that when the digital value of the landing door status exceeds 737, the landing door is considered to be open. Considering the signal fluctuations and noise during the landing door closing process, the closing threshold is set to 120% of the digital value of the fully closed state, that is, 102*120%≈122. This means that when the digital value of the landing door status is lower than 122, the landing door is considered to be closed. The size of the hysteresis interval is equal to the difference between the opening threshold and the closing threshold, that is, 737-122=615. This interval is wide enough to accommodate most of the signal fluctuations and noise, avoiding frequent state switching.

[0037] Determine whether the current digital quantity of the floor door status is within the hysteresis interval. If it is within the hysteresis interval, the current floor door status is maintained unchanged, that is, the open flag and the closed flag value of the previous moment are maintained. This can avoid frequent state switching caused by small fluctuations in the digital quantity and improve the robustness of the system. If the digital quantity of the floor door status is not within the hysteresis interval, the floor door status flag is set according to its size relationship with the open threshold and the closed threshold. When the digital quantity of the floor door status is greater than the floor door open threshold, the floor door open flag of the corresponding floor is set to true (logical 1), and the floor door closed flag is set to false (logical 0). This indicates that the floor door has been opened and is in the open state. Conversely, when the digital quantity of the floor door status is less than the floor door closed threshold, the floor door open flag is set to false and the floor door closed flag is set to true, indicating that the floor door has been closed and is in the closed state.

[0038] Generate the door status parameters according to the set door opening and closing flags, and define the data structure of the door status parameters. The door status parameters should contain the following fields: floor_id (floor number): indicates the floor where the door is located, and the value range is 1~N (N is the total number of floors). For example, for a 10-story building, the floor number can be set to 1, 2, 3, ..., 10. timestamp (timestamp): indicates the time when the parameter is generated, and the Unix timestamp (milliseconds since January 1, 1970) or a custom time format can be used. For example, 12:30:45 on May 20, 2023 can be expressed as 1684560645000. door_open_flag (open flag): indicates whether the door of this floor is in the open state. 1 is open and 0 is closed. door_close_flag (close flag): indicates whether the door of this floor is in the closed state. 1 is closed and 0 is open. Set the open flag and close flag according to the output of the hysteresis comparator. The output of the hysteresis comparator is 0 or 1, 0 means the floor door is closed, 1 means the floor door is open. Then the flag bit can be set by the following logic: If the comparator output is 1, set door_open_flag to 1 and door_close_flag to 0. If the comparator output is 0, set door_open_flag to 0 and door_close_flag to 1. For example, if the hysteresis comparator output of the current floor is 1, then in the floor door status parameter of this floor, door_open_flag should be set to 1 and door_close_flag should be set to 0. Fill in other fields of the floor door status parameter. Assign values ​​to the floor_id and timestamp fields of the parameter according to the current floor number and timestamp information. For example, the floor door status of the 5th floor is currently being processed, and the timestamp is 1684560645000. In the floor door status parameter of this floor, floor_id should be set to 5 and timestamp should be set to 1684560645000. Generate the floor door status parameter. Combine the filled parameter fields into a complete floor door status parameter object and add it to an array or list to represent the floor door status parameters of all floors.In this embodiment, a building has 3 floors. At a certain moment, the floor door status parameters of each floor are as follows: Floor 1: floor_id=1, timestamp=1684560645000, door_open_flag=0, door_close_flag=1; Floor 2: floor_id=2, timestamp=1684560645000, door_open_flag=1, door_close_flag=0; Floor 3: floor_id=3, timestamp=1684560645000, door_open_flag=0, door_close_flag=1; The generated floor door status parameter list should contain the above three parameter objects, indicating that at the current moment, the floor doors of the 1st and 3rd floors are in the closed state, and the floor door of the 2nd floor is in the open state.

[0039] The call state parameters of each floor are obtained by counting the call data through the counter. Specifically, the data structure of the call state parameters is defined. The call state parameters should contain the following fields: floor_id (floor number): indicates the floor where the call occurs, and the value range is 1~N (N is the total number of floors). call_count (number of calls): indicates the cumulative number of calls on the floor, and the initial value is 0. last_call_time (last call time): indicates the timestamp of the most recent call on the floor, and the initial value is 0 or empty. Initialize the call state parameters. Create a call state parameter object for each floor, and initialize the call_count and last_call_time fields to the default values. Process call data. Whenever a new call data is received, perform the following steps: Parse the call data and extract the floor number and timestamp information where the call occurs. Find the corresponding call state parameter object according to the floor number. Add 1 to the call_count field of the floor, indicating that the cumulative number of calls has increased. Update the last_call_time field of the floor to the timestamp of the current call data, indicating the latest call time.

[0040] In this embodiment, a call data is received, showing that a call occurred on the 3rd floor at the timestamp of 1684560645000. The call state parameter object of the 3rd floor is found, its call_count is increased by 1, and the last_call_time is updated to 1684560645000. A call state parameter list is generated. The call state parameter objects of all floors are combined into a list or array in the order of the floor numbers to represent the call state of the entire building. For example, a building has three floors. At a certain moment, the call status parameters of each floor are as follows: Floor 1: floor_id=1, call_count=3, last_call_time=1684560600000; Floor 2: floor_id=2, call_count=0, last_call_time=0; Floor 3: floor_id=3, call_count=5, last_call_time=1684560645000; the generated call status parameter list should contain the above three parameter objects, indicating that at the current moment, the first floor has accumulated 3 calls, and the last call occurred at 1684560600000; there is no call record on the second floor; the third floor has accumulated 5 calls, and the last call occurred at 1684560645000.

[0041] Clear call status parameters regularly. To avoid infinite accumulation of calls, it is necessary to clear the call_count field of all floors and reset the last_call_time field to 0 or empty in each statistical period (such as midnight every day). This can achieve periodic call statistics, which is convenient for data analysis and trend prediction. Use counters to count call data and generate call status parameters that reflect the call situation of each floor. These parameters include the cumulative number of calls and the latest call time for each floor, which can serve as an important basis for subsequent scheduling decisions, such as giving priority to responding to floors with high call frequency and reasonably arranging elevator parking strategies. At the same time, call status parameters can also be used for data visualization and statistical analysis, such as generating a time trend chart of the number of calls, identifying peak call periods, etc., to provide data support for optimizing elevator scheduling.

[0042] like Figure 3As shown in the figure, the door state parameters and call state parameters are used as environmental states, the construction elevator scheduling problem is modeled as a Markov decision process, the scheduling strategy is used as the action, the energy consumption of the construction elevator is used as the reward function, and the scheduling strategy is obtained through reinforcement learning through the Q-learning algorithm; the state space S and the action space A are constructed, and the door state parameters and call state parameters are combined into a state vector as the state representation of the MDP. For example, for an N-story building, the state vector can be expressed as: [f1_door_open, f1_door_close, f1_call_count, f1_last_call_time, ..., fN_door_open, fN_door_close, fN_call_count, fN_last_call_time], where fi_door_open and fi_door_close represent the opening and closing states of the door of the i-th floor, and fi_call_count and fi_last_call_time represent the number of calls and the last call time of the i-th floor. According to the value range of the state vector, the state space S is defined. For example, for the above state vector, each component has a certain range of values: fi_door_open and fi_door_close are 0 or 1; fi_call_count is a non-negative integer; fi_last_call_time is a non-negative real number or 0 (indicating no call). The state space S is the set of all possible state vectors.

[0043] Define action space A according to the user's scheduling needs. For example, the following scheduling actions can be defined: RESPONSE_CALL: respond to a call request on a certain floor and go to that floor; CHANGE_DIRECTION: change the running direction, that is, switch from up to down, or from down to up; STOP_FLOOR_i: stop at the i-th floor; SKIP_FLOOR_i: skip the i-th floor without stopping. Action space A is the set of all possible scheduling actions. Define the state transition probability P and analyze the impact of each scheduling action on state transition. For example: RESPONSE_CALL action will make the elevator go to the floor where the call occurs and may change the running direction; CHANGE_DIRECTION action will change the running direction of the elevator; STOP_FLOOR_i action will make the elevator stop at the i-th floor and may affect the door status and call status; SKIP_FLOOR_i action will not change the state. Estimate the state transition probability based on the current state and scheduling action. The following methods can be used: Based on prior knowledge: According to the general rules of elevator scheduling, the state transition probability is manually set. For example, the RESPONSE_CALL action will cause the elevator to reach the called floor with a high probability. Based on data statistics: By collecting elevator operation data, the frequency of state transitions to other states after taking a certain action in a certain state is counted, and the frequency is normalized to obtain the transition probability. Based on model learning: Use machine learning algorithms (such as maximum likelihood estimation) to fit the state transition probability model from the data.

[0044] Define the reward function R, running distance: the more floors the elevator runs, the greater the energy consumption. Running time: the longer the elevator runs, the greater the energy consumption. Load capacity: the greater the passenger capacity of the elevator, the greater the energy consumption. Acceleration and deceleration times: the more times the elevator starts and stops, the greater the energy consumption. Design the reward function R according to the factors affecting energy consumption. Generally, negative rewards are used, that is, the lower the energy consumption, the greater the reward value. For example: R = -α×running distance-β×running time-γ×load-δ×acceleration and deceleration times, where α, β, γ, δ are weight coefficients, which are adjusted according to actual conditions. Associate the reward value with the state transition. In MDP, the reward is a random variable related to the state transition, that is, after taking a certain action in a certain state, the environment gives a reward value. Therefore, it is necessary to set a corresponding reward value for each state transition.

[0045] Construct a Markov decision process MDP and formally define the five-tuple (S, A, P, R, γ) of the MDP: S: state space, containing all possible state vectors. A: action space, containing all possible scheduling actions. P: state transition probability, indicating the probability of transitioning to other states after taking a certain action in a certain state. R: reward function, indicating the energy consumption reward value obtained after taking a certain action in a certain state. γ: discount factor, indicating the importance of future rewards, with a value range of [0, 1]. Integrate the state space S, action space A, transition probability P, and reward function R into an MDP model to form a complete Markov decision process. The MDP model describes the environmental dynamic characteristics of the elevator scheduling problem and provides a theoretical basis for subsequent optimal decisions. This application constructs a Markov decision process model for elevator scheduling optimization based on the door state parameters and call state parameters. The model comprehensively considers various state information and scheduling actions during the operation of the elevator, and introduces energy consumption as the optimization target, providing a mathematical framework for formulating the optimal scheduling strategy.

[0046] like Figure 4 As shown in the figure, the global scheduling strategy is obtained, including: the input state features of the DQN algorithm come from the state space S of the Markov decision process MDP. The state features include the door state parameters and the call state parameters. The door state parameters represent the open and closed states of the doors on each floor, reflecting the current position and running direction of the elevator. The call state parameters represent the number of calls and the last call time on each floor, reflecting the user's elevator demand and waiting time. The door state parameters and the call state parameters are combined into a state vector as the input of the DQN algorithm, providing comprehensive environmental state information. The DQN algorithm uses a multi-layer feedforward neural network as a Q network to estimate the long-term cumulative reward (i.e., Q value) of taking a certain scheduling action in a certain state. The input of the Q network is the state feature, and the output is the Q value of each scheduling action taken in this state. Scheduling actions include responding to call requests, changing the running direction, stopping or skipping a certain floor, etc. The Q value represents the expected long-term cumulative reward of taking a certain scheduling action in the current state, taking into account the impact of the current action on the future elevator operation. The structure and parameters of the Q network determine the accuracy of the Q value estimation, which in turn affects the quality of the scheduling strategy.

[0047] Specifically, in this embodiment, the representation of the state variable is: the state variable s represents a state in the Markov decision process MDP, which is composed of a door state parameter and a call state parameter. For example, for a 10-story construction elevator, the state variable s can be represented as a 20-dimensional vector: [f1_door_open, f1_door_close, ..., f10_door_open, f10_door_close, f1_call_count, ..., f10_call_count], where fi_door_open and fi_door_close represent the open and closed states of the door of the i-th floor respectively (the values ​​are 0 or 1), and fi_call_count represents the number of calls of the i-th floor (a non-negative integer).

[0048] Representation of action variables: Action variable a represents a scheduling action taken in state s. For example, scheduling actions may include: parking at a certain layer, uplink, downlink, idle, etc. Action variable a can be represented by an integer, and different integers correspond to different scheduling actions. Assuming that parking at the i-th layer is represented by integer i, uplink is represented by integer 11, downlink is represented by integer 12, and idle is represented by integer 0, the value range of action variable a is [0, 1, 2, ..., 10, 11, 12].

[0049] The structure and parameters of the Q network: The Q network is a multi-layer feedforward neural network that takes the state variable s as input and outputs the estimated Q value of each scheduling action taken under state s. The parameters θ of the Q network include the weights and biases of the neurons in each layer, which determine the accuracy of the Q value estimation. For example, a simple Q network can include an input layer (20 neurons, corresponding to the dimension of the state variable), a hidden layer (50 neurons, using the ReLU activation function), and an output layer (13 neurons, corresponding to the range of the action variable). The parameters θ of the Q network can be obtained by random initialization or pre-training, and optimized during the training process.

[0050] Quantify the immediate reward. In the current state, after the construction elevator takes a certain scheduling action, a corresponding immediate reward will be generated. The immediate reward represents the impact of the scheduling action on the energy consumption of the elevator, which can be quantified by the negative value of energy consumption. For example, if the elevator consumes 10 kWh of electricity during operation after taking scheduling action a, the immediate reward can be set to -10. The setting of the immediate reward should take into account the actual energy consumption measurement and normalization processing so that it is compatible with the value range of the Q value. Estimate the maximum Q value of the next state. After taking the current scheduling action, the construction elevator will transfer from the current state to the next state. The maximum Q value of the next state indicates the maximum long-term cumulative reward that can be obtained by taking the optimal scheduling action in the next state. This maximum Q value can be estimated by the target Q network, which is a delayed update version of the Q network parameters. For example, assuming that the current state is s, and after taking action a, it transfers to the next state , the target Q network is in state The maximum Q value output under . Then the maximum Q value of the next state can be expressed as , where, means in the state All possible actions.

[0051] Calculate the target Q value. According to the Bellman equation, the target Q value is the weighted sum of the immediate reward and the maximum Q value of the next state. The calculation formula is: , where y represents the target Q value, r represents the immediate reward, and γ represents the discount factor. The discount factor γ ranges from [0, 1] and represents the importance of future rewards. The larger the γ, the more attention is paid to future long-term rewards; the smaller the γ, the more attention is paid to the current immediate reward. For example, if the immediate reward r = -10, the maximum Q value of the next state , the discount factor γ=0.9, then the target Q value y=-10+0.9*50=35.

[0052] The mean square error loss function and stochastic gradient descent algorithm are used to optimize the Q network parameters. Specifically, according to the expression of the mean square error loss function: L ( θ ) = E [ ( Q ( s , a ; θ ) − y ) 2 ] , calculate the loss value of the Q network. For each training sample , calculate its estimated Q value and target Q value: Estimated Q value It is calculated by inputting the state s and action a into the Q network and is a function of the network parameter θ. The target Q value y is calculated using the Bellman equation: , where r is the immediate reward, γ is the discount factor, is the next state, is the optimal action in the next state, are the parameters of the target network. Calculate the difference between the estimated Q value and the target Q value: , get the mean square error of a single sample. Average the mean square error of all training samples to get the loss function value of the entire training batch .

[0053] According to the loss function Gradient of the Q network parameter θ , use the stochastic gradient descent algorithm to update the parameters. Calculate the loss function for each parameter The partial derivative of , and get the gradient vector . According to the update formula of gradient descent: , update each parameter of the Q network . Where α is the learning rate, which controls the step size of each update. The learning rate is usually set to a small positive number, such as 0.001. Repeat the above steps until the loss function value converges or the preset number of training rounds is reached. In order to improve training stability, a separate target network is usually used to calculate the target Q value. The parameters of the target network Initially, the parameters θ of the Q network are the same, but they are not updated in time during the training process. Every certain number of training steps (such as 1000 steps), the parameters θ of the Q network are copied to the parameters θ' of the target network to achieve delayed update of the target network. This delayed update method can reduce the fluctuation of the target Q value and improve the stability and convergence of training.

[0054] After the training is completed, an optimized Q network is obtained, which can be used to generate the optimal scheduling strategy. For any input state s, it is input into the trained Q network to calculate the estimated Q value of each action taken in this state. The action with the largest estimated Q value is selected as the optimal scheduling action, that is: . The state s and the optimal action Combined together, we get an optimal state-action pair , which represents the optimal scheduling strategy that should be adopted in state s.

[0055] like Figure 5As shown in the figure, the local scheduling strategy includes: extracting feature subsets, extracting feature subsets reflecting the state of a single construction elevator from the state space S of the Markov decision process MDP. The feature subset contains information such as the current floor, running direction, speed, and floor door state of the construction elevator. These features can be represented as a vector or matrix as the input of the policy network. Constructing a policy network, constructing a multi-layer feedforward neural network as the policy network of the PPO algorithm. The policy network takes the feature subset as input and outputs the probability distribution of taking various scheduling actions in the current state. The output layer of the network uses the softmax activation function to ensure that the output is a legal probability distribution. The hidden layer of the network can use activation functions such as ReLU to enhance the nonlinear expression ability of the network. Collecting data, according to the current policy network parameters, a set of state-action-reward data is collected in the state space S and action space A of the MDP. During the data collection process, according to the probability distribution output by the policy network, the scheduling action is selected by random sampling. Specifically, in each state, the output of the policy network is regarded as a discrete probability distribution, and random sampling is performed using methods such as roulette to select a scheduling action. Execute the selected scheduling action, observe the feedback from the environment, obtain the immediate reward and the next state, and form a four-tuple of (state, action, reward, next state). Repeat the above process and collect a sufficient number of data samples for subsequent strategy optimization.

[0056] In this embodiment, a feature subset reflecting the state of a single construction elevator is extracted from the state space S of the Markov decision process MDP. For a 10-story construction elevator, the feature subset may include the following information: Current floor: an integer value indicating the floor where the construction elevator is currently located, with a value range of 1 to 10. Running direction: a binary variable indicating the current running direction of the construction elevator, 1 for upward movement and 0 for downward movement. Speed: a real value indicating the current running speed of the construction elevator, in meters per second. Floor door status: a 10-dimensional binary vector, each element of which corresponds to the door status of a floor, 1 for door open and 0 for door closed. These features are arranged in a fixed order to form a 13-dimensional feature vector as the input of the policy network. A three-layer feedforward neural network is constructed as the policy network of the PPO algorithm. The input layer contains 13 neurons, corresponding to the dimension of the feature subset. The first hidden layer contains 64 neurons and uses the ReLU activation function. The second hidden layer contains 32 neurons and uses the ReLU activation function. The output layer contains 3 neurons, corresponding to three scheduling actions: up, down, and stop. The softmax activation function is used in the output layer to convert the output into a legal probability distribution. The weight parameters of the network are initialized using the Xavier initialization method, and the bias parameters are initialized to 0. According to the current policy network parameters, a set of state-action-reward data is collected in the state space S and action space A of the MDP. The number of episodes for data collection is set to 1000, and the maximum number of steps for each episode is 100. At the beginning of each episode, the state of the construction elevator is randomly initialized, including the current floor, running direction, speed, and floor door state. At each time step: the current state is input into the policy network to obtain the probability distribution of the three scheduling actions. A scheduling action is randomly selected according to the probability distribution using the roulette method. The selected scheduling action is executed, and the immediate reward and the next state are obtained according to the predefined state transition function and reward function. The four-tuple of (current state, selected action, immediate reward, next state) is recorded as a data sample. The next state is updated to the current state, and the above process is repeated until the maximum number of steps of the episode is reached. Repeat the above process until all episodes are collected to obtain a set of state-action-reward data. The collected data is randomly divided into a training set and a validation set for subsequent strategy optimization.

[0057] Calculate the importance weights. According to the collected state-action-reward data, calculate the importance weights between the current policy network and the data collection policy. The importance weights are used to correct the deviation of data distribution and make the optimization process more stable and efficient. The calculation formula of the importance weights is: .in, represents the probability that the current policy network takes action a in state s, It represents the probability that the data collection strategy takes action a in state s. The importance weight can be calculated on each data sample to obtain a set of weight values.

[0058] Construct the objective function. According to the importance weight, construct the objective function of the PPO algorithm to balance the exploration and utilization of the strategy. The objective function consists of three parts: strategy loss, value function loss, and entropy regularization term. The calculation formula of strategy loss is: E [ r ( θ ) × A ( s , a )] ,in: is the importance weight, which represents the probability ratio of the current strategy to the data collection strategy, and can be calculated by the following formula: . is the advantage function, which indicates the advantage value of taking action a in state s, and can be estimated by the following formula: ,in is the state-action value function, is the state value function. In practical calculations, the generalized advantage estimation (GAE) method can be used to estimate the advantage function to reduce variance and bias.

[0059] The calculation formula of the value function loss is: ,in: It is the output of the current value function network in state s, which represents the estimated value function of state s. is the objective value function, which can be calculated as follows: , where r is the immediate reward, γ is the discount factor, is the next state. β is the weight coefficient of the value function loss, which is used to balance the policy loss and the value function loss, and is usually set to 0.5.

[0060] The calculation formula of the entropy regularization term is: ,in: It is the entropy of the action probability distribution output by the policy network in state s, which indicates the randomness and uncertainty of the policy. Entropy can be calculated by the following formula: ,in represents the sum of all possible actions. η is the weight coefficient of the entropy regularization term, which is used to encourage policy exploration and is usually set to 0.01. The policy loss, value function loss, and entropy regularization term are weighted and summed to obtain the complete objective function , represents the target value under the policy network parameters θ.

[0061] Optimize the policy network, using the stochastic gradient ascent algorithm, in the objective function Update the policy network parameters θ in the gradient direction. First, use the collected data samples to calculate the importance weights, advantage functions, and objective value functions. Then, substitute these values ​​into the objective function , calculate the gradient of the objective function with respect to the policy network parameters θ The calculation of gradients can be done efficiently using the back-propagation algorithm. Modern deep learning frameworks such as PyTorch and TensorFlow provide automatic differentiation functions. According to the update formula of gradient ascent: , update the parameters θ of the policy network. Among them, α is the learning rate, which controls the step size of each update and is usually set to a small positive number, such as 0.0003. Repeat the above steps, that is, calculate the gradient and update the parameters, until the objective function converges or reaches the preset number of training rounds, such as 1000 rounds. After each round of training, the performance of the current strategy can be evaluated using the validation set data, and the hyperparameters can be adjusted or early stopping can be performed based on the evaluation results. This application constructs the objective function of the PPO algorithm based on the importance weights, and uses the stochastic gradient ascent algorithm to optimize the parameters of the policy network. The design of the objective function reflects the advantages of the PPO algorithm in balancing exploration and utilization. By introducing importance weights, advantage functions, and entropy regularization terms, the optimal strategy can be learned more efficiently and stably. The optimization process of the policy network demonstrates the effectiveness of the stochastic gradient ascent algorithm in dealing with complex non-convex optimization problems. By continuously iteratively calculating gradients and updating parameters, the performance of the policy network can be gradually improved, and ultimately a local strategy that can intelligently schedule construction elevators is obtained.

[0062] Generate a local scheduling strategy, and use the trained policy network to generate a local scheduling strategy under each feature subset. In this embodiment, several representative feature subsets are selected according to the actual situation of the construction elevator. For example, the following three feature subsets can be selected: Feature subset 1: the number of floors is 10, the current floor is 1, the running direction is upward, the speed is 1.5 meters per second, and all floor doors are closed. Feature subset 2: the number of floors is 10, the current floor is 5, the running direction is downward, the speed is 2.0 meters per second, and the floor doors of the 3rd and 7th floors are open. Feature subset 3: the number of floors is 10, the current floor is 8, the running direction is upward, the speed is 1.0 meters per second, and the floor doors of the 2nd and 9th floors are open. These feature subsets cover typical scenarios during the operation of the construction elevator, including different floor positions, running directions, speeds, and floor door states. For each selected feature subset, a local scheduling strategy is generated using the trained policy network. The feature subset is used as input, and the output of the policy network is calculated through forward propagation to obtain the probability distribution of each scheduling action. According to the probability distribution, a scheduling action is selected by random sampling. Specifically, the numpy.random.choice function can be used to take the probability distribution of the scheduling action as the parameter p, so as to randomly select an action according to the probability. For example, for feature subset 1, the action probability distribution output by the policy network is [0.6, 0.3, 0.1], corresponding to the three actions of up, down, and stop, respectively. Through random sampling, the up action may be selected. For each feature subset, random sampling is repeated multiple times to obtain a series of scheduling actions. These scheduling actions are arranged in the order of sampling to form a scheduling action sequence, which constitutes the local scheduling strategy under the feature subset. The length of the scheduling action sequence can be set according to actual needs, for example, set to 10 actions. For feature subset 1, through 10 random samplings, the following scheduling action sequence may be obtained: [up, up, up, stop, up, up, stop, up, up, up]. This scheduling action sequence indicates that in the current state, the continuous execution of these actions can enable the construction elevator to achieve better energy consumption and operating efficiency. Repeat the above process for different feature subsets to generate corresponding local scheduling strategies. In practical applications, the feature subset that best matches the state of the construction elevator can be selected, and the local scheduling strategy corresponding to the feature subset can be used to make scheduling decisions. These local scheduling strategies are combined to form a complete construction elevator scheduling strategy. For example, during the operation of the construction elevator, if the current state is most similar to feature subset 2, the local scheduling strategy corresponding to feature subset 2 can be used: [down, down, stop, down, down, down, stop, down, down, down] to guide scheduling decisions. By dynamically selecting the best matching local scheduling strategy, the optimal scheduling of the construction elevator in different states can be achieved, improving the energy-saving effect of the entire system.The local scheduling strategies under different feature subsets can cover various typical scenarios in the operation of construction elevators. By dynamically selecting the most matching strategy for scheduling decisions, intelligent energy-saving optimization of construction elevators can be achieved.

[0063] S46, according to the global scheduling strategy and the local scheduling strategy, the final scheduling strategy is obtained through collaborative optimization; specifically, the global scheduling strategy is encoded as a chromosome, and the chromosome is composed of scheduling actions in each state. For example, for a global scheduling strategy containing 10 states, there are 3 optional scheduling actions (up, down, stop) in each state, then the length of the chromosome is 10, and each gene bit can take the value of 0, 1, and 2, respectively representing three scheduling actions. A possible global scheduling strategy chromosome is [0, 1, 2, 0, 1, 1, 2, 0, 1, 0], which represents the scheduling actions taken in 10 states. Multiple local scheduling strategies are regarded as a population, and each local scheduling strategy represents an individual in the population. The chromosome of an individual is composed of a scheduling action sequence under the feature subset of the corresponding local scheduling strategy. For example, for a local scheduling strategy containing 10 scheduling actions, its feature subset is that the number of floors is 10, the current floor is 5, the running direction is down, the speed is 2.0 m / s, and the floor doors of the 3rd and 7th floors are open. The chromosome of this local scheduling strategy can be expressed as [1, 1, 0, 1, 1, 1, 0, 1, 1, 1], which represents 10 consecutive scheduling actions taken under this feature subset.

[0064] Construct a fitness function to evaluate the consistency between the local scheduling strategy and the global scheduling strategy. The fitness function is expressed as: ,in represents the i-th local scheduling strategy, It represents the instant reward obtained by the strategy at the tth time step. The larger the value of the fitness function, the higher the consistency between the local scheduling strategy and the global scheduling strategy, and the better the quality of the local scheduling strategy. For example, for a local scheduling strategy , the immediate reward sequence obtained in 10 time steps is [0.5, 0.8, 0.6, 0.7, 0.9, 0.8, 0.6, 0.7, 0.8, 0.9], and the discount factor γ is 0.9, then its fitness function value is: .

[0065] The local scheduling strategy population is optimized using a genetic algorithm, and the population is updated through selection, crossover, and mutation operations. Selection operation: Use the roulette selection operator to select a group of individuals from the current population as parents with the fitness function value as the weight. The individuals with higher fitness function values ​​have a greater probability of being selected. Crossover operation: Parent individuals are randomly combined in pairs, and chromosomes are cross-recombined through single-point crossover or multi-point crossover operators to generate new offspring individuals. Mutation operation: Randomly mutate the chromosomes of offspring individuals, flip the values ​​of certain gene bits with a certain probability, and introduce new search directions. Iterative update: Repeat the selection, crossover, and mutation operations to continuously update the population until the preset number of iterations or convergence conditions are met.

[0066] When the genetic algorithm meets the iteration termination condition, the individual with the highest fitness function value is selected from the last generation population. The selected individual chromosome is decoded into a local scheduling strategy, which is output as the local scheduling strategy that is most coordinated with the global scheduling strategy. For example, if the chromosome of the optimal individual is [1, 1, 0, 1, 1, 1, 0, 1, 1, 1], the decoded optimal local scheduling strategy is to take 10 consecutive scheduling actions under the feature subset (the number of floors is 10, the current floor is 5, the running direction is down, the speed is 2.0 m / s, and the floor doors of the 3rd and 7th floors are open).

[0067] The final scheduling strategy is composed of a global scheduling strategy and a local scheduling strategy, taking into account both overall optimization and individual control. In the actual scheduling process, the target floor is first selected according to the global scheduling strategy to determine the overall scheduling direction of the construction elevator. Then, according to the state characteristics of the current construction elevator, the local scheduling strategy that best matches it is selected. The construction elevator executes the corresponding scheduling action sequence, such as up, down or stop, according to the selected local scheduling strategy. Through the synergy of the global scheduling strategy and the local scheduling strategy, the optimal scheduling of the construction elevator group in different states is achieved. Specifically, in this embodiment, in the current state, according to the global scheduling strategy chromosome [0, 1, 2, 0, 1, 1, 2, 0, 1, 0], scheduling action 2 should be taken, that is, stop. The current state characteristics of a construction elevator are: the number of floors is 10, the current floor is 5, the running direction is down, the speed is 2.0 m / s, and the floor doors of the 3rd and 7th floors are open. According to the local scheduling strategy chromosome [1, 1, 0, 1, 1, 1, 0, 1, 1, 1], the construction elevator should perform 10 scheduling actions in sequence: down, down, stop, down, down, down, stop, down, down, down. Under the guidance of the global scheduling strategy, the construction elevator first stops running, and then according to the local scheduling strategy, performs a series of down and stop scheduling actions until it reaches the target floor. In this application, a genetic algorithm is used to optimize the local scheduling strategy population, and it is coordinated with the global scheduling strategy to obtain a final scheduling strategy that comprehensively considers global optimization and local control. This method makes full use of the macro-guidance role of the global scheduling strategy and the micro-control ability of the local scheduling strategy, measures the consistency of the two through the fitness function, and uses the genetic algorithm to search for the optimal strategy combination, thereby realizing the intelligent and energy-saving scheduling of the construction elevator.

Claims

1. A smart construction scheduling method based on the Internet of Things, characterized in that: include: Obtain the call extension data of each floor of the construction elevator. The call extension data includes the floor door status data and call data. The floor door status signal reflects the switch status of the floor door, and the call data reflects the call request of the floor. The floor door status data is converted by using a binary processing algorithm to obtain the floor door status parameters of each floor; wherein the floor door status parameters include a floor door opening flag and a floor door closing flag. When the value of the floor door status data is greater than a threshold, the floor door opening flag is set to true, otherwise the floor door closing flag is set to true; The call data is counted by the counter to obtain the call status parameters of each floor, which include the number of calls and the last call time; The door state parameters and call state parameters are used as environmental states, and the construction elevator scheduling problem is modeled as a Markov decision process. The scheduling strategy is used as the action and the energy consumption of the construction elevator is used as the reward function. The scheduling strategy is obtained through reinforcement learning using the Q-learning algorithm. Obtaining a scheduling strategy, including: constructing a state space S according to the state parameters of the floor door and the call state parameters; according to the state space S of the Markov decision process; and setting the action space A according to the user's scheduling needs; wherein the action space A includes responding to call requests, changing the running direction, and stopping or skipping a certain floor; defining the transition probability P between states according to the state space S and the action space A; wherein the transition probability P represents the probability that the environmental state is transferred to the next state after a certain scheduling action is taken in the current state; using the energy consumption of the construction elevator as the reward function R; constructing a Markov decision process MDP according to the state space S, the action space A, the transition probability P and the reward function R; according to the constructed Markov decision process MDP, using a hierarchical reinforcement learning algorithm, performing reinforcement learning at the global scheduling layer and the local scheduling layer, and obtaining a global scheduling strategy and a local scheduling strategy; according to the global scheduling strategy and the local scheduling strategy, obtaining the final scheduling strategy through collaborative optimization; Among them, the global scheduling strategy is obtained, including: taking the state space S of the Markov decision process MDP as the input state feature of the DQN algorithm; wherein the state feature includes the layer door state parameter and the call state parameter; constructing a multi-layer feedforward neural network as the Q network of the DQN algorithm, the Q network takes the state feature as input, and outputs the Q value of each scheduling action taken in the corresponding state; wherein the Q value represents the long-term cumulative return expectation of taking a certain scheduling action in the current state; selecting the scheduling action with the largest Q value in each state as the global scheduling strategy in the corresponding state; Among them, the local scheduling strategy includes: According to the state space S of the Markov decision process MDP, a feature subset reflecting the state of a single construction elevator is extracted; wherein the feature subset includes the current floor, running direction, speed and floor door state of the construction elevator; Construct a multi-layer feedforward neural network as the policy network of the PPO algorithm. The policy network takes the feature subset as input and outputs the probability distribution of each scheduling action taken under the current state. According to the current policy network parameters, a set of state-action-reward data is collected in the state space S and action space A of the MDP. During the data collection process, the scheduling action is selected by random sampling according to the probability distribution of the policy network output. According to the collected state-action-reward data, the importance weight between the current policy network and the data collection strategy is calculated; the data collection strategy refers to the strategy of randomly sampling and scheduling actions according to the action probability distribution; the importance weight is used to correct the deviation of data distribution; According to the importance weight, the objective function of the PPO algorithm is constructed; the objective function includes the policy loss, the value function loss and the entropy regularization term; Use the stochastic gradient ascent algorithm to update the parameters of the policy network in the gradient direction of the objective function; Using the trained policy network, a scheduling action is selected from the probability distribution of scheduling actions output by the policy network by random sampling under each feature subset. The scheduling action sequence selected by multiple samplings constitutes the local scheduling strategy. Through collaborative optimization, the final scheduling strategy is obtained, including: The global scheduling strategy is encoded as a chromosome, and the global scheduling strategy chromosome is composed of scheduling actions under each state; Take multiple local scheduling strategies as a population, each local scheduling strategy represents an individual in the population, and the chromosome of the individual is composed of the scheduling action sequence under the feature subset of the corresponding local scheduling strategy; Construct a fitness function that reflects the consistency between the local scheduling strategy and the global scheduling strategy, and use the cumulative reward obtained by running the local scheduling strategy in the environment as the fitness value; Genetic algorithms are used to select, crossover and mutate individuals in the local scheduling strategy population. A group of individuals are selected from the current population as parents using a roulette wheel selection operator with fitness values ​​as weights. Parent individuals are randomly combined in pairs, and chromosomes are crossovered and recombined using a crossover operator to generate new offspring individuals. When the genetic algorithm meets the preset iteration termination condition, the individual with the highest fitness value is selected from the last generation population, and the selected individual chromosome is decoded into a local scheduling strategy as the local scheduling strategy output that is most coordinated with the global scheduling strategy; The selected local scheduling strategy and the obtained global scheduling strategy are combined to obtain the final scheduling strategy; wherein the global scheduling strategy is responsible for selecting the target floor, and the local scheduling strategy is responsible for controlling the operation of a single construction elevator; Construct a fitness function. The expression of the fitness function is as follows: in, Represents an individual The fitness function of represents the i-th individual in the population; represents the discount factor; t represents the time step variable, which represents the local scheduling strategy The tth time step of running in the environment; Represents the local scheduling strategy The immediate reward obtained at the tth time step.

2. The method for intelligent construction scheduling based on the Internet of Things according to claim 1 is characterized in that: The door status parameters of each floor include: Digitally process the acquired door status data, convert the door status data into digital quantity signals, and obtain the digital quantity of the door status; The floor door opening threshold and the floor door closing threshold are set by the hysteresis comparator, and the floor door hysteresis interval is set according to the floor door opening threshold and the floor door closing threshold; Determine whether the digital value of the floor door status is in the floor door hysteresis range. If yes, maintain the current floor door status unchanged. If not, set the floor door status flag: When the digital value of the floor door status is greater than the floor door opening threshold, the floor door opening flag of the corresponding floor is set to true, and the floor door closing flag is set to false; When the digital value of the floor door status is less than the floor door closing threshold, the floor door opening flag of the corresponding floor is set to false, and the floor door closing flag is set to true; The floor door status parameters of each floor are generated according to the set floor door open flag and floor door closed flag.

3. The smart construction scheduling method based on the Internet of Things according to claim 1 is characterized in that: Get the call status parameters of each floor, including: The number of calls is obtained by accumulating the call data; The last call time is obtained by obtaining the timestamp of the most recent call data.

4. The smart construction scheduling method based on the Internet of Things according to claim 1 is characterized in that: Get the global scheduling strategy, including: During the Q network training process, according to the Bellman equation, the target Q value of the current state-action pair is calculated using the maximum Q value of the next state and the immediate reward. The target Q value is used as a supervisory signal for Q network training and is used to optimize the Q network parameters. The mean square error loss function is used to calculate the difference between the estimated Q value output by the Q network and the target Q value, and the Q network parameters are updated in the direction of the loss function gradient through the stochastic gradient descent algorithm.

5. The method for intelligent construction scheduling based on the Internet of Things according to claim 4 is characterized in that: The mean square error loss function is expressed as follows: in, represents the mean square error loss function; represents the average of the training data; represents the estimated Q value output by the Q network when taking action a in state s; s represents the state variable; a represents the action variable; θ represents the parameter of the Q network; y represents the target Q value, and the expression is: Where: r represents the immediate reward; γ represents the discount factor; represents the next state variable; represents the next action variable; represents the target parameter of the Q network.

6. The method for intelligent construction scheduling based on the Internet of Things according to claim 1 is characterized in that: The objective function of the PPO algorithm is expressed as follows: in, Represents the objective function of the PPO algorithm; Represents the importance weight, and the calculation formula is: in: Represents the probability that the current policy network takes action a in state s; represents the probability that the data collection strategy takes action a in state s; Represents the advantage function, which represents the advantage value of taking action a in state s. The calculation formula is: in: represents the state-action value function; represents the state value function; The weight coefficient representing the loss of the value function; Represents the output of the current value function network in state s; represents the target value function; Represents the weight coefficient of the entropy regularization term; Represents the entropy of the action probability distribution output by the policy network in state s.

Citation Information

Patent Citations

  • Large-scale intraday operation dynamic scheduling method based on hierarchical reinforcement learning

    CN118675716A