Lift shaft self-learning method and system based on reinforcement learning
By constructing the shaft feature matrix and control strategy model based on reinforcement learning, the problem of dynamic changes in the elevator operating environment in the existing technology is solved, and high-precision and safe elevator control are achieved.
Patent Information
- Application Number
- CN202510700491.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-18
AI Technical Summary
The existing shaft self-learning technology is difficult to cope with dynamic changes and complex disturbances in the elevator operating environment, resulting in a decrease in stopping accuracy, insufficient safety margin, and even induced failure risk.
Using a method based on reinforcement learning, a shaft feature matrix is constructed through multimodal perceptual data, combining digital twin technology and elevator industry knowledge graphs, a control strategy model is generated, and online reinforcement learning is carried out on edge computing devices to optimize control strategies in real time.
It significantly improves the accuracy and safety of elevator operation, can dynamically adapt to environmental changes during long-term operation, and improves the stability and safety of the system.
Smart Images

Figure CN120328291A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of elevator control, and in particular, to an elevator shaft self-learning method and system based on reinforcement learning. Background Art
[0002] In modern urban buildings, elevators, as the core facilities for vertical transportation, their operating efficiency and safety are directly related to the use experience of buildings and the safety of people's lives and property. To ensure that elevators can achieve high-precision and high-reliability docking control in variable building structures and usage environments, the so-called "shaft self-learning" technology is commonly used in the industry. Shaft self-learning refers to the elevator system collecting structural parameters and environmental data in the shaft during the installation and commissioning stage, thereby establishing mapping relationships of key parameters such as floor positions, heights, and leveling points to support precise control and intelligent judgment in subsequent operations.
[0003] The existing shaft self-learning technologies generally include the following processes: First, preparation work is carried out, including confirming the status of the elevator safety circuit, switching the system to the maintenance mode, and clearing obstacles in the shaft; secondly, low-speed scanning is performed, controlling the car to run from the bottom floor to the top floor at low speed. During this period, the number of rotations of the traction wire rope is recorded by the encoder, and at the same time, identification marks such as magnets or magnetic isolation plates set in the door areas of each floor are identified; then, it enters the data collection stage, where the system records the positions of the leveling sensors, measures the distances between each floor, and establishes the correspondence between the floor numbers and the physical positions; finally, the data collected is written into the non-volatile memory through the control system, and solidified into parameters such as the operation curve, deceleration points, and braking strategies for the subsequent operation of the elevator.
[0004] Although the existing shaft self-learning solutions have basically achieved the automatic collection of floor information and the automatic generation of operation curves, they still essentially rely on static modeling and single sensing means, and it is difficult to cope with the dynamic changes and complex disturbances in the elevator operating environment. During long-term operation, the elevator system is affected by non-linear factors such as load changes, thermal expansion and contraction of guide rails, mechanical wear, and magnetic field interference. The existing technologies lack sufficient sensing means and intelligent response mechanisms, resulting in a decrease in docking accuracy, insufficient safety margins, and even an increased risk of faults. Summary of the Invention
[0005] This application provides an elevator shaft self-learning method and system based on reinforcement learning, which can continuously perceive environmental changes, autonomously adjust control strategies, and dynamically optimize operation parameters throughout the entire life cycle, significantly improving the operation accuracy, safety, and maintenance intelligence level. This application provides the following technical solutions:
[0006] In a first aspect, this application provides an elevator shaft self-learning method based on reinforcement learning, and the method includes:
[0007] After confirming that the elevator preparation work is completed, obtain multi-modal perception data of the hoistway environment in the low-speed operation mode;
[0008] Generate a hoistway feature matrix based on the laser SLAM technology and the multi-modal perception data, and introduce the digital twin technology on the basis of the hoistway feature matrix to construct a three-dimensional hoistway reference model;
[0009] Load the pre-trained model of the elevator industry knowledge graph, and perform feature transfer and distillation compression based on the constructed three-dimensional hoistway reference model to generate a control strategy model;
[0010] Deploy the distilled and compressed control strategy model to the edge computing device;
[0011] Perform online reinforcement learning based on the edge computing device, and optimize the control strategy model in real time by designing a multi-dimensional reward function.
[0012] In a specific feasible implementation, the obtaining of the multi-modal perception data of the hoistway environment in the low-speed operation mode includes:
[0013] Obtain spatial point cloud data for describing the relative positions and structural contours of the hoistway wall and guide rail through a lidar;
[0014] Record the rotation of the wire rope or traction sheave for calculating the vertical position change of the elevator in real time through an encoder;
[0015] Obtain the car motion attitude and guide rail inclination through an inertial measurement unit;
[0016] Obtain the current environmental temperature change through a temperature sensor;
[0017] Sense the intensity of local magnetic field interference in the hoistway through a magnetometer.
[0018] In a specific feasible implementation, the generating of the hoistway feature matrix based on the laser SLAM technology and the multi-modal perception data includes:
[0019] Use the laser SLAM algorithm to reconstruct the hoistway space and estimate the trajectory, and generate a dense point cloud map of the internal structure of the hoistway in real time, while obtaining the pose trajectory of the sensor moving with the car. According to each spatial position during the movement of the car along the guide rail, extract the following five-dimensional features and organize them in order into a structured hoistway feature matrix M as shown below:
[0020]
[0021] Among them, x i represents the absolute coordinate of the i-th sampling point in the horizontal direction, y irepresents the absolute coordinate of the i-th sampling point in the vertical direction, θ i represents the inclination angle of the guide rail at this position, T i represents the temperature deformation compensation coefficient at the corresponding position, F i represents the magnetic interference intensity value.
[0022] In a specific feasible implementation, introducing digital twin technology based on the shaft feature matrix to construct a three-dimensional shaft reference model includes:
[0023] The three-dimensional shaft reference model includes a geometric structure model, a physical behavior model, and an electrical control model;
[0024] The geometric structure model restores the spatial layout of components based on point cloud data and coordinate features;
[0025] The physical behavior model establishes a physical response relationship by combining physical elements and simulates the operating state of the car under heat and force conditions;
[0026] The electrical control model maps the response parameters of key components to the shaft space nodes.
[0027] In a specific feasible implementation, loading the pre-trained model of the elevator industry knowledge graph and performing feature transfer and distillation compression based on the constructed three-dimensional shaft reference model to generate a control strategy model includes:
[0028] Taking the constructed three-dimensional shaft reference model and its shaft feature matrix as inputs, load the pre-trained model of the elevator industry knowledge graph;
[0029] The pre-trained model of the elevator industry knowledge graph is trained by a large amount of historical elevator operation data. When the pre-trained model of the elevator industry knowledge graph is loaded, it automatically identifies the key dimension features in the shaft feature matrix, and transfers the operation strategy most similar to the current shaft features in the historical data to the current task domain, and compresses and optimizes the pre-trained model of the elevator industry knowledge graph through feature distillation technology to generate a control strategy model.
[0030] In a specific feasible implementation, performing online reinforcement learning based on the edge computing device and optimizing the control strategy model in real time by designing a multi-dimensional reward function includes:
[0031] Introduce a multi-dimensional hierarchical reward function to evaluate and guide control behaviors. The reward function maps the real-time operating state state of the elevator to a reward value Reward, which is used to measure the quality of the current control behavior. The reward function is as follows:
[0032] reward = r1 + r2 + r3
[0033]
[0034] r2 = -0.3 × energy_consumption
[0035]
[0036] Among them, position_error represents the deviation between the current stopping position of the elevator and the target floor position, energy_consumption represents the energy consumption during a single operation of the elevator, and safety_margin represents the safety margin between the current position and the terminal;
[0037] r1 is the position accuracy reward term, r2 is the energy consumption efficiency reward term, and r3 is the safety reward and punishment term.
[0038] In a specific feasible implementation, the online reinforcement learning based on the edge computing device and the real-time optimization of the control policy model by designing a multi-dimensional reward function further include:
[0039] During the online reinforcement learning process, a reinforcement learning algorithm based on policy gradient is adopted, and the multi-dimensional hierarchical reward function is embedded in the feedback path of the control decision;
[0040] The edge computing device collects the operating state of the hoistway environment. The control policy model inputs and outputs control actions based on the state, obtains environmental feedback data after the actual action is executed, and calls the reward function to calculate the reward value corresponding to the current behavior;
[0041] The control policy model takes the state, action, reward, and next state quadruple as input, and asynchronously trains through the main network and the target network in the dual-network structure to continuously fine-tune the parameters of the policy model.
[0042] In a second aspect, the present application provides an elevator hoistway self-learning system based on reinforcement learning, adopting the following technical solutions:
[0043] An elevator hoistway self-learning system based on reinforcement learning, comprising:
[0044] A data acquisition module, configured to obtain multi-modal perception data of the hoistway environment in a low-speed operation mode after confirming that the elevator preparation work is completed;
[0045] A model construction module, configured to generate a hoistway feature matrix based on the laser SLAM technology and the multi-modal perception data, and introduce the digital twin technology on the basis of the hoistway feature matrix to construct a three-dimensional hoistway reference model;
[0046] A model loading module, configured to load a pre-trained model of an elevator industry knowledge graph, and perform feature transfer and distillation compression based on the constructed three-dimensional hoistway benchmark model to generate a control strategy model;
[0047] A model deployment module, configured to deploy the distilled and compressed control strategy model to an edge computing device;
[0048] A reinforcement learning module, configured to perform online reinforcement learning based on the edge computing device, and optimize the control strategy model in real time by designing a multi-dimensional reward function.
[0049] In a third aspect, the present application provides an electronic device, the device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a self-learning method for an elevator hoistway based on reinforcement learning as described in the first aspect.
[0050] In a fourth aspect, the present application provides a computer-readable storage medium, in which a program is stored, and the program is used to implement a self-learning method for an elevator hoistway based on reinforcement learning as described in the first aspect when executed by a processor.
[0051] In summary, the beneficial effects of the present application at least include:
[0052] (1) In the present application, multi-modal perception data such as laser point cloud, IMU attitude, temperature, magnetic field, and displacement are synchronously collected under low-speed operation, and a dense point cloud map and a spatial trajectory reconstruction are realized by using the SLAM algorithm, and a hoistway feature matrix M including position, attitude, temperature compensation factor, and magnetic field disturbance amount is constructed. On this basis, a digital twin mechanism is introduced to respectively model the geometric structure, physical behavior, and electrical response of the hoistway, and a high-precision three-dimensional restoration of the physical structure and operation boundary of the hoistway is realized. By introducing structured features such as temperature compensation coefficients and magnetic field interference factors, this modeling method dynamically reflects the changes in the hoistway state caused by thermal expansion and contraction, magnetic interference, and wear in space, effectively making up for the defect of insufficient recognition of disturbance changes in traditional static modeling methods, significantly improving the expressiveness of the model to changes in the operating environment and the adaptation boundary of the control strategy, and thus realizing high-precision control under disturbance conditions.
[0053] (2) By constructing an online reinforcement learning mechanism, continuous policy updates are carried out based on the control policy model deployed at the shaft edge. During the control process, a multi-dimensional reward function that integrates position error, energy consumption, and safety margin is introduced, enabling the system to not only minimize docking errors during operation but also synchronously reduce energy consumption and actively avoid control paths with low safety margins. Through a dual-network asynchronous update structure, the main network is responsible for policy execution, and the target network calculates the expected return, avoiding gradient oscillation and overfitting problems during the learning process and improving convergence stability and the generalization ability of control decisions. This mechanism breaks the fixity of traditional control policies based on empirical rules and realizes the continuous adaptive evolution of control policies in the face of complex dynamic working conditions such as load changes, equipment aging, and environmental changes, enabling the system to maintain high control performance and system safety during long-term operation.
[0054] First, multi-modal data of sensors such as lidar, IMU, encoder, temperature sensor, and magnetometer are collected in the low-speed mode to establish a shaft feature matrix, and a three-dimensional shaft model is reconstructed through laser SLAM and sensor fusion. Then, combined with the pre-trained knowledge graph model, a control policy model is generated through feature transfer and distillation compression and deployed to edge devices to build an operation platform with local inference capabilities. Finally, a multi-dimensional reward function and a dual-network asynchronous update mechanism are introduced to carry out online reinforcement learning and continuously optimize the control policy parameters. This solution improves the robustness of the model to structural disturbances by introducing high-precision spatial modeling, enhances the real-time perception ability of dynamic disturbances through multi-source sensing perception, and realizes the continuous adaptive update of control policies to environmental non-linear changes through reinforcement learning, effectively solving the defects of existing technologies in dealing with operation dynamic disturbances and safety control, and significantly improving the docking accuracy, safety margin, and operation stability of the system.
[0055] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly and implement it according to the content of the specification, the following takes the preferred embodiments of this application and combines with the drawings to describe in detail as follows. Description of the Drawings
[0056] Figure 1 It is a flowchart of the elevator shaft self-learning method based on reinforcement learning in the embodiment of this application.
[0057] Figure 2 It is a block diagram of the elevator shaft self-learning system based on reinforcement learning in the embodiment of this application.
[0058] Figure 3 It is a block diagram of the electronic device for elevator shaft self-learning based on reinforcement learning in the embodiment of this application. Detailed Embodiments
[0059] The following will further describe in detail the specific implementation manners of the present application with reference to the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0060] Optionally, the present application takes the elevator shaft self-learning method based on reinforcement learning provided in each embodiment as an example for illustration in an electronic device. The electronic device is a terminal or a server. The terminal can be a computer, a tablet computer, etc. The type of the electronic device is not limited in this embodiment.
[0061] Referring to Figure 1 , which is a schematic flowchart of an elevator shaft self-learning method based on reinforcement learning provided by an embodiment of the present application. The method at least includes the following steps:
[0062] Step S101: After confirming that the elevator preparation work is completed, obtain multi-modal perception data of the shaft environment in the low-speed operation mode.
[0063] In step S101, the elevator system needs to be preprocessed first to ensure that the subsequent self-learning process starts in a safe and stable environment. Specifically, in implementation, it should be confirmed that the safety circuit of the elevator control system is closed normally, the elevator is in the maintenance or learning mode, ordinary call instructions are blocked, and relevant mechanical devices (such as carriages, limiters, guide rail lubrication mechanisms, etc.) are in a controllable state. At the same time, on-site personnel should be arranged to conduct a comprehensive inspection of the shaft environment, remove sundries or abnormal components that may affect the sensor accuracy, and ensure that each floor landing area has recognizable magnets or magnetic isolation plates to assist in preliminary positioning. After the preparation work is completed, the perception data collection stage can be entered.
[0064] In implementation, the elevator car is placed at the starting position at the bottom of the hoistway and runs upward along the hoistway in a low-speed operation mode (typical speed range is 0.1 - 0.3 m / s) to the top floor. During the operation, various types of sensors deployed on the car are used to synchronously collect the environmental characteristics along the hoistway. Specifically, the spatial point cloud data obtained by lidar (such as SICK LMS511) is used to describe the relative positions and structural profiles of the hoistway walls and guide rails; the number of rotations of the wire rope or traction sheave is recorded in real time by an encoder to calculate the vertical position change of the elevator; the motion attitude of the car and the inclination angle of the guide rail are obtained by an inertial measurement unit (IMU); the current environmental temperature change is obtained by a temperature sensor for subsequent estimation of the influence of guide rail thermal expansion; and the local magnetic field interference intensity in the hoistway is sensed by a magnetometer to assist in door area identification and sensing stability evaluation. The above-mentioned various sensors scan the hoistway at a certain spatial frequency, and the sampling period is usually between 10 ms and 20 ms to ensure that high-time-resolution and high-precision multi-modal data can be obtained throughout the process of the elevator running along the hoistway. During the data collection process, a corresponding timestamp and current physical position information need to be marked for each piece of sensed data to construct a cross-modal synchronous data set, providing a basis for subsequent structured modeling.
[0065] After this stage is completed, a set of original multi-modal sensed data describing the hoistway environmental state will be obtained. This data set contains spatial structure parameters, dynamic response information, and local interference characteristics, which are the core inputs for constructing a high-precision hoistway model. The successful execution of this step directly determines the boundary conditions and credibility of subsequent model construction and strategy optimization. Therefore, it is necessary to ensure data integrity, the temporal consistency of the collection process, and the effective control of environmental interference.
[0066] Step S102: Generate a hoistway feature matrix based on laser SLAM technology and multi-modal sensed data, and introduce digital twin technology on the basis of the hoistway feature matrix to construct a three-dimensional hoistway reference model.
[0067] In step S102, the collected multi-modal sensed data is jointly processed. Through laser SLAM technology and multi-source sensor fusion, a structured hoistway feature matrix is generated, and on this basis, digital twin technology is introduced to construct a high-precision and highly restored three-dimensional hoistway reference model.
[0068] Specifically, first, the multi-modal sensed data collected in the low-speed operation mode is subjected to spatio-temporal alignment and collaborative fusion. This sensed data includes, but is not limited to, the point cloud information output by lidar, the attitude and angular velocity information provided by the inertial measurement unit (IMU), the vertical position data recorded by the encoder, the environmental temperature collected by the temperature sensor, and the magnetic field interference signal detected by the magnetometer, etc. Through data timestamp synchronization and coordinate system unification, a spatial reference benchmark for the hoistway operation environment is constructed.
[0069] On this basis, the laser SLAM algorithm is used to reconstruct the hoistway space and estimate the trajectory, generating a dense point cloud map of the internal structure of the hoistway in real time, and at the same time obtaining the pose trajectory of the sensor moving with the car. Further, according to each spatial position during the movement of the car along the guide rail, the following five-dimensional features are extracted and organized into a structured hoistway feature matrix M as follows:
[0070]
[0071] where x i represents the absolute coordinate of the i-th sampling point in the horizontal direction, which is obtained by the combined calculation of the lidar and the encoder. y i represents the absolute coordinate of the i-th sampling point in the vertical direction, which is measured by the laser rangefinder. θ i represents the inclination angle of the guide rail at this position, which is measured by the IMU. T i represents the temperature deformation compensation coefficient at the corresponding position, which is estimated by combining the ambient temperature and the material thermal expansion parameters and is used to improve the geometric stability of the model under environmental disturbances. F i represents the magnetic interference intensity value, which is collected by the magnetometer and is used to reflect the distribution of the electromagnetic environment in the hoistway space.
[0072] After the hoistway feature matrix is constructed, a digital twin modeling mechanism is further introduced based on this matrix to perform three-dimensional reconstruction of the hoistway environment. This modeling process comprehensively considers the physical structure, operating parameters, and electrical characteristics in the hoistway, and constructs the following three types of models respectively:
[0073] Geometric structure model: Restore the spatial layout of components such as guide rails, hoistway walls, and buffers based on the point cloud data and coordinate features, and achieve three-dimensional spatial positioning with millimeter-level accuracy;
[0074] Physical behavior model: Establish a physical response relationship by combining elements such as the guide rail inclination angle and temperature compensation coefficient to simulate the operating state of the car under heat and force conditions;
[0075] Electrical control model: Map the response parameters of key components such as encoders and frequency converters to the hoistway space nodes to support the state recognition and strategy playback of subsequent controllers.
[0076] Through the above modeling process, a three-dimensional hoistway reference model with high precision, high integrity, and high scene restoration ability is finally formed. This model not only visually displays the internal structure of the hoistway but also has practical value for subsequent strategy optimization, self-learning deduction, and fault prediction.
[0077] Step S103: Load the pre-trained model of the elevator industry knowledge graph, and perform feature transfer and distillation compression based on the constructed three-dimensional hoistway reference model to generate a control strategy model.
[0078] In step S103, taking the constructed three-dimensional shaft benchmark model and its corresponding structured feature representation (i.e., the shaft feature matrix) as the input, the pre-trained model of the elevator industry knowledge graph is loaded. The pre-trained model of the elevator industry knowledge graph is trained by a large amount of historical elevator operation data (with a scale of more than 100,000 units), contains rich elevator operation rules, fault evolution paths, and typical structure coping strategies, and has strong generalization ability across shaft structures. When loading, the pre-trained model will automatically identify the key dimension features in the shaft feature matrix, and transfer the operation strategy most similar to the current shaft features in the historical data to the current task domain through the feature transfer mechanism, thereby forming a set of candidate control strategy templates.
[0079] In addition, preferably, considering that the deployment scenario is limited by the computing power and energy consumption of the elevator control system, this step also compresses and optimizes the above pre-trained model through feature distillation technology. This process uses the teacher-student model architecture to extract the core strategy logic, response mode, and control structure in the knowledge graph model, and trains a lightweight control strategy network with a smaller volume and a more concise computational graph, enabling it to operate efficiently on an embedded platform. Finally, the model volume is compressed to less than 500MB, and the control strategy model optimized for the current shaft structure and equipment conditions is completed, which has the actual deployment conditions without significantly losing control performance.
[0080] To sum up, this step transfers industry-level knowledge to the current elevator shaft control task through four sub-processes: pre-trained model loading - structural feature matching - feature transfer - model distillation and compression, laying a strategy and structure foundation for subsequent edge deployment and online self-learning, and ensuring that the entire system can smoothly transition to the strategy generation stage after cognitive modeling.
[0081] Step S104: Deploy the control strategy model compressed by distillation to the edge computing device.
[0082] In step S104, further deploy the above control strategy model compressed by distillation to the edge AI computing device at the shaft end to build an operation platform supporting real-time response and adaptive control. The selected edge computing device usually integrates a neural network acceleration module, has medium computing power and high energy efficiency ratio, and can meet the requirements of elevator operation control for response delay, operation stability, and system enclosure.
[0083] During the deployment process, the embedded encapsulation and runtime optimization of the model are first completed, and an edge inference interface layer for the hoistway scenario is constructed to receive real-time status data and use it as the input of the policy model. At the same time, to ensure the operational continuity and security of the system, basic fault tolerance management logic and a remote update interface are also deployed on the edge device, allowing the subsequent optimized model obtained from training to be pushed without interrupting the operation of the main system. Through the edge deployment in this step, the elevator control system has the ability to locally and independently execute the policy model in the hoistway environment, providing an operation and interaction carrier for the subsequent online reinforcement learning stage. At the same time, this step also realizes the transition from policy generation to local execution and feedback, which is an important central link for implementing a closed-loop self-learning system.
[0084] Step S105: Perform online reinforcement learning based on the edge computing device, and optimize the control policy model in real time by designing a multi-dimensional reward function.
[0085] In step S105, based on the AI computing device deployed at the edge of the elevator hoistway, the online reinforcement learning process of the control policy model relying on the edge computing device is officially started. This process takes the real-time perception data of the elevator operation state as the input, and continuously updates the internal parameters of the policy model through the reinforcement learning algorithm to achieve the continuous optimization and adaptive evolution of the control policy. The core goal of the system is to use the actual operation environment as the training sample source, and optimize the elevator control behavior through interactive feedback, so that it has good accuracy, energy consumption and safety performance under different working conditions.
[0086] Specifically, this step introduces a multi-dimensional hierarchical reward function to evaluate and guide the control behavior. This function maps the real-time operation state state of the elevator to a reward value Reward, which is used to measure the quality of the current control behavior. The reward function is as follows:
[0087] reward = r1 + r2 + r3
[0088]
[0089] r2 = -0.3 × energy_consumption
[0090]
[0091] Among them, position_error represents the deviation between the current stopping position of the elevator and the target floor position, energy_consumption represents the energy consumption during a single operation of the elevator, and safety_margin represents the safety margin between the current position and the terminal. r1 is the position accuracy reward term, which encourages precise docking through a negative correlation mechanism, that is, the smaller the error, the higher the reward. r2 is the energy consumption efficiency reward term, which reflects the optimization goal of energy-saving control. By imposing a linear penalty on high-energy-consuming behaviors, it guides the strategy to converge towards low energy consumption. r3 is the safety reward and punishment term. When the safety margin is insufficient, the system will severely punish the current policy behavior to prevent the elevator from running to a dangerous position and ensure the safety of users' lives and the stable operation of the system.
[0092] In implementation, during the online reinforcement learning process, a reinforcement learning algorithm based on policy gradients is adopted, and the above reward function is embedded in the feedback path of control decisions. Its usage process includes: the edge AI device collects the operating state of the hoistway environment, including basic perception data such as position error, energy consumption value, and safety margin. The current control policy model outputs control actions based on the state input. After the actual action is executed, it obtains the environmental feedback data (the next moment state) and calls the reward function to calculate the reward value corresponding to the current behavior. The control policy model takes the quadruple of state, action, reward, and next state as input and asynchronously trains through the main network and the target network in the dual-network structure to continuously fine-tune the parameters of the policy model. With the continuous accumulation of operation data, the system will continuously optimize the control policy and gradually form an adaptive optimal policy under different changing conditions such as load, temperature, and wear.
[0093] In addition, preferably, to achieve an efficient and stable policy optimization mechanism, this step adopts a dual-network asynchronous update architecture: decouple the main policy network and the target valuation network, and iteratively update the policy parameters asynchronously. The main network is used for real-time decision-making to generate control inputs, and the target network is used to calculate the expected return value (i.e., the future cumulative reward). This design can effectively alleviate gradient oscillation, avoid overfitting, and improve the stability of learning convergence.
[0094] In summary, the present application first collects multi-modal data of sensors such as lidar, IMU, encoder, temperature sensor, and magnetometer in the low-speed mode, establishes a shaft feature matrix, and reconstructs a three-dimensional shaft model through laser SLAM and sensor fusion; then combines with a pre-trained knowledge graph model, generates a control strategy model through feature transfer and distillation compression, and deploys it to an edge device to construct an operating platform with local inference ability; finally, introduces a multi-dimensional reward function and a dual-network asynchronous update mechanism to carry out online reinforcement learning and continuously optimize the control strategy parameters. This solution improves the robustness of the model against structural disturbances by introducing high-precision spatial modeling, enhances the real-time perception ability of dynamic disturbances through multi-source sensing, and realizes the continuous adaptive update of the control strategy to the non-linear changes of the environment through reinforcement learning, effectively solving the defects of the existing technology in dealing with operating dynamic disturbances and safety control, and significantly improving the docking accuracy, safety margin, and operating stability of the system.
[0095] Figure 2 FIG. is a structural block diagram of an elevator shaft self-learning system based on reinforcement learning provided by an embodiment of the present application. The system at least includes the following modules:
[0096] A data acquisition module, configured to obtain multi-modal perception data of the shaft environment in the low-speed operation mode after confirming that the elevator preparation work is completed;
[0097] A model construction module, configured to generate a shaft feature matrix based on laser SLAM technology and multi-modal perception data, and introduce digital twin technology on the basis of the shaft feature matrix to construct a three-dimensional shaft reference model;
[0098] A model loading module, configured to load a pre-trained model of an elevator industry knowledge graph, and perform feature transfer and distillation compression based on the constructed three-dimensional shaft reference model to generate a control strategy model;
[0099] A model deployment module, configured to deploy the distilled and compressed control strategy model to an edge computing device;
[0100] A reinforcement learning module, configured to perform online reinforcement learning based on the edge computing device, and optimize the control strategy model in real time by designing a multi-dimensional reward function.
[0101] For related details, refer to the above method embodiment.
[0102] Figure 3 FIG. is a block diagram of an electronic device provided by an embodiment of the present application. The device at least includes a processor 401 and a memory 402.
[0103] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0104] The memory 402 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 402 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 402 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 401 to implement the elevator shaft self-learning method based on reinforcement learning provided in the method embodiments of the present application.
[0105] In some embodiments, the electronic device may optionally further include: a peripheral device interface and at least one peripheral device. The processor 401, the memory 402, and the peripheral device interface may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface through a bus, signal lines, or a circuit board. Schematically, the peripheral devices include but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply, etc.
[0106] Of course, the electronic device may also include fewer or more components, and this embodiment does not limit this.
[0107] Optionally, the present application also provides a computer-readable storage medium, and a program is stored in the computer-readable storage medium, and the program is loaded and executed by the processor to implement the elevator shaft self-learning method based on reinforcement learning in the above method embodiments.
[0108] Optionally, the present application also provides a computer product, which includes a computer-readable storage medium. A program is stored in the computer-readable storage medium and is loaded and executed by a processor to implement the elevator shaft self-learning method based on reinforcement learning in the above method embodiments.
[0109] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0110] The above embodiments only express several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. An elevator shaft self-learning method based on reinforcement learning, characterized in that The method includes: After confirming that the elevator preparation work is completed, obtaining multi-modal perception data of the hoistway environment in the low-speed operation mode; Generating a hoistway feature matrix based on the laser SLAM technology and the multi-modal perception data, and introducing the digital twin technology on the basis of the hoistway feature matrix to construct a three-dimensional hoistway reference model; Loading a pre-trained model of the elevator industry knowledge graph, and performing feature transfer and distillation compression based on the constructed three-dimensional hoistway reference model to generate a control strategy model; Deploying the distilled and compressed control strategy model to an edge computing device; Performing online reinforcement learning based on the edge computing device, and optimizing the control strategy model in real time by designing a multi-dimensional reward function.
2. The self-learning method for an elevator shaft based on reinforcement learning according to claim 1, wherein The obtaining multi-modal perception data of the hoistway environment in the low-speed operation mode includes: Obtaining spatial point cloud data for describing the relative positions and structural contours of the hoistway wall and the guide rail through a lidar; Recording in real time the number of rotations of the wire rope or the traction sheave for calculating the vertical position change of the elevator through an encoder; Obtaining the car motion attitude and the guide rail inclination angle through an inertial measurement unit; Obtaining the current environmental temperature change through a temperature sensor; Sensing the local magnetic field interference intensity in the hoistway through a magnetometer.
3. The elevator shaft self-learning method based on reinforcement learning according to claim 1, characterized in that The generating a hoistway feature matrix based on the laser SLAM technology and the multi-modal perception data includes: Using the laser SLAM algorithm to reconstruct the hoistway space and estimate the trajectory, generating a dense point cloud map of the internal structure of the hoistway in real time, and obtaining the pose trajectory of the sensor moving with the car at the same time. According to each spatial position during the movement of the car along the guide rail, the following five-dimensional features are extracted and organized into a structured hoistway feature matrix M as follows: Among them, x i represents the absolute coordinate of the i-th sampling point in the horizontal direction, and y i represents the absolute coordinate of the i-th sampling point in the vertical direction, and θ i represents the inclination angle of the guide rail at this position, and T i represents the temperature deformation compensation coefficient at the corresponding position, and F i represents the magnetic interference intensity value.
4. The elevator shaft self-learning method based on reinforcement learning according to claim 1, characterized in that The introducing the digital twin technology on the basis of the hoistway feature matrix to construct a three-dimensional hoistway reference model includes: The three-dimensional hoistway reference model includes a geometric structure model, a physical behavior model, and an electrical control model; The geometric structure model restores the spatial layout of the components according to the point cloud data and coordinate features; The physical behavior model establishes a physical response relationship by combining physical elements, and simulates the operating state of the car under heat and force conditions; The electrical control model maps the response parameters of the key components to the hoistway space nodes.
5. The elevator shaft self-learning method based on reinforcement learning according to claim 1, characterized in that The loading a pre-trained model of the elevator industry knowledge graph, and performing feature transfer and distillation compression based on the constructed three-dimensional hoistway reference model to generate a control strategy model includes: Taking the constructed three-dimensional hoistway reference model and its hoistway feature matrix as inputs, and loading a pre-trained model of the elevator industry knowledge graph; The pre-trained model of the elevator industry knowledge graph is trained by a large amount of historical elevator operation data. When the pre-trained model of the elevator industry knowledge graph is loaded, it automatically identifies the key dimension features in the hoistway feature matrix, and transfers the operation strategy most similar to the current hoistway feature in the historical data to the current task domain through a feature transfer mechanism, and compresses and optimizes the pre-trained model of the elevator industry knowledge graph through a feature distillation technology to generate a control strategy model.
6. The elevator shaft self-learning method based on reinforcement learning according to claim 1, characterized in that The performing online reinforcement learning based on the edge computing device, and optimizing the control strategy model in real time by designing a multi-dimensional reward function includes: A multi-dimensional hierarchical reward function is introduced to evaluate and guide control behaviors. The reward function maps the real-time operating state state of the elevator to a reward value Reward, which is used to measure the quality of the current control behavior. The reward function is as follows: reward = r1 + r2 + r3 r2 = -0.3 × energy_consumption where position_error represents the deviation between the current stopping position of the elevator and the target floor position, energy_consumption represents the energy consumption during a single operation of the elevator, and safety_margin represents the safety margin between the current position and the terminal; r1 is the position accuracy reward term, r2 is the energy consumption efficiency reward term, and r3 is the safety reward and punishment term.
7. The elevator shaft self-learning method based on reinforcement learning according to claim 6, characterized in that The online reinforcement learning based on the edge computing device and the real-time optimization of the control policy model by designing a multi-dimensional reward function further includes: During the online reinforcement learning process, a reinforcement learning algorithm based on policy gradients is adopted, and the multi-dimensional hierarchical reward function is embedded in the feedback path of the control decision; The edge computing device collects the operating state of the hoistway environment. The control policy model inputs and outputs control actions based on the state, obtains environmental feedback data after the actual action is executed, and calls the reward function to calculate the reward value corresponding to the current behavior; The control policy model takes a quadruple of state, action, reward, and next state as input, and asynchronously trains through the main network and the target network in the dual-network structure to continuously fine-tune the parameters of the policy model.
8. An elevator shaft self-learning system based on reinforcement learning, characterized in that, It includes: A data collection module, which is used to obtain multi-modal perception data of the hoistway environment in the low-speed operation mode after confirming that the elevator preparation work is completed; A model construction module, which is used to generate a hoistway feature matrix based on the laser SLAM technology and the multi-modal perception data, and introduce the digital twin technology on the basis of the hoistway feature matrix to construct a three-dimensional hoistway reference model; A model loading module, which is used to load the pre-trained model of the elevator industry knowledge graph, and perform feature transfer and distillation compression based on the constructed three-dimensional hoistway reference model to generate a control policy model; A model deployment module, which is used to deploy the distilled and compressed control policy model to the edge computing device; A reinforcement learning module, which is used to perform online reinforcement learning based on the edge computing device, and real-time optimize the control policy model by designing a multi-dimensional reward function.
9. An electronic device, characterized in that, The device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a self-learning method for an elevator hoistway based on reinforcement learning as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A program is stored in the storage medium, and when the program is executed by the processor, it is used to implement a self-learning method for an elevator hoistway based on reinforcement learning as described in any one of claims 1 to 7.
Citation Information
Cited By
Elevator leveling method and system
CN121107201A