Reinforcement learning-based crane sling anti-swing method and system
Through the reinforcement learning algorithm based on deep Q network, a crane spreader anti-sway model was constructed and trained, which solved the problem of limited anti-sway effect of traditional methods under complex working conditions, realized adaptive control of crane spreaders, and improved operational efficiency and safety.
Patent Information
- Application Number
- CN202510988610.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-16
AI Technical Summary
Traditional crane sling anti-sway control methods are difficult to adapt to complex and changeable working conditions, and the anti-sway effect is limited, which cannot meet the needs of efficient and safe operations.
A reinforcement learning algorithm based on deep Q-network is used to build and train a deep Q-network model by defining the state space, action space and reward function. This method can optimize the spreader control strategy in real time and suppress spreader swing.
It realizes adaptive control of the crane under complex working conditions, improves operating efficiency and safety, and effectively suppresses spreader swing.
Smart Images

Figure CN120646681A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of crane control, and in particular relates to a crane spreader anti-sway method and system based on reinforcement learning. Background Art
[0002] Spreader sway is a common and thorny problem in crane operations. It not only reduces crane efficiency and increases loading and unloading time, but can also cause accidents and pose safety risks to operators and the surrounding environment.
[0003] Traditional crane spreader anti-sway control methods, including input shaping and fuzzy control, often rely on precise mathematical models or fixed control laws. However, the crane's operating environment is complex and variable. Factors such as the weight and shape of the cargo being hoisted, the length of the sling, and external interference can affect the spreader's motion characteristics. These traditional methods struggle to adapt to these complex and changing operating conditions, resulting in limited anti-sway effectiveness and failing to meet the demands for higher efficiency and safety. Therefore, a crane spreader anti-sway method that can adapt to these complex operating conditions is urgently needed.
[0004] Chinese invention patent application number "202311619398.2" discloses a method and device for gantry crane anti-sway control based on deep reinforcement learning. The reinforcement learning model is based on a dual-Q network model, separating action selection and value estimation. This separate calculation requires an additional forward pass of the network to separate action selection and value estimation, resulting in high computational complexity and significant computational resources required for both training and deployment. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for anti-sway of crane spreaders based on reinforcement learning. Through the reinforcement learning algorithm, the crane can adapt to different working scenarios and working conditions, optimize the control strategy of the spreader in real time, effectively suppress the swing of the spreader, and improve the working efficiency and safety of the crane.
[0006] In a first aspect, the present invention provides a crane spreader anti-sway method based on reinforcement learning, comprising the following steps:
[0007] S1. Establishing a crane spreader swing model: establishing a dynamic equation of the crane spreader;
[0008] S2. Define the state space, action space, and reward function: including:
[0009] S21. When defining a state space, select a variable describing the current motion state of the crane spreader to obtain a state vector of the crane spreader;
[0010] S22. When defining the action space, controlling the force or acceleration of the crane spreader to obtain an action set of the crane spreader;
[0011] S23, when defining the reward function, by quantifying the quality of the crane spreader behavior, guiding the agent to gradually learn the optimal anti-sway control strategy;
[0012] S3. Build a deep Q network: Use the deep Q network method, using a deep neural network as an approximator of the Q function, with the input being the state variables in the state space and the output being the Q value of each action;
[0013] S4. Training the Deep Q Network: Performing reinforcement learning training by interacting with the crane spreader anti-sway system. During training, the agent selects actions based on the current state, and the actuator module obtains a new state and reward after executing the action.
[0014] S5. Apply the trained deep Q network for anti-sway control: In actual operation, the state vector of the crane spreader is obtained and input into the trained deep Q network. The deep Q network outputs the Q value of each action and selects the action with the largest Q value as the control action to control the operation of the crane and suppress the swing of the crane spreader.
[0015] Preferably, in step S1, the crane trolley is defined to run in the horizontal direction, and the crane spreader is suspended on the crane trolley through a rope; the crane spreader is regarded as a mass point, the length of the rope is defined as 1, and the mass of the rope is neglected; then, according to Newton's law, the dynamic equation of the crane spreader in the horizontal direction is established as formula (1):
[0016]
[0017] Wherein, m is the total mass of the spreader and the load, F is the horizontal force exerted by the crane trolley on the crane spreader, g is the acceleration due to gravity, and θ is the swing angle of the crane spreader.
[0018] Preferably, in step S21, the variables of the current motion state of the crane spreader include the position x, velocity v, swing angle θ, and angular velocity ω of the crane spreader, and a state vector S = (x, v, θ, ω) of the crane spreader is obtained; and the continuous variables of the current motion state of the crane spreader are discretized, and the operating speed and operating angle of the crane trolley are measured by a speed sensor and an angle sensor, respectively.
[0019] In step S22, the actions include acceleration of the crane spreader in different directions, remaining stationary, or different force levels;
[0020] In step S23, a positive reward is given when θ is close to 0, a negative reward is given when θ is too large, a large positive reward is given when the target position is reached, and excessive control actions are penalized; the reward function R is defined as formula (2):
[0021] R=-|θ|-0.1|ω|+0.5v max -|vv desired | (2)
[0022] Wherein, |θ| and |ω| represent the absolute values of the swing angle and the swing angular velocity of the crane spreader, respectively, and are used to penalize the increase of the swing angle and the swing angular velocity, v max is the maximum permissible speed of the crane trolley, 0.5v max Reward for the crane trolley to operate within a reasonable speed range, |vv desired | represents the deviation between the actual speed of the crane trolley and the expected speed, and is used to penalize speed deviation.
[0023] Preferably, in step S21, three actions are defined: positive force, negative force, and no force applied, and the action set A = {-1, 0, 1}; wherein 1 represents a positive force, 0 represents no force applied, and -1 represents a negative force.
[0024] Preferably, in step S3, the deep Q network adopts a three-layer fully connected neural network, and the network structure includes an input layer, a hidden layer and an output layer; the hidden layer adopts a ReLU activation function, and the output layer adopts a linear activation function.
[0025] Preferably, the input layer has 4 neurons, corresponding to the 4 elements of the state vector; there are two hidden layers, the first hidden layer has 32 neurons, and the second hidden layer has 16 neurons; the output layer has 3 neurons, corresponding to the 3 actions in the action space.
[0026] Preferably, step S4 includes:
[0027] S41. Initialize the parameters of the deep Q network and the target network, set the initial state S0, the reward discount factor and the maximum number of iterations;
[0028] S42, obtaining the motion state of the crane spreader at the current moment, determining the action parameters according to the greedy strategy, and the agent executing the set values until the next moment;
[0029] S43, transfer the sample (S t , A t , R t , S t+1 ) is stored in the experience replay buffer, a batch of samples are randomly drawn from the experience replay buffer, the reward obtained is calculated and the Q value is updated;
[0030] S44. Determine whether the learning is completed. If the maximum number of iterations has been reached, the learning is completed and the optimal strategy is obtained. If the maximum number of iterations has not been reached, the learning is not completed and steps S42 and S43 are repeated.
[0031] Preferably, in step S41, the capacity of the experience playback buffer is set.
[0032] Preferably, in step S41, the capacity of the experience playback buffer is set to 10000;
[0033] In step S42, an ε-greedy strategy is used to select actions, with an initial ε of 0.9, which gradually decays to 0.1 as training progresses. For each training round, the maximum time step is 1000. When the swing angle of the crane spreader is less than a set threshold, the crane spreader is considered stable and the round ends early.
[0034] In step S43, at each time step, after executing the action, a new state and reward are obtained, and the transferred samples are stored in the experience replay buffer. When the number of samples in the experience replay buffer exceeds 1000, a batch of samples is randomly sampled for training, and the rewards obtained are calculated and the Q value is updated.
[0035] In step S44, the parameters of the deep Q network are copied to the target network every 100 time steps.
[0036] In a second aspect, the present invention provides a crane spreader anti-sway system based on reinforcement learning, which is used to implement the crane spreader anti-sway method based on reinforcement learning described in the first aspect of the present invention, comprising:
[0037] A sensor module, comprising an angle sensor, an acceleration sensor, and a displacement sensor, for detecting the swing angle, swing acceleration, and displacement of the crane spreader in real time, and obtaining motion state information of the crane trolley;
[0038] A controller module, which uses a computer or microprocessor to run the reinforcement learning algorithm and output control instructions;
[0039] The actuator module includes the motor driver and brake of the crane trolley and is used to adjust the motion state of the crane trolley according to control instructions.
[0040] The crane spreader anti-sway method and system based on reinforcement learning provided by the present invention have the following beneficial effects:
[0041] The reinforcement learning-based crane sling anti-sway method and system of the present invention, by defining appropriate state space, action space and reward function, constructing and training a reinforcement learning model, can effectively suppress sling swing and improve the efficiency and safety of crane operations; it uses the reinforcement learning algorithm to enable the crane to adapt to different working scenarios and working conditions, and optimize the sling control strategy in real time. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is a flowchart of a crane spreader anti-sway method based on reinforcement learning provided by one embodiment of the present invention.
[0043] Figure 2 The figure is a schematic diagram of the structure of a reinforcement learning system of a crane spreader anti-sway system based on reinforcement learning provided by one embodiment of the present invention.
[0044] Figure 3 This is a parameter control flow chart of a crane spreader anti-sway method based on reinforcement learning provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0045] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0046] Example 1
[0047] Please refer to Figure 1 and Figure 2 This embodiment provides a crane sling anti-sway method based on reinforcement learning, comprising the following steps:
[0048] S1. Establishing a crane spreader swing model: Establishing the dynamic equations of the crane spreader;
[0049] S2. Define the state space, action space, and reward function: including:
[0050] S21. When defining the state space, select variables describing the current motion state of the crane spreader to obtain the state vector of the crane spreader;
[0051] S22. When defining the action space, control the force or acceleration of the crane spreader to obtain the action set of the crane spreader;
[0052] S23, when defining the reward function, by quantifying the quality of the crane spreader behavior, guiding the agent to gradually learn the optimal anti-sway control strategy;
[0053] S3. Build a deep Q network: Use the deep Q network method, using a deep neural network as an approximator of the Q function, with the input being the state variables in the state space and the output being the Q value of each action;
[0054] S4. Training the Deep Q Network: This is done through reinforcement learning training by interacting with the crane's anti-sway system. During training, the agent selects actions based on its current state, and the actuator module executes the action to obtain a new state and reward.
[0055] S5. Apply the trained deep Q network for anti-sway control: In actual operation, the state vector of the crane spreader is obtained and input into the trained deep Q network. The deep Q network outputs the Q value of each action and selects the action with the largest Q value as the control action to control the operation of the crane and suppress the swing of the crane spreader.
[0056] Specifically, in step S1, the crane trolley is defined to run in the horizontal direction, and the crane spreader is suspended on the crane trolley through a rope. The crane spreader is regarded as a mass point, the length of the rope is defined as 1, and the mass of the rope is ignored. Then, according to Newton's law, the dynamic equation of the crane spreader in the horizontal direction is established as formula (1):
[0057]
[0058] Where m is the total mass of the spreader and the load, F is the horizontal force exerted by the crane trolley on the crane spreader, g is the acceleration due to gravity, and θ is the swing angle of the crane spreader.
[0059] Specifically, the state space, action space, and reward function are designed for the crane spreader system:
[0060] In step S21, the variables of the crane spreader's current motion state include the crane spreader's position x, velocity v, swing angle θ, and angular velocity ω, and the crane spreader's state vector S = (x, v, θ, ω) is obtained. Furthermore, the continuous variables of the crane spreader's current motion state are discretized, and the operating speed and operating angle of the crane trolley are measured by a speed sensor and an angle sensor, respectively.
[0061] In other words, the state space S includes the crane spreader's position, velocity, swing angle, and angular velocity. These variables describe the crane spreader's current motion state. Assuming the crane spreader's position is x, its velocity is v, its swing angle is θ, and its angular velocity is ω, then the state can be expressed as S = (x, v, θ, ω). In practical applications, these continuous state variables need to be discretized.
[0062] In step S22, the actions include acceleration of the crane spreader in different directions, remaining stationary, or different force levels;
[0063] That is, the action space A includes the forces or accelerations that control the movement of the crane spreader. Actions can include accelerations in different directions, holding still, or different force levels.
[0064] In step S23, a positive reward is given when θ is close to 0, a negative reward is given when θ is too large, a large positive reward is given when the target position is reached, and excessive control actions are penalized; the reward function R is defined as formula (2):
[0065] R=-|θ|-0.1|ω|+0.5v max -|vv desired | (2)
[0066] Among them, |θ| and |ω| represent the absolute values of the swing angle and swing angular velocity of the crane spreader, respectively, which are used to penalize the increase of the swing angle and swing angular velocity, v max is the maximum permissible speed of the crane trolley, 0.5v max Indicates a reward for the crane trolley to operate within a reasonable speed range, |vv desired |Indicates the deviation between the actual speed of the crane trolley and the expected speed, which is used to penalize the speed deviation.
[0067] In other words, the goal of anti-sway is to reduce oscillation while simultaneously reaching the target position as quickly as possible. Therefore, the reward function encourages reducing the oscillation angle θ while also considering the speed at which the target position is reached. Therefore, the reward is designed to give a positive reward when θ is close to 0, a negative reward when θ is too large, and a large positive reward when the target position is reached. At the same time, excessive control actions need to be penalized to prevent instability caused by violent movements.
[0068] Specifically, in step S21, three actions are defined: positive force, negative force, and no force applied, and the action set A = {-1, 0, 1}; where 1 represents a positive force, 0 represents no force applied, and -1 represents a negative force.
[0069] That is, the method of this embodiment designs three actions: positive force, negative force, and no force, so the action space is these three options.
[0070] After defining the state space, action space, and reward function, a reinforcement learning model is selected. The method of this embodiment adopts the deep Q network (DQN) method, using the deep neural network as an approximator of the Q function, with the input being the state variables in the state space and the output being the Q value of each action.
[0071] Specifically, in step S3, the deep Q network adopts a three-layer fully connected neural network, and the network structure includes an input layer, a hidden layer and an output layer; the hidden layer adopts a ReLU activation function, and the output layer adopts a linear activation function.
[0072] Specifically, the input layer has 4 neurons, corresponding to the 4 elements of the state vector; there are two hidden layers, the first hidden layer has 32 neurons, and the second hidden layer has 16 neurons; the output layer has 3 neurons, corresponding to the 3 actions in the action space.
[0073] In step S4, during the training of the constructed deep Q network, the agent selects an action based on the current state, and the actuator module obtains a new state and reward after executing the action. Experience replay and target network technology are used to stabilize the training process; that is, the agent makes decisions and the actuator module executes action decisions.
[0074] Please refer to Figure 3 Specifically, step S4 includes:
[0075] S41. Initialize the parameters of the deep Q network and the target network, set the initial state S0, the reward discount factor and the maximum number of iterations;
[0076] S42, obtaining the motion state of the crane spreader at the current moment, determining the action parameters according to the greedy strategy, and the agent executing the set values until the next moment;
[0077] S43, transfer the sample (S t ,A t ,R t ,S t+1 ) is stored in the experience replay buffer, a batch of samples are randomly drawn from the experience replay buffer, the reward obtained is calculated and the Q value is updated;
[0078] S44. Determine whether the learning is completed. If the maximum number of iterations has been reached, the learning is completed and the optimal strategy is obtained. If the maximum number of iterations has not been reached, the learning is not completed and steps S42 and S43 are repeated.
[0079] Specifically, in step S41, the capacity of the experience playback buffer is set.
[0080] Specifically, in step S41, the capacity of the experience playback buffer is set to 10000;
[0081] In step S42, an ε-greedy strategy is used to select actions, with an initial ε of 0.9, which gradually decays to 0.1 as training progresses. For each training round, the maximum time step is 1000. When the swing angle of the crane spreader is less than the set threshold, the crane spreader is considered stable and the round ends early.
[0082] In step S43, at each time step, after executing the action, the new state and reward are obtained, and the transferred samples are stored in the experience replay buffer. When the number of samples in the experience replay buffer exceeds 1000, a batch of samples are randomly sampled for training, and the rewards obtained are calculated and the Q value is updated.
[0083] In step S44, the parameters of the deep Q network are copied to the target network every 100 time steps.
[0084] Finally, in step S5, the trained reinforcement learning network model is deployed to the actual system to obtain the state information of the spreader in real time, input it into the trained deep Q network, and output the optimal control action to control the operation of the crane trolley, thereby suppressing the swing of the crane spreader.
[0085] In actual operation, the position x, speed v, swing angle θ, and angular velocity ω of the crane spreader are obtained in real time through sensors to form a state vector S, which is input into the trained deep Q network. The network outputs the Q value of each action and selects the action with the largest Q value as the control action to control the operation of the crane trolley, thereby suppressing the swing of the crane spreader.
[0086] Example 2
[0087] Please refer to Figure 1 and Figure 2 This embodiment provides a crane spreader anti-sway system based on reinforcement learning, which is used to implement the crane spreader anti-sway method based on reinforcement learning described in Example 1, including:
[0088] The sensor module includes an angle sensor, an acceleration sensor, and a displacement sensor, which is used to detect the swing angle, swing acceleration, and displacement of the crane spreader in real time and obtain the motion status information of the crane trolley;
[0089] A controller module, which uses a computer or microprocessor to run the reinforcement learning algorithm and output control instructions;
[0090] The actuator module includes the motor driver and brake of the crane trolley, and is used to adjust the motion state of the crane trolley according to the control instructions.
[0091] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0092] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0093] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0094] These computer program instructions can also be loaded onto a computer or other programmable computer device so that a series of operating steps are executed on the computer or other programmable computer device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable computer device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0095] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A crane sling anti-sway method based on reinforcement learning, characterized in that: The following steps are involved: S1. Establishing a crane spreader swing model: establishing a dynamic equation of the crane spreader; S2. Define the state space, action space, and reward function: including: S21. When defining a state space, select a variable describing the current motion state of the crane spreader to obtain a state vector of the crane spreader; S22. When defining the action space, controlling the force or acceleration of the crane spreader to obtain an action set of the crane spreader; S23, when defining the reward function, by quantifying the quality of the crane spreader behavior, guiding the agent to gradually learn the optimal anti-sway control strategy; S3. Build a deep Q network: Use the deep Q network method, using a deep neural network as an approximator of the Q function, with the input being the state variables in the state space and the output being the Q value of each action; S4. Training the Deep Q Network: Performing reinforcement learning training by interacting with the crane spreader anti-sway system. During training, the agent selects actions based on the current state, and the actuator module obtains a new state and reward after executing the action. S5. Apply the trained deep Q network for anti-sway control: In actual operation, the state vector of the crane spreader is obtained and input into the trained deep Q network. The deep Q network outputs the Q value of each action and selects the action with the largest Q value as the control action to control the operation of the crane and suppress the swing of the crane spreader.
2. The crane spreader anti-sway method based on reinforcement learning according to claim 1 is characterized in that: In step S1, the crane trolley is defined to run in the horizontal direction, and the crane spreader is suspended on the crane trolley through a rope; the crane spreader is regarded as a mass point, the length of the rope is defined as 1, and the mass of the rope is neglected; then, according to Newton's law, the dynamic equation of the crane spreader in the horizontal direction is established as formula (1): Wherein, m is the total mass of the spreader and the load, F is the horizontal force exerted by the crane trolley on the crane spreader, g is the acceleration due to gravity, and θ is the swing angle of the crane spreader.
3. The crane sling anti-sway method based on reinforcement learning according to claim 2 is characterized in that: In step S21, the variables of the current motion state of the crane spreader include the position x, velocity v, swing angle θ, and angular velocity ω of the crane spreader, and the state vector S = (x, v, θ, ω) of the crane spreader is obtained. The continuous variables of the current motion state of the crane spreader are discretized, and the operating speed and operating angle of the crane trolley are measured by a speed sensor and an angle sensor, respectively. In step S22, the actions include acceleration of the crane spreader in different directions, remaining stationary, or different force levels; In step S23, a positive reward is given when θ is close to 0, a negative reward is given when θ is too large, a large positive reward is given when the target position is reached, and excessive control actions are penalized; the reward function R is defined as formula (2): R=-|θ|-0.1|ω|+0.5v max -|v-v desired |(2) Wherein, |θ| and |ω| represent the absolute values of the swing angle and the swing angular velocity of the crane spreader, respectively, and are used to penalize the increase of the swing angle and the swing angular velocity, v max is the maximum permissible speed of the crane trolley, 0.5v max Reward for the crane trolley to operate within a reasonable speed range, |vv desired | represents the deviation between the actual speed of the crane trolley and the expected speed, and is used to penalize speed deviation.
4. The crane spreader anti-sway method based on reinforcement learning according to claim 3 is characterized in that: In step S21 , three actions are defined: positive force, negative force, and no force applied, and the action set A={−1, 0, 1}; 1 represents a positive force, 0 represents no force applied, and −1 represents a negative force.
5. The crane spreader anti-sway method based on reinforcement learning according to claim 4 is characterized in that: In step S3, the deep Q network adopts a three-layer fully connected neural network, and the network structure includes an input layer, a hidden layer and an output layer; the hidden layer adopts a ReLU activation function, and the output layer adopts a linear activation function.
6. The crane spreader anti-sway method based on reinforcement learning according to claim 5 is characterized in that: The input layer has 4 neurons, corresponding to the 4 elements of the state vector; there are two hidden layers, the first hidden layer has 32 neurons, and the second hidden layer has 16 neurons; the output layer has 3 neurons, corresponding to the 3 actions in the action space.
7. The crane spreader anti-sway method based on reinforcement learning according to any one of claims 1 to 6, characterized in that: Step S4 includes: S41. Initialize the parameters of the deep Q network and the target network, set the initial state S0, the reward discount factor and the maximum number of iterations; S42, obtaining the motion state of the crane spreader at the current moment, determining the action parameters according to the greedy strategy, and the agent executing the set values until the next moment; S43, transfer the sample (S t , A t , R t , S t+1 ) is stored in the experience replay buffer, a batch of samples are randomly drawn from the experience replay buffer, the reward obtained is calculated and the Q value is updated; S44. Determine whether the learning is completed. If the maximum number of iterations has been reached, the learning is completed and the optimal strategy is obtained. If the maximum number of iterations has not been reached, the learning is not completed and steps S42 and S43 are repeated.
8. The crane spreader anti-sway method based on reinforcement learning according to claim 7 is characterized in that: In step S41, the capacity of the experience playback buffer is set.
9. The crane spreader anti-sway method based on reinforcement learning according to claim 8, characterized in that: In step S41, the capacity of the experience playback buffer is set to 10000; In step S42, an ε-greedy strategy is used to select actions, with an initial ε of 0.9, which gradually decays to 0.1 as training progresses. For each training round, the maximum time step is 1000. When the swing angle of the crane spreader is less than a set threshold, the crane spreader is considered stable and the round ends early. In step S43, at each time step, after executing the action, a new state and reward are obtained, and the transferred samples are stored in the experience replay buffer. When the number of samples in the experience replay buffer exceeds 1000, a batch of samples is randomly sampled for training, and the rewards obtained are calculated and the Q value is updated. In step S44, the parameters of the deep Q network are copied to the target network every 100 time steps.
10. A crane sling anti-sway system based on reinforcement learning, characterized in that: The method for preventing sway of a crane spreader based on reinforcement learning according to any one of claims 1 to 9 comprises: A sensor module, comprising an angle sensor, an acceleration sensor, and a displacement sensor, for detecting the swing angle, swing acceleration, and displacement of the crane spreader in real time, and obtaining motion state information of the crane trolley; A controller module, which uses a computer or microprocessor to run the reinforcement learning algorithm and output control instructions; The actuator module includes the motor driver and brake of the crane trolley and is used to adjust the motion state of the crane trolley according to control instructions.
Citation Information
Patent Citations
Anti-swing control method and device for gantry crane based on deep reinforcement learning
CN117466145A
Cited By
Self-learning control method and system for voltage regulation and speed regulation of crane
CN121386938A
Space truss vibration control method and system based on deep reinforcement learning
CN122239834A
Space truss vibration control method and system based on deep reinforcement learning
CN122239834B