Crane anti-swing positioning control method based on near-end strategy optimization

By applying a deep reinforcement learning control method optimized by proximity strategy in cranes, the problem of limited performance of traditional methods in complex environments is solved, and the stable and high-precision positioning of the crane under complex conditions is achieved.

CN120057753AActive Publication Date: 2025-05-30SPECIAL EQUIP SAFETY SUPERVISION INSPECTION INST OF JIANGSU PROVINCE

Patent Information

Application Number
CN202510518263.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-30
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

Traditional crane control methods have limited performance in complex environments, rely on precise modeling, and have poor adaptability to environmental changes, making it difficult to effectively solve the problem of anti-swing positioning control of cranes.

Method used

The deep reinforcement learning control method based on proximal strategy optimization is adopted to perceive the working environment and motion state of the crane through deep learning, and combine the reinforcement learning of the anti-slope positioning control strategy to directly optimize the control strategy in the actual environment.

Benefits of technology

It significantly improves the stability and adaptability of the crane under different loads, interference and task conditions, reduces swaying, improves positioning accuracy and operation safety, and reduces the threshold for technical implementation and control deviations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120057753A_ABST
    Figure CN120057753A_ABST
Patent Text Reader

Abstract

The invention discloses a crane anti-swing positioning control method based on near-end strategy optimization, which is characterized in that a crane and cable mechanism, a rope length detection module, a position and speed detection module, a swing angle detection module, a deep reinforcement learning controller and a control panel are included; the rope length detection module is used for detecting the rope length # imgabs0 #; the position and speed detection module is used for detecting a trolley position # imgabs 1 #, a cart position # imgabs 2 #, a trolley speed # imgabs 3 # and a cart speed # imgabs 4 #; the swing angle detection module is used for outputting a swing angle # imgabs 5 # and a swing angle velocity # imgabs 6 # of a hanging object; the control panel outputs a target position # imgabs7 # of a hoisted object and a mass # imgabs8 # of the hoisted object; the deep reinforcement learning controller inputs data collected by the above modules as state information, the data is processed and analyzed by the deep reinforcement learning controller to calculate an optimal control signal, and the control signal is transmitted to the large trolley, the small trolley and the cable mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of motion control methods, and particularly to a crane anti-sway positioning control method based on proximal policy optimization. Background Art

[0002] Deep reinforcement learning algorithms have achieved remarkable control effects in fields such as industrial robots and autonomous driving. This algorithm does not require precise modeling of the system. Through continuous interaction between the agent and the environment, it autonomously learns and optimizes the control strategy. Deep reinforcement learning combines the perception ability of deep learning and the decision-making ability of reinforcement learning, and has significant advantages in dealing with complex perception and decision-making problems.

[0003] Traditional crane control methods, such as PID, linear quadratic regulator (LQR), and model predictive control (MPC), have limited performance in complex environments, rely on precise modeling, and have poor adaptability to environmental changes.

[0004] In summary, there is an urgent need to provide a crane anti-sway positioning control method based on proximal policy optimization. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a crane anti-sway positioning control method based on proximal policy optimization, which performs stably under different load, interference, and task conditions and can be transferred to new scenarios. Compared with traditional methods, its adaptability is significantly enhanced.

[0006] The technical solution of the present invention is: A crane anti-sway positioning control method based on proximal policy optimization, including a traveling crane and a cable mechanism, a rope length detection module, a position and speed detection module, a swing angle detection module, a deep reinforcement learning controller, and a control panel;

[0007] The rope length detection module is used to measure the rope length ; The position and speed detection module is used to detect the trolley position , the gantry position , the trolley speed , the gantry speed ; The swing angle detection module is used to output the swing angle of the suspended load and the swing angle angular velocity ; The control panel outputs the target position of the suspended load and the mass of the suspended load ;

[0008] The deep reinforcement learning controller takes the data collected by the above modules as state information inputs. After being processed and analyzed by the deep reinforcement learning controller, the optimal control signal is calculated and the control signal is transmitted to the trolley, gantry, and cable mechanisms.

[0009] Further, the training process of the deep reinforcement learning controller includes:

[0010] Step 1: Initialize the working environment;

[0011] Step 2: Build a deep reinforcement learning controller based on proximal policy optimization, including an action network and a value network;

[0012] Step 3: The crane controller continuously interacts with the working environment to collect an experience dataset;

[0013] Step 4: Train the action network and update the parameters of the action network;

[0014] Step 5: Train the value network and update the parameters of the value network;

[0015] Step 6: Continuously repeat Steps 3 - 5 until the parameter changes of the action network and the value network are less than a certain threshold, or the cumulative reward no longer increases significantly, then stop training and save the parameters of the latest action network and value network.

[0016] Further, in the above Step 1, when initializing the working environment, the target position of the trolley, the target position of the gantry, the position of the trolley, the speed of the trolley, the position of the gantry, the speed of the gantry, the swing angle, the angular velocity of the swing, the length of the cable, and the weight m of the suspended object in the crane working environment at the initial moment are used as the state information of the working environment, denoted as s =[ x r , y r , x , y , x ̇ , y ̇ , θ , θ , ̇ l , m ] 。

[0017] Further, in the above Step 2, the crane controller interacts with the working environment according to the current policy to generate a series of experience data ;

[0018] Among them, is the status information of the working environment at a certain moment, is the new status information generated by the working environment when the crane controller makes an action in the current state;

[0019] a t =[ ω t 1 , ω t 2 , ω t 3 ] , among them, is the trolley motor speed at a certain moment, is the gantry motor speed at a certain moment, is the cable motor speed at a certain moment; Reward is the reward obtained by calculating through the reward function at a certain moment;

[0020] The crane controller continuously interacts with the working environment until a sufficient number of experience data are collected to form an experience dataset .

[0021] Furthermore, the said reward includes a position reward and a swing angle reward , and the total reward is the sum of the position reward and the swing angle reward ;

[0022] For the said position reward, when the crane reaches the end position, a corresponding reward is given according to the error between the end position and the target position. The formula is as follows:

[0023]

[0024] Where is the target position of the lifted object, , is the trolley position at a certain moment, is the gantry position at a certain moment, represents the calculation of the Euclidean distance, is the required position accuracy, , ;

[0025] The swing angle reward is given according to the size of the swing angle, and the formula is as follows:

[0026]

[0027] In the formula, is the swing angle of the suspended object at time the required swing angle accuracy, , .

[0028] Furthermore, step 4 specifically includes the following steps:

[0029] Step 4.1: Randomly sample N pieces of empirical data from the empirical data set ;

[0030] Step 4.2: Evaluate the value of the current state through the value network ;

[0031] Step 4.3: Calculate the estimated value of the advantage function :

[0032]

[0033] where is the discount factor;

[0034] Step 4.4: Calculate the loss of the action network

[0035]

[0036] where is the probability ratio of the actions taken by the new and old action networks at time , are the parameters of the action network before update, is the clipping function, is the clipping parameter used to limit the difference between the new action and the old action;

[0037] This loss function is used to measure the gap between the current policy and the policy that maximizes the cumulative reward.

[0038] Step 4.5: Based on the Adam optimization algorithm, update the parameters of the action network according to the loss function, so that the action network is optimized in the direction of obtaining a greater cumulative reward.

[0039] Furthermore, step 5 specifically includes the following steps:

[0040] Step 5.1: According to the rewards in the sampled data and the value estimation of the next state , calculate the target value ;

[0041] Step 5.2, calculate the loss of the value network :

[0042]

[0043] Step 5.3, through the Adam optimization algorithm, update the parameters of the value network based on the loss function, so that the value network can more accurately evaluate the state value.

[0044] The beneficial technical effects of the present invention are as follows:

[0045] 1. The crane anti-sway positioning control method based on proximal policy optimization proposed by the present invention combines the perception ability of deep learning and the decision-making ability of reinforcement learning. The method of deep learning is used to accurately perceive the working environment of the crane, the characteristics of the lifted object, and the motion state, etc., and extract useful feature information. Reinforcement learning is used to enable the crane to interact with the external working environment to collect information and autonomously learn the anti-sway positioning control strategy. Through continuous trial and error and learning, the crane can gradually learn the optimal anti-sway positioning control strategy, reduce the sway generated during the operation of the crane, and accurately operate to the set position.

[0046] 2. Traditional crane control often relies on the prior establishment of an accurate dynamic mathematical model, but this method uses reinforcement learning to autonomously learn the anti-sway positioning control strategy through the interaction between the crane and the external environment. This autonomous learning ability enables the crane to no longer rely on the accurate dynamic mathematical model of the crane, avoids the problem of difficult accurate modeling due to the complex crane system and variable working conditions, and reduces the technical implementation threshold. At the same time, it also reduces the control deviation caused by model errors, further improving the reliability and stability of the control. And it can adjust the strategy in real time according to the actual working conditions, adapt to different working scenarios, and greatly improve the flexibility and adaptability of the control strategy.

[0047] 3. Through the deep learning method, it is possible to accurately perceive the working environment of the crane, the characteristics of the lifted object, and the motion state, and effectively extract the key feature information. Compared with the traditional method, this greatly improves the cognitive ability of complex working conditions, thus providing solid input data for subsequent control decisions. Combining the autonomous decision-making ability of reinforcement learning, the crane can continuously optimize the anti-sway positioning strategy, significantly improve the positioning accuracy, and ensure that the lifted object can accurately reach the target position.

[0048] 4. Through continuous trial and error and learning, the crane can gradually master the optimal anti-sway positioning control strategy. In actual operation, it can significantly reduce the sway generated during operation and accurately run to the set position, improving the safety and accuracy of the operation, reducing the risk of damage to the hanging objects caused by sway and the problem of low operating efficiency caused by positioning deviation.

[0049] 5. Accurate anti-sway positioning control can reduce the adjustment time of the crane during the lifting process, so that each lifting task can be completed more quickly, thereby improving the overall operation efficiency and saving a lot of time and labor costs in industrial production and other fields.

[0050] 6. This method enables the crane to have intelligent capabilities, reduces the operator's manual intervention in complex lifting tasks, reduces the dependence on the operator's experience and skill level, reduces the operator's workload, and also reduces the possibility of human operational errors.

[0051] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a schematic diagram of the overall structure of the present invention;

[0053] Figure 2 This is a framework diagram of the deep reinforcement learning algorithm of the present invention;

[0054] Figure 3 It is a structural schematic diagram of the action network of the present invention;

[0055] Figure 4 It is a structural schematic diagram of the value network of the present invention. DETAILED DESCRIPTION

[0056] In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the specific implementation methods of the present invention are further described in detail below in conjunction with the drawings and examples. The following examples are used to illustrate the present invention but are not used to limit the scope of the present invention.

[0057] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so as to describe the embodiments of the present application described herein.

[0058] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship recorded in the embodiments and shown in the drawings, or the orientation or positional relationship in which the invention product is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.

[0059] As Figure 1 shown, the present invention specifically relates to a crane anti-sway positioning control method based on proximal policy optimization, including a traveling crane and a cable mechanism, a rope length detection module, a position and speed detection module, a swing angle detection module, a deep reinforcement learning controller, and a control panel.

[0060] The rope length detection module is used to measure the rope length ; the position and speed detection module is used to detect the trolley position , the gantry position , the trolley speed , the gantry speed .

[0061] The swing angle detection module is used to output the swing angle of the suspended load and the swing angle angular velocity .

[0062] The control panel outputs the target position of the suspended load and the mass of the suspended load .

[0063] The deep reinforcement learning controller takes the data collected by the above modules as state information input, calculates the optimal control signal after processing and analysis by the deep reinforcement learning controller, and transmits the control signal to the trolley, gantry, and cable mechanism;

[0064] Further, as Figure 2 shown, the training process of the deep reinforcement learning controller includes:

[0065] Step 1, initialize the working environment;

[0066] Further, in the step 1, the initialization of the working environment is based on the trolley target position in the crane working environment at the initial moment , the gantry target position , the trolley position , the trolley speed , the gantry position , the gantry speed , the swing angle , the swing angle angular velocity , the length of the cable , the weight m of the suspended object, as the status information of the working environment, denoted as s =[ x r , y r , x , y , x ̇ , y ̇ , θ , θ , ̇ l , m ] .

[0067] Step 2: Construct a deep reinforcement learning controller based on Proximal Policy Optimization, including an action network and a value network;

[0068] Among them, as Figure 3 shown, the action network consists of an input layer of a 10-dimensional vector containing status information, a fully connected layer with 256 neurons, layer normalization, a ReLU activation function, and an output layer.

[0069] Among them, as Figure 4 shown, the value network consists of an input layer of a 10-dimensional vector containing status information, a fully connected layer with 256 neurons, layer normalization, a ReLU activation function, a fully connected layer with 1 neuron, and an output layer.

[0070] Furthermore, in the said Step 2, in the said Step 2, the crane controller and the working environment interact according to the current policy to generate a series of experience data ;

[0071] Among them, is the status information of the working environment at time is the action made by the crane controller in the current state , and is the new status information generated by the working environment;

[0072] a t =[ ω t 1 , ω t 2 , ω t 3 ] , among which, is the rotational speed of the trolley motor at time is the rotational speed of the gantry motor at time is the rotational speed of the cable motor at time; Reward is the reward obtained by calculating through the reward function at time

[0073] The crane controller continuously interacts with the working environment until a sufficient number of experience data are collected to form an experience dataset .

[0074] Furthermore, the reward includes a position reward and a swing angle reward , and the total reward is the sum of the position reward and the swing angle reward ;

[0075] For the position reward, when the crane reaches the end position, a corresponding reward is given according to the error between the end position and the target position. The formula is as follows:

[0076]

[0077] where is the target position of the lifted object, , is the trolley position at time is the gantry position at time represents the calculation of the Euclidean distance, is the required position accuracy, , ;

[0078] For the swing angle reward, a corresponding reward is given according to the size of the swing angle. The formula is as follows:

[0079]

[0080] In the formula, is the swing angle of the lifted object at time is the required swing angle accuracy, , .

[0081] By setting different levels of error ranges, it can help the agent to more easily explore the required accuracy.

[0082] Step 3: The crane controller continuously interacts with the working environment to collect the experience dataset;

[0083] Step 4: Train the action network and update the action network parameters;

[0084] Further, step 4 specifically includes the following steps:

[0085] Step 4.1: Randomly sample N pieces of empirical data from the empirical data set ;

[0086] Step 4.2: Evaluate the value at the current state through the value network ;

[0087] Step 4.3: Calculate the estimated value of the advantage function :

[0088]

[0089] where is the discount factor;

[0090] Step 4.4: Calculate the loss of the action network

[0091]

[0092] where is the probability ratio of the actions taken by the new and old action networks at time , is the parameter of the action network before update, is the clipping function, is the clipping parameter, used to limit the difference between the new action and the old action;

[0093] This loss function is used to measure the gap between the current policy and the policy that maximizes the cumulative reward.

[0094] Step 4.5: Through the Adam optimization algorithm, update the parameters of the action network based on the loss function, so that the action network is optimized in the direction of obtaining a greater cumulative reward.

[0095] Step 5: Train the value network and update the value network parameters.

[0096] Further, step 5 specifically includes the following steps:

[0097] Step 5.1: Calculate the target value based on the reward in the sampled data and the value estimate of the next state;

[0098] Step 5.2: Calculate the loss :

[0099]

[0100] Step 5.3: Update the parameters of the value network based on the loss function through the Adam optimization algorithm, so that the value network can more accurately evaluate the state value.

[0101] Step 6: Continuously repeat Steps 3 - 5 until the parameter changes of the action network and the value network are less than a certain threshold, or the cumulative reward no longer increases significantly. Then stop the training and save the parameters of the latest action network and value network.

[0102] The above embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: Any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention.

Claims

1. A crane anti-sway positioning control method based on proximal strategy optimization, characterized in that: It includes a crane and cable mechanism, a rope length detection module, a position and speed detection module, a swing angle detection module, a deep reinforcement learning controller and a control panel; The rope length detection module is used to measure the rope length The position and speed detection module is used to detect the position of the car , Cart position , car speed , vehicle speed The swing angle detection module is used to output the swing angle of the hanging object. and the angular velocity of the swing angle ; The control panel outputs the target position of the hanging object And the quality of the load ; The deep reinforcement learning controller inputs the data collected by the above modules as state information, calculates the optimal control signal after processing and analysis by the deep reinforcement learning controller, and transmits the control signal to the large and small vehicles and cable mechanisms.

2. The crane anti-sway positioning control method based on proximal strategy optimization according to claim 1 is characterized in that: The deep reinforcement learning controller training process includes: Step 1: Initialize the working environment; Step 2: Build a deep reinforcement learning controller based on proximal policy optimization, including action network and value network; Step 3: The crane controller continuously interacts with the working environment to collect experience data sets; Step 4: Train the action network and update the action network parameters; Step 5: Train the value network and update the value network parameters; Step 6: Repeat steps 3-5 until the parameter changes of the action network and value network are less than a certain threshold, or the cumulative reward is no longer significantly improved, stop training, and save the latest action network and value network parameters.

3. The crane anti-sway positioning control method based on proximal strategy optimization according to claim 2 is characterized in that: In step 1, the initial working environment is based on the target position of the trolley in the crane working environment at the initial moment. , the target position of the vehicle , car position , car speed , Cart position , vehicle speed , swing angle , angular velocity of the swing angle , cable length , the weight of the hanging object m, as the state information of the working environment, is recorded as .

4. The crane anti-sway positioning control method based on proximal strategy optimization according to claim 1 is characterized in that: In step 2, the crane controller and the working environment follow the current strategy Interact and generate a series of empirical data ; in, yes Status information of the working environment at all times, is in the current state The lower crane controller makes a move , new state information generated by the working environment; ,in, for The motor speed of the car at the moment, for The motor speed of the trolley at the moment, for Cable motor speed at the moment; reward yes The reward obtained by calculating the reward function at each moment; The crane controller continues to interact with the working environment until a sufficient amount of experience data is collected to form an experience data set .

5. The crane anti-sway positioning control method based on proximal strategy optimization according to claim 4 is characterized in that: The Reward Includes location bonus and swing angle bonus , and the total reward is the sum of the position reward and the swing angle reward ; The position reward, when the crane reaches the final position, is given a corresponding reward based on the error between the final position and the target position. The formula is as follows: ; in is the target position of the load, , for The position of the car at the moment, for The location of the vehicle at the time, Indicates the Euclidean distance. is the required position accuracy, , ; The swing angle reward is given according to the swing angle size. The formula is as follows: ; In the formula, for The swing angle of the hanging object at the moment, For the required swing angle accuracy, , .

6. The crane anti-sway positioning control method based on proximal strategy optimization according to claim 1 is characterized in that: The step 4 specifically comprises the following steps: Step 4.1: From the empirical data set Randomly sample N empirical data; Step 4.2: Through the value network To assess the current status The value of ; Step 4.3: Calculate the advantage function estimate : ; in, is the discount factor; Step 4.4: Calculate the loss of the action network : ; in, Is the new and old action network in the moment The probability ratio of taking an action, is the action network parameter before updating, is the trimming function, is a clipping parameter used to limit the difference between the new action and the old action; This loss function is used to measure the gap between the current strategy and the strategy that maximizes the cumulative reward; Step 4.5: Use the Adam optimization algorithm to update the parameters of the action network based on the loss function, so that the action network is optimized in the direction of obtaining greater cumulative rewards.

7. The crane anti-sway positioning control method based on proximal strategy optimization according to claim 1 is characterized in that: The step 5 specifically comprises the following steps: Step 5.1: Based on the rewards in the sampled data and the value estimate of the next state , calculate the target value ; Step 5.2: Calculate the loss of the value network : ; Step 5.3: Use the Adam optimization algorithm to update the parameters of the value network based on the loss function, so that the value network can more accurately evaluate the state value.

Citation Information

Patent Citations

  • Inter-ship motion compensation method and system

    CN110162048A

  • Automatic crane control method based on reinforcement learning

    CN113343582A

  • Automatic driving anti-swing control method based on video feedback signal reinforcement learning

    CN114265361A

  • Train cooperative operation control method based on reference deep reinforcement learning

    CN114880770A

  • Anti-swing control method of flexible cable type parallel lifting device

    CN115849185A

Cited By

  • Anti-swing active control method based on deep reinforcement learning

    CN122239483A

  • Ship, umbilical cable and underwater robot system prediction control method

    CN122284269A