A crane electronic anti-swing method and system based on markov decision process

CN115594083BActive Publication Date: 2026-09-25QINGDAO HAIXI HEAVY DUTY MASCH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211253572.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-13
Publication Date
2026-09-25
Estimated Expiration
2042-10-13

AI Technical Summary

Technical Problem

[0004]早期的起重机电子闭环控制主要研究方法为经典控制方法,存在一定的局限性,其使用情况并不理想

Benefits of technology

[0019]本发明适用于桥式起重机及相似的起重设备,能够不断训练优化决策模型,不断接近理论最优策略,一直执行已知最优的防摇操作,大大提高了吊具防摇效果和可靠性,特别是经过时间累积,模型训练十分可靠时,执行策略最优,完全实现自动智能防摇。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115594083B_ABST
    Figure CN115594083B_ABST
Patent Text Reader

Abstract

The application provides a crane electronic anti-swing method and system based on a Markov decision process, acquires crane state information and pre-processes based on an artificial intelligence framework; constructs an optimal strategy model based on the Markov decision process, inputs the pre-processed crane state information into the optimal strategy model to output an action strategy, and trains the optimal strategy model; inputs the current acquired crane state information into the trained optimal strategy model to obtain a crane current to-be-executed action strategy. Intelligent electronic anti-swing can be effectively performed under different operating conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of crane anti-sway control, and particularly relates to a crane electronic anti-sway method and system based on Markov decision process. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Currently, my country's ports have a massive and ever-growing container throughput, which inevitably drives terminals to continuously improve their operational efficiency. However, relying on increasing manpower to achieve this efficiency improvement will inevitably lead to significant increases in labor costs. Swaying has a major negative impact on operational efficiency and safety, especially for port cranes. Near the coast, high wind speeds and turbulent currents make it very easy for spreader swaying to occur, severely affecting lifting speed and even causing safety hazards and economic losses. As the biggest factor affecting terminal operating efficiency, the swaying problem of crane spreaders is a core element of automation and intelligent upgrades. Upgrades and modifications at the automation control level depend on the efficiency and reliability of the transfer process.

[0004] Early research on electronic closed-loop control for cranes primarily relied on classical control methods, which had certain limitations and their application was not ideal. These methods were mainly based on traditional pendulum algorithms and applied to yard cranes. However, their installation height was low, and the requirements for anti-sway performance and control accuracy were not high, thus failing to meet the needs of large-scale quay crane anti-sway systems. Summary of the Invention

[0005] To overcome the shortcomings of the prior art, the present invention provides a crane electronic anti-sway method and system based on Markov decision process, which takes into account the influence of wind field and can effectively perform intelligent electronic anti-sway under different operating conditions.

[0006] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solution: an electronic anti-sway method for cranes based on Markov decision processes, comprising the following steps:

[0007] Acquire crane status information and preprocess it based on an artificial intelligence framework;

[0008] An optimal policy model is constructed based on a Markov decision process. The preprocessed crane state information is input into the optimal policy model to output the action policy, and the optimal policy model is trained.

[0009] The current state information of the crane is input into the trained optimal policy model to obtain the current action policy to be executed by the crane.

[0010] A second aspect of the present invention provides an electronic anti-sway system for cranes based on Markov decision processes, comprising:

[0011] The information acquisition module is used to obtain crane status information;

[0012] An artificial intelligence framework is used to preprocess the acquired crane status information to obtain the pose information of the lifting device, the weight information of the lifted object, the wind speed and direction information, the trolley speed information, and the trolley acceleration information.

[0013] An optimal strategy model is constructed based on a Markov decision process. The crane state information is input into the optimal strategy model to output the action strategy. The optimal strategy model is then pre-trained.

[0014] The intelligent anti-sway module is used to input the currently acquired crane state information into the pre-trained optimal strategy model to obtain the crane's current action strategy to be executed.

[0015] The motor control module is used to control the speed change of the motor through the frequency converter according to the action strategy output by the intelligent anti-sway module, so as to realize the electronic anti-sway of the crane.

[0016] A third aspect of the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the steps described in the above method.

[0017] A fourth aspect of the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the steps described in the above method.

[0018] The above one or more technical solutions have the following beneficial effects:

[0019] This invention is applicable to bridge cranes and similar lifting equipment. It can continuously train and optimize the decision-making model, constantly approach the theoretical optimal strategy, and always execute the known optimal anti-sway operation, which greatly improves the anti-sway effect and reliability of the lifting equipment. In particular, after the accumulation of time, when the model training is very reliable, the execution strategy is optimal, and fully automatic intelligent anti-sway is realized.

[0020] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0021] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0022] Figure 1 This is a schematic diagram of the electronic anti-shake method in Embodiment 1 of the present invention;

[0023] Figure 2 This is a schematic diagram of the electronic anti-shake system in Embodiment 2 of the present invention;

[0024] Figure 3-1 This is a schematic diagram of wind speed level classification in Embodiment 1 of the present invention;

[0025] Figure 3-2 This is a schematic diagram of wind direction division in Embodiment 1 of the present invention;

[0026] Figure 3-3 This is a schematic diagram of the offset angle and direction division in Embodiment 1 of the present invention;

[0027] Figure 3-4 This is a schematic diagram of height division in Embodiment 1 of the present invention;

[0028] Figure 3-5 This is a schematic diagram illustrating the division of the weight range of the suspended object in Embodiment 1 of the present invention;

[0029] Figure 3-6 This is a schematic diagram of the state space S in Embodiment 1 of the present invention;

[0030] Figure 4 This is a schematic diagram of the cumulative reward process in Embodiment 1 of the present invention. Detailed Implementation

[0031] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0032] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0033] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0034] Example 1

[0035] like Figures 1-2 As shown, this embodiment discloses an electronic anti-sway method for cranes based on Markov decision processes, including the following steps:

[0036] The system acquires crane status information, processes it using an artificial intelligence framework to obtain offset and spreader height information, and then performs further processing to obtain the crane state s=(sv,r′,h,t,v,a)∈S={s1、s2、s3.......s n};

[0037] Given the crane's state set S, anti-sway operation set A, immediate reward R, decay factor γ, and model state transition probability P, for each s t Each corresponds to a behavior space A t s is determined based on the Markov decision process. t The optimal policy π* at time step is obtained, and a model of the optimal policy is constructed by training on a large dataset.

[0038] The crane's state information s is input into the trained optimal policy model to output the action policy, thus obtaining the crane's current action policy to be executed.

[0039] In this embodiment, external wind field factors, trolley operation status, and lifting device position information are mainly considered as the influencing factors of the lifting device anti-sway status. Wind field factors are obtained using an anemometer, lifting device position information is obtained through image acquisition equipment (most often a camera), trolley speed and acceleration are obtained through a speed encoder, and the weight information of the lifted object is obtained through a weight sensor.

[0040] It should be noted that, depending on the actual situation, the factors affecting the status can be designed freely, and the acquisition methods can also be flexibly adapted based on the existing equipment.

[0041] By leveraging the mainstream open-source AI framework Facebook Torch, one can select or improve existing neural network models to process state information. Facebook TorchTorch is a scientific computing framework that widely supports machine learning algorithms. Torch has a large ecosystem of libraries in the machine learning field, including computer vision packages, signal processing, parallel processing, image, video, audio, and network libraries. Additionally, it has open-sourced deep learning libraries based on Torch, allowing for flexible and easy creation of custom scientific algorithms to process images acquired by cameras and build Markov decision process models. Besides Facebook TorchTorch, other mainstream open-source AI frameworks such as Google TensorFlow and IBM SystemML can also be chosen based on specific needs.

[0042] In this embodiment, the computer vision software package included in the open-source artificial intelligence framework is used to process the pose information of the lifting device acquired by the image acquisition device, and the pose information is converted into information such as swing deviation, offset direction, and lifting device height. Then, the offset range information and height range information are obtained according to the swing deviation, offset direction, and height of the lifting device.

[0043] The entire operation of the crane exhibits Markov properties; the state of the lifting device and the lifted object at the next moment depends only on their current state and actions, and is independent of their previous states and actions. t+1 |s t ]=P[s t+1 |s1,…,s t The processed state set serves as a state element in the state space defined by the MDP, representing the current state information of the spreader. The optimal strategy is found based on the Markov decision process model, and the behavior is determined by the optimal strategy, i.e., the corresponding anti-sway command is generated.

[0044] A Markov decision process consists of a quintuple: (S, A, {p sa Let S be the state space, which is a finite set representing the states of the crane.

[0045] A represents the action space set (actions). Each element in our action space is also a state space, indicating that each state of the crane corresponds to an action space in the action space set. The elements in each action space are the collected anti-sway command training set for that state; P sa It is the state transition probability of the crane. The transition from one state to another in S requires the participation of A.

[0046] P sa This represents the probability distribution of other states that will be transitioned to after the action of a∈A in the current state s∈S (the current state may jump to many states after the action of a).

[0047] γ∈[0,1) is the discount factor. When γ=0, it is equivalent to only considering immediate returns and not long-term returns. When γ=1, long-term returns and immediate returns are considered equally important.

[0048] R: S×A→R, where R is the reward function. The reward function is often written as a function of S (which only depends on S). In this case, R is rewritten as R: S→R.

[0049] In this embodiment, wind speed information, lifting device height, weight of the lifted object, and lifting device offset information are segmented according to levels. The maximum tolerable offset angle is determined based on production safety. Within this offset angle range, similar to the wind speed range, the offset angle range is further divided. The offset direction is then combined to obtain the offset information status of the lifting device and the lifted object. The same method is used to divide the height and weight ranges. For example, wind speed can be divided into levels, such as 0-1 m / s for wind speed state sv1, 1-2 m / s for wind speed state sv2, and so on. The size of the segmented range can be determined according to actual needs.

[0050] The wind speed information, lifting device height, weight of the lifted object, and lifting device offset information are represented by their level representations. The speed and acceleration of the trolley are retained to one decimal place, forming a state set. The entire state set is used as a state element s to describe the state of the lifting device.

[0051] In this embodiment, the anemometer measures an outside wind speed of 3.1 m / s, with an eastward wind direction; the height of the lifting device is 12 m, and the offset angle is 5 degrees, which is the opposite direction of the trolley's movement. The offset angle of the lifting device is divided into two types: the direction of movement and its opposite direction, represented by + and - respectively. Therefore, the offset information here is -5; the height is 12 m; and the trolley speed read from the speed encoder is v. t =0.2m / s², acceleration is a t =0.1m / s2; the weight of the suspended object obtained from the weight sensor is 2.4t.

[0052] according to Figure 3-1 and Figure 3-2 Find the state representation of the wind field factors: sv = sv3 + East; according to Figure 3-3 The spreader offset state was found to be r1′; according to Figure 3-4 The corresponding spreader height status is found to be: h 12 ,according to Figure 3-5 The corresponding weight state is found to be t4; the state s is obtained as [wind: sv3+east, offset: r1′, height: h]. 12 Weight: t4, Car speed: v t The acceleration of the small car: a t ].

[0053] exist Figure 3-6 It is certain that a match s = s can be found in the middle. t ∈S:{s1, s2, s3.......s n Let S be the state space. Due to the range division of the initial parameters, each parameter also has corresponding upper and lower limits according to the actual situation. The state space is a finite set. Each state in the state set has a set of actions that can be taken, namely anti-shake operations, which constitute the behavior space. The behavior space is a function of the state set. Each different s corresponds to a different behavior space. The behavior space of state s∈S is denoted as A(s). All behavior spaces are a finite set of anti-shake operations A.

[0054] Taking an anti-shake operation a∈A(s) in state st∈S will result in a state transition. The transition to s(t+1)∈S has a certain probability (this probability is known through statistical training on a large amount of data), and the probability distribution satisfies:

[0055] 0≤p(s∣st ,a t )≤1 s,s t ∈S,a t ∈A(s t (1-1)

[0056] ∑p(s∣s t ,a t ) = 1 s t ∈S,a t ∈A(s t )s∈S (1-2)

[0057] In a certain state s t Under state ∈S, take an anti-shake operation a∈A(s) in the corresponding action space and transition to state s. (t+1) At that time, a corresponding return R(s) is generated based on the set return function. t ,a (t+1) ,s (t+1) The return is certain, depending on the current state of the spreader, the anti-sway operation performed, and the state it enters after the anti-sway operation.

[0058] Regarding the reward value, a single-step reward function can be customized according to different needs. In this embodiment, after the action is performed, the offset angle is adjusted and improved. The smaller the offset angle, the longer the time spent within the safe range, and the larger the reward value. It can be seen that the reward value is inversely proportional to the offset angle of the next state and directly proportional to the time of maintaining a better posture. Therefore, we set the calculation formula of the reward function as follows:

[0059] R(s t ,a (t+1) ,s (t+1) )=k1|r′|+k2d,(k1<0,k2>0) (2)

[0060] Where k1 is the coefficient of the offset angle |r′|, where the offset angle is a vector representing direction and magnitude, so its absolute value represents the magnitude of the offset; k2 is the coefficient of the duration d, where d is the duration for which the spreader remains within the safe offset range; R(s) t ,a( t+1 ),s( t+1 The return value is calculated based on the return function. Substitute |r′| and d into the function to obtain its value. The magnitude of its value is used to measure the effectiveness of single-step anti-shake. The larger the return value, the better the anti-shake command.

[0061] It should be noted that this embodiment mainly considers using offset angle and maintenance time as measurement factors to set the reward function, but it is not limited to this. According to actual needs, measurement factors can be added, deleted or modified. According to the importance and different meanings of each factor, the magnitude and sign of coefficient k can be dynamically set. As long as the reward function has a linear relationship with the next state and can measure the quality of the strategy, it is acceptable.

[0062] The set of anti-sway operations adopted during the entire lifting process of the quay crane is considered as a strategy, which is represented as: π=(π0,π1,....,π T-1 ), π i It is a mapping from state space to behavior space. When the spreader is in a certain initial state, it performs anti-sway operation according to the strategy, which will generate a... Figure 4 The trajectory shown depicts alternating states and behaviors. Due to the randomness of the spreader state transitions, the trajectory also exhibits a degree of randomness, denoted by τ. The sum of all single-step rewards constitutes the cumulative reward. In the MDP problem, the focus is more on the current period's reward, i.e., on the next spreader state after executing the anti-sway command. Therefore, future rewards are discounted back to the current period using a discount rate γ (0 < γ < 1) to obtain the cumulative reward.

[0063]

[0064] Among them, G t It represents the cumulative reward obtained from step t to the end of the trajectory. The closer γ is to 0, the more one cares about the current reward; the closer it is to 1, the more one values ​​the future reward. k represents the anti-shake measure at step t+1.

[0065] G t The trajectory is sampled from all possible trajectories, thus possessing randomness, and its cumulative reward is not an expected return. Therefore, it cannot be used to evaluate the value of the current spreader state. The value function v(s) eliminates the randomness of the cumulative reward by calculating the expectation, and can accurately measure the value of the spreader at a certain state. Recursively decomposing the value function yields the Bellman expectation equation:

[0066] V(s)=E[G t |S t =s]

[0067]

[0068] P ss′ Let V(s) be the probability of transitioning from state s to the next state s′, and V(s′) be the value function after transitioning to the new state. Starting from the current state, the expected cumulative reward of the lifting device, based on the current anti-sway strategy π, is the state value function of this MDP, and its Bellman expectation equation is:

[0069]

[0070] The action value function of MDP is the value obtained by the spreader starting from state s, taking anti-sway measure a, and subsequently acting according to the current anti-sway strategy π. In other words, the action value function allows for actions taken at the current moment with a deviation from the strategy. Its Bellman equation is:

[0071]

[0072] π(a|s) represents the probability that the anti-sway strategy π assigns an anti-sway command to the spreader in each state s. If the given strategy π is deterministic, then strategy π assigns a deterministic anti-sway command in each state s.

[0073] The solution process for the optimal anti-shake strategy is as follows:

[0074] Optimal anti-shake strategy π π* It is certain that the optimal value function can be obtained: v π* (s)=v * (s), optimal action value function: q π* (s, a) = q * (s, a). From formula (2) and formulas (5-1) and (5-2), we get:

[0075]

[0076]

[0077] The optimal anti-shake strategy can be achieved by adjusting q * (s, a) The anti-sway operation command a, given the maximum value under the known lifting gear conditions, is obtained as follows:

[0078]

[0079] Using Bellman's optimal equation:

[0080] v * (s)=max a q * (s, a) (8-1)

[0081]

[0082] q can be obtained by solving the optimal state value function (8-1). * (s, a) are used to find the optimal strategy π.

[0083] The Bellman optimal equation is a nonlinear equation with no closed-form solution, but it can be solved using an iterative method.

[0084] The following example illustrates an iterative method for solving the value:

[0085] S1: Initialize the state space S, action space A, and state transition probabilities. Immediate reward function Damping coefficient γ, initialization function v(s) = 0, iteration number k = 0;

[0086] S2: Solve using the Gauss-Seidel iterative algorithm based on function (5.1). The iterative formula is:

[0087]

[0088] S3: Calculate formula (9) for each state s;

[0089] S4: Increment the iteration count k by 1;

[0090] S5: Repeat steps S3 and S4 until v t+1 =v t ;

[0091] S6: Obtain the current state value function value and the anti-shake action a.

[0092] It is important to note that during each iteration, the state space needs to be scanned once, and the anti-shake instruction set corresponding to each state space s needs to be scanned in order to obtain a greedy strategy.

[0093] The optimal strategy π is obtained using the above methods. * Then, the state s of the crane is input into the optimal strategy model, and the corresponding anti-sway command action and the anti-sway strategy in the optimal model can be matched to achieve intelligent electronic anti-sway.

[0094] The optimal strategy is given by reorganizing the execution actions based on the known set of strategies. Updating and improving the set of strategies under different states is more conducive to giving a more reasonable optimal strategy.

[0095] The Markov decision process model is the currently known optimal decision model obtained through training with a large amount of data accumulation. As the crane's working time increases and the collected training set increases, the optimal decision model is continuously updated and optimized, and the known optimal strategy in the model is constantly approaching the theoretical optimal decision.

[0096] Example 2

[0097] like Figure 2 As shown, the purpose of this embodiment is to provide an electronic anti-sway system for cranes based on Markov decision processes, including:

[0098] The information acquisition module is used to obtain crane status information;

[0099] An artificial intelligence framework is used to preprocess the acquired crane status information to obtain the pose information of the lifting device, the weight information of the lifted object, the wind speed and direction information, the trolley speed information, and the trolley acceleration information.

[0100] An optimal strategy model is constructed based on a Markov decision process. The crane state information is input into the optimal strategy model to output the action strategy. The optimal strategy model is then pre-trained.

[0101] The intelligent anti-sway module is used to input the currently acquired crane state information into the pre-trained optimal strategy model to obtain the crane's current action strategy to be executed.

[0102] The motor control module is used to control the speed change of the motor through the frequency converter according to the action strategy output by the intelligent anti-sway module, so as to realize the electronic anti-sway of the crane.

[0103] Example 3

[0104] The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the above-described method.

[0105] Example 4

[0106] The purpose of this embodiment is to provide a computer-readable storage medium.

[0107] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the above method.

[0108] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0109] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0110] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A crane electronic anti-sway method based on Markov decision process, characterized in that, Includes the following steps: The crane status information is acquired and preprocessed based on an artificial intelligence framework, specifically as follows: The pose information of the lifting device is converted into swing deviation and lifting device height using an artificial intelligence framework. The corresponding offset range and offset direction are obtained based on the swing deviation. The offset state information of the lifting device is formed based on the offset range and offset direction. The wind speed and direction information, the weight information of the suspended object, the offset information of the lifting device, and the lifting device height are divided and then combined with the trolley acceleration information and trolley speed information to form a state space. An optimal policy model is constructed based on a Markov decision process. Preprocessed crane state information is input into the optimal policy model to output action policies. The optimal policy model is then trained, where: The Markov decision process includes a reward function, which is set based on the offset angle of the spreader's next state and the time the spreader remains within the safe offset range. The reward function value is inversely proportional to the offset angle of the spreader's next state and directly proportional to the time the spreader remains within the safe offset range. The formula for calculating the reward function is as follows: Where k1 is the offset angle |r ′ | is the coefficient, where the offset angle is a vector representing direction and magnitude, and its absolute value represents the magnitude of the offset; k2 is the coefficient of the duration d, where d is the duration for which the spreader stays within the safe offset range; The current state information of the crane is input into the trained optimal policy model to obtain the current action policy to be executed by the crane.

2. The electronic anti-sway method for cranes based on Markov decision processes as described in claim 1, characterized in that, The crane status information includes wind speed information, trolley speed, trolley acceleration information, lifting device position information, and the weight of the object being lifted.

3. The electronic anti-sway method for cranes based on Markov decision processes as described in claim 1, characterized in that, In the state space, each state corresponds to a different action, and the different actions constitute the behavior space. The crane is in a certain state... s When an action is taken in the action space under state (t), and the state transitions to state s(t+1), a corresponding reward is generated based on the reward function. The reward is related to the current state, the action taken, and the next state entered after taking the action. The optimal action is determined based on the generated reward to generate an anti-shake command.

4. The electronic anti-sway method for cranes based on Markov decision processes as described in claim 1, characterized in that, The anti-sway operations taken by the crane throughout the entire lifting process are considered as a strategy. The spreader is in a certain initial state, and anti-sway operations are performed according to the strategy, resulting in a trajectory of alternating states and behaviors. The expected cumulative reward obtained by the spreader from the initial state to the end of the trajectory is used as a value function to measure the spreader in a certain state.

5. The electronic anti-sway method for cranes based on Markov decision processes as described in claim 1, characterized in that, Training the optimal policy model includes: S1: Initialize the state space S, action space A, and state transition probabilities. Immediate return function Damping coefficient γ, initialization function v(s) = 0, iteration number k = 0; S2: Solved using the Gauss-Seidel iterative algorithm, the iterative formula is: S3: Calculate for each state s using the formula in S2; S4: Increment the iteration count by one; S5: Repeat S3-S4 until... Or it may reach the expected convergence value; S6: Obtain the state value function and anti-shake action.

6. A crane electronic anti-sway system based on Markov decision process, characterized in that, include: The information acquisition module is used to obtain crane status information; An artificial intelligence framework is used to preprocess the acquired crane status information to obtain the pose information of the lifting device, the weight information of the lifted object, the wind speed and direction information, the trolley speed information, and the trolley acceleration information, specifically: The pose information of the lifting device is converted into swing deviation and lifting device height using an artificial intelligence framework. The corresponding offset range and offset direction are obtained based on the swing deviation. The offset state information of the lifting device is formed based on the offset range and offset direction. The wind speed and direction information, the weight information of the suspended object, the offset information of the lifting device, and the lifting device height are divided and then combined with the trolley acceleration information and trolley speed information to form a state space. An optimal policy model is constructed based on a Markov decision process. The crane's state information is input into the optimal policy model to output the action policy. The optimal policy model is then pre-trained, wherein: The Markov decision process includes a reward function, which is set based on the offset angle of the spreader's next state and the time the spreader remains within the safe offset range. The reward function value is inversely proportional to the offset angle of the spreader's next state and directly proportional to the time the spreader remains within the safe offset range. The formula for calculating the reward function is as follows: Where k1 is the offset angle |r ′ | is the coefficient, where the offset angle is a vector representing direction and magnitude, and its absolute value represents the magnitude of the offset; k2 is the coefficient of the duration d, where d is the duration for which the spreader stays within the safe offset range; The intelligent anti-sway module is used to input the currently acquired crane state information into the pre-trained optimal strategy model to obtain the crane's current action strategy to be executed. The motor control module is used to control the speed change of the motor through the frequency converter according to the action strategy output by the intelligent anti-sway module, so as to realize the electronic anti-sway of the crane.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the electronic anti-sway method for cranes based on Markov decision processes as described in any one of claims 1-5.

8. A processing apparatus, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the electronic anti-sway method for cranes based on Markov decision processes as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Automatic driving anti-swing control method based on video feedback signal reinforcement learning

    CN114265361A

  • Dynamic flex compensation, coordinated hoist control, and Anti-sway control for load handling machines

    FI20215750A1