Intelligent crane anti-swing control method based on reinforcement learning
Through reinforcement learning methods, combined with multi-sensor data acquisition and reasonable state space and action space design, the shortcomings of anti-slope control in traditional cranes are solved, and high-precision and robust anti-slope control effect is achieved, and the operation efficiency and safety of cranes are improved under complex working conditions.
Patent Information
- Application Number
- CN202510767744.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Traditional crane anti-rock control methods are difficult to achieve high accuracy and robustness under complex working conditions. The existing intelligent control technology has problems such as insufficient definition of state space, unreasonable reward function design, low model training efficiency, and imperfect control action verification mechanism, resulting in uneven anti-rock effects.
Using a method based on reinforcement learning, the state space and action space are reasonably defined through multi-sensor fusion data acquisition, reward functions are designed, reinforcement learning models are trained, and the precise matching and verification of control actions and state space are ensured through convergence criteria and secondary verification mechanism.
It realizes high-precision anti-swing control of the crane under complex working conditions, improves the system's perception of environmental changes, enhances control performance and robustness, reduces the risk of malfunction, and improves operating efficiency and safety.
Smart Images

Figure CN120270909A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of crane control, and particularly to an intelligent anti-sway control method for cranes based on reinforcement learning. Background Art
[0002] As a key device in fields such as industrial production and logistics transportation, the problem of load sway during the operation of a crane has always been an important challenge affecting operation efficiency, safety, and accuracy. Traditional crane anti-sway control methods mainly rely on classical control theories such as PID control. These methods usually require the establishment of an accurate mathematical model. However, the crane system has complex characteristics such as nonlinearity, strong coupling, and time-variation. It is difficult to obtain an accurate model in practical applications, resulting in limited anti-sway control effects. Especially when facing complex working conditions or load changes, the control accuracy and robustness are insufficient.
[0003] With the development of industrial automation and intelligence, the control requirements for cranes are increasing day by day, and the limitations of traditional control methods are gradually emerging. For example, in scenarios such as port loading and unloading and large equipment installation, it is required that the crane can quickly and accurately complete the handling and positioning of the load, while effectively suppressing sway to avoid collision accidents and improve operation efficiency. Traditional methods often have difficulty achieving an ideal anti-sway effect when dealing with complex environments with multiple variables and strong interference, and require operators to have rich experience for manual adjustment, increasing labor costs and operation risks.
[0004] In recent years, the rise of intelligent control technology has provided new ideas for crane anti-sway control. As a machine learning method based on trial and error, reinforcement learning can learn the optimal control strategy through interaction with the environment without an accurate mathematical model, and is suitable for solving the control problems of complex systems. However, applying reinforcement learning to the field of crane intelligent anti-sway control still faces many challenges. For example, how to reasonably define the state space, action space, and reward function to accurately describe the operating state of the crane system and the effect of control actions; how to efficiently collect and process the motion state data and load sway data of the crane to ensure the accuracy and real-time nature of the data; how to design an effective model training and update mechanism to improve the convergence speed and control performance of the reinforcement learning model; and how to achieve the precise matching and verification of control actions and the state space to ensure the reliability and effectiveness of anti-sway control operations, etc.
[0005] In the prior art, although there have been some application studies on intelligent control in the anti-sway of cranes, most of them have problems such as incomplete definition of the state space, unreasonable design of the reward function, low model training efficiency, and imperfect control action verification mechanism, resulting in uneven anti-sway effects in actual applications and being difficult to meet the requirements of the industrial site for intelligent and high-precision control of cranes. Therefore, there is an urgent need for a crane intelligent anti-sway control method based on reinforcement learning, which can make full use of the advantages of reinforcement learning, combine the characteristics of the crane system, solve the deficiencies of traditional control methods, and improve the anti-sway control performance and operation efficiency of the crane under complex working conditions. Summary of the Invention
[0006] The purpose of the present invention is to provide a crane intelligent anti-sway control method based on reinforcement learning to solve the problems raised in the above background technology.
[0007] To achieve the above purpose, the present invention provides the following technical solution: A crane intelligent anti-sway control method based on reinforcement learning, the method includes: Collect the motion state data and load swing data of the crane, and the motion state data includes position and speed information; Based on the motion state data and swing data, combined with a preset reinforcement learning model, judge whether each control action can effectively reduce the load swing; If the judgment is effective, perform an anti-sway control operation based on the control action; Among them, judging whether each control action is effective includes defining a state space, an action space, and a reward function, and training a reinforcement learning model; The state space includes the crane position, the load swing angle and speed; The action space includes the direction and amplitude of the control instruction; The reward function is calculated based on the change in swing amplitude and energy consumption.
[0008] Preferably, the motion state data is real-time data collected by sensors, including data collected by an inertial measurement unit arranged on the crane boom and an angle sensor on the load; the position information is extracted based on global positioning system data, and the extraction method is filtering and smoothing.
[0009] Preferably, the method for defining the state space is as follows: Calculate the dimension of the state space in data collection; set a feature subset for the state space in data collection, and uniformly select k feature points as the associated feature points of the state space in the feature subset, where k is a positive integer; among them, the feature subset of the state space consists of all feature points whose distance from the boundary of the state space is not less than r and not greater than s, r is a preset first boundary threshold, and s is a preset second boundary threshold.
[0010] Preferably, the training of the reinforcement learning model includes using associated feature points as input for model update, and the steps are as follows: S401: Check whether each input feature point meets the convergence criterion. If it does, mark the feature point as a training sample and add it to the current model; The specific convergence criterion is as follows: If the Euclidean distance between the feature point and the model prediction in the error space is less than a preset error threshold, the feature point meets the convergence criterion; otherwise, the feature point does not meet the convergence criterion. S402: Mark the input that has completed the convergence criterion check for all feature points as processed. S403: Repeat S401 to S402. The input marked as processed will no longer participate in the check until all inputs are marked as processed, and then stop model update; record the parameter range of the current model as the updated reinforcement learning model.
[0011] Preferably, the associated feature points are data points corresponding to the state space; the method for determining the associated data points of each feature point is as follows: Extract m data points updated based on j associated feature points in the state space; remove duplicates from the m data points and delete completely overlapping data points; if only one data point remains after deduplication, the remaining data point is the associated data point of the feature point; if n data points remain after deduplication, where n is a positive integer greater than 1, count the number of associated feature points contained in each data point, and the data point containing the most associated feature points is the associated data point of the feature point.
[0012] Preferably, the load swing data is extracted based on real-time data collected by an angle sensor; the method for marking the position of each control action in the action space is as follows: Detect and identify each control action in the action space, and extract the identification information of the control action; based on the identification information, find the state points in the state space corresponding to each control action in the action space; the control actions corresponding in the action space and the state space have the same identification information; obtain the number of each control action and mark the position of each control action in the action space.
[0013] Preferably, after determining the associated data points of each control action in the action space, perform a secondary verification on each control action and its corresponding associated data point, including the following steps: Segment the associated data points of the control action from the state space and mark them as the first associated point set; segment the associated data points of the control action from the action space and mark them as the second associated point set; Based on feature matching, determine whether the first associated point set and the second associated point set belong to the same set of control sequences; if so, perform secondary verification; if not, re-determine the associated data points of the control action; belonging to the same set of control sequences means that the first associated point set and the second associated point set are representations of the same control action in different spaces.
[0014] Preferably, the method for re-determining the associated data points of the control action is as follows: Partition t data points after duplicate removal from the state space, and all are marked as state data points; partition u data points after duplicate removal from the action space, and all are marked as action data points; both t and u are positive integers; Based on feature matching, group the state data points and action data points; any group contains one state data point belonging to the same set of control sequences and the corresponding action data point; delete the state data points and action data points that are not grouped. If there is only one group, the state data points and action data points in the group are both the associated data points of the control action; if there are more than one group, count the total number of associated feature points contained in each group, and the state data points and action data points in the group with the most associated feature points are both the associated data points of the control action; mark the state data points in the associated data points as the first associated point set, and mark the action data points in the associated data points as the second associated point set.
[0015] Preferably, the control action includes direction adjustment and amplitude adjustment; the execution of the anti-sway control operation includes direction control verification and amplitude control verification; The method for the direction control verification is as follows: Compare and match the identification information of the control action with the action templates in the preset action database to identify the control direction; if the identified control direction is consistent with the control direction obtained from defining the action space, the direction control verification is passed. The method for the amplitude control verification is as follows: Extract the amplitude information of all control actions, and calculate the total amplitude of all actions, denoted as the command total amplitude; calculate the actual total amplitude executed through sensor data, denoted as the actual total amplitude; calculate the amplitude difference between the actual total amplitude and the command total amplitude.
[0016] Preferably, the method for generating the identification information of the control action is as follows: Assign a unique action code to each control action; Bind the action code to the position coordinates in the state space; Generate an identification string containing the direction type and amplitude level based on the action code; The identification string is stored in the index field of the action database for matching state points and action positions.
[0017] Compared with the prior art, the beneficial effects of the present invention are: At the data acquisition and processing level, through the inertial measurement unit installed on the crane boom, the angle sensor on the load, and the global positioning system, accurate acquisition of multi-source real-time data such as the position, speed, and load swing angle of the crane is achieved. And the position information is processed through methods such as filtering and smoothing to ensure the accuracy and reliability of the input data, providing a solid data foundation for subsequent control decisions. This multi-sensor fusion data acquisition method can comprehensively and real-time reflect the operating state of the crane system, overcome the limitations of single-sensor data, and improve the system's perception ability of environmental changes.
[0018] In terms of constructing the reinforcement learning model, by reasonably defining the state space, action space, and reward function, the model can accurately describe the dynamic characteristics of the crane system and the effects of control actions. The state space covers key parameters such as the crane position, load swing angle, and speed, comprehensively reflecting the state of the system; the action space clarifies the direction and amplitude of control commands, making the description of control actions more refined; the reward function is calculated based on the change in swing amplitude and energy consumption, taking into account both the core goal of anti-sway control - reducing load swing, and the energy consumption optimization of the system, achieving a balance between control performance and economy. In addition, by calculating the dimension of the state space, setting feature subsets, and selecting associated feature points, etc., the structure of the state space is optimized, improving the training efficiency and generalization ability of the model.
[0019] During the model training process, a feature point processing mechanism based on convergence criteria is adopted. By judging whether the Euclidean distance between the feature point and the model prediction in the error space is less than the preset threshold, it is ensured that only feature points meeting the accuracy requirements are included in the training samples, thereby improving the training quality of the model. By repeating the inspection and marking of the processed inputs, the convergence verification of all feature points is gradually completed, ensuring the comprehensiveness and accuracy of model updates, enabling the model to continuously approach the optimal control strategy, and enhancing control accuracy and robustness.
[0020] In terms of the association and verification between control actions and data points, through methods such as extracting associated feature points, deduplication processing, and grouping and statistics based on feature matching, it is ensured that control actions can be accurately associated with data points in the state space, avoiding data confusion and incorrect matching. The secondary verification mechanism further improves the consistency and reliability between control actions and states by comparing whether the associated point sets in the state space and action space belong to the same control sequence, ensuring the effectiveness of anti-sway control operations. This multi-level data verification and association mechanism effectively reduces the risk of misoperation of the system and improves the accuracy of control decisions.
[0021] In the execution and verification of control actions, through the direction control verification and amplitude control verification mechanisms, the correct execution of control instructions is ensured. The direction control verification ensures the accuracy of the control direction by comparing the identification information of the control action with the preset action template; the amplitude control verification realizes the quantitative evaluation of the execution effect of the control action by calculating the difference between the actual total amplitude and the instruction total amplitude, facilitating the timely discovery and adjustment of control deviations, and improving the stability and controllability of the control process. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is the working principle diagram of the intelligent anti-sway control method for cranes based on reinforcement learning according to the present invention; Figure 2 is the flowchart of the motion state data acquisition; Figure 3 is the flowchart of the method for defining the state space; Figure 4 is the flowchart of marking the position of the control action in the action space. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] Please refer to Figures 1-4 , the intelligent anti-sway control method for cranes based on reinforcement learning involved in the present invention is specifically implemented as follows: Collect the motion state data and load sway data of the crane through sensors. Among them, the motion state data includes position and speed information, and the real-time data is specifically collected by the inertial measurement unit arranged on the crane boom. The position information is extracted based on the global positioning system data and processed by the filtering and smoothing method; the load sway data is collected by the angle sensor arranged on the load in real time.
[0025] Based on the collected motion state data and sway data, combined with the preset reinforcement learning model, it is judged whether each control action can effectively reduce the load sway. Specifically, it is realized by defining the state space, action space and reward function, and training the reinforcement learning model.
[0026] State space: includes the crane position, load sway angle and speed.
[0027] Action space: includes the direction and amplitude of the control instruction.
[0028] Reward function: Based on the change in swing amplitude and energy consumption calculation, it is used to evaluate the effectiveness of control actions.
[0029] If the control action is determined to be effective, anti-sway control operations are performed based on this control action, including direction control verification and amplitude control verification. Direction control verification is carried out by comparing and matching the identification information of the control action with the action templates in the preset action database to identify the control direction and verify the consistency; amplitude control verification calculates the total command amplitude by extracting the amplitude information of the control action, combines the sensor data to calculate the actual total amplitude, and verifies the amplitude difference between the two.
[0030] The present invention will be further described below in conjunction with Embodiments 1 to 5: Embodiment 1: In the data acquisition link, the acquisition of motion state data and load swing data needs to rely on specific sensors to ensure the real-time and accuracy of the data. Among them, the motion state data includes position and speed information, which is specifically collected in real time through an inertial measurement unit set on the crane boom. The inertial measurement unit can sensitively capture parameters such as acceleration and angular velocity of the crane boom during the movement process, and then obtain the speed information through data processing. The position information is extracted based on the Global Positioning System (GPS) data. Since the GPS signal may be affected by factors such as multipath effect and atmospheric interference in the actual environment, resulting in noise in the positioning data, it is necessary to use a filtering and smoothing method to process the original GPS data to remove noise interference and improve the accuracy and stability of the position information. The load swing data is collected in real time through an angle sensor set on the load. The angle sensor can directly measure the swing angle of the load relative to the crane, providing key data support for subsequent anti-sway control.
[0031] Regarding the definition of the state space, a series of rigorous steps need to be implemented. First, calculate the dimension of the state space in data acquisition. The dimension of the state space is determined by the physical quantities it contains. In this solution, the state space includes the crane position, load swing angle, and speed. Therefore, corresponding data for these three physical quantities need to be obtained respectively during data acquisition to determine that the state space is a three-dimensional space. Next, set a feature subset for the state space in data acquisition. The selection of the feature subset needs to meet specific conditions, that is, it consists of all feature points whose distance from the boundary of the state space is not less than the preset first boundary threshold r and not greater than the preset second boundary threshold s. Here, the boundary thresholds r and s are preset according to factors such as the actual working range of the crane, the physical characteristics of the load, and the control accuracy requirements. By setting the feature subset, the data acquisition range can be limited to a reasonable interval, avoiding the adverse impact of data instability near the boundary on subsequent model training and control effects.
[0032] In the feature subset, it is necessary to uniformly select k feature points as the associated feature points of the state space, where k is a positive integer. The purpose of uniformly selecting feature points is to ensure that all regions of the state space can be fully covered and avoid data sampling bias. For example, in a three-dimensional state space, sampling can be performed at certain intervals within the value ranges of the crane position, load swing angle, and speed, so that the k selected feature points can be evenly distributed in the feature subset. These associated feature points will serve as the input data for the subsequent training of the reinforcement learning model, and the rationality of their selection directly affects the model's ability to represent the crane's motion state and load swing situation.
[0033] In practical applications, the working environment of the crane is complex and variable, and factors such as the weight, shape, and lifting height of the load will affect the swing characteristics of the load. Therefore, when setting the boundary thresholds r and s of the feature subset, factors such as the maximum working range of the crane, the maximum swing angle and speed that the load may exhibit need to be comprehensively considered. For example, if the maximum working radius of the crane is R, then when setting the boundary threshold of the crane position, r can be set to 0.1R and s can be set to 0.9R, so that the feature subset is limited to between 10% and 90% of the crane's working radius, which can not only avoid abnormal data that may exist near the crane base and the boundary of the maximum working radius, but also cover most of the normal working areas. For the boundary thresholds of the load swing angle and speed, they can be determined based on statistical analysis of the physical characteristics of the load and historical operation data to ensure that the data within the feature subset can reflect the swing situation of the load under normal working conditions.
[0034] When uniformly selecting k feature points, the limitations of computing resources and model training efficiency also need to be considered. If the value of k is too large, it may lead to a sharp increase in the amount of data, increasing the computational cost and time of model training; if the value of k is too small, it may not be able to fully represent the characteristics of the state space, affecting the accuracy of the model. Therefore, it is necessary to reasonably select the value of k according to the actual situation. For example, in a three-dimensional state space, each dimension can be divided into m equally spaced intervals, and then one feature point is selected from each dimension's interval, thus obtaining k = m×m×m feature points. By adjusting the value of m, a balance can be achieved between the amount of data and model performance.
[0035] In addition, attention should also be paid to the installation position and accuracy calibration of sensors during the data acquisition process. The inertial measurement unit should be installed near the center of gravity of the crane boom to ensure that the motion state of the boom can be accurately measured; the angle sensor should be firmly installed on the load, and the installation direction should be consistent with the swing direction of the load to ensure the accuracy of the measurement results. At the same time, regularly calibrate and maintain the sensors to ensure that their measurement accuracy meets the requirements, and avoid inaccurate state space data caused by sensor errors, which in turn affects the training of the reinforcement learning model and the anti-swing control effect.
[0036] Through the above methods of data acquisition and state space definition, high-quality input data can be provided for the intelligent anti-sway control of cranes based on reinforcement learning, enabling the reinforcement learning model to accurately perceive the motion state of the crane and the sway of the load, thus laying a solid foundation for judging the effectiveness of control actions and performing anti-sway control operations. In the subsequent model training and control process, based on the data of these associated feature points, the model can continuously learn and optimize the control strategy to effectively suppress the sway of the crane load and improve the safety and efficiency of crane operations.
[0037] Example 2: When training the reinforcement learning model, the associated feature points need to be used as inputs for model update. The specific implementation method is as follows: The core process of model update revolves around the convergence criterion. First, it is necessary to clarify the definition and judgment logic of the convergence criterion. The convergence criterion measures the consistency between the feature points and the model prediction through the Euclidean distance in the error space. Specifically, if the Euclidean distance between the feature points and the model prediction in the error space is less than the preset error threshold, the feature points meet the convergence criterion; otherwise, they do not. Among them, the error space refers to the space formed by taking the difference between the actual value and the model prediction value of the feature points as coordinate components, and the Euclidean distance is used to quantify the overall deviation degree between the two. Let the actual value vector of the feature points be , and the model prediction value vector be , then the Euclidean distance calculation formula is:
[0038] In the formula, represents the Euclidean distance, is the dimension of the feature points (i.e., the number of parameters of the crane position, load sway angle, and speed in the state space. Here ), is the actual value of the th feature parameter, and is the model prediction value of the th feature parameter. The preset error threshold is a constant preset according to the accuracy requirements of crane anti-sway control, used to determine whether the feature points are close enough to the model prediction to decide whether to include them in the training samples.
[0039] The specific steps of model update are as follows: Step S401: Perform convergence criterion checks on each input feature point. First, input the actual value vector of the feature points into the current reinforcement learning model, and the model outputs the prediction value vector through its internal neural network structure or algorithm logic. Then, calculate according to the above Euclidean distance formula, and compare with the preset error threshold Compare. If , it indicates that the deviation between the feature point and the model prediction is within the allowable range, meeting the convergence criterion. At this time, mark this feature point as a training sample and add it to the current model to update the model's parameters (such as the weights and biases of the neural network, etc.); if , then this feature point does not meet the convergence criterion and is not included in the training sample for the time being, and subsequent processing is required.
[0040] In actual operation, the input order of feature points may affect the efficiency and stability of model update. Therefore, a batch processing method can be adopted. A certain number of feature points are grouped into a batch, and the convergence criterion is checked and processed for each feature point in the batch in turn. For example, each time feature points are input ( is the batch size, which can be set according to computing resources and model training requirements), calculate the Euclidean distance of each feature point one by one and determine whether it meets the convergence criterion. For feature points that meet the conditions, immediately update the model parameters with them as training samples; for feature points that do not meet the conditions, record their status and perform the next operation uniformly after all feature points in the batch are processed.
[0041] Step S402: When all feature points in the batch have completed the convergence criterion check, mark the input of this batch as "processed". The purpose of marking "processed" is to avoid reprocessing the feature points in the same batch and ensure the orderliness and efficiency of the model update process. The processed input will no longer participate in the convergence criterion check in subsequent model update iterations, thus reducing the waste of computing resources.
[0042] Step S403: Repeat steps S401 to S402 until all inputs are marked as "processed". During the repeated iteration process, each time new unmarked batches of feature points are processed, the model parameters have been updated through the training samples of the previous batches. Therefore, the model's prediction ability for subsequent feature points will gradually improve. As the number of iterations increases, more and more feature points will meet the convergence criterion and be included in the training sample until all input feature points are processed. At this time, stop the model update and record the parameter range of the current model as the updated reinforcement learning model.
[0043] During the model training process, the following points need to be noted: Setting of the error threshold: The size of the error threshold directly affects the convergence speed and accuracy of the model. If is set too small, a large number of feature points will not meet the convergence criterion, and the model will require more iteration times to complete training, and may even fall into overfitting; if If the setting is too large, the fitting accuracy of the model for feature points is insufficient, resulting in inaccurate judgment of the effectiveness of control actions and affecting the anti-sway control effect. Therefore, it is necessary to reasonably set the value according to the actual control requirements and historical data of the crane through experiments or simulation means. For example, for the crane position parameter, can be set to 0.1 meters to ensure that the model can accurately capture the position changes of the crane at the meter-level accuracy; for the load sway angle parameter, can be set to 1 degree to meet the requirements of anti-sway control for angle accuracy.
[0044] Preprocessing of feature points: Before inputting the feature points, it is necessary to standardize them, converting the feature parameters of different dimensions into a unified numerical range to avoid model training deviation caused by the dimensional differences of feature parameters. For example, the unit of crane position is meters, the unit of load sway angle is degrees, and the unit of speed is meters per second. The feature parameters of each dimension can be converted into values within the range of [0,1] through the normalization method. The specific formula is:
[0045] where is the standardized feature value, and are the historical minimum and maximum values of the th feature parameter respectively. The standardization process enables the model to more efficiently learn the correlation relationships between different feature parameters, improving the training efficiency and prediction accuracy.
[0046] Model parameter update strategy: When adding feature points that meet the convergence criterion to the current model, it is necessary to adopt appropriate parameter update strategies, such as Stochastic Gradient Descent (SGD), Adaptive Moment Estimation (Adam), etc. These strategies can automatically adjust the update step size of model parameters according to the gradient information of training samples, ensuring that the model parameters are optimized in the direction of reducing prediction errors. Taking Stochastic Gradient Descent as an example, each time the model parameters are updated, the update amount of the parameters is calculated according to the gradient of the current batch of training samples. The formula is:
[0047] where is the model parameter vector at the th iteration, is the learning rate (a preset constant used to control the update step size), is the loss function at The gradient vector. Through continuous iterative updates, the model parameters gradually converge to the vicinity of the optimal solution, enabling the model to accurately judge the effectiveness of control actions based on the input feature points.
[0048] Prevention of Overfitting: To avoid overfitting in the training process of the model (i.e., the model overfits the training data, resulting in a decline in generalization ability on new data), regularization methods such as L1 regularization and L2 regularization can be used. Add a regularization term to the loss function to penalize the complexity of the model parameters. For example, the expression of the loss function for L2 regularization is:
[0049] In the formula, is the number of training samples, is the loss function of the th sample (such as mean squared error), is the regularization parameter (a preset constant used to control the regularization strength), is the number of model parameters. By introducing the regularization term, it can force the absolute values of the model parameters to be as small as possible, thereby reducing the complexity of the model and improving its generalization ability.
[0050] Through the above implementation methods, the reinforcement learning model can perform iterative training based on the associated feature point data, gradually optimize the model parameters, and enable it to accurately judge whether each control action can effectively reduce the load swing.
[0051] Example 3: The implementation method for determining the associated data points of each feature point and marking the position of the control action in the action space is as follows: When determining the associated data points of the feature points, first, it is necessary to extract the data points related to the associated feature points from the state space. Specifically, extract m data points updated based on j associated feature points in the state space, where j is the number of associated feature points preset to participate in data point update, and m is the total number of generated data points after update. Due to possible duplicate sampling or noise interference during data collection and processing, there may be completely overlapping data points among the m data points. Therefore, it is necessary to perform deduplication on these data points and delete the completely overlapping data points to ensure the uniqueness and effectiveness of the subsequent processed data points.
[0052] After the deduplication operation, it is necessary to judge the determination method of the associated data points according to the number of remaining data points. If only one data point remains after deduplication, it means that this data point is unique and effective, and it can be directly determined as the associated data point of the feature point. For example, when the update process of the j associated feature points is relatively stable and the data collection accuracy is high, only one unique data point may be generated, and at this time, this data point can accurately reflect the associated information of the feature point.
[0053] If there are n data points remaining after deduplication (n is a positive integer greater than 1), it is necessary to determine the associated data points by counting the number of associated feature points contained in each data point. The specific method is as follows: Analyze each data point and count the number of feature points that actually participate in data generation among the j associated feature points it contains. The data point with the most associated feature points is determined as the associated data point of the feature point. This is because the data point with more associated feature points can more comprehensively reflect the comprehensive influence of each feature parameter in the state space, has a stronger correlation with the feature point, and is more suitable as the basis for model training and control action judgment. For example, in a three-dimensional state space, if a data point contains the complete information of three associated feature points, namely the position of the crane, the load swing angle, and the speed, while other data points only contain the information of two or one of these feature points, then this data point will be preferentially selected as the associated data point.
[0054] When marking the position of each control action in the action space, it is first necessary to detect and identify the control actions. Specifically, the control instructions of the crane are monitored in real time through sensors or the control system to identify the type and parameters of each control action, such as the direction of the control instruction (such as lifting, lowering, translation, etc.) and the amplitude (such as the magnitude of the speed, the magnitude of the acceleration, etc.), and the identification information of each control action is extracted. The identification information is the feature data used to uniquely identify each control action, and its generation method needs to ensure that the control actions in the action space and the state space can correspond one by one.
[0055] Based on the extracted identification information, it is necessary to find the corresponding state points in the state space for each control action in the action space. Since the corresponding control actions in the action space and the state space have the same identification information, the mapping between the two can be achieved by matching the identification information. For example, a unique code is assigned to each control action as the identification information, and this code is simultaneously associated with the state parameters such as the position of the crane, the load swing angle, and the speed in the state space when the control action is executed, thereby establishing a one-to-one correspondence between the action space and the state space.
[0056] After completing the mapping between the control actions and the state points, it is necessary to obtain the number of each control action and mark its position in the action space. The number of the control action is a unique identifier assigned to each control action according to a certain logical order, such as sorting and numbering according to the type (direction) and amplitude level of the control action, for the convenience of subsequent management and invocation. When marking the position of the control action in the action space, using the number of the control action as the index, and taking its corresponding direction and amplitude parameters as the coordinate values, to determine the specific position in the coordinate system of the action space. For example, in a two-dimensional action space, taking the control direction as the abscissa and the control amplitude as the ordinate, each control action corresponds to a point in the coordinate system, and the coordinates of this point are the position of the control action in the action space.
[0057] In practical applications, the control actions of a crane may be affected by various factors, such as load weight, crane boom length, operating environment, etc. Therefore, when determining the associated data points and marking the positions of control actions, the dynamic changes of these factors need to be fully considered. For example, when the load weight changes, the same control action may result in different load swing angles and speeds. At this time, it is necessary to re-extract the data points in the state space and update the associated data points and the position marks in the action space according to the new data points to ensure that the model can adapt to the changes in the crane operating conditions in real time.
[0058] In addition, the accuracy of data acquisition and processing is crucial for determining the associated data points and marking the positions of control actions. The measurement errors of sensors may lead to inaccurate data points in the state space, which in turn affects the selection of associated data points and the mapping accuracy between the action space and the state space. Therefore, it is necessary to regularly calibrate and maintain the sensors to ensure the accuracy of their measurement results. At the same time, in the data processing process, algorithms such as filtering and smoothing can be used to preprocess the original data to remove noise interference and improve the reliability of the data.
[0059] Example 4: After determining the associated data points of each control action in the action space, it is necessary to perform a secondary verification on the control action and the corresponding associated data points to ensure that the mapping relationship between the action space and the state space is accurate. The following describes the implementation method in detail with specific examples: Suppose the crane performs a translation control action in a certain operating scenario. The data points collected in the state space related to this action include the crane position, load swing angle, and speed. The corresponding control action in the action space includes the direction (horizontal to the right) and amplitude (speed of 0.5 m / s). At this time, the associated data points of this control action are segmented from the state space and marked as the first associated point set (for example, it includes three data points: A (position 1, angle 5°, speed 0.2 m / s), B (position 2, angle 3°, speed 0.1 m / s), C (position 3, angle 4°, speed 0.15 m / s)); the associated data points of this control action are segmented from the action space and marked as the second associated point set (for example, it includes two data points: X (direction right, amplitude 0.5 m / s), Y (direction right, amplitude 0.4 m / s)).
[0060] Based on feature matching, determine whether two sets of associated point sets belong to the same control sequence, that is, determine whether the data points in the state space and the data points in the action space are representations of the same control action in different spaces. The basis for feature matching includes the logical relationships between the direction type and amplitude level of the control action and the position, angle, and speed in the state space. For example, in a translational control action, if the control amplitude in the action space is 0.5 m / s, the speed parameter in the state space should be close to this amplitude value, and the load swing angle should have a reasonable dynamic relationship with the change in the crane position (such as the angle may increase first and then decrease due to inertia as the position moves farther).
[0061] If the first associated point set and the second associated point set belong to the same control sequence after feature matching (such as the speed parameters of data points B and X are close, and the angle change conforms to the physical law of the translational action), then through secondary verification, confirm that the associated data points of this control action are valid. If the judgment is negative (such as the amplitude of Y in the action space is 0.4 m / s, which is quite different from the speed parameters of all data points in the state space and there is no reasonable physical association), then it is necessary to re-determine the associated data points of the control action.
[0062] When re-determining the associated data points, first split out t de-duplicated data points from the state space (for example, t = 5, including data points D, E, F, G, H), all marked as state data points; split out u de-duplicated data points from the action space (for example, u = 4, including data points M, N, O, P), all marked as action data points. Next, group the state data points and action data points based on feature matching, and each group should contain one state data point belonging to the same control sequence and the corresponding action data point.
[0063] The specific methods of feature matching include: Direction consistency matching: The direction type of the action data point needs to be consistent with the crane movement direction in the state data point. For example, if the direction of the action data point is "horizontal to the right", the change in the crane position in the state data point should be reflected as a rightward movement (such as an increase in the coordinate value).
[0064] Amplitude correlation matching: The amplitude level (such as the speed magnitude) of the action data point needs to have a reasonable numerical association with the speed parameter in the state data point. For example, when the action amplitude is 0.6 m / s, the speed of the state data point should fluctuate within the range of 0.6 m / s ± 0.1 m / s, and the specific fluctuation range can be preset according to the mechanical characteristics and control accuracy of the crane.
[0065] Dynamic Logic Matching: The change in the load swing angle in the status data point should conform to the physical laws of the control actions corresponding to the action data points. For example, in the lifting control action, the load swing angle usually generates a small swing due to inertia at the beginning of lifting and then gradually stabilizes; in the translation control action, the angle change may show a trend of first increasing and then decreasing, and the peak value is related to the translation speed and distance.
[0066] Taking the re - determination of the associated data points of the above translation control action as an example: Status data point D (position 4, angle 6°, speed 0.6 m / s), action data point M (direction right, amplitude 0.6 m / s): The directions are the same, the speed values match, and the angle of 6° may be the inertial swing at the start of translation, which conforms to the dynamic logic and is grouped into one set.
[0067] Status data point E (position 5, angle 2°, speed 0.55 m / s), action data point N (direction right, amplitude 0.55 m / s): The directions are the same, the speed values are close, and the angle of 2° may be the stable state during translation, and is grouped into one set.
[0068] Status data point F (position 3, angle 8°, speed 0.7 m / s), action data point O (direction left, amplitude 0.7 m / s): The directions are not the same, so it is excluded from grouping.
[0069] Status data point G (position 6, angle 3°, speed 0.58 m / s), action data point P (direction right, amplitude 0.58 m / s): The directions are the same, the speed values match, and the angle of 3° conforms to the characteristics of the stable stage, and is grouped into one set.
[0070] Status data point H (position 2, angle 5°, speed 0.4 m / s): There is no matching action data point and it is not grouped.
[0071] After grouping, delete the status data point H and the action data point O that are not grouped. At this time, there are three groups (D - M, E - N, G - P). Count the total number of associated feature points in each group (each status data point contains 3 associated feature points: position, angle, speed; each action data point contains 2 associated feature points: direction, amplitude). The total number of each group is 5 (3 + 2). If there are multiple groups and the total numbers are the same, any one of the groups can be selected as the associated data points, or further screened according to additional conditions such as time - series continuity. For example, in the order of data acquisition time, select the earliest - acquired group D - M as the associated data points, mark the status data point D in it as the first associated point set, and the action data point M as the second associated point set.
[0072] In another scenario, if there is only one group after re-segmentation (for example, the state data point Q and the action data point R are completely matched), then directly determine the state data point and the action data point in this group as the associated data points of the control action, without further statistics.
[0073] The process of secondary verification and re-determination of associated data points needs to ensure that the mapping relationship between the action space and the state space conforms to the actual motion law and control logic of the crane. For example, in the hoisting control action, if the amplitude in the action space is "low-speed hoisting", the speed parameter in the state space should be in the low-speed range (such as 0.1 - 0.3 m / s), and the load swing angle should be small; if there is a situation where the speed parameter is high-speed (such as 0.8 m / s) but matches the "low-speed hoisting" action, it can be identified as abnormal through feature matching and trigger the re-determination process.
[0074] In practical applications, the control actions of the crane may have composite operations (such as hoisting and translation at the same time). At this time, the verification and determination of associated data points need to be carried out for each independent control action respectively. For example, the composite action can be decomposed into a hoisting action and a translation action, and the state data points and action data points of each are extracted and verified and matched step by step according to the above steps.
[0075] Example 5: The generation process of the identification information of the control action and the execution process of the anti-sway control operation are as follows, which will be described in detail with specific examples: I. Generation and mapping of identification information When assigning a unique action code to each control action, the combination of "type - amplitude - timestamp" is used to ensure uniqueness. For example, for a certain hoisting control action with type 01 (defined as the hoisting direction), amplitude level 03 (preset low-speed gear, corresponding to a speed of 0.2 m / s), and timestamp 20250604143001, the action code is "01 - 03 - 20250604143001". Bind this code to the position coordinates in the state space. Assuming that the initial position coordinates of the crane when performing this action are (X = 10 m, Y = 5 m, Z = 3 m), then the identification information includes this coordinate as the spatial positioning reference.
[0076] When generating an identification string based on an action code, the direction type and amplitude level need to be parsed. For example, in the action code "01-03-20250604143001", "01" corresponds to the direction type of "lifting", and "03" corresponds to the amplitude level of the low speed gear. Therefore, the identification string is "QS-D3" (QS represents lifting, and D3 represents the third low speed gear). This string is stored in the index field of the action database for subsequent matching with status points. When the crane executes a translation control action (direction to the right, amplitude 0.5 m / s), the action code is "02-05-20250604143510" (02 is the translation direction, 05 is the medium speed gear), the identification string is "PY-Z5", and the initial position of the bound state space is (X = 12 m, Y = 5 m, Z = 3 m).
[0077] II. Direction Control Verification Example Suppose the preset lifting action template in the action database contains the direction type "lifting" and its corresponding identification string prefix "QS-". When the crane executes a certain control action, extract its identification string "QS-D3", match it with the "QS-" prefix in the action template, and identify the control direction as "lifting". At the same time, the control direction obtained when defining the action space is also "lifting", and the two are consistent, passing the direction control verification.
[0078] If an abnormal situation occurs, such as the identification string of a certain control action is "PY-D3" (the prefix is the translation direction PY), but the crane actually moves in the downward direction during execution. At this time, by comparing the identification string with the action template, it is found that the direction type is "translation", while the actual movement direction is "downward", and the two are inconsistent. The direction control verification fails, and the system will trigger an alarm and stop the execution of this action to avoid exacerbating the load swing caused by misjudgment of the direction.
[0079] III. Amplitude Control Verification Example Taking a set of translation control actions as an example, suppose it includes three consecutive actions: Action 1 (amplitude 0.3 m / s), Action 2 (amplitude 0.4 m / s), Action 3 (amplitude 0.5 m / s). Then the total commanded amplitude is 0.3 + 0.4 + 0.5 = 1.2 m / s. Real-time data is collected through a speed sensor installed on the crane drive motor to calculate the actual executed total amplitude: the actual speed of Action 1 is 0.29 m / s, the actual speed of Action 2 is 0.41 m / s, and the actual speed of Action 3 is 0.50 m / s. The actual total amplitude is 0.29 + 0.41 + 0.50 = 1.20 m / s. Calculate the amplitude difference as 1.20 - 1.2 = 0 m / s, indicating that the actual executed amplitude is consistent with the command, and the amplitude control verification passes.
[0080] In another example, if the total instruction amplitude is 0.8 m / s (composed of two actions with an amplitude of 0.4 m / s each), and the sensor data shows that the actual total amplitude is 0.75 m / s, the amplitude difference is 0.05 m / s. At this time, it is necessary to determine whether this difference is within the preset allowable range (such as ±0.1 m / s). If it is within the range, the verification passes; if it exceeds the range, the compensation mechanism is triggered to adjust the amplitude of subsequent actions to correct the deviation. For example, if the allowable range is ±0.1 m / s, a difference of 0.05 m / s is within the allowable range, and the verification passes; if the difference is 0.15 m / s, it fails, and the system automatically increases the amplitude of the next action by 0.05 m / s to balance the cumulative error.
[0081] IV. Verification of Multi-Action Collaboration Scenarios When the crane executes compound control actions (such as lifting and translation simultaneously), it is necessary to verify the direction and amplitude of each independent action separately. For example, the lifting action (identification string QS-D3, amplitude 0.2 m / s) and the translation action (identification string PY-Z5, amplitude 0.5 m / s) are executed simultaneously: Direction verification: The prefix "QS-" of the lifting action identification string matches the lifting direction in the action template, and the prefix "PY-" of the translation action matches the translation direction, both of which are consistent with the direction obtained from the defined action space, and the verification passes.
[0082] Amplitude verification: The instruction amplitude of the lifting action is 0.2 m / s, and the actual measurement by the sensor is 0.19 m / s; the instruction amplitude of the translation action is 0.5 m / s, and the actual measurement is 0.51 m / s. The lifting amplitude difference is -0.01 m / s, and the translation amplitude difference is +0.01 m / s, both of which are within the allowable range, and the verification passes.
[0083] If the verification of a certain action fails, such as the actual amplitude of the translation action is 0.6 m / s, and the amplitude difference is 0.1 m / s (exactly the upper limit of the allowable range), although the system passes the verification, it will record this abnormality and focus on monitoring the execution parameters of this action in the future. If deviations close to the upper limit of the allowable range occur continuously for multiple times, the system will automatically adjust the amplitude control parameters of this action, such as reducing the preset value of the instruction amplitude, to reduce the actual execution deviation.
[0084] V. Application of Identification Information and Database Index The index fields of the motion database are implemented with identification strings for fast retrieval. For example, when it is necessary to query the corresponding state points of a certain lifting motion in the state space, by inputting the identification string "QS-D3", the database can directly locate the bound position coordinates (X = 10m, Y = 5m, Z = 3m) and the corresponding historical state data (such as the curve of the load swing angle change and the speed fluctuation data when performing this motion). This mapping relationship ensures the real-time linkage between control actions and state data, facilitating the system to quickly call relevant information during the anti-sway control process and optimize the control strategy.
[0085] In actual operations, the execution effect of the same control action of the crane may change due to factors such as load changes and mechanical wear. For example, for the same "PY-Z5" horizontal movement (amplitude 0.5m / s), after the load weight increases, the actual execution speed may decrease due to the increased motor load, resulting in the amplitude difference exceeding the allowable range. At this time, the system updates the associated state points of this action in the motion database with the real-time collected state data, and regenerates the binding relationship between the identification information and the state coordinates, ensuring that the model can adapt to the changes in the dynamic characteristics of the crane and maintain the effectiveness of the anti-sway control.
[0086] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0087] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A crane intelligent anti-sway control method based on reinforcement learning, characterized in that, It includes the following steps: Collect the motion state data and load swing data of the crane, where the motion state data includes position and speed information; Based on the motion state data and swing data, combined with a preset reinforcement learning model, judge whether each control action can effectively reduce the load swing; If it is judged to be effective, perform an anti-swing control operation based on the control action; Among them, judging whether each control action is effective includes defining a state space, an action space, and a reward function, and training a reinforcement learning model; The state space includes the crane position, the load swing angle, and the speed; The action space includes the direction and amplitude of the control instruction; The reward function is calculated based on the swing amplitude change and energy consumption.
2. The intelligent anti-sway control method for a crane based on reinforcement learning according to claim 1, wherein: The motion state data is real-time data collected by sensors, including data collected by an inertial measurement unit arranged on the crane boom and an angle sensor on the load; the position information is extracted based on global positioning system data, and the extraction method is filtering and smoothing.
3. The intelligent anti-sway control method for a crane based on reinforcement learning according to claim 2, characterized in that, The method for defining the state space is as follows: calculate the dimension of the state space in data collection; set a feature subset for the state space in data collection, and uniformly select k feature points as the associated feature points of the state space in the feature subset, where k is a positive integer; among them, the feature subset of the state space consists of all feature points whose distance from the boundary of the state space is not less than r and not greater than s, where r is a preset first boundary threshold and s is a preset second boundary threshold.
4. The intelligent anti-sway control method for a crane based on reinforcement learning according to claim 3, wherein, Training the reinforcement learning model includes using the associated feature points as inputs for model update, and the steps are as follows: S401: Check whether each input feature point meets the convergence criterion. If it meets, mark the feature point as a training sample and add it to the current model; The convergence criterion is specifically as follows: if the Euclidean distance between the feature point and the model prediction in the error space is less than a preset error threshold, the feature point meets the convergence criterion; otherwise, the feature point does not meet the convergence criterion; S402: Mark the input that has completed the convergence criterion check for all feature points as processed; S403: Repeat S401 to S402, and the input marked as processed will no longer participate in the check until all inputs are marked as processed, and stop model update; record the parameter range of the current model as the updated reinforcement learning model.
5. The intelligent anti-sway control method for a crane based on reinforcement learning according to claim 4, wherein: The associated feature points are data points corresponding to the state space; the method for determining the associated data points of each feature point is as follows: Extract m data points updated based on j associated feature points of the state space; remove duplicates from the m data points and delete completely overlapping data points; if only one data point remains after deduplication, the remaining data point is the associated data point of the feature point; if n data points remain after deduplication, where n is a positive integer greater than 1, count the number of associated feature points contained in each data point, and the data point containing the most associated feature points is the associated data point of the feature point.
6. The intelligent anti-sway control method for a crane based on reinforcement learning according to claim 5, wherein: The load swing data is extracted based on the real-time data collected by the angle sensor; the method for marking the position of each control action in the action space is as follows: Detect and identify each control action in the action space, and extract the identification information of the control action; Based on the identification information, find the corresponding state points of each control action in the state space in the action space; The control actions corresponding in the action space and the state space have the same identification information; Obtain the number of each control action and mark the position of each control action in the action space.
7. The intelligent anti-sway control method for a crane based on reinforcement learning according to claim 6, wherein: After determining the associated data points of each control action in the action space, perform a secondary verification on each control action and the corresponding associated data points, including the following steps: Split the associated data points of the control action from the state space and mark them as the first associated point set; Split the associated data points of the control action from the action space and mark them as the second associated point set; Based on feature matching, determine whether the first associated point set and the second associated point set belong to the same set of control sequences; If so, pass the secondary verification; If not, re-determine the associated data points of the control action; Belonging to the same set of control sequences means that the first associated point set and the second associated point set are the representations of the same control action in different spaces.
8. The intelligent anti-sway control method for a crane based on reinforcement learning according to claim 7, characterized in that: The method for re-determining the associated data points of the control action is as follows: Split t de-duplicated data points from the state space, all marked as state data points; Split u de-duplicated data points from the action space, all marked as action data points; Both t and u are positive integers; Based on feature matching, group the state data points and the action data points; Any group contains a state data point belonging to the same set of control sequences and the corresponding action data point; Delete the state data points and action data points that are not included in the group; If there is only one group, the state data points and action data points in the group are both the associated data points of the control action; If there are more than one group, count the total number of associated feature points contained in each group, and the state data points and action data points in the group containing the most associated feature points are both the associated data points of the control action; Mark the state data points in the associated data points as the first associated point set, and mark the action data points in the associated data points as the second associated point set.
9. The intelligent anti-sway control method for a crane based on reinforcement learning according to claim 8, characterized in that: The control action includes direction adjustment and amplitude adjustment; The execution of the anti-sway control operation includes direction control verification and amplitude control verification; The method for direction control verification is as follows: Compare and match the identification information of the control action with the action template in the preset action database to identify the control direction; If the identified control direction is consistent with the control direction obtained by defining the action space, pass the direction control verification; The method for amplitude control verification is as follows: Extract the amplitude information of all control actions, and calculate the total amplitude of all actions, denoted as the command total amplitude; Calculate the actual total amplitude executed through the sensor data, denoted as the actual total amplitude; Calculate the amplitude difference between the actual total amplitude and the command total amplitude.
10. The intelligent anti-sway control method for a crane based on reinforcement learning according to claim 6, characterized in that: The method for generating the identification information of the control action is as follows: Assign a unique action code to each control action; Bind the action code to the position coordinates in the state space; Generate an identification string including the direction type and the amplitude level based on the action code; The identification string is stored in the index field of the action database and is used to match the state point with the action position.
Citation Information
Patent Citations
Anti-swing control method and device for gantry crane based on deep reinforcement learning
CN117466145A
Positioning and anti-swing linear active disturbance rejection control method for double-swing bridge crane
CN118151535A
Intelligent tower crane path planning method based on reinforcement learning and imitation learning
CN120069250A
Intelligent anti-swing and automatic control portal crane
CN211644380U
Abnormality detecting device for stacker crane
JP2010018426A