Reinforcement learning-based laminated slab hoisting pose intelligent control method and system

By using a reinforcement learning-based approach, the control parameters for hoisting high-rise composite slabs are dynamically optimized, which solves the problem of dynamic coupling risks between wind disturbance and obstacles, improves the stability and safety of hoisting posture control, adapts to complex working conditions, and increases operational efficiency.

CN122380233APending Publication Date: 2026-07-14HANGZHOU CHILDRENS HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU CHILDRENS HOSPITAL
Filing Date
2026-04-28
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing technologies fail to adequately consider the coupling risks of wind disturbance and dynamic changes in obstacles during the hoisting of high-rise composite slabs, resulting in delayed adjustment of control parameters and an inability to achieve dynamic coordination between risk perception and control parameters, thus affecting the stability and safety of hoisting posture control.

Method used

By employing a reinforcement learning-based approach, the raw data of the hoisting time window is acquired, the sling swing sequence and obstacle position sequence are extracted, and time-series regression analysis is performed to dynamically optimize control parameters. The weights of the multi-objective reward function are corrected using a pre-trained reinforcement learning model, thereby achieving real-time perception and smooth adjustment of wind disturbance and obstacle risks.

Benefits of technology

It improves the accuracy and foresight of wind disturbance and obstacle risk calculations, ensures the stability and safety of the hoisting process, avoids sudden changes in control commands, adapts to complex working conditions, and improves the safety and efficiency of hoisting operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122380233A_ABST
    Figure CN122380233A_ABST
Patent Text Reader

Abstract

The application discloses a laminated slab hoisting pose intelligent control method and system based on reinforcement learning, relates to the technical field of laminated slab hoisting pose intelligent control, and acquires original sling obstacle data of a hoisting time window and extracts a sequence, obtains a wind disturbance risk value and an obstacle risk value through time series regression analysis; according to a comparison result of the risk value and a corresponding threshold value, the weight of a multi-target reward function is dynamically corrected or smoothly recovered, a pose control instruction set is solved and output. The application realizes accurate intelligent control of high-rise laminated slab hoisting pose, improves hoisting stability, safety and control precision, and is suitable for complex hoisting working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control technology for the hoisting posture of composite slabs, and more specifically, this application relates to a method and system for intelligent control of the hoisting posture of composite slabs based on reinforcement learning. Background Technology

[0002] The hoisting of high-rise composite slabs is a core component of high-rise prefabricated building construction. The stability and safety of positional control under the complex working conditions of high-rise hoisting directly affect construction efficiency and operational safety. As flexible components, the slings used in the hoisting of high-rise composite slabs are easily affected by natural wind disturbances, and the positions of obstacles in the work area may also change dynamically. This often results in wind-related risks and obstacle-related risks being coupled and superimposed.

[0003] Existing technologies, when addressing such dynamically coupled risks, fail to fully consider the temporal accumulation characteristics and dynamic changes of these risks. Control parameter adjustments often employ fixed patterns or simple linear adjustment logic, failing to dynamically optimize the priority of multi-objective control based on real-time risk levels. This results in existing technologies either lagging in control parameter adjustments, making it unable to promptly suppress sling sway caused by wind disturbances. Essentially, existing technologies fail to achieve dynamic coordination between risk perception and control parameter adjustment, ignoring the vibration inertia of flexible slings and the response characteristics of actuators, and are unable to construct closed-loop control logic adapted to dynamic risks.

[0004] To address the aforementioned issues, the field requires a technical solution that can accurately adapt to the dynamic coupling risks during the hoisting of high-rise composite slabs, achieve dynamic optimization and smooth adjustment of control parameters, and thereby improve the accuracy and stability of hoisting posture control. Summary of the Invention

[0005] To address the aforementioned technical problems, this paper provides a method and system for intelligent control of the hoisting posture of composite slabs based on reinforcement learning. This technical solution solves the problems mentioned in the background section.

[0006] In a first aspect, embodiments of this application provide an intelligent control method for the lifting posture of composite slabs based on reinforcement learning, comprising the following steps: acquiring the original sling obstacle data of the target composite slab within the lifting time window and extracting the sling swing sequence and obstacle position sequence; performing time-series regression analysis on the sling swing sequence and obstacle position sequence respectively to obtain wind disturbance risk value and obstacle risk value; if the wind disturbance risk value is greater than a preset wind disturbance risk threshold and the obstacle risk value is greater than a preset collision risk threshold, then normalizing the first difference between the wind disturbance risk value and the preset wind disturbance risk threshold, and the second difference between the obstacle risk value and the preset collision risk threshold. The first and second risk coefficients are obtained, and the preset hoisting stability weights and preset obstacle avoidance safety weights of the multi-objective reward function of the pre-trained reinforcement learning model are modified accordingly to obtain the first-level multi-objective reward function and solve it, thereby obtaining and outputting the first pose control command set. If the wind disturbance risk value is not greater than the preset wind disturbance risk threshold or the obstacle risk value is less than or equal to the preset collision risk threshold, the current hoisting stability weights and current obstacle avoidance safety weights are smoothly restored to the preset hoisting stability weights and preset obstacle avoidance safety weights by a preset linear step size, thereby obtaining the second-level multi-objective reward function and solving it, thereby obtaining and outputting the second pose control command set.

[0007] Secondly, embodiments of this application provide an intelligent control system for the lifting posture of composite slabs based on reinforcement learning, including: a data acquisition module: used to acquire the original sling obstacle data of the target composite slab within the lifting time window and extract the sling swing sequence and obstacle position sequence; a time-series regression analysis module: used to perform time-series regression analysis on the sling swing sequence and obstacle position sequence respectively to obtain wind disturbance risk value and obstacle risk value; and a first posture control instruction set module: used to, if the wind disturbance risk value is greater than a preset wind disturbance risk threshold and the obstacle risk value is greater than a preset collision risk threshold, then set the first difference value between the wind disturbance risk value and the preset wind disturbance risk threshold, and the second difference value between the obstacle risk value and the preset collision risk threshold... The difference values ​​are normalized to obtain the first risk coefficient and the second risk coefficient. Based on these, the preset hoisting stability weights and preset obstacle avoidance safety weights of the multi-objective reward function of the pre-trained reinforcement learning model are respectively corrected to obtain the first-level multi-objective reward function and solve it, thereby obtaining and outputting the first pose control command set. The second pose control command set module is used to smoothly restore the current hoisting stability weights and current obstacle avoidance safety weights to the preset hoisting stability weights and preset obstacle avoidance safety weights by a preset linear step size if the wind disturbance risk value is not greater than the preset wind disturbance risk threshold or the obstacle risk value is less than or equal to the preset collision risk threshold. This yields the second-level multi-objective reward function, which is then solved, and the second pose control command set is obtained and output.

[0008] Thirdly, this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described reinforcement learning-based intelligent control method for the hoisting posture of composite slabs.

[0009] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0010] 1. Effectively improves the accuracy and foresight of wind disturbance risk value and obstacle risk value calculation. By combining the sling swing sequence and obstacle position sequence of historical hoisting time window for time series fitting to obtain the predicted value, and weighted summation with the initial risk value of the current hoisting time window, it solves the problem of lagging risk perception and inability to predict risk change trend in existing technology, provides accurate basis for subsequent weight adjustment, and ensures hoisting safety.

[0011] 2. Achieve dynamic adaptation and smooth adjustment of lifting posture control weights. Based on the risk coefficient obtained by normalization of the first difference between the wind disturbance risk value and the preset wind disturbance risk threshold, and the second difference between the obstacle risk value and the preset collision risk threshold, the preset lifting stability weights and preset obstacle avoidance safety weights are corrected. After the risk is mitigated, the weights are smoothly restored in linear steps to avoid abrupt changes in control commands. This solves the problem of unscientific weight adjustment in existing technologies, which can easily lead to control instability, and balances lifting safety and control accuracy.

[0012] 3. Comprehensively improve the stability and adaptability of the hoisting process. By adapting to sudden wind disturbances and adjusting the hoisting time window, suppressing resonance and adjusting the preset sampling frequency, and correcting the weight accumulation deviation and adjusting the preset risk threshold, the potential defects of the core control logic are made up for. It effectively deals with problems such as sudden wind disturbances, sling resonance, and long-term hoisting weight deviation accumulation. It further solves the shortcomings of existing technologies that are difficult to adapt to complex high-rise hoisting conditions, and ensures that hoisting operations are carried out efficiently and safely. Attached Figure Description

[0013] Figure 1 A schematic diagram illustrating the steps of the reinforcement learning-based intelligent control method for the hoisting posture of composite slabs provided in this application embodiment;

[0014] Figure 2 A schematic diagram of the logic flow of the intelligent control method for the hoisting posture of composite slabs based on reinforcement learning provided in the embodiments of this application;

[0015] Figure 3 This is a schematic diagram of the structure of the reinforcement learning-based intelligent control system for the hoisting posture of composite slabs provided in an embodiment of this application. Detailed Implementation

[0016] This application's embodiments address the technical problem in the prior art where neglecting the vibration inertia of flexible slings and the response characteristics of actuators leads to insufficient robustness in constructing closed-loop control logic adapted to dynamic risks, through a reinforcement learning-based intelligent control method and system for the hoisting posture of composite slabs.

[0017] In the hoisting of high-rise composite slabs, the slings, as flexible components, are susceptible to wind disturbance and swaying, while the positions of obstacles may dynamically change. These two risks are coupled and superimposed. Existing technologies cannot achieve dynamic coordination between risk perception and control parameters. The control parameter adjustment mode is fixed, ignoring the vibration inertia of the slings and the response characteristics of the actuators, and cannot construct a closed-loop control adapted to dynamic risks. Therefore, this solution starts with risk perception and gradually derives data processing and control logic to achieve precise adaptation and stable control. First, to solve the problem of lagging risk perception, it is necessary to obtain the original sling and obstacle data within the hoisting time window. From this, sequences that can intuitively reflect the wind disturbance and obstacle status are extracted. By spatially registering and temporally synchronizing the original data, the sling sway sequence and obstacle position sequence are obtained, providing basic data support for subsequent risk analysis. Next, in order to fully consider the time-series accumulation characteristics and dynamic change patterns of risks, it is necessary to perform time-series fitting by combining the current hoisting time window with the historical hoisting time window sequence to obtain the sling swing attenuation trend and obstacle movement approach trend, calculate the risk prediction value within the preset advance time, and at the same time perform time-series regression on the current hoisting time window sequence to obtain the initial risk value. The two are weighted and summed to obtain the accurate wind disturbance risk value and obstacle risk value, thus solving the problem of incomplete and lagging risk perception in existing technologies.

[0018] Based on this, the priority of control parameters needs to be dynamically optimized according to the risk level to avoid the defects of a fixed adjustment mode. When both risks exceed the corresponding thresholds, the difference between the risk and the threshold is calculated and normalized to obtain a risk coefficient reflecting the real-time risk proportion. Combined with the exponential moving average of historical risks and the influencing factors of the hoisting stage, the risk coefficient is ensured to accurately match the actual risk state. Based on this, the relevant weights of the multi-objective reward function are adjusted, taking into account the vibration inertia of the sling and the response lag characteristics of the actuator, to achieve dynamic adaptation of control parameters. When the risk does not exceed the standard, the preset weights are smoothly restored by linear steps to avoid additional excitation of the sling caused by sudden changes in control commands and to ensure hoisting stability. To further improve closed-loop control and address potential problems in sudden working conditions and long-term hoisting, this scheme adds supplementary processing logic: in the event of sudden wind disturbance, the time window length is adjusted by detecting the instantaneous change coefficient of the sling swing to ensure the real-time nature of risk perception; when sling resonance risk occurs, the sampling frequency is adjusted to avoid data distortion; after long-term hoisting, the cumulative deviation of the weights is calculated and the preset risk threshold is adjusted to maintain control accuracy.

[0019] The core innovation of this solution lies in constructing a complete closed-loop logic of "data acquisition - sequence extraction - risk prediction and calculation - dynamic weight adjustment - supplementary adaptation". It realizes the dynamic coordination of risk perception and control parameters, fully considers the vibration inertia of the slings and the response characteristics of the actuators, and accurately adapts to the coupled superposition of wind disturbance and obstacles in the hoisting of high-rise composite slabs. It solves the unique and detailed technical problems of unscientific adjustment of control parameters, lagging risk perception, and inability to construct dynamic closed-loop control in existing technologies, which lead to insufficient hoisting stability and accuracy, thus ensuring safe and efficient hoisting operations.

[0020] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0021] Figure 1 This is a schematic diagram illustrating the steps of the reinforcement learning-based intelligent control method for the lifting posture of composite slabs provided in this application embodiment. The reinforcement learning-based intelligent control method for the lifting posture of composite slabs includes the following steps: acquiring the original sling obstacle data of the target composite slab within the lifting time window and extracting the sling swing sequence and obstacle position sequence; performing time-series regression analysis on the sling swing sequence and obstacle position sequence respectively to obtain wind disturbance risk value and obstacle risk value; if the wind disturbance risk value is greater than a preset wind disturbance risk threshold and the obstacle risk value is greater than a preset collision risk threshold, then the first difference value between the wind disturbance risk value and the preset wind disturbance risk threshold, and the first difference value between the obstacle risk value and the preset collision risk threshold are calculated. The second difference value is normalized to obtain the first risk coefficient and the second risk coefficient. Based on these, the preset hoisting stability weight and the preset obstacle avoidance safety weight of the multi-objective reward function of the pre-trained reinforcement learning model are respectively corrected to obtain the first-level multi-objective reward function and solve it. The first pose control command set is obtained and output. If the wind disturbance risk value is not greater than the preset wind disturbance risk threshold or the obstacle risk value is less than or equal to the preset collision risk threshold, the current hoisting stability weight and the current obstacle avoidance safety weight are smoothly restored to the preset hoisting stability weight and the preset obstacle avoidance safety weight by a preset linear step size. The second-level multi-objective reward function is obtained and solved. The second pose control command set is obtained and output.

[0022] Figure 2 A schematic diagram of the logic flow of the intelligent control method for the hoisting posture of composite slabs based on reinforcement learning provided in the embodiments of this application;

[0023] The preset wind disturbance risk threshold is based on the safe allowable range of sling swing in the "Safety Technical Specification for Lifting and Hoisting Engineering in Building Construction" JGJ276-2012. It is obtained by calibration through linear mapping method and mapped to a threshold in the range of 0 to 1, with a value range of 0.3 to 0.7 and a typical value of 0.5. It is used to determine whether the wind disturbance risk exceeds the standard.

[0024] The preset collision risk threshold is based on the safety distance requirements for hoisting operations in the "Technical Specification for Safety of High-Altitude Operations in Building Construction" JGJ80-2016. It is obtained by calibration through linear mapping method and mapped to a threshold in the range of 0 to 1, with a value range of 0.3 to 0.7 and a typical value of 0.5. It is used to determine whether the collision risk exceeds the standard.

[0025] The preset linear step size is obtained through joint calibration of sling swing dynamics simulation test and tower crane mechanism step response test. The value ranges from 0.05 to 0.2 for each hoisting time window, with a typical value of 0.1. It is used for the smooth recovery of weights after risk elimination to avoid secondary excitation of sling caused by sudden changes in control commands.

[0026] Time series regression analysis fits the changing trend of a sequence through linear regression and maps it to a risk value in the range of 0 to 1, which is used to quantify the level of wind disturbance and collision risk.

[0027] Normalization processes map the two differences to the interval between 0 and 1 through min-max normalization, ensuring that the sum of the two risk coefficients is always 1, which conforms to the weight constraints of the multi-objective reward function.

[0028] The pre-trained reinforcement learning model used in this scheme is a dual-deep Q-network, which includes an online network and a target network. The two networks have completely identical structures, both being 3-layer fully connected neural networks.

[0029] 1. Input layer: The number of neurons is 12. The input vector is multi-source real-time sensing data after time synchronization, including the six-degree-of-freedom pose of the composite plate, the swing amplitude of the sling, the relative distance to obstacles, wind speed, and the tension of the suspension point.

[0030] 2. Hidden layers: 2 layers, each with 64 neurons, using the ReLU activation function to extract depth features of the hoisting state.

[0031] 3. Output layer: The number of neurons is 9, corresponding to the combination of 3 actions of the tower crane's slewing, luffing, and hoisting mechanisms.

[0032] The following is an example of the pre-training process for a reinforcement learning model:

[0033] 1. Training environment setup: A virtual simulation environment for high-rise composite slab hoisting was built based on the Unity engine to simulate different wind disturbance levels, obstacle movement scenarios, and working conditions during the hoisting stage, and to construct 10,000 sets of training samples.

[0034] 2. Pre-training parameter settings: learning rate is 0.001, discount factor γ is 0.95, experience replay pool capacity is 100,000, batch size is 32, target network update frequency is 100 steps, and total training steps are 500,000.

[0035] 3. Pre-trained reward function: A multi-objective reward function with initial fixed weights is adopted, with preset weights of 0.2 for hoisting stability, 0.2 for obstacle avoidance safety, 0.3 for pose accuracy, and 0.3 for hoisting efficiency. The training model learns the basic hoisting pose control strategy.

[0036] 4. Pre-training convergence criteria: The model achieves a first-time placement success rate of ≥99% and a placement accuracy of ≤±5mm in the simulation environment. The training loss converges to below 0.01. The weights of the converged model are saved as the pre-trained model.

[0037] This solution addresses the scenario of coupled wind disturbance and obstacle risks during the hoisting of high-rise prefabricated composite slabs. It accurately quantifies the risk level through time-series regression analysis, adaptively corrects the core weights of the multi-objective reward function based on the dual-risk exceeding state, and smoothly restores the weights in the non-exceeding state. This achieves dynamic coordination between risk perception and control parameter adjustment, effectively solving the problems of existing fixed-weight mode being unable to adapt to dynamic working conditions, control lag, and easy instability. It significantly improves the safety, positioning accuracy, and work efficiency of hoisting operations, perfectly meeting the needs of complex hoisting conditions in high-rise buildings.

[0038] Furthermore, the specific extraction process of the sling swing sequence and obstacle position sequence is as follows: Obtain the preset unified spatial coordinate system where the target composite plate is located and the preset sampling frequency of the hoisting time window; within the hoisting time window, synchronously collect the original data of the sling tilt angle and the original spatial coordinates of the obstacle during the hoisting process of the target composite plate at the preset sampling frequency, as the original sling obstacle data; map the original sling tilt angle data and the original spatial coordinates of the obstacle to the unified spatial coordinate system, perform temporal synchronization and spatial registration, and obtain the registered sling tilt angle data and the registered obstacle spatial coordinates; based on the registered sling tilt angle data at the preset sampling frequency, extract the sling swing sequence within the hoisting time window; based on the registered obstacle spatial coordinates, extract the obstacle position sequence.

[0039] In this embodiment, the preset unified spatial coordinate system is the geodetic coordinate system corresponding to the BIM model at the construction site. It is calibrated before construction by using a total station on-site calibration method. It is a fixed coordinate system used to unify the spatial reference of multiple sensors.

[0040] The preset sampling frequency of the hoisting time window is obtained by combining the modal test of the sling's natural frequency with the Nyquist sampling theorem for joint calibration. The value ranges from 10Hz to 50Hz, with a typical value of 20Hz, to ensure that the sampling data can completely reproduce the dynamic characteristics of the sling swing and obstacle movement.

[0041] The raw data of the sling tilt angle is obtained by installing a dual-axis tilt sensor at the connection point between the sling and the lifting device to collect the lateral and longitudinal tilt angle data of the sling.

[0042] The raw obstacle spatial coordinate data is obtained by collecting point cloud data of obstacles in the working area using a binocular vision camera or LiDAR. The centroid coordinates of the obstacles are extracted by the target detection algorithm. The input is point cloud / image data, and the output is time-series data of the three-dimensional spatial coordinates of the obstacles.

[0043] The example process of timing synchronization and spatial registration is as follows: using a unified spatial coordinate system as a reference, the external parameter matrices of the tilt sensor and the vision / radar sensor are solved through the hand-eye calibration algorithm, and all raw data are mapped to the same coordinate system; using the hardware trigger clock of the acquisition device as a reference, the sampling data of different sensors are linearly interpolated and aligned to eliminate the sampling time difference and ensure that the spatial data under the same timestamp correspond one-to-one.

[0044] Sequence extraction can slice the registered time series data according to the start and end times of the hoisting time window, arrange them in ascending order of time, and form sequence data with equal time intervals.

[0045] This embodiment addresses the inherent spatiotemporal misalignment problem in multi-sensor data during high-rise hoisting by constructing a unified spatiotemporal reference through time synchronization and spatial registration. This solves the sequence distortion problem caused by data mismatch between different sensors in existing technologies, providing accurate basic data for subsequent risk value calculation and weight correction, and avoiding control deviations caused by inconsistent data references.

[0046] This embodiment achieves synchronous acquisition and spatiotemporal registration of raw data from multiple sensors by constructing a unified spatial coordinate system and sampling frequency. Finally, it extracts a sling swing sequence and obstacle position sequence with consistent spatiotemporal reference, effectively eliminating spatiotemporal misalignment errors of multi-sensor data and ensuring the accuracy of data for subsequent risk analysis and control decisions.

[0047] Furthermore, the specific processing procedure for wind disturbance risk values ​​and obstacle risk values ​​is as follows: Prior to obtaining the current hoisting time window, the continuous... The sling swing sequence and obstacle position sequence for each historical hoisting time window. Set as a preset positive integer; obtain the preset lead time; for continuous... Time series fitting was performed on the sling swing sequence and obstacle position sequence of each historical lifting time window and the current lifting time window to obtain the sling swing attenuation trend curve and the obstacle movement approaching trend curve. Based on the sling swing attenuation trend curve, the predicted wind disturbance risk value within a preset lead time was calculated, and based on the obstacle movement approaching trend curve, the predicted obstacle risk value within a preset lead time was calculated. Time series regression analysis was performed on the sling swing sequence of the current lifting time window to obtain the initial value of wind disturbance risk. Time series regression analysis was performed on the obstacle position sequence of the current lifting time window to obtain the initial value of obstacle risk. The initial value of wind disturbance risk and the predicted value of wind disturbance risk were weighted and summed to obtain the wind disturbance risk value. The initial value of obstacle risk and the predicted value of obstacle risk were weighted and summed to obtain the obstacle risk value.

[0048] In this embodiment, N, which represents N consecutive historical hoisting time windows, is obtained through a combination of time series fitting goodness-of-fit test and wind field cycle statistical analysis. It is a preset positive integer with a value range of 3 to 10, and a typical value of 5, to ensure that a stable trend curve can be fitted, while avoiding prediction deviations caused by outdated historical data.

[0049] The preset lead time is obtained through joint calibration by the tower crane actuator step response test and the sling swing period modal analysis. The value ranges from 0.2s to 1.0s, with a typical value of 0.5s. This ensures that the prediction time can cover the response lag of the tower crane action and achieve early prediction of risks.

[0050] This embodiment breaks through the limitations of existing technologies that calculate risks based solely on static data from a single window. By fitting historical sequences to obtain trend curves and predicted values, it weights and integrates measured and predicted risks, achieving advanced risk perception and solving the core problem of lagging risk calculation in existing technologies. It can adapt to the temporal superposition of wind disturbance and obstacle risks in advance.

[0051] This embodiment obtains the risk prediction value by fitting the historical and current sequences over time, and then sums it with the measured initial risk value to obtain the final risk value. This achieves advanced and accurate quantification of wind disturbance risk and obstacle risk, effectively making up for the lag of single-window static analysis and providing a scientific risk basis for subsequent adaptive weight adjustment.

[0052] Furthermore, the specific process for obtaining the first and second risk coefficients is as follows: The exponential moving average of the wind disturbance risk value and the exponential moving average of the obstacle risk value are calculated by performing exponential moving averages on the wind disturbance risk value and obstacle risk value of the N consecutive historical lifting time windows preceding the current lifting time window, respectively. The specific calculation formulas for the first and second risk coefficients are as follows:

[0053] ;

[0054] ;

[0055] in, As the first risk factor, As the second risk factor, The preset real-time risk weight coefficient, The first difference value, The second difference value, This is the exponential moving average of the wind disturbance risk value. This is the exponential moving average of the obstacle risk value. These are the preset influencing factors for the hoisting stage.

[0056] In this embodiment, the preset real-time risk weight coefficient is obtained through optimization and calibration by combining multiple sets of orthogonal experiments under various operating conditions with the response surface methodology. The value ranges from 0.5 to 0.9, with a typical value of 0.7. This coefficient is used to balance the weight ratio of real-time risk and historical accumulated risk, ensuring priority response to real-time risk.

[0057] The preset impact factors for the hoisting stage were obtained by combining risk level classification tests for different hoisting stages with the analytic hierarchy process. The value range for the transfer stage is 0.8 to 1.0, with a typical value of 0.9; the value range for the placement stage is 1.0 to 1.2, with a typical value of 1.1. The placement stage has a narrow working space and a lower risk tolerance, so the impact factor for the corresponding stage is larger, which is used to increase the weight of the risk coefficient under the corresponding working conditions.

[0058] Exponential moving average (EM) is used to calculate a weighted moving average of historical risk values, where more recent historical data has a higher weight.

[0059] This embodiment breaks through the limitation of existing technologies that calculate weight coefficients based solely on instantaneous risk difference values. It introduces the EMA cumulative term of historical risks and the impact factor of the hoisting stage, realizing the multi-dimensional integration of instantaneous risk, historical cumulative risk, and scenario constraints. This allows the risk coefficient to accurately match the risk level of the actual hoisting scenario, solving the problem of excessive fluctuations in risk coefficients and mismatch with the scenario in existing technologies.

[0060] This embodiment calculates the first and second risk coefficients, which are precisely matched with the actual working conditions, by integrating real-time risk difference values, historical cumulative risk values, and influencing factors during the hoisting stage. This effectively reduces the randomness of instantaneous data in a single window, achieves stable and accurate calculation of risk coefficients, and provides a reliable core basis for subsequent weight correction.

[0061] Furthermore, the specific correction process for the preset hoisting stability weights and preset obstacle avoidance safety weights is as follows: Obtain the preset hoisting stability weights and preset obstacle avoidance safety weights of the multi-objective reward function of the pre-trained reinforcement learning model; obtain the preset marginal effect coefficient and preset hysteresis compensation coefficient; after the tower crane execution equipment receives any pose control command set, obtain the actual response time of the tower crane execution equipment; calculate the continuous response time before the current hoisting time window... The response lag time of the tower crane's operating equipment within a historical hoisting time window. The specific formula for calculating the response lag time of the tower crane's actuators is: (As a preset positive integer) ,in, The time output for any pose control command set within the k-th historical hoisting time window. Let be the actual response time of the tower crane's operating equipment in the k-th historical lifting time window; using the first risk coefficient as the core correction coefficient, combined with the preset marginal effect coefficient and the preset lag compensation coefficient, the preset lifting stability weight is corrected to obtain the corrected preset lifting stability weight. The specific calculation formula is as follows: ,in, This indicates the corrected preset hoisting stability weight. This indicates the preset hoisting stability weight. Represents the natural constant. The marginal utility coefficient is the preset value. The preset lag compensation coefficient, The response lag time of the tower crane's operating equipment is used as the core correction coefficient. Combined with the preset marginal effect coefficient and the tower crane response lag compensation coefficient, the preset obstacle avoidance safety weight is corrected to obtain the corrected preset obstacle avoidance safety weight. The specific calculation formula is as follows: ,in, This indicates the revised preset obstacle avoidance safety weights. This indicates the preset obstacle avoidance safety weights.

[0062] In this embodiment, the preset hoisting stability weight and preset obstacle avoidance safety weight are obtained by combining a multi-objective optimization algorithm with Pareto optimal solution analysis and calibration of pre-trained simulation conditions. The values ​​range from 0.2 to 0.8, with a typical value of 0.5, to ensure the weight balance of the multi-objective reward function under risk-free conditions.

[0063] The preset marginal effect coefficient is obtained through weight correction response characteristic test combined with nonlinear fitting calibration, and the value ranges from 1.0 to 5.0, with a typical value of 2.0. The larger the preset marginal effect coefficient, the faster the marginal effect of weight correction decreases under high risk coefficient, which can avoid the failure of other control objectives due to excessive weight increase.

[0064] The preset lag compensation coefficient is obtained through tower crane actuator lag characteristic test combined with advance compensation control simulation calibration. The value ranges from 0.1 to 0.5, with a typical value of 0.3, and is used to compensate for the control delay caused by the mechanical lag of the tower crane.

[0065] The K value for K consecutive historical hoisting time windows is set based on the statistical period of tower crane response lag, with a range of 2 to 5, a typical value of 3, and K ≤ N.

[0066] The actual response time of the tower crane's actuators can be obtained by reading the encoder feedback data of the slewing, luffing, and hoisting mechanisms from the tower crane's PLC control system.

[0067] This embodiment breaks through the limitations of the existing technology of linear proportional weight correction, introduces a marginal effect coefficient to avoid excessive weight increase, and introduces tower crane response lag time for advance compensation, realizing accurate and stable weight correction. It not only ensures the principle that the higher the risk, the greater the weight, but also avoids multi-objective imbalance and control oscillation, and solves the problems of unscientific weight correction and easy control instability caused by the existing technology.

[0068] This embodiment combines risk coefficients, marginal effect constraints, and tower crane lag compensation to adaptively adjust the core weights of the multi-objective reward function. This ensures the achievement of safety priorities under high-risk conditions, maintains the balance of multi-objective control, and compensates for the control delay caused by tower crane mechanical lag, effectively improving the stability and accuracy of hoisting control.

[0069] Furthermore, after obtaining the first-level multi-objective reward function, the process also includes a sudden wind disturbance adaptation step: obtaining a preset window adjustment coefficient; calculating the instantaneous change rate of the sling swing sequence in the current hoisting time window, and calculating the instantaneous abrupt change coefficient of the sling swing through the relative abrupt change rate; if the instantaneous abrupt change coefficient of the sling swing is not greater than a preset abrupt change threshold, no adjustment is made; if the instantaneous abrupt change coefficient of the sling swing is greater than the preset abrupt change threshold, the window length of the hoisting time window is dynamically adjusted through the exponential decay method, and the specific calculation formula is as follows: ,in, The adjusted hoisting time window length. The length of the hoisting time window. The preset window adjustment factor. The instantaneous change coefficient of the sling swing is given by the following: The preset mutation threshold is used; based on the adjusted hoisting time window, the sling swing sequence is re-extracted, and a new first-level multi-objective reward function is repeatedly calculated accordingly, which is denoted as the third-level multi-objective reward function and solved to obtain and output the third pose control instruction set.

[0070] In this embodiment, the preset window adjustment coefficient is obtained by combining the simulation test of sudden gust of wind with the window length optimization analysis and calibration. The value range is from 0.5 to 2.0, and the typical value is 1.0. This ensures that the larger the instantaneous change coefficient of the sling swing, the more obvious the shortening of the hoisting time window length, thus improving the real-time performance of data sampling.

[0071] The preset mutation threshold is obtained through statistical analysis of the 3σ criterion of the sling swing data under normal working conditions. The value ranges from 1.5 to 3.0, with a typical value of 2.0. When the instantaneous mutation coefficient of the sling swing exceeds this threshold, it is determined to be a sudden gust of wind condition.

[0072] The instantaneous rate of change of the sling swing sequence is calculated by performing a first-order difference calculation on the sling swing sequence to obtain the instantaneous rate of change of the swing at each sampling point, which characterizes the instantaneous increase in the sling swing.

[0073] The relative mutation rate calculation is the ratio of the absolute difference between the instantaneous change rate of the current hoisting time window and the previous window to the average of the instantaneous change rates of N consecutive historical hoisting time windows. This yields the instantaneous mutation coefficient of the sling swing, eliminating the influence of the sling swing reference value.

[0074] The exponential decay method for adjusting the window length involves using an exponential function to achieve a non-linear adjustment of the window length. The larger the mutation coefficient, the shorter the window length, thus avoiding data fluctuations caused by abrupt changes in the window length.

[0075] This embodiment addresses the scenario of sudden changes in sling swing caused by sudden gusts of wind. By detecting sudden changes in the instantaneous change coefficient, it dynamically adjusts the length of the hoisting time window, solving the problems of the original fixed window being unable to adapt to sudden wind disturbances and the lag in risk calculation. It achieves real-time improvement in risk perception and rapid adaptation of control strategies under sudden conditions, and is a supplementary optimization to the core control logic, possessing independent creativity.

[0076] This embodiment detects sudden gusts of wind by using the instantaneous change coefficient of sling swing, nonlinearly adjusts the length of the hoisting time window, and recalculates the multi-objective reward function and control instruction set to adapt to the sudden conditions. This effectively improves the real-time performance of risk perception under sudden wind disturbances, avoids the risk of sling swing going out of control, and ensures hoisting safety under extreme conditions.

[0077] Furthermore, after obtaining the sling swing sequence and obstacle position sequence, a resonance suppression processing step is included: spectral analysis of the sling swing sequence within the current hoisting time window is performed using Fast Fourier Transform to extract the natural frequency of the sling swing; a preset sampling frequency and a preset sampling frequency adjustment coefficient are obtained, and the fit ratio between the sampling frequency and the natural frequency is calculated; if the fit ratio is not within the preset resonance range, no adjustment is made; if the fit ratio is within the preset resonance range, it is determined to be a risk of sling swing resonance, and the preset sampling frequency is dynamically adjusted using a nonlinear proportional adjustment method. The specific calculation formula is as follows: ,in, The adjusted preset sampling frequency, The preset sampling frequency before adjustment. For the matching ratio between the sampling frequency and the natural frequency, The preset sampling frequency adjustment coefficient is used; based on the adjusted preset sampling frequency, the original sling obstacle data of the current hoisting time window is re-acquired and a new sling swing sequence and obstacle position sequence are extracted.

[0078] In this embodiment, the preset sampling frequency adjustment coefficient is obtained by combining the sampling resonance condition simulation test with the frequency adjustment sensitivity analysis and calibration. The value range is 0.1 to 0.3, and the typical value is 0.2. This ensures that the closer the matching ratio between the sampling frequency and the natural frequency is to 2, the greater the sampling frequency adjustment amplitude, and the faster it can deviate from the resonance range.

[0079] The preset resonance interval is obtained by combining the Nyquist sampling criterion with statistical analysis of the probability of sampling resonance occurrence. The value range is [1.8, 2.2]. When the matching ratio between the sampling frequency and the natural frequency is within this interval, sampling resonance is likely to occur, resulting in distortion of the sequence data.

[0080] Fast Fourier Transform (FFT) spectral analysis extracts the frequency point with the largest amplitude in the spectrum, i.e., the natural frequency of the sling swing, by performing a frequency domain transformation on the sling swing sequence. The input is time-domain sequence data, and the output is the natural frequency value.

[0081] The adaptation ratio calculation is the ratio of the preset sampling frequency to the natural frequency of the sling swing, used to determine whether it is in the resonance range.

[0082] This embodiment addresses the sampling resonance scenario caused by the proximity of the natural frequency of the sling swing to the sampling frequency. By detecting resonance risk through spectrum analysis and dynamically adjusting the sampling frequency, it solves the problems of sequence data distortion, risk calculation deviation, and control instability caused by the original fixed sampling frequency. It avoids control failure caused by resonance from the source of data acquisition, and represents a fundamental optimization of the core technical solution, demonstrating independent creativity.

[0083] This embodiment extracts the natural frequency of sling swing through spectrum analysis, detects the risk of sampling resonance, dynamically adjusts the sampling frequency, and re-acquires data to extract sequences. This effectively avoids data distortion caused by sampling resonance, ensures the accuracy of risk calculation, suppresses the exacerbation of sling swing and control instability caused by resonance, and improves the stability of long-distance hoisting processes.

[0084] Furthermore, after outputting any pose control instruction set, it also includes a weighted cumulative deviation correction step: obtaining the continuous values ​​before the current hoisting time window. The corrected preset lifting stability weights and corrected preset obstacle avoidance safety weights for each historical lifting time window are calculated. The cumulative deviations of the lifting stability weights and obstacle avoidance safety weights are calculated using the average absolute deviation method. If the cumulative deviation of any weight exceeds the preset weight deviation threshold, the preset wind disturbance risk threshold and preset collision risk threshold are adaptively adjusted using the exponential decay method. The specific calculation formula is as follows: ; ;in, The adjusted preset wind disturbance risk threshold. The preset wind disturbance risk threshold before adjustment The adjusted preset collision risk threshold. The preset collision risk threshold before adjustment. This is the preset threshold adjustment coefficient. The cumulative deviation of the hoisting stability weight, To mitigate the cumulative deviation of obstacle avoidance safety weights, the adjusted preset wind disturbance risk threshold and the adjusted preset collision risk threshold will be used in the next hoisting time window.

[0085] In this embodiment, the preset weight deviation threshold is obtained by long-term hoisting condition simulation test combined with control accuracy attenuation characteristic analysis and calibration. The value range is 0.05 to 0.2, and the typical value is 0.1. When the cumulative weight deviation exceeds this threshold, it is determined that the cumulative weight deviation exceeds the standard, and the risk threshold needs to be adjusted.

[0086] The preset threshold adjustment coefficient is obtained through simulation experiments of weight deviation closed-loop control combined with threshold adjustment sensitivity analysis and calibration. The value range is from 0.5 to 2.0, with a typical value of 1.0. This ensures that the greater the cumulative deviation of the weight, the more obvious the adjustment of the risk threshold, thus reducing the frequency of over-correction of the weight.

[0087] The mean absolute deviation calculation is to calculate the average of the absolute differences between the corrected weights and the preset weights over K consecutive historical hoisting time windows, thereby obtaining the cumulative weight deviation and quantifying the long-term deviation of the weights.

[0088] The exponential decay method adjusts the risk threshold by using an exponential function to achieve non-linear adjustment of the risk threshold. The larger the cumulative deviation, the smaller the risk threshold, thereby improving the sensitivity of risk detection and reducing the frequency of over-correction of weights.

[0089] This embodiment addresses the issue of accumulated minor deviations in weight correction during long-term continuous hoisting. It quantifies the accumulated deviation by means of average absolute deviation, adaptively adjusts the risk threshold, and forms a long-term closed-loop optimization logic. This solves the problems of existing technologies, such as the lack of a long-term deviation correction mechanism, imbalance of multi-objective reward functions, and continuous decline in control accuracy. It achieves stable control accuracy during long-term continuous hoisting and demonstrates independent innovation.

[0090] This embodiment calculates the cumulative deviation of weights and adaptively adjusts the preset risk threshold for risk assessment in the next hoisting time window. This effectively suppresses the accumulation of weight deviations during long-term continuous hoisting, avoids imbalance of multi-objective reward functions, ensures the control accuracy and stability of long-term hoisting operations, and adapts to the engineering scenario requirements of continuous hoisting in multiple flow sections.

[0091] Figure 3 This is a schematic diagram of the intelligent control system for the lifting posture of composite slabs based on reinforcement learning provided in this application embodiment. The intelligent control system for the lifting posture of composite slabs based on reinforcement learning includes: a data acquisition module: used to acquire the original sling obstacle data of the target composite slab within the lifting time window and extract the sling swing sequence and obstacle position sequence; a time-series regression analysis module: used to perform time-series regression analysis on the sling swing sequence and obstacle position sequence respectively to obtain the wind disturbance risk value and obstacle risk value; and a first posture control instruction set module: used to, if the wind disturbance risk value is greater than a preset wind disturbance risk threshold and the obstacle risk value is greater than a preset collision risk threshold, then set the first difference value between the wind disturbance risk value and the preset wind disturbance risk threshold, and the obstacle risk value and the first difference value between the wind disturbance risk value and the preset wind disturbance risk threshold, and the first difference value between the obstacle risk value and the preset collision risk threshold. The second difference value of the preset collision risk threshold is normalized to obtain the first risk coefficient and the second risk coefficient. Based on these, the preset hoisting stability weight and the preset obstacle avoidance safety weight of the multi-objective reward function of the pre-trained reinforcement learning model are respectively corrected to obtain the first-level multi-objective reward function and solve it, and obtain and output the first pose control command set. The second pose control command set module is used to smoothly restore the current hoisting stability weight and the current obstacle avoidance safety weight to the preset hoisting stability weight and the preset obstacle avoidance safety weight by a preset linear step size if the wind disturbance risk value is not greater than the preset wind disturbance risk threshold or the obstacle risk value is less than or equal to the preset collision risk threshold, thereby obtaining the second-level multi-objective reward function and solving it, and obtaining and outputting the second pose control command set.

[0092] This application also provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements a reinforcement learning-based intelligent control method for the hoisting posture of composite slabs.

[0093] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0094] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0095] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0096] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0097] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0098] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A reinforcement learning-based intelligent control method for the hoisting posture of composite slabs, characterized in that, Includes the following steps: Obtain the original sling obstacle data of the target composite plate within the hoisting time window and extract the sling swing sequence and obstacle position sequence; Time-series regression analysis was performed on the sling swing sequence and obstacle position sequence to obtain wind disturbance risk value and obstacle risk value; If the wind disturbance risk value is greater than the preset wind disturbance risk threshold and the obstacle risk value is greater than the preset collision risk threshold, then the first difference between the wind disturbance risk value and the preset wind disturbance risk threshold and the second difference between the obstacle risk value and the preset collision risk threshold are normalized to obtain the first risk coefficient and the second risk coefficient. Based on these, the preset hoisting stability weight and the preset obstacle avoidance safety weight of the multi-objective reward function of the pre-trained reinforcement learning model are respectively corrected to obtain the first-level multi-objective reward function and solve it to obtain and output the first posture control command set. If the wind disturbance risk value is not greater than the preset wind disturbance risk threshold or the obstacle risk value is less than or equal to the preset collision risk threshold, then the current hoisting stability weight and the current obstacle avoidance safety weight are smoothly restored to the preset hoisting stability weight and the preset obstacle avoidance safety weight by a preset linear step size. The second-level multi-objective reward function is obtained and solved, and the second pose control instruction set is obtained and output.

2. The intelligent control method for the hoisting posture of composite slabs based on reinforcement learning according to claim 1, characterized in that, The specific extraction process for the sling swing sequence and obstacle position sequence is as follows: Obtain the preset unified spatial coordinate system where the target composite plate is located and the preset sampling frequency of the hoisting time window; Within the hoisting time window, the original data of the sling tilt angle and the original data of the obstacle spatial coordinates during the hoisting process of the target composite plate are synchronously collected at the preset sampling frequency, and used as the original sling obstacle data; The original data of the sling tilt angle and the original data of the obstacle spatial coordinates are mapped to the unified spatial coordinate system, and time synchronization and spatial registration are performed to obtain the registered sling tilt angle data and the registered obstacle spatial coordinate original data. Based on the registered sling tilt angle data at a preset sampling frequency, the sling swing sequence within the hoisting time window is extracted. Based on the original spatial coordinate data of registered obstacles at a preset sampling frequency, the obstacle position sequence is extracted.

3. The intelligent control method for the hoisting posture of composite slabs based on reinforcement learning according to claim 1, characterized in that, The specific processing procedure for the wind disturbance risk value and obstacle risk value is as follows: Get the continuous time before the current hoisting time window The sling swing sequence and obstacle position sequence for each historical hoisting time window. It is a preset positive integer; Get the preset lead time; For continuous The sling swing sequence and obstacle position sequence of each historical hoisting time window and the current hoisting time window are time-series fitted to obtain the sling swing attenuation trend curve and the obstacle movement approach trend curve. Based on the sling swing attenuation trend curve, the wind disturbance risk prediction value within a preset lead time is calculated, and based on the obstacle movement approach trend curve, the obstacle risk prediction value within a preset lead time is calculated. A time-series regression analysis was performed on the sling swing sequence within the current hoisting time window to obtain the initial value of wind disturbance risk; A time-series regression analysis was performed on the sequence of obstacle locations within the current hoisting time window to obtain initial values ​​for obstacle risk. The wind disturbance risk value is obtained by weighted summing the initial wind disturbance risk value and the predicted wind disturbance risk value. The obstacle risk value is obtained by weighted summing of the initial obstacle risk value and the predicted obstacle risk value.

4. The intelligent control method for the hoisting posture of composite slabs based on reinforcement learning according to claim 1, characterized in that, The specific process for obtaining the first risk coefficient and the second risk coefficient is as follows: The exponential moving average of the wind disturbance risk value and the exponential moving average of the obstacle risk value are calculated by performing exponential moving averages on the wind disturbance risk value and the obstacle risk value of the N consecutive historical lifting time windows before the current lifting time window, respectively. The specific formulas for calculating the first and second risk coefficients are as follows: ; ; in, As the first risk factor, As the second risk factor, The preset real-time risk weight coefficient, The first difference value, The second difference value, This is the exponential moving average of the wind disturbance risk value. This is the exponential moving average of the obstacle risk value. These are the preset influencing factors for the hoisting stage.

5. The intelligent control method for the hoisting posture of composite slabs based on reinforcement learning according to claim 1, characterized in that, The specific correction process for the preset hoisting stability weight and the preset obstacle avoidance safety weight is as follows: Obtain the preset hoisting stability weights and preset obstacle avoidance safety weights of the multi-objective reward function of the pre-trained reinforcement learning model; Obtain the preset marginal effect coefficient and the preset lag compensation coefficient; After receiving any set of posture control commands, the tower crane actuator obtains the actual response time of the tower crane actuator. Calculate the consecutive time windows before the current hoisting time window The response lag time of the tower crane's operating equipment within a historical hoisting time window. The specific formula for calculating the response lag time of the tower crane's actuators is: (As a preset positive integer) ,in, The time output for any pose control command set within the k-th historical hoisting time window. The actual response time of the tower crane's execution equipment during the k-th historical hoisting time window; Using the first risk coefficient as the core correction coefficient, and combining it with the preset marginal effect coefficient and the preset lag compensation coefficient, the preset hoisting stability weight is corrected to obtain the corrected preset hoisting stability weight. The specific calculation formula is as follows: ,in, This indicates the corrected preset hoisting stability weight. This indicates the preset hoisting stability weight. Represents the natural constant. The marginal utility coefficient is the preset value. The preset lag compensation coefficient, The response lag time of the tower crane's operating equipment; Using the second risk coefficient as the core correction coefficient, and combining it with the preset marginal effect coefficient and the tower crane response lag compensation coefficient, the preset obstacle avoidance safety weight is corrected to obtain the corrected preset obstacle avoidance safety weight. The specific calculation formula is as follows: ,in, This indicates the revised preset obstacle avoidance safety weights. This indicates the preset obstacle avoidance safety weights.

6. The intelligent control method for the hoisting posture of composite slabs based on reinforcement learning according to claim 5, characterized in that, After obtaining the first-level multi-objective reward function, the process also includes a sudden wind disturbance adaptation step: Get the preset window adjustment factor; The instantaneous rate of change of the sling swing sequence in the current hoisting time window is calculated, and the instantaneous abrupt change coefficient of the sling swing is calculated through the relative abrupt change rate; If the instantaneous change coefficient of the sling swing is not greater than the preset change threshold, no adjustment is made; If the instantaneous change coefficient of the sling swing is greater than the preset change threshold, the length of the hoisting time window is dynamically adjusted using the exponential decay method. The specific calculation formula is as follows: ,in, The adjusted hoisting time window length. The length of the hoisting time window. The preset window adjustment factor. The instantaneous change coefficient of the sling swing is given by the following: The preset mutation threshold; Based on the adjusted hoisting time window, the sling swing sequence is re-extracted, and a new first-level multi-objective reward function is repeatedly calculated accordingly. This is denoted as the third-level multi-objective reward function and solved to obtain and output the third pose control instruction set.

7. The intelligent control method for the hoisting posture of composite slabs based on reinforcement learning according to claim 1, characterized in that, After obtaining the sling swing sequence and obstacle position sequence, the process also includes a resonance suppression step: The natural frequency of the sling swing is extracted by performing a spectral analysis on the sling swing sequence within the current hoisting time window using Fast Fourier Transform. Obtain the preset sampling frequency and the preset sampling frequency adjustment coefficient, and calculate the matching ratio between the sampling frequency and the inherent frequency; If the adaptation ratio is not within the preset resonance range, no adjustment will be made; If the fitting ratio is within the preset resonance range, it is determined to be a risk of sling oscillation resonance. The preset sampling frequency is then dynamically adjusted using a nonlinear proportional adjustment method. The specific calculation formula is as follows: ,in, The adjusted preset sampling frequency, The preset sampling frequency before adjustment. For the matching ratio between the sampling frequency and the natural frequency, This is the preset sampling frequency adjustment coefficient; Based on the adjusted preset sampling frequency, the original sling obstacle data of the current hoisting time window is re-acquired and new sling swing sequences and obstacle position sequences are extracted.

8. The intelligent control method for the hoisting posture of composite slabs based on reinforcement learning according to claim 7, characterized in that, After outputting any pose control instruction set, a weight accumulation deviation correction step is also included: Get the continuous time before the current hoisting time window The corrected preset hoisting stability weight and the corrected preset obstacle avoidance safety weight for each historical hoisting time window; The cumulative deviations of the hoisting stability weight and the obstacle avoidance safety weight are calculated separately using the mean absolute deviation calculation method. If the cumulative deviation of any weight exceeds the preset weight deviation threshold, the preset wind disturbance risk threshold and the preset collision risk threshold are adaptively adjusted using the exponential decay method. The specific calculation formula is as follows: ; ; in, The adjusted preset wind disturbance risk threshold. The preset wind disturbance risk threshold before adjustment The adjusted preset collision risk threshold. The preset collision risk threshold before adjustment. This is the preset threshold adjustment coefficient. The cumulative deviation of the hoisting stability weight, Accumulated deviation for obstacle avoidance safety weights; The adjusted preset wind disturbance risk threshold and the adjusted preset collision risk threshold will be used in the next hoisting time window.

9. A reinforcement learning-based intelligent control system for the lifting posture of composite slabs, characterized in that, include: Data acquisition module: used to acquire the original sling obstacle data of the target composite plate during the hoisting time window and extract the sling swing sequence and obstacle position sequence; The time-series regression analysis module is used to perform time-series regression analysis on the sling swing sequence and obstacle position sequence to obtain wind disturbance risk value and obstacle risk value. The first posture control instruction set module is used to normalize the first difference between the wind disturbance risk value and the preset wind disturbance risk threshold, and the second difference between the obstacle risk value and the preset collision risk threshold, if the wind disturbance risk value is greater than the preset wind disturbance risk threshold and the obstacle risk value is greater than the preset collision risk threshold, to obtain the first risk coefficient and the second risk coefficient. Based on these, the preset hoisting stability weight and the preset obstacle avoidance safety weight of the multi-objective reward function of the pre-trained reinforcement learning model are respectively corrected to obtain the first-level multi-objective reward function and solve it, and obtain and output the first posture control instruction set. The second pose control instruction set module is used to smoothly restore the current hoisting stability weight and the current obstacle avoidance safety weight to the preset hoisting stability weight and the preset obstacle avoidance safety weight by a preset linear step size if the wind disturbance risk value is not greater than the preset wind disturbance risk threshold or the obstacle risk value is less than or equal to the preset collision risk threshold. This process yields a secondary multi-objective reward function, which is then solved to obtain and output the second pose control instruction set.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.