Horizontal rotary table dynamic leveling control system based on reinforcement learning

CN122593425APending Publication Date: 2026-08-18SUZHOU FURUTA AUTOMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610673645.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-15
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0003]然而,在卧式转台旋转工况下,负载偏心、支撑刚度差异、摩擦滞回、间隙与地基扰动等因素往往呈现随转角变化的非平稳特性,传统控制方法通常采用统一控制律或少量工况分段补偿,难以在全转角范围内兼顾快速性与稳定性,容易出现姿态误差周期性波动、超调明显、恢复时间长等问题,尤其在扰动强、工况切换频繁的情况下,固定参数控制依赖人工整定,存在调参成本高、跨工况泛化差的问题;基于简单模型的前馈补偿又容易受模型失配影响,导致调平动作不够准确或产生反向补偿,进一步引发振动与执行器冲击

Benefits of technology

本发明通过对转台运行状态数据进行同步处理并引入转角扰动域的门控判别,使控制系统能够在卧式转台旋转过程中将随转角变化的偏载、摩擦与支撑差异等扰动进行结构化区分,从而在不同转角工况下调用对应的局部策略响应生成候选调平动作,提升动态调平控制对非平稳扰动的适应能力,降低姿态误差在全转角范围内的周期性波动。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122593425A_ABST
    Figure CN122593425A_ABST
Patent Text Reader

Abstract

The application discloses a horizontal rotary table dynamic leveling control system based on reinforcement learning, comprising the following steps: a state input module for generating state input data; a disturbance domain generation module for generating a rotation angle disturbance domain; a prediction information module for performing attitude evolution prediction and generating prediction information; a candidate action module for performing Black-DROPS strategy search, calling a local strategy response to generate a candidate leveling action; a corrected action module for performing amplitude reduction, direction retention and change rate blunting on the candidate leveling action to generate a corrected leveling action; a real machine response module for obtaining real machine response data; and a strategy updating module for fragmenting a cache, value discrimination screening, re-executing Black-DROPS strategy search and obtaining an updated leveling strategy result. The application improves the adaptability and stability of dynamic leveling control of the horizontal rotary table under a rotating working condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control technology, and in particular to a dynamic leveling control system for a horizontal turntable based on reinforcement learning. Background Technology

[0002] Horizontal rotary tables are widely used in precision machining, inspection and calibration, assembly alignment, and other scenarios. During workpiece rotation or continuous platform rotation, it is necessary to maintain the platform's horizontal orientation or control the orientation error within a small range to ensure machining accuracy and measurement consistency. Existing horizontal rotary table leveling systems typically consist of an orientation sensor, an angle detection device, a leveling actuator, and a controller. The controller outputs leveling commands based on the orientation error and angle position, driving the support mechanism to adjust displacement or support force, achieving static or quasi-dynamic leveling. Common control methods include fixed-parameter PID control, feedforward compensation, and simple adaptive control. Some solutions introduce compensation terms based on angle position or load estimation to reduce orientation fluctuations caused by off-center loading.

[0003] However, under the rotating condition of a horizontal turntable, factors such as load eccentricity, support stiffness differences, frictional hysteresis, clearance, and foundation disturbance often exhibit non-stationary characteristics that vary with the rotation angle. Traditional control methods usually employ a unified control law or piecewise compensation for a small number of operating conditions, which makes it difficult to balance speed and stability across the entire rotation angle range. This can easily lead to problems such as periodic fluctuations in attitude error, significant overshoot, and long recovery time. Especially under conditions of strong disturbance and frequent switching of operating conditions, fixed parameter control relies on manual tuning, resulting in high parameter tuning costs and poor generalization across operating conditions. Feedforward compensation based on a simple model is also easily affected by model mismatch, leading to inaccurate leveling actions or reverse compensation, which further causes vibration and actuator impact.

[0004] Therefore, how to provide a dynamic leveling control system for horizontal turntable based on reinforcement learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a dynamic leveling control system for a horizontal turntable based on reinforcement learning. This invention generates leveling actions collaboratively through angle disturbance domain gating, attitude evolution prediction, and Black-DROPS policy search. It also combines amplitude reduction, direction maintenance, and rate of change passivation to suppress overshoot and oscillation. Furthermore, it utilizes fragmented buffering of real machine response and value discrimination screening to achieve rapid policy updates. This improves the dynamic leveling accuracy, convergence speed, and operational stability under load changes and disturbances, while reducing actuator impact and debugging costs.

[0006] A horizontal turntable dynamic leveling control system based on reinforcement learning according to an embodiment of the present invention includes: The status input module is used to collect the operating status data of the horizontal rotary table during operation and perform time synchronization processing to generate status input data. The disturbance domain generation module is used to extract corner information from the state input data, perform periodic homeomorphic mapping on the corner information to generate periodic representation data, and perform gating discrimination on the periodic representation data to generate the corner disturbance domain. The prediction information module is used to perform attitude evolution prediction based on state input data and the rotation perturbation domain, and generate prediction information. The candidate action module is used to perform Black-DROPS policy search based on state input data, corner perturbation domain and prediction information, call the local policy response corresponding to the corner perturbation domain, and generate candidate leveling actions; The correction action module is used to perform amplitude reduction processing, direction preservation processing, and rate of change passivation processing on candidate leveling actions based on prediction information, and generate corrected leveling actions. The actual machine response module is used to input the correction and leveling action into the horizontal turntable leveling actuator, drive the support mechanism to perform displacement adjustment and support force adjustment, and obtain actual machine response data; The policy update module is used to perform fragmented caching based on the actual machine response data and the corner disturbance domain, generate actual machine running fragment data and perform value discrimination and filtering, update the attitude evolution prediction process and local policy response, and re-execute the Black-DROPS policy search based on the updated local prediction unit to obtain the updated leveling policy result.

[0007] Optionally, the status input module includes: By sampling sequences from different paths, the operating status data of the horizontal rotary table during operation is collected to form the original dataset; The synchronization start time is determined from the original dataset, the sampling period is controlled, a synchronization time series is constructed, time synchronization processing is performed on each sampling sequence, and the synchronization sampling value is calculated by linear interpolation to form a synchronization dataset. The synchronization data set is framed according to the synchronization time sequence. At each synchronization time point, the synchronization sample values ​​of each channel are arranged in a fixed order to form a state vector. The state vector is used as a data frame to generate state input data.

[0008] Optionally, the disturbance domain generation module includes: The current control phase angle information is read from the status input data. The angle information is a single angle value. The angle value is normalized based on a preset main value interval. When the angle value is greater than the upper limit of the main value interval, the angle value is subtracted by an integer angle. When the angle value is less than the lower limit of the main value interval, the angle value is added by an integer angle to obtain the normalized angle. Perform periodic homeomorphism mapping on the normalized rotation angle, calculate the sine and cosine representation values ​​of the normalized rotation angle, and form two-dimensional periodic representation data in a fixed order. Use the two-dimensional periodic representation data as the gating discrimination input data. Gating discrimination is performed on the input data of the gating discrimination. A set of numbers and a gating reference vector and discrimination threshold corresponding to each number in the set are configured. The similarity score between the periodic representation data and each gating reference vector is calculated one by one based on the two-dimensional periodic representation data. Numbers with similarity scores less than the discrimination threshold are removed. The number with the largest similarity score is selected from the remaining numbers and the selected number is determined as the corner perturbation domain.

[0009] Optionally, the prediction information module includes: Based on the corner perturbation domain, the corresponding prediction parameter group is retrieved in the parameter storage area. Dimension matching verification is performed in combination with the state input data. If the dimension matching verification fails, the prediction parameter group corresponding to the corner perturbation domain is retrieved again. If the dimension matching verification passes, a loading completion flag is output. Based on the loading completion marker, the prediction calculation is initiated. The state input data and the prediction parameter group are multiplied by matrix and vector. Each row of the matrix is ​​multiplied element by element with the state input data and accumulated to obtain an intermediate prediction vector. The parameters for bias superposition are taken from the prediction parameter group. The intermediate prediction vector and the bias parameters are added element by element to obtain the attitude prediction value. The uncertainty parameter is taken from the prediction parameter group. The state input data is subjected to a quadratic operation. The quadratic operation calculates the intermediate vector and the inner product of the intermediate vector and the state input data to obtain the uncertainty value. Create a prediction information data structure, write the angle disturbance domain, attitude prediction value and uncertainty value into the corresponding fields of the prediction information data structure, and output the prediction information.

[0010] Optionally, the candidate action module includes: The local policy response is determined based on the corner perturbation domain, policy search input data is generated, the number of candidate policies, action prediction time domain length, number of iterations, perturbation scale parameter, weight generation parameter are set, and the current policy parameter record is read. Initiate Black-DROPS policy search, generate candidate policy parameter record numbers based on the current policy parameter record and perturbation scale parameters, generate random perturbations in sequence, superimpose the random perturbations onto the current policy parameter record to obtain candidate policy parameter records, and summarize all candidate policy parameter records to obtain the candidate policy parameter set; The policy search input data and the current candidate policy parameter record are input into the local policy response to obtain the first time-domain step candidate leveling action. The first time-domain step candidate leveling action is combined with the policy search input data to generate the second time-domain step policy search input data. The time-domain step policy search input data is generated recursively according to the time-domain step number, and the corresponding time-domain step candidate leveling action is obtained. This process continues until the action prediction time-domain length ends, resulting in a candidate leveling action sequence. For each candidate leveling action sequence, rolling prediction evaluation and assessment calculation is performed. The current attitude quantity is extracted from the state input data as the initial rolling value. The initial rolling value, the candidate leveling action of the current time step, and the prediction information are input into the attitude recursion process according to the time step number to obtain the predicted attitude of the current time step. The initial rolling value is then updated until the end of the action prediction time domain length to obtain the predicted attitude sequence. The attitude reference value is read and the attitude error between the predicted attitude and the attitude reference value, the amplitude of the candidate leveling action, and the change of the candidate leveling action in adjacent time steps are calculated according to the time step number. These are weighted according to the corresponding weight coefficients and accumulated within the action prediction time domain length to obtain the evaluation value. The Black-DROPS policy search and update is completed based on the evaluation value. The candidate policy parameter records are sorted from best to worst according to the evaluation value, and normalized weight coefficients are generated. The candidate policy parameter records are weighted and summed according to the normalized weight coefficients to obtain the updated policy parameter records. The sampling, rolling generation, rolling evaluation, evaluation calculation, and weighted update are performed in a loop according to the number of iterations until the number of iterations ends. Based on the final updated policy parameter records and the policy search input data, the local policy response is called and the candidate leveling action is output.

[0011] Optionally, the correction action module includes: Read candidate leveling actions and prediction information and perform amplitude reduction processing. Use preset mapping rules to convert uncertainty index into reduction factor. Input the candidate leveling action into amplitude to calculate the candidate amplitude. Perform amplitude limit calculation with the candidate amplitude and reduction factor to obtain the target amplitude. The candidate leveling action is subjected to direction preservation processing, and the candidate leveling action is subjected to normalized direction calculation to obtain the direction vector. When the candidate amplitude is zero, the direction vector is set to zero vector. Based on the direction vector and the target amplitude, the amplitude assignment calculation is performed to obtain the intermediate leveling action. The intermediate leveling action is subjected to rate passivation processing. The difference vector between the intermediate leveling action and the candidate leveling action is calculated to obtain the change vector. The magnitude of the change vector is calculated to obtain the change magnitude. When the change magnitude is less than the upper limit of the rate of change, the intermediate leveling action is determined as the corrected leveling action. When the change magnitude is greater than the upper limit of the rate of change, the change vector is subjected to amplitude limiting scaling to obtain the amplitude limiting change vector. The candidate leveling action and the amplitude limiting change vector are superimposed to obtain the corrected leveling action.

[0012] Optionally, the actual response module includes: Receive correction and leveling actions, allocate the correction and leveling actions according to the number of control channels of the support mechanism, and generate a set of displacement adjustment amount and a set of support force adjustment amount; Read the current displacement feedback data and current support force feedback data of the support mechanism, perform superposition calculation on the set of displacement adjustment amount and the current displacement feedback data to generate displacement target data, perform superposition calculation on the set of support force adjustment amount and the current support force feedback data to generate support force target data, input the displacement target data and support force target data into the horizontal turntable leveling actuator, drive the support mechanism to perform displacement adjustment and support force adjustment, and generate an execution completion mark; When the completion mark meets the valid conditions, the displacement feedback data and support force feedback data are combined in a fixed order to generate the actual machine response data.

[0013] Optionally, the policy update module includes: The actual machine response data and the corner disturbance domain are written into the fragmented buffer in the order of arrival. The fragment length and fragment step size are set, and the buffer content within the fragment length range is captured by moving according to the fragment step size to generate the actual machine running fragment data. Value judgment calculation is performed on each segment of the actual machine operation data. The displacement feedback difference between adjacent sampling points within the segment is calculated and accumulated to obtain the displacement change index. The support force feedback difference between adjacent sampling points within the segment is calculated and accumulated to obtain the force change index. The consistency index is obtained by counting the number of times the displacement change direction and the force change direction are consistent within the segment. The value index is obtained by weighting and summing the displacement change index, force change index, and consistency index according to preset weights. Set a value threshold, compare the value index with the value threshold, and determine the real machine operation segment data that meets the threshold condition as high value segment data. Classify the high value segment data according to the corner disturbance domain and write it into the feedback buffer to generate a high value segment set. The high-value fragment set is input into the attitude evolution prediction process update flow, the displacement feedback data and support force feedback data in the high-value fragment set are extracted and corresponding records with the rotation disturbance domain are established, the prediction parameters are updated and written to the parameter storage area to complete the attitude evolution prediction process update, the high-value fragment set is input into the local policy response update flow, the policy parameters are updated and written to the policy parameter storage area to complete the local policy response update. Based on the updated attitude evolution prediction process and local policy response, the Black-DROPS policy search is re-executed to generate updated leveling policy results.

[0014] The beneficial effects of this invention are: This invention enables the control system to structurally distinguish disturbances such as off-center load, friction, and support differences that vary with the rotation angle during the rotation of the horizontal turntable by synchronously processing the turntable's operating status data and introducing gating discrimination of the rotation angle disturbance domain. This allows the system to call corresponding local strategy responses to generate candidate leveling actions under different rotation angle conditions, thereby improving the adaptability of dynamic leveling control to non-stationary disturbances and reducing the periodic fluctuation of attitude error across the entire rotation angle range.

[0015] This invention introduces attitude evolution prediction to obtain prediction information before generating candidate leveling actions, and uses Black-DROPS strategy search to realize rolling evaluation and optimization of candidate actions, so that the generation of leveling actions is transformed from a single error feedback to a strategy search process oriented towards prediction effect, thereby improving convergence speed and control stability under disturbance and load changes.

[0016] This invention performs amplitude reduction, direction preservation, and rate of change passivation processing on candidate leveling actions in the action output stage. This can suppress the risks of overshoot, oscillation, and actuator impact caused by excessively large action amplitude and excessively rapid changes while ensuring the consistency of the leveling correction direction, thereby improving the system's operational safety and stability under continuous rotation conditions.

[0017] This invention utilizes real-machine response data and the corner perturbation domain for fragmented caching and value discrimination screening, and updates the attitude evolution prediction process and local policy response accordingly. The Black-DROPS policy search is then performed again to obtain the updated leveling policy results, which transforms the policy update from full-data-driven to high-value fragment-driven, improving online update efficiency and effective sample utilization, reducing debugging costs, and enhancing the continuous optimization capability across operating conditions. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a horizontal turntable dynamic leveling control system based on reinforcement learning proposed in this invention. Figure 2 This is a schematic diagram of the module structure of a horizontal turntable dynamic leveling control system based on reinforcement learning proposed in this invention. Figure 3 This diagram illustrates the strategy update and refeedback mechanism based on Black-DROPS strategy search in a horizontal turntable dynamic leveling control system based on reinforcement learning proposed in this invention. Detailed Implementation

[0019] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0020] refer to Figures 1-3 A dynamic leveling control system for a horizontal turntable based on reinforcement learning, comprising: The status input module is used to collect the operating status data of the horizontal rotary table during operation and perform time synchronization processing to generate status input data. The disturbance domain generation module is used to extract corner information from the state input data, perform periodic homeomorphic mapping on the corner information to generate periodic representation data, and perform gating discrimination on the periodic representation data to generate the corner disturbance domain. The prediction information module is used to perform attitude evolution prediction based on state input data and the rotation perturbation domain, and generate prediction information. The candidate action module is used to perform Black-DROPS policy search based on state input data, corner perturbation domain and prediction information, call the local policy response corresponding to the corner perturbation domain, and generate candidate leveling actions; The correction action module is used to perform amplitude reduction processing, direction preservation processing, and rate of change passivation processing on candidate leveling actions based on prediction information, and generate corrected leveling actions. The actual machine response module is used to input the correction and leveling action into the horizontal turntable leveling actuator, drive the support mechanism to perform displacement adjustment and support force adjustment, and obtain actual machine response data; The policy update module is used to perform fragmented caching based on the actual machine response data and the corner disturbance domain, generate actual machine running fragment data and perform value discrimination and filtering, update the attitude evolution prediction process and local policy response, and re-execute the Black-DROPS policy search based on the updated local prediction unit to obtain the updated leveling policy result.

[0021] In this embodiment, the status input module includes: By sampling sequences from different paths, the operating status data of the horizontal turntable during operation is collected. The turntable operating status data includes attitude information, rotation angle information, support mechanism status information, and leveling execution information, forming a raw data set. The sampling sequence includes sampled values ​​and timestamps. The synchronization start time and sampling period are determined from the original dataset. A synchronization time series is constructed. The synchronization time series is determined by the synchronization start time and the sampling period. Time synchronization processing is performed on each sampling sequence. The time synchronization processing includes selecting a pair of adjacent sampling records in the current sampling sequence at each synchronization time point, and using linear interpolation to calculate the synchronization sampling value to form a synchronization dataset. The synchronization data set is framed according to the synchronization time sequence. At each synchronization time point, the synchronization sample values ​​of each channel are arranged in a fixed order to form a state vector. The state vector is used as a data frame to generate state input data.

[0022] In this embodiment, the disturbance domain generation module includes: The current control phase angle information is read from the status input data. The angle information is a single angle value. The angle value is normalized based on the preset main value range. When the angle value is greater than the upper limit of the main value range, the angle value is subtracted by an integer angle. When the angle value is less than the lower limit of the main value range, the angle value is added by an integer angle to obtain the normalized angle. Perform periodic homeomorphism mapping on the normalized rotation angle, calculate the sine and cosine representation values ​​of the normalized rotation angle, and form two-dimensional periodic representation data in a fixed order. Use the two-dimensional periodic representation data as the gating discrimination input data. Gating discrimination is performed on the input data. A set of numbers and a gate reference vector and discrimination threshold are configured for each number in the set. The gate reference vector is a two-dimensional vector with the same dimension as the periodic representation data. The discrimination threshold is a similarity discrimination threshold. The similarity score between the periodic representation data and each gate reference vector is calculated one by one based on the two-dimensional periodic representation data. The similarity score is calculated by the inner product of the periodic representation data and the gate reference vector. Numbers with similarity scores less than the discrimination threshold are removed. The number with the largest similarity score is selected from the remaining numbers and the selected number is determined as the corner perturbation domain.

[0023] In this embodiment, the prediction information module includes: Based on the corner perturbation domain, the corresponding prediction parameter set is retrieved in the parameter storage area. Dimension matching verification is performed in conjunction with the state input data. If the dimension matching verification fails, the prediction parameter set corresponding to the corner perturbation domain is retrieved again. If the dimension matching verification passes, a loading completion flag is output. The parameter storage area stores the correspondence between the prediction parameter set and the corner perturbation domain. Based on the loading completion marker, the prediction calculation is initiated. The state input data and the prediction parameter group are multiplied by matrix and vector. Each row of the matrix is ​​multiplied element by element with the state input data and the results are accumulated to obtain the intermediate prediction vector. The parameters for bias superposition are taken from the prediction parameter group. The intermediate prediction vector and the bias parameters are added element by element to obtain the attitude prediction value. The uncertainty parameter is taken from the prediction parameter group. The state input data is subjected to a quadratic operation. The quadratic operation calculates the intermediate vector and the inner product of the intermediate vector and the state input data to obtain the uncertainty value. Create a prediction information data structure, write the angle disturbance domain, attitude prediction value and uncertainty value into the corresponding fields of the prediction information data structure, and output the prediction information.

[0024] In this embodiment, the candidate action module includes: The local policy response is determined based on the corner perturbation domain, policy search input data is generated, and the number of candidate policies, action prediction time domain length, number of iterations, perturbation scale parameter, and weight generation parameter are set. The current policy parameter record is read, the local policy response is the policy mapping function, the policy mapping function is established in correspondence with the corner perturbation domain, and the corner perturbation domain determines the policy mapping function selection result. Initiate Black-DROPS policy search, generate candidate policy parameter record numbers based on the current policy parameter record and perturbation scale parameters, generate random perturbations in sequence, superimpose the random perturbations onto the current policy parameter record to obtain candidate policy parameter records, and summarize all candidate policy parameter records to obtain the candidate policy parameter set; The policy search input data and the current candidate policy parameter record are input into the local policy response to obtain the first time-domain step candidate leveling action. The first time-domain step candidate leveling action is combined with the policy search input data to generate the second time-domain step policy search input data. The time-domain step policy search input data is generated recursively according to the time-domain step number, and the corresponding time-domain step candidate leveling action is obtained. This process continues until the action prediction time-domain length ends, resulting in a candidate leveling action sequence. For each candidate leveling action sequence, rolling prediction evaluation and assessment calculation is performed. The current attitude quantity is extracted from the state input data as the initial rolling value. The initial rolling value, the candidate leveling action of the current time step, and the prediction information are input into the attitude recursion process according to the time step number to obtain the predicted attitude of the current time step. The initial rolling value is then updated until the end of the action prediction time domain length to obtain the predicted attitude sequence. The attitude reference value is read and the attitude error between the predicted attitude and the attitude reference value, the amplitude of the candidate leveling action, and the change of the candidate leveling action in adjacent time steps are calculated according to the time step number. These are weighted according to the corresponding weight coefficients and accumulated within the action prediction time domain length to obtain the evaluation value. The Black-DROPS policy search and update is completed based on the evaluation value. The candidate policy parameter records are sorted from best to worst according to the evaluation value, and normalized weight coefficients are generated. The candidate policy parameter records are weighted and summed according to the normalized weight coefficients to obtain the updated policy parameter records. The sampling, rolling generation, rolling evaluation, evaluation calculation, and weighted update are performed in a loop according to the number of iterations until the number of iterations ends. Based on the final updated policy parameter records and the policy search input data, the local policy response is called and the candidate leveling action is output.

[0025] This invention selects local policy responses through the corner perturbation domain and introduces a candidate perturbation sampling and weighted update mechanism based on Black-DROPS policy search. Combined with rolling prediction evaluation and multiple cost accumulation evaluation methods, it achieves rapid optimization and stable output of candidate leveling actions under continuous rotation and changing operating conditions. This improves the adaptability, convergence speed and anti-disturbance capability of dynamic leveling, reduces the risk of overshoot and oscillation, and reduces the cost of manual tuning and trial and error.

[0026] In this embodiment, the correction action module includes: Read candidate leveling actions and prediction information and perform amplitude reduction processing. Use preset mapping rules to convert uncertainty index into reduction factor. Input the candidate leveling action into amplitude to calculate the candidate amplitude. Perform amplitude limit calculation with the candidate amplitude and reduction factor to obtain the target amplitude. The candidate leveling action is subjected to direction preservation processing, and the candidate leveling action is subjected to normalized direction calculation to obtain the direction vector. When the candidate amplitude is zero, the direction vector is set to zero vector. Based on the direction vector and the target amplitude, the amplitude assignment calculation is performed to obtain the intermediate leveling action. The intermediate leveling action is subjected to rate passivation processing. The difference vector between the intermediate leveling action and the candidate leveling action is calculated to obtain the change vector. The magnitude of the change vector is calculated to obtain the change magnitude. When the change magnitude is less than the upper limit of the rate of change, the intermediate leveling action is determined as the corrected leveling action. When the change magnitude is greater than the upper limit of the rate of change, the change vector is subjected to amplitude limiting scaling to obtain the amplitude limiting change vector. The candidate leveling action and the amplitude limiting change vector are superimposed to obtain the corrected leveling action.

[0027] This invention reduces the amplitude of candidate leveling actions by utilizing the uncertainty in the predicted information and keeps the correction direction consistent. At the same time, it uses the rate of change to passivate and limit the magnitude of sudden changes in the action. This makes the leveling output automatically tend to be conservative and continuous under the conditions of model uncertainty or sudden disturbance, reducing the risk of overshoot, oscillation and actuator impact, and improving the stability, safety and control accuracy of dynamic leveling under the rotating condition of horizontal turntable.

[0028] In this embodiment, the actual response module includes: Receive correction and leveling actions, allocate the correction and leveling actions according to the number of control channels of the support mechanism, and generate a set of displacement adjustment amount and a set of support force adjustment amount; Read the current displacement feedback data and current support force feedback data of the support mechanism, perform superposition calculation on the set of displacement adjustment amount and the current displacement feedback data to generate displacement target data, perform superposition calculation on the set of support force adjustment amount and the current support force feedback data to generate support force target data, input the displacement target data and support force target data into the horizontal turntable leveling actuator, drive the support mechanism to perform displacement adjustment and support force adjustment, and generate an execution completion mark; When the completion mark meets the valid conditions, the displacement feedback data and support force feedback data are combined in a fixed order to generate the actual machine response data.

[0029] In this embodiment, the policy update module includes: The actual machine response data and the corner disturbance domain are written into the fragmented buffer in the order of arrival. The fragment length and fragment step size are set, and the buffer content within the fragment length range is captured by moving according to the fragment step size to generate the actual machine running fragment data. Value judgment calculation is performed on each segment of the actual machine operation data. The displacement feedback difference between adjacent sampling points within the segment is calculated and accumulated to obtain the displacement change index. The support force feedback difference between adjacent sampling points within the segment is calculated and accumulated to obtain the force change index. The consistency index is obtained by counting the number of times the displacement change direction and the force change direction are consistent within the segment. The value index is obtained by weighting and summing the displacement change index, force change index, and consistency index according to preset weights. Set a value threshold, compare the value index with the value threshold, and determine the real machine operation segment data that meets the threshold condition as high value segment data. Classify the high value segment data according to the corner disturbance domain and write it into the feedback buffer to generate a high value segment set. The high-value fragment set is input into the attitude evolution prediction process update flow, the displacement feedback data and support force feedback data in the high-value fragment set are extracted and corresponding records with the rotation disturbance domain are established, the prediction parameters are updated and written to the parameter storage area to complete the attitude evolution prediction process update, the high-value fragment set is input into the local policy response update flow, the policy parameters are updated and written to the policy parameter storage area to complete the local policy response update. Based on the updated attitude evolution prediction process and local policy response, the Black-DROPS policy search is re-executed to generate updated leveling policy results.

[0030] This invention organizes real-world response data into fragments according to the corner perturbation domain and performs value discrimination screening, only re-updating high-value fragments. This improves the effective information density and learning efficiency under limited real-world sample conditions, reduces the interference of invalid data on prediction and policy, promotes the rapid convergence of attitude evolution prediction process and local policy response to corner-related perturbations, and obtains more robust updated leveling policy results when re-executing Black-DROPS policy search, thereby improving dynamic leveling accuracy and operational stability.

[0031] Example 1: To verify the feasibility of this invention in practice, it was applied to a dynamic leveling control scenario for a horizontal rotary table under continuous rotation. This rotary table is used to carry medium-to-large tooling fixtures and workpiece assemblies for processing and inspection. During operation, the rotary table needs to rotate continuously at a constant angular velocity, while maintaining the platform tilt angle error within a small range to avoid machining trajectory deviation and measurement reference drift. The prominent problem in this scenario is that the rotary table load eccentricity exhibits periodic disturbances with changes in rotation angle. Frictional hysteresis and clearance in the support mechanism introduce nonlinearity, and ground micro-vibrations and slight drift of the load center of gravity cause continuous fluctuations in attitude error. Traditional fixed parameter control is prone to overshoot and secondary oscillations in certain rotation angle ranges. Furthermore, to avoid oscillations, the control gain must be reduced, resulting in a longer attitude recovery time. It is difficult to simultaneously meet the requirements of speed and stability under different load and speed combinations.

[0032] In this scenario, the state input module of this invention continuously collects turntable operating status data and performs time synchronization processing to form state input data for subsequent modules to call. The state input data covers attitude information, angle information, support mechanism status information, and leveling execution information, ensuring that the control link uses consistent data frames in the same control phase. The disturbance domain generation module extracts angle information from the state input data and performs periodic homeomorphism mapping to obtain periodic characterization data. Then, it obtains the angle disturbance domain through gating discrimination, enabling the system to explicitly represent the disturbance structure that changes with the angle and provide a basis for subsequent local strategy response selection.

[0033] The prediction information module performs attitude evolution prediction under the constraints of state input data and rotation perturbation domain, and outputs prediction information containing attitude prediction values ​​and uncertainty values. The prediction information is used for both rolling evaluation of candidate action modules and correction of amplitude reduction and rate of change passivation modulation of action modules. The candidate action module performs Black-DROPS policy search under the settings of the number of candidate policies, the length of the action prediction time domain, and the number of iterations. It generates a set of candidate policy parameters by superimposing random perturbations on the current policy parameter records, and then calls the local policy response corresponding to the corner perturbation domain to generate a sequence of candidate leveling actions. It performs rolling prediction evaluation and assessment calculations in combination with prediction information, obtains the evaluation value, and then weights and sums the candidate policy parameters according to the normalized weight coefficients and updates the policy parameter records. Finally, it outputs the candidate leveling actions. The correction action module uses the uncertainty index in the prediction information to reduce the amplitude of the candidate leveling actions while keeping the direction of the candidate leveling actions unchanged. It then applies a rate-of-change passivation constraint to the change of intermediate leveling actions relative to the candidate leveling actions to generate correction leveling actions, thereby avoiding overshoot and shock caused by excessively large or rapid changes in actions when the model is uncertain or the perturbation changes abruptly.

[0034] The real-world response module inputs the leveling action into the leveling actuator to drive the support mechanism to adjust displacement and support force, thus obtaining real-world response data. The strategy update module fragments and caches the real-world response data and the angle disturbance domain, and performs value discrimination and screening. Only high-value fragments are used to update the attitude evolution prediction process and local strategy response. After the update, the Black-DROPS strategy search is re-executed to obtain the updated leveling strategy result, enabling the system to continuously improve strategy adaptability under the condition of limited real-world samples.

[0035] To demonstrate the problem solved by this invention and verify its beneficial effects, a traditional control baseline scheme was set as the comparison scheme. The baseline scheme uses a simplified feedforward compensation with fixed-parameter PID superposition. The feedforward term provides the compensation amount based on the estimated rotation angle and static eccentricity. The controller outputs support displacement correction according to the attitude error. The control parameters are obtained through manual tuning. The specific comparative experiments are shown in Table 1. Table 1. Comparison of Dynamic Leveling Performance of Horizontal Turntables under Continuous Rotation Conditions

[0036] Table 1 shows the statistical mean comparison. Under continuous rotation and angle-related disturbance conditions, the proposed solution reduces the peak-to-peak tilt error from 36 arcseconds to 16 arcseconds and the root mean square error from 12 arcseconds to 5 arcseconds. Simultaneously, it reduces the maximum overshoot from 14 arcseconds to 5 arcseconds and shortens the recovery time to ±5 arcsecond bandwidth from 3.3 seconds to 1.6 seconds, demonstrating a simultaneous improvement in dynamic leveling accuracy and convergence speed. The peak value of the actuator command change rate decreases from 0.98 to 0.60, the peak value of the support force fluctuation decreases from 310N to 185N, and the peak value of the drive current decreases from 8.3A to 6.5A, indicating smoother action and less execution impact. Regarding strategy updates, only an average of 28 high-value fragments are needed to reach a stable strategy threshold, reflecting the high sample efficiency of the fragmented value judgment and refeeding update mechanism under limited real-world data conditions.

[0037] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A dynamic leveling control system for a horizontal rotary table based on reinforcement learning, characterized in that, include: The status input module is used to collect the operating status data of the horizontal rotary table during operation and perform time synchronization processing to generate status input data. The disturbance domain generation module is used to extract corner information from the state input data, perform periodic homeomorphic mapping on the corner information to generate periodic representation data, and perform gating discrimination on the periodic representation data to generate the corner disturbance domain. The prediction information module is used to perform attitude evolution prediction based on state input data and the rotation perturbation domain, and generate prediction information. The candidate action module is used to perform Black-DROPS policy search based on state input data, corner perturbation domain and prediction information, call the local policy response corresponding to the corner perturbation domain, and generate candidate leveling actions; The correction action module is used to perform amplitude reduction processing, direction preservation processing, and rate of change passivation processing on candidate leveling actions based on prediction information, and generate correction leveling actions. The actual machine response module is used to input the correction and leveling action into the horizontal turntable leveling actuator, drive the support mechanism to perform displacement adjustment and support force adjustment, and obtain actual machine response data; The policy update module is used to perform fragmented caching based on the actual machine response data and the corner disturbance domain, generate actual machine running fragment data and perform value discrimination and filtering, update the attitude evolution prediction process and local policy response, and re-execute the Black-DROPS policy search based on the updated local prediction unit to obtain the updated leveling policy result.

2. The dynamic leveling control system for horizontal rotary tables based on reinforcement learning according to claim 1, characterized in that, The status input module includes: By sampling sequences from different paths, the operating status data of the horizontal rotary table during operation is collected to form the original dataset; The synchronization start time is determined from the original dataset, the sampling period is controlled, a synchronization time series is constructed, time synchronization processing is performed on each sampling sequence, and the synchronization sampling value is calculated by linear interpolation to form a synchronization dataset. The synchronization data set is framed according to the synchronization time sequence. At each synchronization time point, the synchronization sample values ​​of each channel are arranged in a fixed order to form a state vector. The state vector is used as a data frame to generate state input data.

3. The dynamic leveling control system for horizontal rotary tables based on reinforcement learning according to claim 1, characterized in that, The disturbance domain generation module includes: The current control phase angle information is read from the status input data. The angle information is a single angle value. The angle value is normalized based on a preset main value interval. When the angle value is greater than the upper limit of the main value interval, the angle value is subtracted by an integer angle. When the angle value is less than the lower limit of the main value interval, the angle value is added by an integer angle to obtain the normalized angle. Perform periodic homeomorphism mapping on the normalized rotation angle, calculate the sine and cosine representation values ​​of the normalized rotation angle, and form two-dimensional periodic representation data in a fixed order. Use the two-dimensional periodic representation data as the gating discrimination input data. Gating discrimination is performed on the input data of the gating discrimination. A set of numbers and a gating reference vector and discrimination threshold corresponding to each number in the set are configured. The similarity score between the periodic representation data and each gating reference vector is calculated one by one based on the two-dimensional periodic representation data. Numbers with similarity scores less than the discrimination threshold are removed. The number with the largest similarity score is selected from the remaining numbers and the selected number is determined as the corner perturbation domain.

4. The dynamic leveling control system for horizontal rotary tables based on reinforcement learning according to claim 1, characterized in that, The prediction information module includes: Based on the corner perturbation domain, the corresponding prediction parameter group is retrieved in the parameter storage area. Dimension matching verification is performed in combination with the state input data. If the dimension matching verification fails, the prediction parameter group corresponding to the corner perturbation domain is retrieved again. If the dimension matching verification passes, a loading completion flag is output. Based on the loading completion marker, the prediction calculation is initiated. The state input data and the prediction parameter group are multiplied by matrix and vector. Each row of the matrix is ​​multiplied element by element with the state input data and accumulated to obtain an intermediate prediction vector. The parameters for bias superposition are taken from the prediction parameter group. The intermediate prediction vector and the bias parameters are added element by element to obtain the attitude prediction value. The uncertainty parameter is taken from the prediction parameter group. The state input data is subjected to a quadratic operation. The quadratic operation calculates the intermediate vector and the inner product of the intermediate vector and the state input data to obtain the uncertainty value. Create a prediction information data structure, write the angle disturbance domain, attitude prediction value and uncertainty value into the corresponding fields of the prediction information data structure, and output the prediction information.

5. The dynamic leveling control system for horizontal rotary tables based on reinforcement learning according to claim 1, characterized in that, The candidate action module includes: The local policy response is determined based on the corner perturbation domain, policy search input data is generated, the number of candidate policies, action prediction time domain length, number of iterations, perturbation scale parameter, weight generation parameter are set, and the current policy parameter record is read. Initiate Black-DROPS policy search, generate candidate policy parameter record numbers based on the current policy parameter record and perturbation scale parameters, generate random perturbations in sequence, superimpose the random perturbations onto the current policy parameter record to obtain candidate policy parameter records, and summarize all candidate policy parameter records to obtain the candidate policy parameter set; The policy search input data and the current candidate policy parameter record are input into the local policy response to obtain the first time-domain step candidate leveling action. The first time-domain step candidate leveling action is combined with the policy search input data to generate the second time-domain step policy search input data. The time-domain step policy search input data is generated recursively according to the time-domain step number, and the corresponding time-domain step candidate leveling action is obtained. This process continues until the action prediction time-domain length ends, resulting in a candidate leveling action sequence. For each candidate leveling action sequence, rolling prediction evaluation and assessment calculation is performed. The current attitude quantity is extracted from the state input data as the initial rolling value. The initial rolling value, the candidate leveling action of the current time step, and the prediction information are input into the attitude recursion process according to the time step number to obtain the predicted attitude of the current time step. The initial rolling value is then updated until the end of the action prediction time domain length to obtain the predicted attitude sequence. The attitude reference value is read and the attitude error between the predicted attitude and the attitude reference value, the amplitude of the candidate leveling action, and the change of the candidate leveling action in adjacent time steps are calculated according to the time step number. These are weighted according to the corresponding weight coefficients and accumulated within the action prediction time domain length to obtain the evaluation value. The Black-DROPS policy search and update is completed based on the evaluation value. The candidate policy parameter records are sorted from best to worst according to the evaluation value, and normalized weight coefficients are generated. The candidate policy parameter records are weighted and summed according to the normalized weight coefficients to obtain the updated policy parameter records. The sampling, rolling generation, rolling evaluation, evaluation calculation, and weighted update are performed in a loop according to the number of iterations until the number of iterations ends. Based on the final updated policy parameter records and the policy search input data, the local policy response is called and the candidate leveling action is output.

6. The dynamic leveling control system for a horizontal rotary table based on reinforcement learning of claim 1, wherein, The correction action module includes: Read candidate leveling actions and prediction information and perform amplitude reduction processing. Use preset mapping rules to convert uncertainty index into reduction factor. Input the candidate leveling action into amplitude to calculate the candidate amplitude. Perform amplitude limit calculation with the candidate amplitude and reduction factor to obtain the target amplitude. The candidate leveling action is subjected to direction preservation processing, and the candidate leveling action is subjected to normalized direction calculation to obtain the direction vector. When the candidate amplitude is zero, the direction vector is set to zero vector. Based on the direction vector and the target amplitude, the amplitude assignment calculation is performed to obtain the intermediate leveling action. The intermediate leveling action is subjected to rate passivation processing. The difference vector between the intermediate leveling action and the candidate leveling action is calculated to obtain the change vector. The magnitude of the change vector is calculated to obtain the change magnitude. When the change magnitude is less than the upper limit of the rate of change, the intermediate leveling action is determined as the corrected leveling action. When the change magnitude is greater than the upper limit of the rate of change, the change vector is subjected to amplitude limiting scaling to obtain the amplitude limiting change vector. The candidate leveling action and the amplitude limiting change vector are superimposed to obtain the corrected leveling action.

7. The dynamic leveling control system for a horizontal rotary table based on reinforcement learning according to claim 1, wherein, The actual machine response module includes: Receive correction and leveling actions, allocate the correction and leveling actions according to the number of control channels of the support mechanism, and generate a set of displacement adjustment amount and a set of support force adjustment amount; Read the current displacement feedback data and current support force feedback data of the support mechanism, perform superposition calculation on the set of displacement adjustment amount and the current displacement feedback data to generate displacement target data, perform superposition calculation on the set of support force adjustment amount and the current support force feedback data to generate support force target data, input the displacement target data and support force target data into the horizontal turntable leveling actuator, drive the support mechanism to perform displacement adjustment and support force adjustment, and generate an execution completion mark; When the completion mark meets the valid conditions, the displacement feedback data and support force feedback data are combined in a fixed order to generate the actual machine response data.

8. A horizontal turntable dynamic leveling control system based on reinforcement learning according to claim 1, characterized in that, The policy update module includes: The actual machine response data and the corner disturbance domain are written into the fragmented buffer in the order of arrival. The fragment length and fragment step size are set, and the buffer content within the fragment length range is captured by moving according to the fragment step size to generate the actual machine running fragment data. Value judgment calculation is performed on each segment of the actual machine operation data. The displacement feedback difference between adjacent sampling points within the segment is calculated and accumulated to obtain the displacement change index. The support force feedback difference between adjacent sampling points within the segment is calculated and accumulated to obtain the force change index. The consistency index is obtained by counting the number of times the displacement change direction and the force change direction are consistent within the segment. The value index is obtained by weighting and summing the displacement change index, force change index, and consistency index according to preset weights. Set a value threshold, compare the value index with the value threshold, and determine the real machine operation segment data that meets the threshold condition as high value segment data. Classify the high value segment data according to the corner disturbance domain and write it into the feedback buffer to generate a high value segment set. The high-value fragment set is input into the attitude evolution prediction process update flow, the displacement feedback data and support force feedback data in the high-value fragment set are extracted and corresponding records with the rotation disturbance domain are established, the prediction parameters are updated and written to the parameter storage area to complete the attitude evolution prediction process update, the high-value fragment set is input into the local policy response update flow, the policy parameters are updated and written to the policy parameter storage area to complete the local policy response update. Based on the updated attitude evolution prediction process and local policy response, the Black-DROPS policy search is re-executed to generate updated leveling policy results.