Method and device for adjusting motor control parameters of absolute gravimeter

By introducing a reinforcement learning framework into the absolute gravity meter motor, adaptively adjusting the PID controller parameters, the problem that the motor control parameters cannot be adjusted adaptively in the prior art is solved, and the accuracy of measuring gravity acceleration values ​​is improved.

CN120185480AActive Publication Date: 2025-06-20NATIONAL INSTITUTE OF METROLOGY CHINA

Patent Information

Application Number
CN202510644794.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-20
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The PID controller parameters of existing absolute gravity meter motors mainly rely on experience settings and cannot achieve adaptive adjustment, resulting in unstable motorcycle movement and inaccurate fall of the body, which in turn reduces the accuracy of measuring gravity acceleration values.

Method used

Introducing a reinforcement learning framework, by designing agents and reward functions for each control stage, accumulating empirical data and training agents, dynamically generating accurate PID parameter adjustments, and realizing adaptive adjustment of motor control parameters.

Benefits of technology

By adaptively adjusting the motor control parameters, the accuracy and stability of the motorcycle and the falling movement are significantly improved, thereby improving the accuracy of gravity acceleration value measurement, breaking through the limitations of traditional empirical parameter adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120185480A_ABST
    Figure CN120185480A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of motor control, and discloses an absolute gravimeter motor control parameter adjusting method and device. The method comprises the following steps: determining a current control stage based on position information, motion information and motion duration of a trailer and a falling body; the intelligent agent corresponding to the stage is used for processing the corresponding state data, and action data is generated to adjust motor control parameters; calculating a reward value by utilizing a stage reward function, storing the empirical data to a corresponding playback buffer area, circulating to a receiving section, enabling the trailer and the falling body to have the same speed, and enabling the distance to be 0; if all the intelligent agents are trained, action data are generated through all the intelligent agents; and if an untrained agent exists, corresponding empirical data is extracted to continue training, and the absolute gravimeter is reset to continue iteration. By introducing an interactive learning mechanism of reinforcement learning, the gravimeter motor control parameters can be adaptively adjusted, the accuracy of trailer and falling body movement is improved, and then the measurement accuracy of gravitational acceleration is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of motor control, and in particular to a method and device for adjusting control parameters of an absolute gravimeter motor. Background Art

[0002] An absolute gravimeter is a precision measuring instrument used to directly measure the gravitational acceleration value on the earth's surface. In the process of using an absolute gravimeter to test the gravitational acceleration value, the motor's control of the trolley is divided into three stages: separation stage, free fall stage, and receiving stage. In the separation stage, it is necessary to ensure that the trolley and the falling object are separated smoothly; in the free fall stage, it is necessary to ensure that the trolley and the falling object are as relatively still as possible; in the receiving stage, it is necessary to ensure that the trolley "just" receives the falling object. The high-precision operating characteristics of the absolute gravimeter determine the need to design a rigorous control strategy for its motor.

[0003] Usually, the motion control of the absolute gravimeter motor is achieved through a classic PID (proportional-integral-differential) controller. At present, in the prior art, the absolute gravimeter motor PID controller parameters are mainly set based on experience, and adaptive adjustment cannot be achieved. Therefore, the existing absolute gravimeter motor control parameter adjustment method cannot ensure the stability of the tractor movement, the accuracy of the free fall of the falling body, etc., resulting in poor measurement accuracy of the gravity acceleration value. Therefore, it is urgent to solve this technical problem. Summary of the invention

[0004] In view of the above situation, the embodiments of the present disclosure provide a method and device for adjusting the control parameters of an absolute gravimeter motor, aiming to solve the above problem or at least partially solve the above problem.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for adjusting control parameters of an absolute gravimeter motor, the method comprising: Determine the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object in the target absolute gravimeter; According to the current control stage, the corresponding intelligent agent and state data are obtained, and the state data is processed by the intelligent agent to generate action data; the action data is a set of motor controller parameter adjustment values; Using the action data, adjusting the motor control parameters of the target absolute gravimeter; using the reward function corresponding to the current control stage, calculating the reward value; using the reward value, the action data, the state data and the acquired state data of the next time step as an experience data, storing them in the playback buffer corresponding to the current control stage; returning to the step of determining the current control stage based on the acquired position information, motion information, and motion duration of the trailer and the falling object in the target absolute gravimeter, until the current control stage is the takeover stage, and the speed of the trailer is the same as that of the falling object, and the interval distance is 0; According to the preset training termination condition, generating a judgment result of whether the training of each of the intelligent agents has been completed; If the judgment result is that the training of each of the intelligent agents is completed, then using each of the intelligent agents to process the state data of the target absolute gravimeter to generate action data to adjust the motor control parameters of the target absolute gravimeter; If the judgment result is that there is a target intelligent agent that has not been trained, and the number of experience data in the target playback buffer corresponding to the target intelligent agent is greater than a preset threshold, then sample experience is extracted from the target playback buffer, and the target intelligent agent is trained using the sample experience to obtain a new target intelligent agent; the target absolute gravimeter is reset, and the step of determining the current control stage is returned based on the position information, motion information, and motion duration of the trailer and the falling object obtained in the target absolute gravimeter.

[0006] In a second aspect, the present disclosure also provides a device for adjusting motor control parameters of an absolute gravimeter, the device comprising: A phase determination module, used to determine the current control phase based on the position information, motion information, and motion duration of the trailer and the falling object acquired in the target absolute gravimeter; An experience collection module is used to obtain the corresponding intelligent agent and state data according to the current control stage, and use the intelligent agent to process the state data to generate action data; the action data is a set of motor controller parameter adjustment amounts; the motor control parameters of the target absolute gravimeter are adjusted using the action data; the reward value is calculated using the reward function corresponding to the current control stage; the reward value, the action data, the state data and the acquired state data of the next time step are stored as an experience data in the playback buffer corresponding to the current control stage; return to the step of determining the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object in the acquired target absolute gravimeter, until the current control stage is a takeover stage, and the speed of the trailer and the falling object is the same and the interval distance is 0; A training module, configured to generate a judgment result on whether each of the agents has been trained according to a preset training termination condition; if the judgment result indicates that there is a target agent that has not been trained yet, and the number of pieces of experience data in the target replay buffer corresponding to the target agent is greater than a preset threshold, then sample experiences are extracted from the target replay buffer, and the target agent is trained using the sample experiences to obtain a new target agent; reset the target absolute gravimeter, and return to the step of determining the current control stage based on the position information, motion information, and motion duration of the tow truck and the falling body obtained by the target absolute gravimeter. An adjustment module, configured to, if the judgment result indicates that each of the agents has been trained, use each of the agents to process the state data of the target absolute gravimeter to generate action data, so as to adjust the motor control parameters of the target absolute gravimeter.

[0007] By means of the above technical solution, the method and device for adjusting the motor control parameters of an absolute gravimeter provided by the embodiments of the present disclosure introduce a reinforcement learning framework. For each stage when the motor controls the tow truck, first, using the agents and reward functions designed for each stage respectively, a large amount of experience data including successes / failures is accumulated. Using this experience data, the agents in each control stage can be trained to gradually optimize the controller parameter adjustment strategy of the agents. Finally, the trained mature agents can dynamically generate accurate PID parameter adjustment amounts according to the real-time state data of the target absolute gravimeter. Using these PID parameters, accurate adjustment of the motor control parameters of the gravimeter can be achieved. Through this interactive learning mechanism, adaptive adjustment of the motor control parameters of the absolute gravimeter can be realized, breaking through the limitations of traditional empirical parameter tuning, significantly improving the accuracy and stability of the tow truck and falling body movements, and thus improving the accuracy of measuring the gravitational acceleration value.

[0008] The above description is only an overview of the technical solution of the present disclosure. In order to be able to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present disclosure more obvious and understandable, the specific embodiments of the present disclosure are specifically given below. Description of the Drawings

[0009] The drawings described herein are used to provide a further understanding of the present disclosure, and constitute a part of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings: Figure 1 A flowchart showing the method for adjusting the motor control parameters of the absolute gravimeter provided by the embodiments of the present disclosure is shown; Figure 2 A structural diagram showing the free-fall device of the absolute gravimeter provided by the embodiments of the present disclosure is shown; Figure 3 The structural schematic diagram of the motor control parameter adjustment device for an absolute gravimeter provided by an embodiment of the present disclosure is shown. Detailed implementation manners

[0010] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the technical solutions of the present disclosure will be clearly and completely described below in conjunction with specific embodiments of the present disclosure and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.

[0011] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0012] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such use can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the term "including" and its variants should be interpreted as open-ended terms meaning "including but not limited to".

[0013] As introduced above, the parameters of the PID controller of the absolute gravimeter motor mainly rely on experience for setting and cannot achieve adaptive adjustment. Therefore, the existing methods for adjusting the motor control parameters of the absolute gravimeter cannot ensure the smoothness of the trolley movement, the accuracy of the free fall of the falling body, etc., resulting in poor measurement accuracy of the gravitational acceleration value. Based on this, the present invention proposes a method and device for adjusting the motor control parameters of the absolute gravimeter, and the present disclosure will be described in detail below through specific embodiments.

[0014] Before introducing the specific embodiments, the professional terms involved in the embodiments of the present disclosure will be explained first: 1) Environment: In reinforcement learning, the environment is the external world where the agent is located. The agent receives the state information of the environment at each time step, selects an appropriate action, and then the environment gives feedback on this action, including rewards and new states.

[0015] 2) State: The current description of the environment.

[0016] 3) Action: An action is one of the decisions that an agent can take in a given state, used to maximize the reward in a certain state. In the present disclosure, the action is a control signal output by the agent for controlling the motor of the absolute gravimeter.

[0017] 4) Reward: A reward is the feedback given by the environment after the agent executes an action, usually used to measure the quality of that action.

[0018] 5) Time step: A basic unit for the interaction between the agent and the environment. After each time step, the agent takes an action based on the current state, and the environment gives a new state and a reward.

[0019] For ease of understanding of this embodiment, first, a method for adjusting the control parameters of the motor of an absolute gravimeter disclosed in the embodiments of the present disclosure will be introduced in detail. The execution subject of the method for adjusting the control parameters of the motor of the absolute gravimeter provided in the embodiments of the present disclosure is generally a computer device with certain computing capabilities. Such a computer device includes, for example: a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, etc. In some possible implementation manners, the method for adjusting the control parameters of the motor of the absolute gravimeter may be implemented by the processor invoking computer-readable instructions stored in the memory.

[0020] Figure 1 The flowchart of the method for adjusting the control parameters of the motor of the absolute gravimeter provided in the embodiments of the present disclosure is shown. From Figure 1 it can be seen that the embodiments of the present disclosure at least include steps S101 - S106: S101: Based on the obtained position information, motion information, and motion duration of the trolley and the falling body in the target absolute gravimeter, determine the current control phase.

[0021] S102: According to the current control phase, obtain the corresponding agent and state data, and use the agent to process the state data to generate action data; the action data is a set of motor controller parameter adjustment amounts.

[0022] S103: Using the action data, adjusting the motor control parameters of the target absolute gravimeter; using the reward function corresponding to the current control stage, calculating the reward value; using the reward value, the action data, the state data and the acquired state data of the next time step as an experience data, storing them in the playback buffer corresponding to the current control stage; returning to the steps of determining the current control stage based on the acquired position information, motion information, and motion duration of the trolley and the falling object in the target absolute gravimeter, until the current control stage is the taking-over stage, and the speeds of the trolley and the falling object are the same and the interval distance is 0.

[0023] S104: According to a preset training termination condition, a judgment result is generated as to whether the training of each of the intelligent agents is completed.

[0024] S105: If the judgment result is that the training of each of the intelligent agents is completed, each of the intelligent agents is used to process the state data of the target absolute gravimeter to generate action data to adjust the motor control parameters of the target absolute gravimeter.

[0025] S106: If the judgment result is that there is a target intelligent agent that has not been trained, and the number of experience data in the target playback buffer corresponding to the target intelligent agent is greater than a preset threshold, then sample experience is extracted from the target playback buffer, and the target intelligent agent is trained using the sample experience to obtain a new target intelligent agent; the target absolute gravimeter is reset, and the step of returning to the step of determining the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object obtained in the target absolute gravimeter is returned.

[0026] The absolute gravimeter motor control parameter adjustment method provided in this embodiment introduces a reinforcement learning framework. For each stage when the motor controls the trailer, first, the intelligent agent and reward function designed for each stage are used to accumulate a large amount of experience data containing successes / failures. Using these experience data, the intelligent agent in each control stage can be trained to gradually optimize the controller parameter adjustment strategy of the intelligent agent. Finally, the maturely trained intelligent agent can dynamically generate accurate PID parameter adjustment amounts based on the real-time status data of the target absolute gravimeter. Using these PID parameters, accurate adjustment of the gravimeter motor control parameters can be achieved. Through this interactive learning mechanism, adaptive adjustment of the absolute gravimeter motor control parameters can be achieved, breaking through the limitations of traditional experience-based parameter adjustment, significantly improving the accuracy and stability of trailer and falling motion, and thus improving the accuracy of gravity acceleration value measurement.

[0027] The above S101-S106 are described in detail below.

[0028] Regarding S101 above: In this embodiment, training samples, i.e., empirical data, are first collected. In specific implementation, in this step, the position information of the trailer, the motion information of the trailer, the position information of the falling object, the motion information of the falling object, and the duration of the motion in the target absolute gravimeter are first obtained. Then, based on the above information, the current control stage is determined. Here, the position information is used to describe the spatial position of the trailer and the falling object, which can be the displacement, ordinate value, etc. of the trailer and the falling object in the target absolute gravimeter. The motion information includes instantaneous velocity value, instantaneous acceleration value, etc. The duration of the motion is the time that the target absolute gravimeter lasts from the start of the test. As described in the background technology, the current control stage includes a separation section, a free fall section, and a receiving section.

[0029] This embodiment does not limit the method for obtaining the position information, motion information, and motion duration of the trailer and the falling object. During implementation, you can choose according to the actual situation. For example, the execution subject of this embodiment can connect to the internal sensor of the target absolute gravimeter (such as laser interferometer to measure displacement, accelerometer to measure acceleration, etc.) through a USB or Ethernet interface, etc., and read the position information and motion information of the trailer and the falling object in real time. The timer module or the built-in clock of the gravimeter can be used to obtain the duration of the movement.

[0030] Then, based on the position information, motion information and motion duration of the trailer and the falling object, the current control stage is determined. Specifically, in some embodiments, the motion information includes: the instantaneous acceleration value of the trailer, the instantaneous speed value of the trailer and the instantaneous speed value of the falling object; the current control stage is determined based on the position information, motion information and motion duration of the trailer and the falling object in the target absolute gravimeter, including: Calculating the distance between the trailer and the falling object according to the position information of the trailer and the falling object; If the separation distance is 0, and the instantaneous acceleration value of the trailer is less than or equal to the gravity acceleration value, then the current control stage is determined to be the separation stage; If the movement duration is less than or equal to a preset time threshold, and the instantaneous acceleration value of the trailer is less than or equal to the gravity acceleration value, then the current control stage is determined to be a free fall stage; If the movement duration is greater than the time threshold, and the instantaneous speed value of the trailer is less than or equal to the instantaneous speed value of the falling object, the current control stage is determined as the taking-over stage.

[0031] After the target absolute gravimeter starts testing, it enters the separation stage, the trailer accelerates and falls, and when the acceleration is greater than the gravitational acceleration g, the falling object and the trailer smoothly separate. It can be seen that in this stage, the interval distance between the trailer and the falling object is 0, and the instantaneous acceleration value of the trailer should be less than or equal to the gravitational acceleration value. Therefore, during implementation, the interval distance between the trailer and the falling object can be calculated based on the position information of the trailer and the falling object. Exemplarily, the interval distance between the trailer and the falling object can be obtained by subtracting the longitudinal coordinate value of the trailer from the longitudinal coordinate value of the falling object. Then determine whether the interval distance between the trailer and the falling object is 0, and whether the instantaneous acceleration value of the trailer is less than or equal to the gravitational acceleration value. If so, determine that the current control stage is the separation stage.

[0032] Then it enters the free fall stage. In this stage, it is necessary to ensure that the trailer cannot give external force to the falling object, so that the falling object can fall freely, and the trailer and the falling object should remain relatively still as much as possible. After the falling object falls freely for a certain distance, it enters the receiving stage. In this stage, the trailer slows down and catches the falling object steadily. When the trailer catches the falling object, the instantaneous speed values ​​of the two should be as close as possible to avoid the falling object hitting the trailer. It can be seen that in the free fall stage, in order to keep the trailer and the falling object as relatively still as possible, it is necessary to ensure that the initial speed and acceleration of the trailer and the falling object are the same. Therefore, in the early stage of the free fall stage, the instantaneous acceleration value of the trailer needs to be less than the gravity acceleration value first, so as to be the same as the speed of the falling object. Then, the instantaneous acceleration value of the trailer is made the same as the instantaneous acceleration value of the falling object, that is, both are equal to the gravity acceleration value, so that the relative stillness of the trailer and the falling object can be achieved. When entering the receiving section, in order to make the instantaneous speed values ​​of the tractor and the falling object as close as possible, it is necessary to first slow down the tractor and approach the falling object, and then speed up so that the tractor speed is infinitely close to the speed of the falling object.

[0033] Therefore, during implementation, it can be determined whether the instantaneous acceleration value of the trailer is less than or equal to the gravity acceleration value, and whether the motion duration is less than or equal to the preset time threshold. If so, the current control stage is determined to be the free fall stage. It can be determined whether the motion duration is greater than the time threshold, and whether the instantaneous speed value of the trailer is less than or equal to the instantaneous speed value of the falling object. If so, the current control stage is determined to be the receiving stage.

[0034] Here, if the judgment basis of "the relationship between the duration of the movement and the preset time threshold" is not added, misjudgment may occur, because in the receiving section, the instantaneous acceleration value of the trailer may be less than or equal to the gravity acceleration value (the trailer is decelerating, which will reduce the acceleration value). In the free fall section, the instantaneous speed value of the trailer may be equal to the instantaneous speed value of the falling object (in order to remain relatively still with the falling object, the trailer needs to slow down first to be equal to the speed of the falling object). Therefore, in this embodiment, the time threshold is used to achieve a unique division of stages through the constraints of the time dimension. During implementation, the time threshold can be set according to actual conditions, and this embodiment is not limited to this. For example, the time threshold can be the duration of the test at the end of the free fall section in theory.

[0035] Regarding S102-S103 above: In this disclosure, each control stage corresponds to its own agent, state data, reward function, and replay buffer.

[0036] For the intelligent agent, the intelligent agent in the present disclosure can adopt an actor-critic structure. When implemented, the lightweight of the intelligent agent can be achieved by reducing the number of hidden layers or neurons in the actor network, sharing some underlying parameters of the two critic networks, etc., so as to meet the real-time requirements of the absolute gravimeter motor control.

[0037] For state data and reward function: In the separation stage, the motor speed must be controlled to accelerate the trolley's fall. Studies have shown that if the trolley's acceleration is too large, it will easily cause the falling object to move horizontally or rotate; if the trolley's acceleration is too small, the falling object cannot be separated from the trolley.

[0038] Based on this, when the current control stage is the separation stage, in some embodiments, the state data includes: the position information of the tractor, the position information of the falling object, and the instantaneous acceleration value of the tractor.

[0039] In some implementations, the reward function is:

[0040] Among them, R 分离 is the separation segment reward value, T1 is the total number of time steps of the preset separation segment, and a t is the instantaneous acceleration value of the trailer, a 目标 is the ideal instantaneous acceleration value of the trailer in the separation section, 落体 is the angular velocity of the falling object, v 落体水平 is the instantaneous horizontal velocity value of the falling object.

[0041] Here, a 目标It is preset according to actual needs. In this embodiment, the means for obtaining the independent variable value of the reward function is prior art and will not be elaborated here.

[0042] The separation segment reward function provided in this embodiment can guide the intelligent agent in the separation segment to accurately control the motor speed by introducing penalty terms such as the deviation between the instantaneous acceleration value of the tow truck and the ideal value, the falling body's rotational angular velocity, and the horizontal velocity, so as to ensure that the acceleration of the tow truck meets the standard and the falling body detaches smoothly, avoiding separation failure or motion deviation caused by improper acceleration.

[0043] In the free fall segment, it is necessary to control the motor speed to make the tow truck and the falling body relatively stationary as much as possible. Based on this, when the current control stage is the free fall segment, in some embodiments, the state data includes: the position information of the tow truck, the position information of the falling body, the instantaneous speed value of the tow truck, the instantaneous speed value of the falling body, and the instantaneous acceleration value of the tow truck.

[0044] In some embodiments, the reward function is:

[0045] where R 自由下落 is the free fall segment reward value, T2 - T1 - 1 is the total number of preset time steps in the free fall segment, x 托车 is the position information of the tow truck, x 落体 is the position information of the falling body, is the ideal interval distance preset between the tow truck and the falling body, v t is the instantaneous speed value of the tow truck, v 目标 is the ideal instantaneous speed value of the tow truck preset.

[0046] During implementation, A can be set according to actual needs, and this embodiment does not make any limitations in this regard. For example, it can be set to 3mm. The means for obtaining the independent variable value in the free fall segment reward function is prior art and will not be elaborated here.

[0047] The free fall segment reward function provided in this embodiment can guide the intelligent agent in the free fall segment to accurately adjust the motor speed by introducing penalty terms such as the deviation between the interval distance between the tow truck and the falling body and the ideal interval distance, and the speed deviation between the tow truck and the falling body, so as to keep the tow truck and the falling body relatively stationary, avoiding the two from deviating from the ideal motion state due to excessive position or speed differences, thus meeting the high-precision control requirements.

[0048] In the receiving stage, it is necessary to control the rotational speed of the motor to achieve a slow deceleration of the towing vehicle, so that the towing vehicle "just" catches the falling object. The instantaneous speeds of the towing vehicle and the falling object should be as close as possible to avoid the situation of the falling object hitting the towing vehicle. Based on this, when the current control stage is the receiving stage, in some embodiments, the state data includes: the position information of the towing vehicle, the position information of the falling object, the instantaneous speed value of the towing vehicle, the instantaneous speed value of the falling object, and the instantaneous acceleration value of the towing vehicle.

[0049] In some embodiments, the reward function is:

[0050] where R 承接 is the receiving stage reward value, T3 - T2 - 1 is the total number of preset time steps in the receiving stage, v 托车承接时瞬时 represents the instantaneous speed value when the towing vehicle catches the falling object, and v 落体承接时瞬时 represents the instantaneous speed value when the falling object is caught.

[0051] Here, the means for obtaining the independent variable value in the receiving stage reward function is a prior art and will not be elaborated here.

[0052] The receiving stage reward function provided in this embodiment can guide the intelligent agent in the receiving stage to accurately adjust the rotational speed of the motor by introducing penalty terms such as the instantaneous speed difference between the towing vehicle and the falling object during reception (ensuring that the speeds are close to avoid impact) and the deviation between the towing vehicle speed and the ideal speed (controlling the smoothness of the deceleration process), so that the towing vehicle decelerates at an ideal speed curve and achieves the control goal of "just catching" the falling object.

[0053] It has been found through research that the self - vibration of the drive mechanism of the absolute gravimeter caused by excitations such as mechanical friction will cause the movement trajectory of the towing vehicle to deviate from the ideal state. In addition, as a key transmission component driving the movement of the towing vehicle, the small deformation of the steel belt caused by temperature changes will also cause a systematic deviation in the movement path of the towing vehicle after accumulation. Both of the above - mentioned deviations will cause adjustment deviations in the motor feedback control system, ultimately leading to a decrease in control accuracy. It can be seen that the self - vibration characteristics of the drive mechanism and the temperature stability of the steel belt are two key factors affecting the motor control accuracy.

[0054] Based on this, in some embodiments of the present disclosure, when the current control stage is the separation stage, the status data include: the position information of the trolley, the position information of the falling body, the instantaneous acceleration value of the trolley, and at least one of the following: the peak value of the main frequency of the self-oscillation of the transmission mechanism in the target absolute gravimeter, and the temperature change data of the steel belt in the target absolute gravimeter; when the current control stage is the free fall stage or the receiving stage, the status data include: the position information of the trolley, the position information of the falling body, the instantaneous speed value of the trolley, the instantaneous speed value of the falling body, the instantaneous acceleration value of the trolley, and at least one of the following: the peak value of the main frequency of the self-oscillation of the transmission mechanism in the target absolute gravimeter, and the temperature change data of the steel belt in the target absolute gravimeter.

[0055] In this embodiment, for the transmission mechanism of the target absolute gravimeter, for example, Figure 2 The schematic diagram of the structure of a free-fall device of an absolute gravimeter provided by the present disclosure is shown. Figure 2 As shown, the free fall device is mainly composed of a trolley, a steel belt, a vacuum transmission shaft and a motor. Among them, the steel belt, the vacuum transmission shaft and the motor constitute a transmission mechanism.

[0056] In order to obtain the peak value of the self-oscillation main frequency of the transmission mechanism, during implementation, an acceleration sensor (such as a piezoelectric acceleration sensor) can be installed on the transmission mechanism (such as key parts such as steel belts, trailers, and vacuum transmission shafts); the analog vibration signal output by the acceleration sensor can be converted into a digital signal using a data acquisition card; the collected time domain signal is then pre-processed by filtering and amplifying it through software (such as MATLAB, LabVIEW), and the time domain vibration signal is converted into a frequency domain signal using a fast Fourier transform algorithm to generate a spectrum diagram. In the spectrum diagram, the frequency corresponding to the peak point with the largest amplitude is the peak value of the self-oscillation main frequency.

[0057] In order to obtain the temperature change data of the steel strip, a temperature sensor can be fixed on the steel strip during implementation. The temperature data can be obtained in real time through the temperature sensor, and the temperature change data can be calculated based on each adjacent temperature data.

[0058] This embodiment introduces the natural vibration main frequency peak value of the transmission mechanism and / or the temperature change data of the steel belt into the state data of each control stage, so that each state space covers the mechanical vibration characteristics and thermal characteristics of the gravimeter. This provides accurate environmental feedback for the intelligent agent, enables the intelligent agent to learn the optimal control strategy under the interaction of multiple factors, reduces the movement deviation of the tractor caused by mechanical friction excitation and temperature cumulative deformation, and significantly improves the response accuracy and robustness of the motor feedback control.

[0059] In some embodiments, the above-mentioned reward function further includes: a peak value of a main frequency of natural vibration of a transmission mechanism in the target absolute gravimeter.

[0060] Specifically, if the current control stage is the separation stage, the reward function is:

[0061] Among them, A 自振 It is the peak value of the natural main frequency of the transmission mechanism.

[0062] If the current control stage is the free fall stage, the reward function is:

[0063] If the current control stage is the continuation stage, the reward function is:

[0064] In this embodiment, in addition to paying attention to its own core control objectives, the reward function of each stage also pays attention to the self-oscillation interference of the transmission mechanism, which enables each intelligent agent to generate a more accurate PID parameter adjustment strategy and further improves the reliability of the absolute gravimeter motor control.

[0065] After determining the corresponding agent, state data and reward function according to the current control stage, the agent can be used to process the corresponding state data to generate action data. Here, the action data, that is, the PID controller parameter adjustment amount, is the weight of the actor network in the agent. The action data can be used to adjust the motor control parameters. Then, the reward value is calculated using the corresponding reward function. Then the state data of the next time step is obtained. The reward value, action data, state data of the current time step and state data of the next time step constitute an experience data, which is stored in the playback buffer corresponding to the current control stage. By analogy, at each time step, the above steps S101-S103 are repeated to collect experience data until the current control stage is the takeover stage, and the speed of the trailer and the falling object is the same and the interval distance is 0, that is, the takeover stage ends. At this point, the collection of experience data for one round is completed.

[0066] Regarding S104-S106 above: Here, the training termination condition may be, for example, the reward function converges, or the number of training rounds reaches a preset number of training rounds. The training termination condition may be set according to actual needs, and the embodiment of the present disclosure does not limit this.

[0067] It can be determined whether each intelligent agent meets the preset training termination conditions. If so, it indicates that the performance of each intelligent agent meets the requirements. The state data of the target absolute gravimeter can be input into the corresponding intelligent agent to generate a motor controller parameter adjustment set. The motor controller parameter adjustment set can be used to adjust the motor control parameters of the target absolute gravimeter.

[0068] If it is determined that there is an untrained target agent and the number of experience data in the target replay buffer corresponding to the target agent is greater than a preset threshold, then sample experiences are extracted from the target replay buffer. During implementation, random sampling can be performed in the target replay buffer to obtain multiple sample experiences, and the target agent is trained using the extracted sample experiences to obtain a new target agent. Here, the preset threshold can be set according to actual needs. Then, the target absolute gravimeter is reset (the reset operation includes: resetting the trolley and the falling body to the initial position, clearing the control parameters and status data, etc.), so that the system restarts from the unified initial conditions to ensure the consistency of training. Then, return to step S101 and iterate continuously until each agent meets the preset training termination conditions.

[0069] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0070] It should be noted that in practical applications, all the above possible implementation manners can be combined in any combination to form possible embodiments of the present disclosure, which will not be elaborated herein one by one.

[0071] Based on the same concept, the embodiment of the present disclosure also provides an absolute gravimeter motor control parameter adjustment device, and the absolute gravimeter motor control parameter adjustment device corresponds one-to-one with the absolute gravimeter motor control parameter adjustment method in the above embodiment. Figure 3 The structural schematic diagram of the absolute gravimeter motor control parameter adjustment device provided by the embodiment of the present disclosure is shown. Refer to Figure 3 As shown, the absolute gravimeter motor control parameter adjustment device 300 provided by the embodiment of the present disclosure includes: A stage determination module 301, configured to determine the current control stage based on the obtained position information, motion information, and motion duration of the trolley and the falling body in the target absolute gravimeter; The experience collection module 302 is used to obtain the corresponding intelligent agent and state data according to the current control stage, and use the intelligent agent to process the state data to generate action data; the action data is a set of motor controller parameter adjustment amounts; the motor control parameters of the target absolute gravimeter are adjusted using the action data; the reward value is calculated using the reward function corresponding to the current control stage; the reward value, the action data, the state data and the acquired state data of the next time step are stored as an experience data in the playback buffer corresponding to the current control stage; return to the step of determining the current control stage based on the acquired position information, motion information, and motion duration of the trailer and the falling object in the target absolute gravimeter, until the current control stage is a takeover stage, and the speed of the trailer and the falling object is the same and the interval distance is 0; The training module 303 is used to generate a judgment result of whether the training of each of the intelligent agents has been completed according to a preset training termination condition; if the judgment result is that there is a target intelligent agent that has not been trained, and the number of experience data in the target playback buffer corresponding to the target intelligent agent is greater than a preset threshold, then sample experience is extracted from the target playback buffer, and the target intelligent agent is trained with the sample experience to obtain a new target intelligent agent; the target absolute gravimeter is reset, and the step of determining the current control stage is returned based on the position information, motion information, and motion duration of the trailer and the falling object obtained in the target absolute gravimeter; The adjustment module 304 is used to process the state data of the target absolute gravimeter using each of the intelligent agents to generate action data to adjust the motor control parameters of the target absolute gravimeter if the judgment result is that the training of each of the intelligent agents is completed.

[0072] In some embodiments, the motion information includes: an instantaneous acceleration value of the trailer, an instantaneous speed value of the trailer and an instantaneous speed value of the falling object; in the above-mentioned device, the stage determination module is specifically used to: calculate the interval distance between the trailer and the falling object according to the position information of the trailer and the falling object; if the interval distance is 0, and the instantaneous acceleration value of the trailer is less than or equal to the gravity acceleration value, then determine that the current control stage is a separation segment; if the motion duration is less than or equal to a preset time threshold, and the instantaneous acceleration value of the trailer is less than or equal to the gravity acceleration value, then determine that the current control stage is a free fall segment; if the motion duration is greater than the time threshold, and the instantaneous speed value of the trailer is less than or equal to the instantaneous speed value of the falling object, then determine the current control stage as a taking-over segment.

[0073] In some embodiments, in the above-mentioned device, if the current control stage is the separation stage, the state data includes: the position information of the tractor, the position information of the falling object, and the instantaneous acceleration value of the tractor.

[0074] In some embodiments, in the above-mentioned device, if the current control stage is a free fall stage or a receiving stage, the status data includes: the position information of the trolley, the position information of the falling object, the instantaneous speed value of the trolley, the instantaneous speed value of the falling object, and the instantaneous acceleration value of the trolley.

[0075] In some embodiments, in the above device, if the current control stage is a separation stage, the reward function is:

[0076] Among them, R 分离 is the separation segment reward value, T1 is the total number of time steps of the preset separation segment, and a t is the instantaneous acceleration value of the trailer, a 目标 is the ideal instantaneous acceleration value of the trailer in the separation section, 落体 is the angular velocity of the falling object, v 落体水平 is the instantaneous horizontal velocity value of the falling object.

[0077] In some embodiments, in the above device, if the current control stage is a free fall stage, the reward function is:

[0078] Among them, R 自由下落 is the free fall segment reward value, T2-T1-1 is the total number of time steps of the preset free fall segment, x 托车 is the position information of the trailer, x 落体 is the position information of the falling object, is the preset ideal distance between the trailer and the falling object, v t is the instantaneous speed of the trailer, v 目标 It is the preset ideal instantaneous speed value of the trailer.

[0079] In some embodiments, in the above device, if the current control stage is a takeover stage, the reward function is:

[0080] Among them, R 承接 is the reward value of the succession segment, T3-T2-1 is the total number of time steps of the preset succession segment, v 托车承接时瞬时 represents the instantaneous speed value of the trailer when it receives the falling object, v 落体承接时瞬时Represents the instantaneous velocity value when the falling object is received, v t Is the instantaneous velocity value of the carrier vehicle, v 目标 Is the preset ideal instantaneous velocity value of the carrier vehicle.

[0081] In some embodiments, in the above device, if the current control stage is the separation stage, the state data includes: the position information of the carrier vehicle, the position information of the falling object, the instantaneous acceleration value of the carrier vehicle, and at least one of the following: the peak value of the natural vibration frequency of the transmission mechanism in the target absolute gravimeter, the temperature change data of the steel strip in the target absolute gravimeter; if the current control stage is the free fall stage or the receiving stage, the state data includes: the position information of the carrier vehicle, the position information of the falling object, the instantaneous velocity value of the carrier vehicle, the instantaneous velocity value of the falling object, the instantaneous acceleration value of the carrier vehicle, and at least one of the following: the peak value of the natural vibration frequency of the transmission mechanism in the target absolute gravimeter, the temperature change data of the steel strip in the target absolute gravimeter.

[0082] In some embodiments, in the above device, the reward function further includes the peak value of the natural vibration frequency of the transmission mechanism in the target absolute gravimeter.

[0083] The present invention provides an absolute gravimeter motor control parameter adjustment device, which introduces a reinforcement learning framework. For each stage of the motor controlling the carrier vehicle, first, using the agents and reward functions designed for each stage respectively, a large amount of experience data containing success / failure is accumulated. Using these experience data, the agents in each control stage can be trained to gradually optimize the controller parameter adjustment strategy of the agents. Finally, the trained mature agents can dynamically generate accurate PID parameter adjustment amounts according to the real-time state data of the target absolute gravimeter. Using these PID parameters, the accurate adjustment of the absolute gravimeter motor control parameters can be realized. Through this interactive learning mechanism, the adaptive adjustment of the absolute gravimeter motor control parameters can be achieved, breaking through the limitations of traditional empirical parameter adjustment, significantly improving the accuracy and stability of the movement of the carrier vehicle and the falling object, and thus improving the accuracy of the measured value of the gravitational acceleration.

[0084] For the specific limitations of the absolute gravimeter motor control parameter adjustment device, reference can be made to the limitations of the absolute gravimeter motor control parameter adjustment method in the above text, which will not be elaborated here. Each module in the above absolute gravimeter motor control parameter adjustment device can be implemented in whole or in part by software, hardware and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.

[0085] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of another identical element in the process, method, commodity or device comprising said element.

[0086] Those skilled in the art should understand that the embodiments of the present disclosure may be provided as a method, a system or a computer program product. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0087] The above are only the embodiments of the present disclosure and are not intended to limit the present disclosure. For those skilled in the art, various changes and modifications can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the scope of the claims of the present disclosure.

Claims

1. A method for adjusting control parameters of an absolute gravimeter motor, characterized in that: The method comprises: Determine the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object in the target absolute gravimeter; According to the current control stage, the corresponding intelligent agent and state data are obtained, and the state data is processed by the intelligent agent to generate action data; the action data is a set of motor controller parameter adjustment values; Using the action data, adjusting the motor control parameters of the target absolute gravimeter; using the reward function corresponding to the current control stage, calculating the reward value; using the reward value, the action data, the state data and the acquired state data of the next time step as an experience data, storing them in the playback buffer corresponding to the current control stage; returning to the step of determining the current control stage based on the acquired position information, motion information, and motion duration of the trailer and the falling object in the target absolute gravimeter, until the current control stage is the takeover stage, and the speed of the trailer is the same as that of the falling object, and the interval distance is 0; According to the preset training termination condition, generating a judgment result of whether the training of each of the intelligent agents has been completed; If the judgment result is that the training of each of the intelligent agents is completed, then using each of the intelligent agents to process the state data of the target absolute gravimeter to generate action data to adjust the motor control parameters of the target absolute gravimeter; If the judgment result is that there is a target intelligent agent that has not been trained, and the number of experience data in the target playback buffer corresponding to the target intelligent agent is greater than a preset threshold, then sample experience is extracted from the target playback buffer, and the target intelligent agent is trained using the sample experience to obtain a new target intelligent agent; the target absolute gravimeter is reset, and the step of determining the current control stage is returned based on the position information, motion information, and motion duration of the trailer and the falling object obtained in the target absolute gravimeter.

2. The method according to claim 1, characterized in that The motion information includes: the instantaneous acceleration value of the trailer, the instantaneous speed value of the trailer and the instantaneous speed value of the falling object; the current control stage is determined based on the position information, motion information and motion duration of the trailer and the falling object in the target absolute gravimeter obtained, including: Calculating the distance between the trailer and the falling object according to the position information of the trailer and the falling object; If the separation distance is 0, and the instantaneous acceleration value of the trailer is less than or equal to the gravity acceleration value, then the current control stage is determined to be the separation stage; If the movement duration is less than or equal to a preset time threshold, and the instantaneous acceleration value of the trailer is less than or equal to the gravity acceleration value, then the current control stage is determined to be a free fall stage; If the movement duration is greater than the time threshold, and the instantaneous speed value of the trailer is less than or equal to the instantaneous speed value of the falling object, the current control stage is determined as the taking-over stage.

3. The method according to claim 1, characterized in that: If the current control stage is the separation stage, the state data includes: the position information of the trailer, the position information of the falling object, and the instantaneous acceleration value of the trailer.

4. The method according to claim 1, characterized in that: If the current control stage is a free fall stage or a receiving stage, the status data includes: the position information of the trailer, the position information of the falling object, the instantaneous speed value of the trailer, the instantaneous speed value of the falling object, and the instantaneous acceleration value of the trailer.

5. The method according to claim 1, characterized in that: If the current control stage is a separation stage, the reward function is: Among them, R 分离 is the separation segment reward value, T1 is the total number of time steps of the preset separation segment, and a t is the instantaneous acceleration value of the trailer, a 目标 is the ideal instantaneous acceleration value of the trailer in the separation section, 落体 is the angular velocity of the falling object, v 落体水平 is the instantaneous horizontal velocity value of the falling object.

6. The method according to claim 1, characterized in that If the current control stage is the free fall stage, the reward function is: Among them, R 自由下落 is the free fall segment reward value, T2-T1-1 is the total number of time steps of the preset free fall segment, x 托车 is the position information of the trailer, x 落体 is the position information of the falling object, is the preset ideal distance between the trailer and the falling object, v t is the instantaneous speed of the trailer, v 目标 It is the preset ideal instantaneous speed value of the trailer.

7. The method according to claim 1, characterized in that If the current control stage is a succession stage, the reward function is: Among them, R 承接 is the reward value of the succession segment, T3-T2-1 is the total number of time steps of the preset succession segment, v 托车承接时瞬时 represents the instantaneous speed value of the trailer when it receives the falling object, v 落体承接时瞬时 represents the instantaneous velocity value of the falling object when it is caught, v t is the instantaneous speed of the trailer, v 目标 It is the preset ideal instantaneous speed value of the trailer.

8. The method according to claim 3 or 4, characterized in that: The state data also includes at least one of the following: the peak value of the main frequency of the self-oscillation of the transmission mechanism in the target absolute gravimeter, and the temperature change data of the steel belt in the target absolute gravimeter.

9. The method according to any one of claims 5 to 7, characterized in that: The reward function also includes the peak value of the self-oscillation main frequency of the transmission mechanism in the target absolute gravimeter.

10. An absolute gravimeter motor control parameter adjustment device, characterized in that: The device comprises: A phase determination module, used to determine the current control phase based on the position information, motion information, and motion duration of the trailer and the falling object acquired in the target absolute gravimeter; An experience collection module is used to obtain the corresponding intelligent agent and state data according to the current control stage, and use the intelligent agent to process the state data to generate action data; the action data is a set of motor controller parameter adjustment amounts; the motor control parameters of the target absolute gravimeter are adjusted using the action data; the reward value is calculated using the reward function corresponding to the current control stage; the reward value, the action data, the state data and the acquired state data of the next time step are stored as an experience data in the playback buffer corresponding to the current control stage; return to the step of determining the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object in the acquired target absolute gravimeter, until the current control stage is a takeover stage, and the speed of the trailer and the falling object is the same and the interval distance is 0; The training module is used to generate a judgment result of whether the training of each of the intelligent agents has been completed according to a preset training termination condition; if the judgment result is that there is a target intelligent agent that has not been trained, and the number of experience data in the target playback buffer corresponding to the target intelligent agent is greater than a preset threshold, then sample experience is extracted from the target playback buffer, and the target intelligent agent is trained with the sample experience to obtain a new target intelligent agent; the target absolute gravimeter is reset, and the step of determining the current control stage is returned based on the position information, motion information, and motion duration of the trailer and the falling object obtained in the target absolute gravimeter; The adjustment module is used to use each of the intelligent agents to process the state data of the target absolute gravimeter and generate action data to adjust the motor control parameters of the target absolute gravimeter if the judgment result is that the training of each of the intelligent agents is completed.

Citation Information

Patent Citations

  • Method for platform gravimeter to reduce influence of horizontal acceleration

    CN111123381A

  • Underwater vehicle attitude control system and method based on reinforcement learning compensator

    CN116449856A

  • Absolute gravimeter falling body and trailer separation distance measurement system and measurement method

    CN118795562A

  • Balanced falling mechanism and gravimeter

    WO2021244426A1

Cited By

  • Motor motion curve optimization method and device based on iterative learning

    CN122386728A

  • Motor motion curve optimization method and device based on iterative learning

    CN122386728B