Method and device for adjusting motor control parameters of an absolute gravimeter
By introducing an interactive learning mechanism of reinforcement learning framework and agent reward function, the absolute gravity meter motor control parameters are optimized, and the problem of insufficient adaptive adjustment in the existing technology is solved, the accuracy and stability of motorcycle and fall motion are improved, and the measurement accuracy of gravity acceleration values is improved.
Patent Information
- Application Number
- CN202510644794.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The existing absolute gravity meter motor control parameter adjustment method cannot achieve adaptive adjustment, resulting in unstable movement of the motorcycle and inaccurate free fall, which in turn affects the measurement accuracy of gravity acceleration value.
Introduce a reinforcement learning framework, use agents and reward functions to accumulate empirical data, optimize motor control parameters, and realize adaptive adjustment through interactive learning mechanisms.
It significantly improves the accuracy and stability of the motorcycle and falling movement, thereby improving the measurement accuracy of gravity acceleration values.
Smart Images

Figure CN120185480B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of motor control technology, and in particular to a method and device for adjusting control parameters of an absolute gravimeter motor. Background Art
[0002] An absolute gravimeter is a precision measuring instrument used to directly measure the acceleration of gravity at the Earth's surface. During gravity acceleration testing using an absolute gravimeter, motor control of the trolley is divided into three phases: separation, free fall, and catch-up. During the separation phase, the trolley must smoothly separate from the falling object; during the free fall phase, the trolley and the falling object must remain as relatively stationary as possible; and during the catch-up phase, the trolley must precisely catch the falling object. The high-precision operation of the absolute gravimeter necessitates a rigorous control strategy for its motor.
[0003] Typically, motion control of an absolute gravimeter motor is achieved using a classic PID (Proportional-Integral-Derivative) controller. Currently, existing techniques for adjusting the PID controller parameters rely primarily on empirical experience, lacking adaptive adjustment. Consequently, existing methods for adjusting the control parameters of absolute gravimeter motors fail to ensure smooth trolley motion or accurate free-fall measurements, resulting in poor gravitational acceleration measurement accuracy. Therefore, a solution to this technical problem is urgently needed. Summary of the Invention
[0004] In view of the above situation, the embodiments of the present disclosure provide a method and device for adjusting the control parameters of an absolute gravimeter motor, aiming to solve the above problem or at least partially solve the above problem.
[0005] In a first aspect, an embodiment of the present disclosure provides a method for adjusting control parameters of an absolute gravimeter motor, the method comprising:
[0006] Determine the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object obtained from the target absolute gravimeter;
[0007] According to the current control stage, the corresponding intelligent agent and state data are obtained, and the state data is processed by the intelligent agent to generate action data; the action data is a set of motor controller parameter adjustment values;
[0008] Using the action data, adjusting motor control parameters of the target absolute gravimeter; using a reward function corresponding to the current control stage, calculating a reward value; storing the reward value, the action data, the state data, and the acquired state data for the next time step as a piece of experience data in a playback buffer corresponding to the current control stage; returning to the step of determining the current control stage based on the acquired position information, motion information, and motion duration of the tractor and the falling object in the target absolute gravimeter, until the current control stage is a takeover stage, and the tractor and the falling object have the same speed and the separation distance is 0;
[0009] According to the preset training termination conditions, generating a judgment result on whether the training of each of the intelligent agents is completed;
[0010] If the judgment result is that the training of each of the intelligent agents is completed, then using each of the intelligent agents to process the state data of the target absolute gravimeter to generate action data to adjust the motor control parameters of the target absolute gravimeter;
[0011] If the judgment result is that there is a target intelligent agent that has not been trained, and the number of experience data in the target replay buffer corresponding to the target intelligent agent is greater than a preset threshold, then sample experience is extracted from the target replay buffer, and the target intelligent agent is trained using the sample experience to obtain a new target intelligent agent; the target absolute gravimeter is reset, and the process returns to the step of determining the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object obtained in the target absolute gravimeter.
[0012] In a second aspect, an embodiment of the present disclosure further provides a device for adjusting motor control parameters of an absolute gravimeter, the device comprising:
[0013] A phase determination module is used to determine the current control phase based on the position information, motion information, and motion duration of the trailer and the falling object obtained from the target absolute gravimeter;
[0014] An experience collection module is configured to obtain, according to the current control stage, a corresponding intelligent agent and state data, and process the state data using the intelligent agent to generate action data; the action data is a set of motor controller parameter adjustment values; the motor control parameters of the target absolute gravimeter are adjusted using the action data; a reward value is calculated using a reward function corresponding to the current control stage; the reward value, the action data, the state data, and the acquired state data for the next time step are stored as a piece of experience data in a playback buffer corresponding to the current control stage; and the step of determining the current control stage based on the acquired position information, motion information, and motion duration of the trailer and the falling object in the target absolute gravimeter is returned until the current control stage is a takeover stage, and the speeds of the trailer and the falling object are the same and the separation distance is zero.
[0015] A training module is configured to generate a determination result as to whether all of the agents have completed training based on a preset training termination condition; if the determination result indicates that a target agent has not been trained and the number of experience data items in the target replay buffer corresponding to the target agent is greater than a preset threshold, extract sample experience from the target replay buffer, use the sample experience to train the target agent, and obtain a new target agent; reset the target absolute gravimeter, and return to the step of determining the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object obtained from the target absolute gravimeter;
[0016] The adjustment module is used to use each of the intelligent agents to process the state data of the target absolute gravimeter and generate action data to adjust the motor control parameters of the target absolute gravimeter if the judgment result is that the training of each of the intelligent agents is completed.
[0017] By leveraging the aforementioned technical solution, the method and apparatus for adjusting the motor control parameters of an absolute gravimeter provided in the disclosed embodiments introduce a reinforcement learning framework. For each stage of motor control of a trailer, a large amount of empirical data containing successes and failures is first accumulated using an agent and reward function designed for each stage. This empirical data can be used to train the agent for each control stage, gradually optimizing the agent's controller parameter adjustment strategy. Ultimately, a maturely trained agent can dynamically generate precise PID parameter adjustments based on the real-time status data of the target absolute gravimeter. These PID parameters can be used to precisely adjust the gravimeter motor control parameters. This interactive learning mechanism enables adaptive adjustment of the absolute gravimeter motor control parameters, overcoming the limitations of traditional empirical parameter adjustment methods and significantly improving the accuracy and smoothness of trailer and falling object motion, thereby increasing the accuracy of gravitational acceleration measurement.
[0018] The above description is only an overview of the technical solution of the present disclosure. In order to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present disclosure more obvious and understandable, the specific embodiments of the present disclosure are specifically exemplified below. Description of the Drawings
[0019] The drawings described herein are used to provide a further understanding of the present disclosure and form a part of the present disclosure. The schematic embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:
[0020] Figure 1 A flowchart showing the method for adjusting the motor control parameters of the absolute gravimeter provided by the embodiment of the present disclosure is shown;
[0021] Figure 2 A schematic structural diagram of the free-fall device of the absolute gravimeter provided by the embodiment of the present disclosure is shown;
[0022] Figure 3 A schematic structural diagram of the device for adjusting the motor control parameters of the absolute gravimeter provided by the embodiment of the present disclosure is shown. Detailed Embodiments
[0023] To make the purpose, technical solution and advantages of the present disclosure clearer, the technical solution of the present disclosure will be clearly and completely described below in conjunction with the specific embodiments of the present disclosure and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0024] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0025] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such use can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the term "including" and its variants should be interpreted as open-ended terms meaning "including but not limited to".
[0026] As introduced above, the parameters of the PID controller for the absolute gravimeter motor mainly rely on empirical settings and cannot achieve adaptive adjustment. Therefore, the existing methods for adjusting the control parameters of the absolute gravimeter motor cannot ensure the smoothness of the trolley movement, the accuracy of the free fall of the falling body, etc., resulting in poor measurement accuracy of the gravitational acceleration value. Based on this, the present invention proposes a method and device for adjusting the control parameters of the absolute gravimeter motor, and the following will describe the present disclosure in detail through specific embodiments.
[0027] Before introducing the specific embodiments, the professional terms involved in the embodiments of the present disclosure will be explained first:
[0028] 1) Environment: In reinforcement learning, the environment is the external world where the agent is located. The agent receives the state information of the environment at each time step, selects an appropriate action, and then the environment gives feedback on this action, including rewards and new states.
[0029] 2) State: The current description of the environment.
[0030] 3) Action: An action is one of the decisions that the agent can take in a given state, used to maximize the reward in a certain state. In the present disclosure, the action is the control signal output by the agent, used to control the motor of the absolute gravimeter.
[0031] 4) Reward: A reward is the feedback given by the environment after the agent executes a certain action, usually used to measure the quality of this action.
[0032] 5) Time step: A basic unit for the interaction between the agent and the environment. Every time a time step passes, the agent will take an action according to the current state, and the environment will give a new state and a reward.
[0033] For ease of understanding of this embodiment, first, a method for adjusting the control parameters of the absolute gravimeter motor disclosed in the embodiments of the present disclosure will be introduced in detail. The execution subject of the method for adjusting the control parameters of the absolute gravimeter motor provided in the embodiments of the present disclosure is generally a computer device with certain computing capabilities. Such a computer device includes, for example: a terminal device or a server or other processing devices. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, etc. In some possible implementation manners, the method for adjusting the control parameters of the absolute gravimeter motor can be implemented by a processor calling computer-readable instructions stored in a memory.
[0034] Figure 1 The flowchart of the method for adjusting the control parameters of the absolute gravimeter motor provided in the embodiments of the present disclosure is shown. From Figure 1 it can be seen that the embodiments of the present disclosure at least include steps S101 - S106:
[0035] S101: Determine the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object obtained from the target absolute gravimeter.
[0036] S102: According to the current control stage, a corresponding intelligent agent and state data are obtained, and the state data is processed by the intelligent agent to generate action data; the action data is a set of motor controller parameter adjustment values.
[0037] S103: Using the action data, adjust the motor control parameters of the target absolute gravimeter; using the reward function corresponding to the current control stage, calculate the reward value; store the reward value, the action data, the state data and the acquired state data of the next time step as an experience data in the playback buffer corresponding to the current control stage; return to the steps of determining the current control stage based on the position information, motion information, and motion duration of the tractor and the falling object acquired in the target absolute gravimeter, until the current control stage is the takeover stage, and the speed of the tractor and the falling object is the same and the interval distance is 0.
[0038] S104: Based on a preset training termination condition, a judgment result is generated as to whether the training of each of the intelligent agents is completed.
[0039] S105: If the judgment result is that the training of each of the intelligent agents is completed, each of the intelligent agents is used to process the state data of the target absolute gravimeter to generate action data to adjust the motor control parameters of the target absolute gravimeter.
[0040] S106: If the judgment result is that there is a target intelligent agent that has not been trained, and the number of experience data in the target replay buffer corresponding to the target intelligent agent is greater than the preset threshold, then sample experience is extracted from the target replay buffer, and the target intelligent agent is trained using the sample experience to obtain a new target intelligent agent; the target absolute gravimeter is reset, and the process returns to the step of determining the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object obtained in the target absolute gravimeter.
[0041] The absolute gravimeter motor control parameter adjustment method provided in this embodiment introduces a reinforcement learning framework. For each stage of motor control of a trailer, a large amount of empirical data containing successes and failures is first accumulated using an intelligent agent and reward function designed for each stage. This empirical data can be used to train the intelligent agent for each control stage, gradually optimizing the intelligent agent's controller parameter adjustment strategy. Ultimately, a maturely trained intelligent agent can dynamically generate precise PID parameter adjustments based on the real-time status data of the target absolute gravimeter. These PID parameters can be used to precisely adjust the gravimeter motor control parameters. This interactive learning mechanism enables adaptive adjustment of the absolute gravimeter motor control parameters, breaking through the limitations of traditional empirical parameter adjustment methods and significantly improving the accuracy and smoothness of trailer and falling object motion, thereby increasing the accuracy of gravitational acceleration measurement.
[0042] The above S101-S106 are described in detail below.
[0043] Regarding S101 above:
[0044] In this embodiment, training samples, i.e., empirical data, are first collected. In specific implementation, in this step, the position information of the trailer, the motion information of the trailer, the position information of the falling object, the motion information of the falling object, and the duration of the motion in the target absolute gravimeter are first obtained. Then, based on the above information, the current control stage is determined. Here, the position information is used to describe the spatial position of the trailer and the falling object, which can be the displacement, vertical coordinate value, etc. of the trailer and the falling object in the target absolute gravimeter. The motion information includes instantaneous velocity value, instantaneous acceleration value, etc. The duration of the motion is the time that the target absolute gravimeter lasts from the start of the test. As described in the background technology, the current control stage includes a separation section, a free fall section, and a takeover section.
[0045] This embodiment does not limit the method for obtaining the position information, motion information, and motion duration of the trailer and falling object. During implementation, this method can be selected based on actual circumstances. For example, the execution entity of this embodiment can connect to the internal sensors of the target absolute gravimeter (such as a laser interferometer for displacement or an accelerometer for acceleration) via a USB or Ethernet interface to read the position and motion information of the trailer and falling object in real time. The motion duration can be obtained using a timer module or the gravimeter's built-in clock.
[0046] Then, based on the position information, motion information, and motion duration of the trailer and the falling object, the current control stage is determined. Specifically, in some embodiments, the motion information includes: the trailer's instantaneous acceleration value, the trailer's instantaneous velocity value, and the falling object's instantaneous velocity value; the determination of the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object obtained from the target absolute gravimeter includes:
[0047] Calculating the distance between the trailer and the falling object based on the position information of the trailer and the falling object;
[0048] If the separation distance is 0, and the instantaneous acceleration value of the trailer is less than or equal to the gravity acceleration value, then the current control stage is determined to be the separation stage;
[0049] If the motion duration is less than or equal to a preset time threshold, and the instantaneous acceleration value of the trailer is less than or equal to the gravity acceleration value, then the current control stage is determined to be a free fall stage;
[0050] If the movement duration is greater than the time threshold, and the instantaneous speed value of the trailer is less than or equal to the instantaneous speed value of the falling object, the current control stage is determined to be the taking-over stage.
[0051] After the target absolute gravimeter starts testing, it enters the separation stage, and the trailer accelerates to fall. When the acceleration is greater than the acceleration of gravity g, the falling object and the trailer are smoothly separated. It can be seen that in this stage, the distance between the trailer and the falling object is 0, and the instantaneous acceleration value of the trailer should be less than or equal to the acceleration of gravity. Therefore, during implementation, the distance between the trailer and the falling object can be calculated based on the position information of the trailer and the falling object. For example, the vertical coordinate value of the falling object can be subtracted from the vertical coordinate value of the trailer to obtain the distance between the trailer and the falling object. Then determine whether the distance between the trailer and the falling object is 0, and whether the instantaneous acceleration value of the trailer is less than or equal to the acceleration of gravity. If so, it is determined that the current control stage is the separation stage.
[0052] Then it enters the free fall stage. In this stage, it is necessary to ensure that the trailer cannot exert any external force on the falling object, so that the falling object can fall freely, and the trailer and the falling object should remain relatively still as much as possible. After the falling object has fallen freely for a certain distance, it enters the receiving stage. In this stage, the trailer slows down and steadily catches the falling object. When the trailer catches the falling object, the instantaneous speed values of the two should be as close as possible to avoid the falling object hitting the trailer. It can be seen from this that in the free fall stage, in order to keep the trailer and the falling object as relatively still as possible, it is necessary to ensure that the initial speed and acceleration of the trailer and the falling object are the same. Therefore, in the early stage of the free fall stage, the instantaneous acceleration value of the trailer needs to be less than the acceleration due to gravity first, so as to be the same as the speed of the falling object. Then, the instantaneous acceleration value of the trailer is made the same as the instantaneous acceleration value of the falling object, that is, both are equal to the acceleration due to gravity. In this way, the relative stillness of the trailer and the falling object can be achieved. When entering the receiving section, in order to make the instantaneous speed values of the tractor and the falling object as close as possible, it is necessary to first slow down the tractor and approach the falling object, and then speed up to make the tractor speed as close as possible to the falling object speed.
[0053] Therefore, during implementation, it is possible to determine whether the trailer's instantaneous acceleration is less than or equal to the acceleration due to gravity, and whether the duration of the motion is less than or equal to a preset time threshold. If so, the current control phase is determined to be the free-fall phase. It is also possible to determine whether the duration of the motion is greater than the time threshold, and whether the trailer's instantaneous speed is less than or equal to the instantaneous speed of the falling object. If so, the current control phase is determined to be the takeover phase.
[0054] Here, if the judgment basis of "the relationship between the duration of the movement and the preset time threshold" is not added, misjudgment may occur, because in the connection section, the instantaneous acceleration value of the trailer may be less than or equal to the acceleration value of gravity (the trailer is decelerating, which will reduce the acceleration value). In the free fall section, the instantaneous speed value of the trailer may be equal to the instantaneous speed value of the falling object (in order to remain relatively still with the falling object, the trailer needs to slow down first to be equal to the speed of the falling object). Therefore, in this embodiment, the time threshold is used to achieve a unique division of the stages through the constraint of the time dimension. During implementation, the time threshold can be set according to actual conditions, and this embodiment does not limit this. For example, the time threshold can be the duration of the test at the end of the free fall section in theory.
[0055] Regarding S102-S103 above:
[0056] In this disclosure, each control stage corresponds to its own agent, state data, reward function, and replay buffer.
[0057] The intelligent agent in this disclosure can adopt an actor-critic architecture. This can be achieved by reducing the number of hidden layers or neurons in the actor network and sharing some underlying parameters between the two critic networks, thereby achieving a lightweight intelligent agent that meets the real-time requirements of absolute gravimeter motor control.
[0058] For state data and reward function:
[0059] During the separation phase, the motor speed must be controlled to accelerate the trolley's descent. Research has shown that excessive acceleration can easily cause the falling object to move horizontally or rotate, while excessive slow acceleration can prevent the falling object from separating from the trolley.
[0060] Based on this, when the current control stage is the separation stage, in some embodiments, the state data includes: the position information of the trailer, the position information of the falling object, and the instantaneous acceleration value of the trailer.
[0061] In some implementations, the reward function is:
[0062]
[0063] Among them, R 分离is the separation segment reward value, T1 is the total number of preset time steps in the separation segment, a t is the instantaneous acceleration value of the tow truck, a 目标 is the ideal instantaneous acceleration value of the tow truck in the separation segment, 落体 is the rotational angular velocity of the falling object, v 落体水平 is the instantaneous horizontal velocity value of the falling object.
[0064] Here, a 目标 is preset according to actual needs. In this embodiment, the means for obtaining the independent variable values of the reward function is the prior art and will not be elaborated here.
[0065] The separation segment reward function provided in this embodiment can guide the intelligent body in the separation segment to accurately control the motor speed by introducing penalty terms such as the deviation between the instantaneous acceleration value of the tow truck and the ideal value, the rotational angular velocity of the falling object, and the horizontal velocity, so as to ensure that the acceleration of the tow truck meets the standard and the falling object detaches smoothly, and avoid separation failure or motion deviation caused by improper acceleration.
[0066] In the free fall segment, it is necessary to control the speed of the motor to make the tow truck and the falling object as relatively stationary as possible. Based on this, when the current control stage is the free fall segment, in some embodiments, the state data includes: the position information of the tow truck, the position information of the falling object, the instantaneous velocity value of the tow truck, the instantaneous velocity value of the falling object, and the instantaneous acceleration value of the tow truck.
[0067] In some embodiments, the reward function is:
[0068]
[0069] wherein, R 自由下落 is the free fall segment reward value, T2 - T1 - 1 is the total number of preset time steps in the free fall segment, x 托车 is the position information of the tow truck, x 落体 is the position information of the falling object, is the ideal interval distance preset between the tow truck and the falling object, v t is the instantaneous velocity value of the tow truck, v 目标 is the ideal instantaneous velocity value preset for the tow truck.
[0070] During implementation, A can be set according to actual needs, and this embodiment does not limit this. For example, it can be set to 3mm. The means for obtaining the independent variable values in the free fall segment reward function is the prior art and will not be elaborated here.
[0071] The free-fall segment reward function provided in this embodiment can guide the free-fall segment intelligent agent to accurately adjust the motor speed by introducing penalty items such as the deviation between the distance between the trolley and the falling object and the ideal distance, and the speed deviation between the trolley and the falling object. This allows the trolley and the falling object to remain relatively stationary, avoiding the two from deviating from the ideal motion state due to excessive differences in position or speed, thereby meeting high-precision control requirements.
[0072] During the catch phase, the motor speed must be controlled to slowly decelerate the tractor so that it "just" catches the falling object. The instantaneous speeds of the tractor and the falling object are as close as possible to avoid the falling object striking the tractor. Therefore, when the current control phase is the catch phase, in some embodiments, the state data includes: the tractor's position information, the falling object's position information, the tractor's instantaneous speed, the falling object's instantaneous speed, and the tractor's instantaneous acceleration.
[0073] In some embodiments, the reward function is:
[0074]
[0075] Among them, R 承接 is the reward value of the succession segment, T3-T2-1 is the total number of time steps of the preset succession segment, v 托车承接时瞬时 represents the instantaneous speed of the vehicle when it receives the falling object, v 落体承接时瞬时 It represents the instantaneous velocity value when the falling object is caught.
[0076] Here, the means for obtaining the independent variable value in the succession segment reward function is the existing technology and will not be described in detail here.
[0077] The reward function for the receiving segment provided in this embodiment introduces penalty terms such as the instantaneous speed difference between the trolley and the falling object (ensuring that the speeds are close to avoid collision) and the deviation between the trolley speed and the ideal speed (controlling the smoothness of the deceleration process). This can guide the intelligent agent in the receiving segment to accurately adjust the motor speed so that the trolley decelerates along the ideal speed curve, achieving the control goal of "just catching" the falling object.
[0078] Research has found that the self-vibration of the absolute gravimeter's transmission mechanism, caused by excitations such as mechanical friction, can cause the tractor's trajectory to deviate from its ideal state. Furthermore, as a key transmission component driving the tractor, the accumulated micro-deformations of the steel belt due to temperature fluctuations can also cause systematic deviations in the tractor's motion path. Both of these deviations can cause regulation errors in the motor feedback control system, ultimately leading to a decrease in control accuracy. Therefore, the self-vibration characteristics of the transmission mechanism and the temperature stability of the steel belt are the two key factors affecting motor control accuracy.
[0079] Based on this, in some embodiments of the present disclosure, when the current control stage is the separation stage, the status data include: the position information of the tractor, the position information of the falling object, the instantaneous acceleration value of the tractor, and at least one of the following: the peak value of the main frequency of the natural vibration of the transmission mechanism in the target absolute gravimeter, and the temperature change data of the steel belt in the target absolute gravimeter; when the current control stage is the free fall stage or the taking-over stage, the status data include: the position information of the tractor, the position information of the falling object, the instantaneous speed value of the tractor, the instantaneous speed value of the falling object, the instantaneous acceleration value of the tractor, and at least one of the following: the peak value of the main frequency of the natural vibration of the transmission mechanism in the target absolute gravimeter, and the temperature change data of the steel belt in the target absolute gravimeter.
[0080] In this embodiment, for the transmission mechanism of the target absolute gravimeter, for example, Figure 2 The schematic diagram of the structure of the free-fall device of the absolute gravimeter provided by the present disclosure is shown. Figure 2 As shown in FIG, the free fall device mainly consists of a trolley, a steel belt, a vacuum transmission shaft and a motor, wherein the steel belt, the vacuum transmission shaft and the motor constitute a transmission mechanism.
[0081] To obtain the peak value of the transmission mechanism's main natural frequency, accelerometers (such as piezoelectric accelerometers) can be installed on the transmission mechanism (e.g., key locations such as the steel belt, trolley, and vacuum drive shaft). A data acquisition card can be used to convert the analog vibration signal output by the accelerometer into a digital signal. Software (such as MATLAB or LabVIEW) can then be used to preprocess the acquired time-domain signal by filtering and amplifying it. The fast Fourier transform algorithm is then used to convert the time-domain vibration signal into a frequency-domain signal, generating a spectrum. In the spectrum, the frequency corresponding to the peak point with the largest amplitude is the peak value of the main natural frequency.
[0082] To obtain temperature variation data of the steel strip, a temperature sensor can be fixed on the steel strip. The temperature sensor can obtain temperature data in real time, and the temperature variation data can be calculated based on the adjacent temperature data.
[0083] This embodiment introduces the natural vibration main frequency peak value of the transmission mechanism and / or the temperature change data of the steel belt into the state data of each control stage, so that each state space covers the mechanical vibration characteristics and thermal characteristics of the gravimeter. This provides accurate environmental feedback for the intelligent agent, enables the intelligent agent to learn the optimal control strategy under the interaction of multiple factors, reduces the tractor motion deviation caused by mechanical friction excitation and temperature cumulative deformation, and significantly improves the response accuracy and robustness of the motor feedback control.
[0084] In some embodiments, the above-mentioned reward function further includes: a peak value of the natural vibration main frequency of the transmission mechanism in the target absolute gravimeter.
[0085] Specifically, if the current control stage is the separation stage, the reward function is:
[0086]
[0087] Among them, A 自振 is the peak value of the natural frequency of the transmission mechanism.
[0088] If the current control stage is the free fall stage, the reward function is:
[0089]
[0090] If the current control stage is the continuation stage, the reward function is:
[0091]
[0092] In this embodiment, in addition to focusing on its own core control objectives, the reward function at each stage also pays attention to the self-oscillation interference of the transmission mechanism. This enables each intelligent agent to generate a more accurate PID parameter adjustment strategy, further improving the reliability of the absolute gravimeter motor control.
[0093] After determining the corresponding intelligent agent, state data, and reward function according to the current control stage, the intelligent agent can be used to process the corresponding state data to generate action data. Here, the action data, i.e., the PID controller parameter adjustment amount, is the weight of the actor network in the intelligent agent. The action data can be used to adjust the motor control parameters. Next, the reward value is calculated using the corresponding reward function. The state data of the next time step is then obtained. The reward value, action data, state data of the current time step, and state data of the next time step constitute a piece of experience data, which is stored in the playback buffer corresponding to the current control stage. Similarly, at each time step, the above steps S101-S103 are repeated to collect experience data until the current control stage is the takeover stage, and the speed of the trailer and the falling object is the same, and the distance between them is 0, that is, the takeover stage ends. At this point, the collection of experience data for one round is completed.
[0094] Regarding S104-S106 above:
[0095] Here, the training termination condition may be, for example, the convergence of the reward function, or the number of training rounds reaching a preset number of training rounds. The training termination condition may be set according to actual needs and is not limited in the present embodiment.
[0096] It is possible to determine whether each intelligent agent meets the preset training termination conditions. If all of them are met, it indicates that the performance of each intelligent agent has reached the requirements at this time. Then, the state data of the target absolute gravimeter can be input into the corresponding intelligent agent to generate a set of motor controller parameter adjustment amounts, and the motor control parameters of the target absolute gravimeter can be adjusted by using the set of motor controller parameter adjustment amounts.
[0097] If it is determined that there are target intelligent agents that have not completed training, and the number of experience data in the target replay buffer corresponding to the target intelligent agent is greater than the preset threshold, then sample experiences are extracted from the target replay buffer. During implementation, random sampling can be performed in the target replay buffer to obtain multiple sample experiences, and the target intelligent agent is trained using the extracted sample experiences to obtain a new target intelligent agent. Here, the preset threshold can be set according to actual needs. Then, the target absolute gravimeter is reset (the reset operation includes: resetting the trolley and the falling body to the initial position, clearing the control parameters and state data, etc.), so that the system restarts from the unified initial conditions to ensure the consistency of training. Then, it returns to step S101 and iterates continuously until each intelligent agent meets the preset training termination conditions.
[0098] Those skilled in the art can understand that in the above method of the specific embodiment mode, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0099] It should be noted that in practical applications, all the above possible implementation manners can be combined in any combination to form possible embodiments of the present disclosure, which will not be elaborated herein one by one.
[0100] Based on the same concept, the embodiment of the present disclosure also provides an absolute gravimeter motor control parameter adjustment device, which corresponds one-to-one to the absolute gravimeter motor control parameter adjustment method in the above embodiment. Figure 3 The structural schematic diagram of the absolute gravimeter motor control parameter adjustment device provided by the embodiment of the present disclosure is shown. Refer to Figure 3 As shown, the absolute gravimeter motor control parameter adjustment device 300 provided by the embodiment of the present disclosure includes:
[0101] A stage determination module 301, configured to determine the current control stage based on the obtained position information, motion information, and motion duration of the trolley and the falling body in the target absolute gravimeter;
[0102] The experience collection module 302 is configured to obtain the corresponding intelligent agent and state data according to the current control stage, and use the intelligent agent to process the state data to generate action data; the action data is a set of motor controller parameter adjustment values; the motor control parameters of the target absolute gravimeter are adjusted using the action data; a reward value is calculated using a reward function corresponding to the current control stage; the reward value, the action data, the state data, and the acquired state data of the next time step are stored as a piece of experience data in a playback buffer corresponding to the current control stage; and the step of determining the current control stage based on the acquired position information, motion information, and motion duration of the trailer and the falling object in the target absolute gravimeter is returned until the current control stage is a takeover stage, and the speed of the trailer and the falling object is the same and the distance between them is zero.
[0103] The training module 303 is configured to generate a determination result as to whether all agents have been trained, based on a preset training termination condition. If the determination result indicates that a target agent has not been trained, and the number of experience data items in the target replay buffer corresponding to the target agent is greater than a preset threshold, sample experience data is extracted from the target replay buffer, and the target agent is trained using the sample experience data to obtain a new target agent. The target absolute gravimeter is reset, and the process returns to the step of determining the current control phase based on the position information, motion information, and motion duration of the trailer and the falling object obtained from the target absolute gravimeter.
[0104] The adjustment module 304 is used to process the state data of the target absolute gravimeter using each of the intelligent agents if the judgment result is that the training of each of the intelligent agents is completed, and generate action data to adjust the motor control parameters of the target absolute gravimeter.
[0105] In some embodiments, the motion information includes: the instantaneous acceleration value of the trailer, the instantaneous speed value of the trailer and the instantaneous speed value of the falling object; in the above-mentioned device, the stage determination module is specifically used to: calculate the interval distance between the trailer and the falling object based on the position information of the trailer and the falling object; if the interval distance is 0, and the instantaneous acceleration value of the trailer is less than or equal to the gravity acceleration value, then the current control stage is determined to be the separation segment; if the motion duration is less than or equal to a preset time threshold, and the instantaneous acceleration value of the trailer is less than or equal to the gravity acceleration value, then the current control stage is determined to be the free fall segment; if the motion duration is greater than the time threshold, and the instantaneous speed value of the trailer is less than or equal to the instantaneous speed value of the falling object, then the current control stage is determined to be the taking-over segment.
[0106] In some embodiments, in the above-mentioned device, if the current control stage is the separation stage, the state data includes: position information of the trailer, position information of the falling object, and instantaneous acceleration value of the trailer.
[0107] In some embodiments, in the above-mentioned device, if the current control stage is a free fall stage or a receiving stage, the status data includes: the position information of the tractor, the position information of the falling object, the instantaneous speed value of the tractor, the instantaneous speed value of the falling object, and the instantaneous acceleration value of the tractor.
[0108] In some embodiments, in the above device, if the current control stage is the separation stage, the reward function is:
[0109]
[0110] Among them, R 分离 is the separation segment reward value, T1 is the total number of time steps of the preset separation segment, a t is the instantaneous acceleration value of the trailer, a 目标 is the ideal instantaneous acceleration value of the trailer in the separation section, 落体 is the angular velocity of the falling object, v 落体水平 is the instantaneous horizontal velocity of the falling object.
[0111] In some embodiments, in the above device, if the current control stage is a free-fall stage, the reward function is:
[0112]
[0113] Among them, R 自由下落 is the free-fall segment reward value, T2-T1-1 is the total number of time steps of the preset free-fall segment, x 托车 is the position information of the trailer, x 落体 is the position information of the falling object, is the preset ideal distance between the trailer and the falling object, v t is the instantaneous speed of the trailer, v 目标 is the preset ideal instantaneous speed value of the trailer.
[0114] In some embodiments, in the above device, if the current control stage is a takeover stage, the reward function is:
[0115]
[0116] Among them, R 承接 is the reward value of the succession segment, T3-T2-1 is the total number of time steps of the preset succession segment, v托车承接时瞬时 represents the instantaneous velocity value of the carrier when it catches the falling object, v 落体承接时瞬时 represents the instantaneous velocity value of the falling object when it is caught, v t is the instantaneous velocity value of the carrier, v 目标 is the preset ideal instantaneous velocity value of the carrier.
[0117] In some embodiments, in the above device, if the current control stage is the separation stage, the state data includes: the position information of the carrier, the position information of the falling object, the instantaneous acceleration value of the carrier, and at least one of the following: the peak value of the natural vibration frequency of the transmission mechanism in the target absolute gravimeter, the temperature change data of the steel belt in the target absolute gravimeter; if the current control stage is the free fall stage or the catching stage, the state data includes: the position information of the carrier, the position information of the falling object, the instantaneous velocity value of the carrier, the instantaneous velocity value of the falling object, the instantaneous acceleration value of the carrier, and at least one of the following: the peak value of the natural vibration frequency of the transmission mechanism in the target absolute gravimeter, the temperature change data of the steel belt in the target absolute gravimeter.
[0118] In some embodiments, in the above device, the reward function further includes the peak value of the natural vibration frequency of the transmission mechanism in the target absolute gravimeter.
[0119] The present invention provides an absolute gravimeter motor control parameter adjustment device, which introduces a reinforcement learning framework. For each stage of the motor controlling the carrier, first, using the agents and reward functions designed for each stage respectively, a large amount of experience data including a large number of successes / failures is accumulated. Using this experience data, the agents for each control stage can be trained to gradually optimize the controller parameter adjustment strategy of the agents. Finally, the trained mature agents can dynamically generate accurate PID parameter adjustment amounts according to the real-time state data of the target absolute gravimeter. Using these PID parameters, the accurate adjustment of the gravity meter motor control parameters can be realized. Through this interactive learning mechanism, the adaptive adjustment of the absolute gravimeter motor control parameters can be achieved, breaking through the limitations of traditional empirical parameter adjustment, significantly improving the accuracy and stability of the movement of the carrier and the falling object, and thus improving the accuracy of the measured value of the gravitational acceleration.
[0120] For the specific limitations of the absolute gravimeter motor control parameter adjustment device, reference can be made to the limitations of the absolute gravimeter motor control parameter adjustment method in the above text, which will not be elaborated here. Each module in the above absolute gravimeter motor control parameter adjustment device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0121] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0122] Those skilled in the art will appreciate that embodiments of the present disclosure may be provided as methods, systems, or computer program products. Thus, the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0123] The above are merely examples of the present disclosure and are not intended to limit the present disclosure. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure are intended to be included within the scope of the claims of the present disclosure.
Claims
1. A method for adjusting the motor control parameters of an absolute gravimeter, characterized in that, The method comprises: Determine the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object obtained from the target absolute gravimeter; According to the current control stage, the corresponding intelligent agent and state data are obtained, and the state data is processed by the intelligent agent to generate action data; the action data is a set of motor controller parameter adjustment values; Using the action data, adjusting motor control parameters of the target absolute gravimeter; using a reward function corresponding to the current control stage, calculating a reward value; storing the reward value, the action data, the state data, and the acquired state data for the next time step as a piece of experience data in a playback buffer corresponding to the current control stage; returning to the step of determining the current control stage based on the acquired position information, motion information, and motion duration of the tractor and the falling object in the target absolute gravimeter, until the current control stage is a takeover stage, and the tractor and the falling object have the same speed and the separation distance is 0; According to the preset training termination conditions, generating a judgment result on whether the training of each of the intelligent agents is completed; If the judgment result is that the training of each of the intelligent agents is completed, then using each of the intelligent agents to process the state data of the target absolute gravimeter to generate action data to adjust the motor control parameters of the target absolute gravimeter; If the judgment result is that there is a target intelligent agent that has not been trained, and the number of experience data in the target replay buffer corresponding to the target intelligent agent is greater than a preset threshold, then sample experience is extracted from the target replay buffer, and the target intelligent agent is trained using the sample experience to obtain a new target intelligent agent; the target absolute gravimeter is reset, and the process returns to the step of determining the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object obtained in the target absolute gravimeter.
2. The method according to claim 1, characterized in that, The motion information includes: the instantaneous acceleration value of the trailer, the instantaneous speed value of the trailer, and the instantaneous speed value of the falling object; the determining of the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object obtained from the target absolute gravimeter includes: Calculating the distance between the trailer and the falling object based on the position information of the trailer and the falling object; If the separation distance is 0, and the instantaneous acceleration value of the trailer is less than or equal to the gravity acceleration value, then the current control stage is determined to be the separation stage; If the motion duration is less than or equal to a preset time threshold, and the instantaneous acceleration value of the trailer is less than or equal to the gravity acceleration value, then the current control stage is determined to be a free fall stage; If the movement duration is greater than the time threshold, and the instantaneous speed value of the trailer is less than or equal to the instantaneous speed value of the falling object, the current control stage is determined to be the taking-over stage.
3. The method according to claim 1, wherein If the current control stage is the separation stage, the state data includes: the position information of the trailer, the position information of the falling object, and the instantaneous acceleration value of the trailer.
4. The method according to claim 1, wherein If the current control stage is the free fall stage or the receiving stage, the status data includes: the position information of the trailer, the position information of the falling object, the instantaneous speed value of the trailer, the instantaneous speed value of the falling object, and the instantaneous acceleration value of the trailer.
5. The method according to claim 1, wherein If the current control stage is the separation stage, the reward function is: Among them, R 分离 is the separation segment reward value, T1 is the total number of time steps of the preset separation segment, a t is the instantaneous acceleration value of the trailer, a 目标 is the ideal instantaneous acceleration value of the trailer in the separation segment, 落体 is the rotational angular velocity of the falling object, v 落体水平 is the instantaneous horizontal velocity value of the falling object.
6. The method according to claim 1, wherein If the current control phase is the free fall phase, the reward function is: Among them, R 自由下落 is the reward value for the free fall section, T2 - T1 - 1 is the total number of time steps preset for the free fall section, x 托车 is the position information of the trolley, x 落体 is the position information of the falling object, is the ideal distance preset between the trolley and the falling object, v t is the instantaneous velocity value of the trolley, v 目标 is the ideal instantaneous velocity value preset for the trolley.
7. The method according to claim 1, characterized in that, If the current control stage is the continuation stage, the reward function is: Among them, R 承接 is the reward value of the succession segment, T3-T2-1 is the total number of time steps of the preset succession segment, v 托车承接时瞬时 represents the instantaneous speed of the vehicle when it receives the falling object, v 落体承接时瞬时 represents the instantaneous velocity of the falling object when it is caught, v t is the instantaneous speed of the trailer, v 目标 is the preset ideal instantaneous speed value of the trailer.
8. The method according to claim 3 or 4, characterized in that The state data further includes at least one of the following: a peak value of a main frequency of natural vibration of a transmission mechanism in the target absolute gravimeter, and temperature change data of a steel belt in the target absolute gravimeter.
9. The method according to any one of claims 5-7, characterized in that The reward function also includes the peak value of the natural vibration main frequency of the transmission mechanism in the target absolute gravimeter.
10. An apparatus for adjusting the motor control parameters of an absolute gravimeter, characterized in that, The device comprises: A phase determination module is used to determine the current control phase based on the position information, motion information, and motion duration of the trailer and the falling object obtained from the target absolute gravimeter; An experience collection module is configured to obtain, according to the current control stage, a corresponding intelligent agent and state data, and process the state data using the intelligent agent to generate action data; the action data is a set of motor controller parameter adjustment values; the motor control parameters of the target absolute gravimeter are adjusted using the action data; a reward value is calculated using a reward function corresponding to the current control stage; the reward value, the action data, the state data, and the acquired state data for the next time step are stored as a piece of experience data in a playback buffer corresponding to the current control stage; and the step of determining the current control stage based on the acquired position information, motion information, and motion duration of the trailer and the falling object in the target absolute gravimeter is returned until the current control stage is a takeover stage, and the speeds of the trailer and the falling object are the same and the separation distance is zero. A training module is configured to generate a determination result as to whether all of the agents have completed training based on a preset training termination condition; if the determination result indicates that a target agent has not been trained and the number of experience data items in the target replay buffer corresponding to the target agent is greater than a preset threshold, extract sample experience from the target replay buffer, use the sample experience to train the target agent, and obtain a new target agent; reset the target absolute gravimeter, and return to the step of determining the current control stage based on the position information, motion information, and motion duration of the trailer and the falling object obtained from the target absolute gravimeter; The adjustment module is used to use each of the intelligent agents to process the state data of the target absolute gravimeter and generate action data to adjust the motor control parameters of the target absolute gravimeter if the judgment result is that the training of each of the intelligent agents is completed.
Citation Information
Patent Citations
Method for platform gravimeter to reduce influence of horizontal acceleration
CN111123381A
Underwater vehicle attitude control system and method based on reinforcement learning compensator
CN116449856A