Optimization Method, Device, Water Movable Device and Storage Medium of Controller

By training the probability dynamic model and optimizing the controller parameters using particle swarm algorithm, the problem of cumbersome controller parameter setting in intelligent ship motion control is solved, and rapid and automated parameter setting is achieved, reducing labor costs.

CN119882459BActive Publication Date: 2025-07-01DONGGUAN EPROPULSION INTELLIGENCE TECH LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510380244.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-01
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

In intelligent ship motion control, controller parameter setting usually requires engineers to adjust repeatedly through trial and error, which consumes a lot of time and energy, and parameter adjustments based on human intuition may not produce better parameter values.

Method used

By obtaining the actual driving data of movable devices in the water, training the probability dynamics model, and optimizing the controller parameters using the particle swarm algorithm to achieve automated parameter tuning.

Benefits of technology

This method can effectively reduce the number of interactions with movable devices in actual waters, realize the rapid adjustment of the auxiliary driving function controller, reduce labor costs, and quickly obtain controller parameters that meet actual needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119882459B_ABST
    Figure CN119882459B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides an optimization method, device, waterborne movable device, and storage medium for a controller. Among them, the method includes the following steps: obtaining the actual driving data of the waterborne movable device; training a probabilistic dynamics model of the waterborne movable device according to the actual driving data; using a particle swarm algorithm and the probabilistic dynamics model to optimize the controller parameters of the controller; using the currently optimized controller parameters to control the operation of the waterborne movable device, and obtaining the actual driving data generated by the waterborne movable device under the control of the optimized controller; when the optimization process of the controller meets a first predetermined condition, determining the currently optimized controller parameters as the controller parameters of the target controller; when the optimization process of the controller does not meet the first predetermined condition, returning to the step of training the probabilistic dynamics model according to the actual driving data. The technical solution provided by the embodiment of the present application can reduce labor costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of control technologies, and particularly to an optimization method and device for a controller, a waterborne movable device, and a storage medium. Background Art

[0002] In the motion control of intelligent ships, the controller completes a series of autonomous driving tasks according to the ship navigation data, such as course keeping and path tracking. The control effect is directly related to the controller parameters. Therefore, the tuning of controller parameters is an indispensable process in the industry.

[0003] However, in practical applications, this work usually requires engineers to repeatedly adjust by trial and error based on the theoretical knowledge of the controller and expert experience. This is a cumbersome task for engineers, consuming a large amount of time and effort, and the parameter adjustment based on human intuition may not produce an optimal value of the parameters. Summary of the Invention

[0004] In view of the above problems, the present application is proposed to provide an optimization method and device for a controller, a waterborne movable device, and a storage medium that solve the above problems or at least partially solve the above problems.

[0005] In a first aspect of the present application, an optimization method for a controller is provided. The controller is used to control the auxiliary driving function of a waterborne movable device. The method includes:

[0006] Obtain the actual driving data of the waterborne movable device, where the actual driving data is generated when the waterborne movable device moves under the control of the controller;

[0007] Train a probabilistic dynamics model of the waterborne movable device according to the actual driving data;

[0008] Optimize the controller parameters of the controller by using the particle swarm optimization algorithm and the probabilistic dynamics model;

[0009] Control the operation of the waterborne movable device by using the currently optimized controller parameters, and obtain the actual driving data generated by the waterborne movable device under the control of the optimized controller;

[0010] When the optimization process of the controller meets a first predetermined condition, determine the currently optimized controller parameters as the controller parameters of the target controller;

[0011] When the optimization process of the controller does not meet the first predetermined condition, return to the step of training a probabilistic dynamics model of the waterborne movable device according to the actual driving data.

[0012] In a second aspect of the present application, an electronic device is provided. The electronic device includes: a memory and a processor, where

[0013] the memory is used to store programs;

[0014] the processor is coupled to the memory and is used to execute the programs stored in the memory to implement the method described above.

[0015] In a third aspect of the present application, a waterborne mobile device is provided, and the device includes the above-mentioned electronic device.

[0016] In a fourth aspect of the present application, a computer-readable storage medium storing a computer program is provided, and when the computer program is executed by a computer, it can implement the method described in any one of the above.

[0017] In the technical solution provided by the embodiments of the present application, only a small amount of actual driving data of the waterborne mobile device needs to be collected for training the probabilistic dynamics model. Subsequently, in the process of optimizing the controller parameters using the particle swarm algorithm, only this probabilistic dynamics model is needed to simulate the movement of the waterborne mobile device, without the need to interact with the actual waterborne mobile device. It can be seen that the technical solution provided by the embodiments of the present application can effectively reduce the number of interactions with the actual waterborne mobile device, and can achieve the rapid tuning of the auxiliary driving function controller of the waterborne mobile device under the condition of very few interactions. It can be seen that the technical solution provided by the embodiments of the present application can not only reduce the labor cost, but also quickly obtain the controller parameters that meet the actual auxiliary driving function requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1 It is a schematic flowchart of an optimization method provided by an embodiment of the present application;

[0020] Figure 2 It is a structural example diagram of a probabilistic dynamics model provided by an embodiment of the present application;

[0021] Figure 3 It is a structural block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To enable those skilled in the art to better understand the solution of this application, the technical solution in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the scope of protection of this application.

[0023] In addition, in some processes described in the specification, claims, and the above-mentioned accompanying drawings of this application, a plurality of operations that appear in a specific order are included. These operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish each different operation, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., do not represent the sequence, and do not limit that "first" and "second" are different types.

[0024] Figure 1 It is a schematic flowchart of an optimization method for a controller provided in an embodiment of this application. The execution subject of this method may be an electronic device, and this electronic device may be a central controller of a watercraft propeller (a controller integrated into the watercraft propeller) or a cloud server. The controller is used to control the assisted driving function of a waterborne mobile device, that is, the controller is a controller for completing the assisted driving function. The controller may be a proportional, integral, and derivative (PID) controller, a neural network controller, etc. For a PID controller, the controller parameters are the proportional coefficient (Kp), the integral coefficient (Ki), and the derivative coefficient (Kd). For a neural network controller, the controller parameters are the weights (weight) and biases (bias) of the neural network, etc. The assisted driving function may include heading keeping, position keeping, speed keeping, path following, trajectory tracking, etc.

[0025] As Figure 1 shown, the optimization method of this controller may include the following steps:

[0026] 101. Obtain the actual driving data of the waterborne mobile device.

[0027] Among them, the waterborne mobile device may be a ship, an underwater robot, a submarine, etc.

[0028] Among them, the actual driving data is generated by the movement of the waterborne mobile device under the control of the controller.

[0029] In some embodiments, the actual driving data in step 101 may be generated by the movement of the waterborne mobile device under the control of the initial controller. The controller parameters of the initial controller are the initial controller parameters. That is to say, the actual driving data in step 101 is generated by the initial controller controlling the movement of the waterborne mobile device according to the initial controller parameters. The initial controller parameters may be set based on actual experience or randomly initialized.

[0030] Exemplarily, the actual driving data includes the historical motion state quantities and historical motion control quantities of the waterborne mobile device at multiple historical moments. Among them, the historical motion control quantity at any historical moment is determined by the controller based on the controller parameters it has at that historical moment and the historical motion state quantity of the waterborne mobile device at that historical moment.

[0031] Among them, the motion control quantity refers to the parameter used to control the motion state of the waterborne mobile device. This parameter is transmitted by the controller to the actuator to achieve the control of the motion state of the waterborne mobile device. Taking the waterborne mobile device as a ship as an example, the actuator may be a propeller and a rudder, and the motion control quantity may include propeller power data, rudder angle data for controlling the rudder angle steering, or steering wheel rotation angle data.

[0032] Among them, the motion state quantity may include at least one of the position and speed of the waterborne mobile device in the three degrees of freedom of surge, sway, and yaw. The position and speed of the waterborne mobile device in the three degrees of freedom of surge, sway, and yaw include: the position in surge, the linear velocity in surge, the position in sway, the linear velocity in sway, the position in yaw, and the angular velocity in yaw. Exemplarily, the motion state quantity may be the motion state quantity in the global coordinate system (i.e., the geodetic coordinate system) or the local coordinate system of the waterborne mobile device.

[0033] The driving data (i.e., the data that needs to be collected) concerned by different assisted driving functions is different. For example: The driving data concerned by course keeping is mainly: the angular velocity of yaw, the position of yaw, the linear velocity of sway, the linear velocity of surge; The driving data concerned by position keeping is mainly: the position of sway, the linear velocity of sway, the position of surge, the linear velocity of surge; The driving data concerned by speed keeping is mainly: the linear velocity of surge. Of course, in actual applications, in order to make the probabilistic neural network model more fitting to the real ship, for each assisted driving function, the six kinds of data mentioned above can be used.

[0034] In practical applications, the actual driving data of the waterborne mobile device within a preset time period can be obtained. The duration of the preset time period can be set according to actual needs, and the embodiments of the present application do not make specific limitations thereto. For example, it can be 20 seconds. In some embodiments, a data acquisition device can be provided on the waterborne mobile device. The data acquisition device acquires the actual driving data of the waterborne mobile device within the preset time period and sends the acquired data to an electronic device for the electronic device to execute Figure 1 the optimization method shown.

[0035] 102. Train a probabilistic dynamics model of the waterborne mobile device according to the actual driving data.

[0036] Among them, the probabilistic dynamics model of the waterborne mobile device is a mathematical framework used to describe and predict the motion behavior of the waterborne mobile device in an uncertain environment. Such models combine classical dynamics equations with probability theory to cope with the inherent randomness and uncertainty factors in the marine environment, such as the influence of waves, ocean currents, wind power, etc. on the motion of the waterborne mobile device. Moreover, as a prediction model constructed based on probability theory and mathematical statistics methods, the probabilistic dynamics model can also exhibit good accuracy and reliability when facing small samples or data containing random errors.

[0037] Training a probabilistic dynamics model of the waterborne mobile device according to the actual driving data can, to a certain extent, truly simulate the motion process of the waterborne mobile device.

[0038] In some embodiments, the probabilistic dynamics model is used to predict the predicted state change amount of the waterborne mobile device, and the predicted state change amount is represented in the form of a probability distribution. Exemplarily, the predicted state change amount follows a normal distribution, and the probabilistic dynamics model outputs the variance and mean of the normal distribution followed by the predicted state change amount.

[0039] Exemplarily, the probabilistic dynamics model is used to output a predicted state change amount represented in the form of a probability distribution based on the motion state amount and motion control amount of the waterborne mobile device at the previous moment. In this way, based on the motion state amount of the waterborne mobile device at the previous moment and the predicted state change amount represented in the form of a probability distribution output by the probabilistic dynamics model, the predicted motion state amount of the waterborne mobile device at the next moment represented in the form of a probability distribution can be obtained. The probabilistic dynamics model outputs the variance and mean of the normal distribution followed by the predicted state change amount. Add the motion state amount of the waterborne mobile device at the previous moment to the mean value output by the probabilistic dynamics model to obtain the target mean value. Determine the predicted motion state amount represented in the form of a probability distribution according to the target mean value and the variance output by the probabilistic dynamics model. The predicted motion state amount follows a normal distribution determined by the target mean value and the variance output by the probabilistic dynamics model.

[0040] In some other embodiments, a probabilistic dynamics model is used to predict the predicted motion state quantity of the waterborne movable device, and the predicted motion state quantity is represented in the form of a probability distribution. Exemplarily, the predicted state change quantity follows a normal distribution. What the probabilistic dynamics model outputs are the variance and mean of the normal distribution followed by the predicted motion state quantity.

[0041] 103. Use the particle swarm optimization algorithm and the probabilistic dynamics model to optimize the controller parameters of the controller.

[0042] Wherein, in the particle swarm optimization algorithm, the position of a particle is used to represent the controller parameters of a controller, and the controller parameters can be regarded as candidate controller parameters. The fitness of the particle is determined based on the state difference between the predicted operating state and the reference motion state, and the predicted operating state is predicted by using the probabilistic dynamics model.

[0043] Wherein, the reference motion state can be understood as the desired motion state, and the reference motion state can be set according to the preset assisted driving function to be achieved by the controller. It should be noted that when the preset assisted driving functions to be achieved by the controller are different, the reference motion states are also different.

[0044] Wherein, optimizing the controller parameters of the controller is equivalent to optimizing the controller.

[0045] 104. Use the currently optimized controller parameters to control the operation of the waterborne movable device, and obtain the actual driving data generated by the waterborne movable device under the control of the optimized controller.

[0046] The currently optimized controller parameters refer to the controller parameters optimized through step 103. The controller parameters of the optimized controller are the currently optimized controller parameters. Among them, the optimized controller can be considered as the controller parameters optimized through step 103.

[0047] Taking the waterborne movable device as a ship as an example, the optimized controller parameters can be sent back to the ship control module, and the control module uses the optimized controller to control the intelligent assisted driving function.

[0048] 105. When the optimization process of the controller meets the first predetermined condition, determine the currently optimized controller parameters as the controller parameters of the target controller.

[0049] 106. When the optimization process of the controller does not meet the first predetermined condition, return to the step of training the probabilistic dynamics model of the waterborne movable device according to the actual driving data.

[0050] When the optimization process of the controller does not meet the first predetermined condition, new actual driving data is determined according to the actual driving data generated by the waterborne mobile device under the control of the optimized controller, and a probabilistic dynamics model of the waterborne mobile device is trained and returned based on the new actual driving data.

[0051] When the optimization process of the controller does not meet the first predetermined condition, the i-th iteration is started. In the i-th iteration, the actual driving data generated by controlling the movement of the waterborne mobile device by using the controller obtained in the (i - 1)-th iteration is acquired. Based on this actual driving data, a probabilistic dynamics model of the waterborne mobile device is trained, and the controller parameters of the controller are optimized by using a particle swarm algorithm and the probabilistic dynamics model. Here, i is an integer greater than or equal to 2. It should be noted that the actual driving data used in the first iteration can be generated by controlling the waterborne mobile device by using an initial controller.

[0052] In an alternative manner, in the i-th iteration, not only the actual driving data generated by controlling the movement of the waterborne mobile device by using the controller obtained in the (i - 1)-th iteration can be used to train the model, but also the actual driving data used to train the model in the (i - 1)-th iteration can be used to train the model. Here, i is an integer greater than or equal to 2. That is to say, the total data set used in the i-th iteration is the set of the actual driving data collected during the i iterations. For example, if 200 pairs of data are collected in the first iteration and 200 pairs of data are collected in the second iteration, then the total data set used in the second iteration is 200 + 200 = 400 pairs, and so on.

[0053] In some embodiments, the first predetermined condition may include: the number of iterations of the optimization process of the controller reaches a sixth preset number, and / or, in seven consecutive preset iterations of the optimization process of the controller, the difference between the actual cumulative costs corresponding to the controller parameters obtained in any two iterations is less than a second preset value, and the actual cumulative cost corresponding to the controller parameter obtained in any one iteration is calculated based on the actual driving data generated by controlling the movement of the waterborne mobile device by using the controller obtained in this iteration and a preset cost function.

[0054] Exemplarily, the preset cost function is:

[0055] (1)

[0056] In the formula, are the historical motion state and the reference state at time t respectively, and L is a pre-specified weighted coefficient matrix, which determines the cost magnitudes corresponding to different state dimensions.

[0057] In practical applications, the actual driving data generated by controlling the water area movable device using the controller obtained from this optimization process includes the historical motion states at multiple historical moments. Among them, the historical motion state is characterized by motion state variables. For the historical motion state at each moment, the cost corresponding to that moment can be calculated according to the above formula (1), and the costs corresponding to multiple historical moments are accumulated to obtain the actual cumulative cost corresponding to the controller parameters obtained from this optimization process.

[0058] Among them, the sixth preset number of times can be the maximum number of iterations set in advance, that is: the maximum number of executions of the optimization process. In practical applications, the maximum number of iterations can be set according to actual needs, and the embodiments of the present application do not make specific limitations in this regard. For example: the maximum number of iterations is 5 times.

[0059] Among them, the seventh preset number of times can be set according to actual needs, and the embodiments of the present application do not make specific limitations in this regard. For example: the seventh preset number of times is 2 times.

[0060] Among them, the sixth preset number of times can be greater than the seventh preset number of times.

[0061] Exemplarily, when the number of executions of the optimization process reaches the maximum number of iterations or the actual cumulative cost corresponding to the controller parameters obtained from consecutive optimization processes tends to be stable, the iteration stops.

[0062] In some other embodiments, the first preset condition may include: the optimization process of the controller reaches the sixth preset number of times, and / or when the actual cumulative cost corresponding to the controller parameters obtained from a certain optimization process is less than the third preset value. Exemplarily, in the embodiments of the present application, when the number of executions of the optimization process reaches the maximum number of iterations or the actual cumulative cost corresponding to the controller parameters obtained from a certain optimization process is less than the third preset value, the iteration stops.

[0063] In the technical solution provided by the embodiments of the present application, only a small amount of actual driving data of the water area movable device needs to be collected for training the probabilistic dynamics model. Subsequently, in the process of optimizing the controller parameters using the particle swarm algorithm, only this probabilistic dynamics model needs to be used to simulate the motion of the water area movable device, without the need to interact with the actual water area movable device. It can be seen that the technical solution provided by the embodiments of the present application can effectively reduce the number of interactions with the actual water area movable device, and can quickly tune the controller of the assisted driving function of the water area movable device under the condition of very few interactions. It can be seen that the technical solution provided by the embodiments of the present application can not only reduce the labor cost, but also quickly obtain the controller parameters that meet the actual assisted driving function requirements.

[0064] In some embodiments, the actual driving data includes the historical motion state quantities and historical motion control quantities of the waterborne mobile device at multiple historical moments. The step "training the probabilistic dynamics model of the waterborne mobile device according to the actual driving data" in step 102 can be implemented by the following steps:

[0065] 1021. According to the historical motion state quantity and historical motion control quantity at the first historical moment among the multiple historical moments, use the probabilistic dynamics model to obtain a first predicted state change quantity represented in the form of a probability distribution.

[0066] Exemplarily, the probabilistic dynamics model outputs the mean and variance of the normal distribution followed by the first predicted state change quantity.

[0067] In some embodiments, input the historical motion state quantity and historical motion control quantity at the first historical moment among the multiple historical moments into the probabilistic dynamics model to obtain a first predicted state change quantity represented in the form of a probability distribution.

[0068] 1022. Use a preset loss function to calculate the state change difference between the first predicted state change quantity and the expected state change quantity.

[0069] Among them, the state change difference refers to the difference between the first predicted state change quantity and the expected state change quantity.

[0070] Among them, the expected state change quantity is the state change quantity between the historical motion state quantity at the second historical moment and the historical motion state quantity at the first historical moment among the multiple historical moments. The first historical moment and the second historical moment are any two adjacent moments among the multiple historical moments, and the first historical moment is earlier than the second historical moment.

[0071] Among them, the state change quantity between the historical motion state quantity at the second moment and the historical motion state quantity at the first historical moment may include at least one of the position change quantities in the surge, sway, and yaw degrees of freedom and the speed change quantities in the surge, sway, and yaw degrees of freedom.

[0072] In an alternative embodiment, the above preset loss function is the negative log-likelihood function, and the function value of the preset loss function is the state change difference.

[0073] In practical applications, given a training data set of size N, denoted as , where is the training input set, is the corresponding training target, , where,​​ The positions and velocities at time n in the surge, sway, and yaw degrees of freedom are the propeller power and rudder angle at time n, and the training objective is the state change amount between adjacent times , where are the position change amounts at time n + 1 and time n in the sway, surge, and yaw degrees of freedom, respectively, and the velocity change amounts at time n + 1 and time n in the sway, surge, and yaw degrees of freedom.

[0074] The probability dynamics model is trained by minimizing the negative log-likelihood function:

[0075] (2)

[0076] In the formula, , are the mean and variance of the model prediction / output when the input is respectively.

[0077] Among them, the negative log-likelihood function incorporates the uncertainty of the output, which helps to improve the prediction accuracy of the model.

[0078] 1023. Optimize the model parameters of the probability dynamics model according to the state change difference.

[0079] According to the state change difference, use an optimizer to optimize the model parameters of the probability dynamics model. Among them, the optimizer can optimize the model parameters of the probability dynamics model based on optimization algorithms such as gradient descent method, stochastic gradient descent method, momentum method, etc.

[0080] In some embodiments, as Figure 2 shown, the probability dynamics model may include multiple sub-models 200 (illustrated as B sub-models). Exemplarily, the probability dynamics model can be established using the Probabilistic Ensemble (PE) method, which includes two parts: a probabilistic neural network model and an idol ensemble.

[0081] In the above step 103, each sub-model is used to predict the state change amount, and the output of the probability dynamics model is obtained based on the outputs of multiple sub-models. That is, the optimized probability dynamics model is used to predict the state change amount between the second moment and the first moment based on the motion state amount and motion control amount at the first moment, where the motion state amount and motion control amount at the first moment are used as the input of each sub-model, and the state change amount is obtained based on the outputs of the multiple sub-models.

[0082] The mean and variance of the probability distribution output by the probability dynamics model are respectively:

[0083] (3)

[0084] (4)

[0085] In the formula, are respectively the mean and variance of the prediction / output of the b-th sub-model, are respectively the mean and variance of the prediction / output of the probability dynamics model.

[0086] Among them, the sub-model can be a probabilistic neural network.

[0087] Given the state quantity and control quantity at time t, the predicted state distribution at time t + 1 can be expressed as:

[0088] (5)

[0089] Among them,

[0090] (6)

[0091] (7)

[0092] In some embodiments, the initial model parameters of each of the sub-models are obtained by randomization; the training sample set of each of the sub-models is randomly sampled from the actual driving data for a first preset number of times. For each of the sub-models, the training sample set corresponding to the sub-model includes training samples, and the training sample includes: the historical motion state quantity and historical motion control quantity at a first historical moment, and the state change quantity between the historical motion state quantity at a second historical moment and the historical motion state quantity at the first historical moment. In this training sample, the historical motion state quantity and historical motion control quantity at the first historical moment are used as the input of the sub-model, and the state change quantity between the historical motion state quantity at the second historical moment and the historical motion state quantity at the first historical moment is used as the expected state change quantity (i.e., the training label) of the sub-model. In this way, during training, for each sub-model, the historical motion state quantity and historical motion control quantity of the training samples in the training sample set of the sub-model can be used as the input of the sub-model to obtain the first predicted state change quantity output by the sub-model. Then, using a preset loss function, calculate the state change difference between the first predicted state change quantity and the expected state change quantity in the training sample, and optimize the model parameters of the sub-model based on this state change difference. In the embodiments of the present application, the training sample sets (or training data sets) corresponding to different sub-models are different, that is, different sub-models are trained using different training sample sets.

[0093] In some other embodiments, during training, the historical motion state quantity and the historical motion control quantity at the first moment among multiple historical moments are used as the input of each of the sub-models. The first predicted state change quantity is obtained based on the outputs of the multiple sub-models. To obtain the first predicted state change quantity based on the outputs of the multiple sub-models, the above formulas (3) and (4) can be specifically used for implementation. Then, using a preset loss function, the state change difference between the first predicted state change quantity and the desired state change quantity is calculated, and based on this state change difference, the model parameters of the multiple sub-models are optimized.

[0094] Each time a training sample is randomly sampled, the first preset number of times can be designed according to actual needs, and the embodiments of the present application do not make specific limitations on this. For example: 200 times.

[0095] Optionally, during the i-th iteration in the optimization process of the controller, the actual driving data generated by controlling the water area movable device using the controller obtained in the (i - 1)-th iteration is sampled the first preset number of times to obtain new added training samples; the new added training samples and the training samples used in the (i - 1)-th iteration are used as the training samples required for the i-th iteration, and based on the training samples required for the i-th iteration, the probabilistic dynamics model of the water area movable device is trained. Wherein, i is an integer greater than or equal to 2.

[0096] Among them, each sub-model can specifically be a bootstrap model, for example, it can be a probabilistic neural network model.

[0097] The model parameters of each sub-model are randomly initialized (i.e., the parameters of each probabilistic neural network model may be different), and each sub-model has its own independent training data set, which is obtained by sampling the total data set collected from the real environment N times (data can be repeated), where N is the size of the total data set. Specifically, assuming that 200 pairs of data are collected within 20s, where each pair of data includes the motion state quantity at time t as the model input and the change quantity between the motion state quantity at time t+1 and the motion state quantity at time t as the training label, then each model will sample 200 times from these 200 pairs of data to obtain the training data set unique to this model. Among them, it is possible that the sampled data is the same data, that is, the data can be repeated. Using different training data sets for different sub-models helps to capture epistemic uncertainty (currently there are only 200 pairs of data, the data samples are relatively few, and the lack of understanding of the real environment by the model leads to epistemic uncertainty), making the model more fitting to the real ship. It should be noted that since the optimization of the controller parameters requires multiple iterations, therefore, the total data set used in the i-th iteration is the set of all actual driving data collected during the i-th iteration. For example, 200 pairs of data are collected in the first iteration, and 200 pairs of data are collected again at the beginning of the second iteration, then the total data set used in the second iteration is 200+200=400 pairs, and so on.

[0098] In some embodiments, for the step 103 of "optimizing the controller parameters of the controller by using the particle swarm algorithm and the probabilistic dynamics model", the following steps can be used to implement:

[0099] 1031. Based on multiple particles in the particle swarm algorithm and the probabilistic dynamics model, predict the predicted motion state of the waterborne mobile device.

[0100] Among the multiple particles, the position of each particle serves as the controller parameter of a controller. It can be considered that the position of each particle serves as a candidate controller parameter of the controller.

[0101] Among them, the velocity of each particle represents the direction and speed of position update.

[0102] For each particle, the probabilistic dynamics model and the controller parameters represented by the particle can be used to predict the predicted motion state of the waterborne mobile device.

[0103] In an alternative embodiment, the predicted motion state is characterized by a predicted motion state quantity. For the step 1031 of "based on multiple particles in the particle swarm algorithm and the probabilistic dynamics model, predict the predicted motion state of the waterborne mobile device", the following steps can be used to implement:

[0104] S11. For each of the said particles, calculate the motion control quantity at the first moment based on the motion state quantity at the first moment and the particle.

[0105] Wherein, the motion state quantity at the first moment can be a randomly generated motion state quantity, or a historical motion state quantity of the water area movable device at a certain historical moment (the motion state quantity is an actually collected motion state quantity), or a motion state quantity predicted by a probability dynamics model.

[0106] S12. Input the motion state quantity and the motion control quantity at the first moment into the probability dynamics model to obtain a second predicted state change quantity represented in the form of a probability distribution.

[0107] Exemplarily, input the motion state quantity and the motion control quantity at the first moment into the probability dynamics model, and the probability dynamics model outputs the mean and variance of the normal distribution followed by the second predicted state change quantity.

[0108] S13. Extract a predicted state change quantity from the second predicted state change quantity as the target predicted state change quantity.

[0109] Sample from the normal distribution followed by the second predicted state change quantity to obtain a predicted state change quantity, and use this predicted state change quantity as the target predicted state change quantity. Exemplarily, the Monte Carlo method can be used to sample from the normal distribution followed by the second predicted state change quantity to obtain a predicted state change quantity.

[0110] S14. Calculate the predicted motion state quantity at the second moment according to the motion state quantity at the first moment and the target predicted state change quantity.

[0111] When the target predicted state change quantity and the motion state quantity at the first moment are expressed based on the same coordinate system (such as the global coordinate system), the motion state quantity at the first moment can be added to the target predicted state change quantity to obtain the predicted motion state quantity at the second moment.

[0112] When the target predicted state change quantity is expressed based on the first coordinate system (such as the local coordinate system of the water area movable device) and the motion state quantity at the first moment is expressed based on the second coordinate system (such as the global coordinate system), the target predicted state change quantity can be transformed to the second coordinate system, and then added to the motion state quantity at the first moment to obtain the predicted motion state quantity at the second moment.

[0113] 1032. Determine the fitness of the multiple particles based on the state difference between the predicted motion state and the reference motion state.

[0114] For each particle, the fitness of the particle can be determined according to the state difference between the predicted motion state quantity calculated for the example and the reference motion state quantity.

[0115] 1033. When the fitness does not meet the second predetermined condition, update the positions and velocities of the multiple particles, and return to the step of predicting the predicted motion state of the water area movable device based on the multiple particles in the particle swarm algorithm and the probability dynamics model.

[0116] In some embodiments, the second predetermined condition includes at least one of the following: the number of update iterations of the fitness of the multiple particles reaches a fourth preset number; for any particle, during the update iterations of a continuous fifth preset number, the difference between the fitness values obtained in any two update iterations is less than a first preset value.

[0117] Wherein, the fourth preset number refers to the maximum number of iterations, and the magnitude of the maximum number of iterations can be set according to actual needs, and the embodiments of the present application do not make specific limitations in this regard. It should be noted that in this embodiment, the magnitude of the maximum number of iterations may be different from the magnitude of the maximum number of iterations involved in the first predetermined condition in the above embodiment.

[0118] The magnitude of the above first preset value can be set according to actual needs, and the embodiments of the present application do not make specific limitations in this regard. The first preset value in the embodiments of the present application may be different from the second preset value in the above embodiment.

[0119] For any particle, during the update iterations of a continuous fifth preset number, the difference between the fitness values obtained in any two update iterations is less than a first preset value, indicating that the fitness of any particle tends to be stable.

[0120] During the process of the particle searching for the optimal solution, the personal best position and the global best position are recorded, and the position and velocity of the particle are updated according to these two positions. Let the position and velocity of particle j be and . Let (k) represent the best position of this particle after the k-th iteration, and let represent the current global best position after the k-th iteration. The position and velocity of the particle are updated according to the following formulas:

[0121] (8)

[0122] (9)

[0123] In the formula, represents the number of iterations, , is the learning factor, usually set to , and are two random numbers between 0 and 1. is the inertia factor, reflecting the ability of the particle to maintain its previous velocity. Optionally, a linearly decreasing inertia factor is adopted, and its value at the k-th iteration is determined by the following formula:

[0124] (10)

[0125] wherein, and respectively represent the initial inertia factor and the inertia factor at the maximum number of iterations S.

[0126] 1034. When the fitness meets the second predetermined condition, select the position corresponding to the particle with the minimum fitness from the multiple particles as the optimized controller parameter.

[0127] In some embodiments, the position of the particle with the minimum fitness in the last iteration process can be used as the optimized controller parameter.

[0128] In other embodiments, the position of the particle with the minimum fitness in the multiple iteration processes of the fitness of the multiple particles can be used as the optimized controller parameter.

[0129] In an alternative embodiment, in order to improve the calculation accuracy of the fitness of each particle, for "determining the fitness of the multiple particles based on the state difference between the predicted motion state and the reference state" in step 1032 above, the following steps can be adopted to implement:

[0130] S21. For each of the particles, use the predicted motion state quantity and the reference motion state quantity as the input of a preset cost function to calculate the instantaneous cost at the first moment.

[0131] wherein, the instantaneous cost is used to measure the state difference between the predicted motion state quantity and the reference motion state quantity.

[0132] Exemplarily, the preset cost function is:

[0133] (11)

[0134] wherein, are respectively the predicted motion state and the reference state at time t, and L is a pre-specified weighted coefficient matrix, determining the cost magnitude corresponding to different state dimensions.

[0135] S22. Take the second moment as the updated first moment, and take the predicted motion state quantity as the motion state quantity at the updated first moment.

[0136] S23. Return to the step of calculating the motion control amount at the first moment based on the motion state amount at the first moment and the particle until the second preset number of iterations is completed.

[0137] That is: after taking the second moment as the updated first moment and the predicted motion state amount as the motion state amount at the updated first moment, return to execute the step of calculating the motion control amount at the first moment based on the motion state amount at the first moment and the particle until the second preset number of iterations is completed.

[0138] Through the second preset number of iterations, the predicted motion state amounts at multiple moments can be obtained, and the predicted motion state amounts at these multiple moments form a predicted / simulated motion trajectory.

[0139] S24. Calculate the cumulative cost based on all the instantaneous costs during the second preset number of iterations.

[0140] Exemplarily, the sum of all the instantaneous costs during the second preset number of iterations can be used as the cumulative cost.

[0141] S25. Determine the fitness based on the cumulative cost.

[0142] In an alternative embodiment, the cumulative cost can be used as the fitness of the particle.

[0143] In practical applications, in order to further improve the calculation accuracy of the fitness, multiple motion state amount samples can be correspondingly set for each particle, and the multiple motion state amount samples are sampled from a normal distribution with a preset motion state amount as the mean. Among them, the preset motion state amount can be a randomly generated motion state amount, or the historical motion state amount of the water area movable device at a certain historical moment (the motion state amount is an actually collected motion state amount). In the above step S24, "calculating the cumulative cost based on all the instantaneous costs during the second preset number of iterations" may include:

[0144] S241. For each motion state amount sample, calculate the cumulative cost corresponding to the motion state amount sample based on all the instantaneous costs during the second preset number of iterations.

[0145] Among them, for each motion state amount sample, the motion state amount at the first moment in the first iteration process during the second preset number of iterations is the motion state amount sample.

[0146] That is to say, for each motion state amount sample, execute one second preset number of iteration processes.

[0147] In the above step S25, "determining the fitness according to the cumulative cost" can be implemented by the following steps:

[0148] S251. Determine the fitness according to the cumulative cost corresponding to each of the multiple motion state quantity samples.

[0149] The average value of the cumulative costs corresponding to the multiple motion state quantity samples can be used as the fitness.

[0150] (12)

[0151] Among them, is the cumulative cost corresponding to the m-th motion state quantity sample. Among them, T is the second preset number of times, which can also be called the formulated prediction step number.

[0152] In the formula, represents the parameter to be updated of the controller. The smaller the cumulative cost, the closer the predicted running trajectory is to the reference trajectory, and thus the better the current controller parameters. The reference trajectory includes the reference motion state quantities at the multiple moments.

[0153] Taking the water area movable device as a ship as an example, the technical solution provided by the embodiment of the present application realizes the automatic adjustment of the intelligent auxiliary driving function controller by using the ship historical driving data, minimizes the number of interactions between the engineer and the actual ship environment, and can complete the rapid tuning of the intelligent auxiliary driving function controller for ships of different sizes and different types under the condition of very few interactions. Specifically, this solution adopts a model-based reinforcement learning method. Among them, the historical driving data of the ship is first used to train the model to make the model close to the real ship, so as to use the model to replace the real ship, and then the model is used to optimize the controller parameters, so as to obtain the controller parameters that conform to the real ship. Compared with the method in the background technology, this solution only needs to collect a small amount of driving data of the ship, that is, the number of interactions with the actual ship environment is small, so that the rapid tuning of the intelligent auxiliary driving function controller for ships of different sizes and different types can be completed under the condition of very few interactions. And, this solution adjusts the controller parameters through several iteration cycles in the initial operation stage of the ship, and the controller parameters do not need to be adjusted during the subsequent operation process.

[0154] Figure 3 shows a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 3As shown, the electronic device includes a memory 1101 and a processor 1102. The memory 1101 can be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application or method operating on the electronic device. The memory 1101 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks.

[0155] The memory 1101 is used to store programs;

[0156] The processor 1102, coupled to the memory 1101, is used to execute the programs stored in the memory 1101 to implement the methods provided in the above method embodiments.

[0157] Furthermore, as Figure 3 shown, the electronic device further includes other components such as a communication component 1103, a display 1104, a power supply component 1105, an audio component 1106, etc. Figure 3 Only some components are schematically shown in Figure 3 the figure, which does not mean that the electronic device only includes

[0158] the components shown in the figure.

[0159] The embodiments of the present application further provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the steps or functions of the methods provided in the above method embodiments.

[0160] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0161] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM (Read Only Memory), RAM (Random Access Memory), magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A controller optimization method, characterized in that: The controller is used to control the auxiliary driving function of a mobile device in water area, and the method includes: Acquire actual driving data of the movable device in the water area, where the actual driving data is generated by the movable device in the water area moving under the control of the controller; According to the actual driving data, a probabilistic dynamics model of the movable device for water area is obtained by training, wherein the probabilistic dynamics model is used to output a predicted state change amount expressed in a probability distribution form based on the motion state amount and motion control amount of the movable device for water area at a previous moment; Optimizing controller parameters of the controller using a particle swarm algorithm and the probabilistic dynamics model; Using the currently optimized controller parameters to control the operation of the mobile device in the water area, and obtaining actual driving data generated by the mobile device in the water area under the control of the optimized controller; When the optimization process of the controller satisfies a first predetermined condition, determining the currently optimized controller parameters as controller parameters of the target controller; When the optimization process of the controller does not satisfy the first predetermined condition, the process returns to the step of training and obtaining a probabilistic dynamic model of the movable device in the water area according to the actual driving data.

2. The method according to claim 1, characterized in that The actual driving data includes historical motion state quantities and historical motion control quantities of the movable device in the water area at multiple historical moments; the probabilistic dynamics model of the movable device in the water area is trained based on the actual driving data, including: According to the historical motion state quantity and the historical motion control quantity of the first historical moment among the multiple historical moments, using the probabilistic dynamics model, a first predicted state change quantity represented in the form of probability distribution is obtained; Using a preset loss function, calculating a state change difference between the first predicted state change amount and an expected state change amount, wherein the expected state change amount is a state change amount between a historical motion state amount at a second historical moment among the multiple historical moments and a historical motion state amount at the first historical moment, wherein the first historical moment and the second historical moment are any two adjacent moments among the multiple historical moments, and the first historical moment is earlier than the second historical moment; The model parameters of the probabilistic dynamics model are optimized according to the state change difference.

3. The method according to claim 2, characterized in that The probabilistic dynamics model includes a plurality of sub-models, and the initial model parameters of each sub-model are obtained by randomization; The training sample set of each sub-model is obtained by randomly sampling the actual driving data a first preset number of times, and the training sample sets of different sub-models are different. Each sub-model is trained based on the training sample set of the sub-model.

4. The method according to any one of claims 1 to 3, characterized in that The method of optimizing the controller parameters of the controller by using the particle swarm algorithm and the probabilistic dynamics model includes: Based on the plurality of particles in the particle swarm algorithm and the probabilistic dynamics model, predicting the predicted motion state of the movable device in the water area, wherein the position of each particle in the plurality of particles is used as a candidate controller parameter of the controller; determining the fitness of the plurality of particles based on a state difference between the predicted motion state and a reference motion state; When the fitness does not satisfy the second predetermined condition, updating the positions and velocities of the plurality of particles, and returning to the step of predicting the motion state of the movable device in the water area based on the plurality of particles in the particle swarm algorithm and the probabilistic dynamics model; When the fitness satisfies the second predetermined condition, a position corresponding to a particle with the smallest fitness is selected from the multiple particles as the optimized controller parameter.

5. The method according to claim 4, characterized in that The predicted motion state is characterized by a predicted motion state quantity; the predicted motion state of the movable device in the water area is predicted based on the plurality of particles in the particle swarm algorithm and the probabilistic dynamics model, including: For each of the particles, calculating a motion control amount at the first moment based on the motion state amount at the first moment and the particle; Inputting the motion state quantity and the motion control quantity at the first moment into the probabilistic dynamics model to obtain a second predicted state change quantity represented in the form of probability distribution; Extracting a predicted state change amount from the second predicted state change amount as a target predicted state change amount; The predicted motion state quantity at the second moment is calculated based on the motion state quantity at the first moment and the target predicted state change quantity.

6. The method according to claim 5, characterized in that The reference motion state is represented by a reference motion state quantity, and determining the fitness of the plurality of particles based on the state difference between the predicted motion state and the reference state includes: For each of the particles, the predicted motion state quantity and the reference motion state quantity are used as inputs of a preset cost function to calculate an instantaneous cost at the first moment, wherein the instantaneous cost is used to measure the state difference between the predicted motion state quantity and the reference motion state quantity; Taking the second moment as the updated first moment, and taking the predicted motion state quantity as the updated motion state quantity at the first moment; Returning to the step of calculating the motion control amount at the first moment based on the motion state amount at the first moment and the particle, until a second preset number of iterations is completed; Calculate the cumulative cost based on all the instantaneous costs during the second preset number of iterations; The fitness is determined according to the accumulated cost.

7. The method according to claim 6, characterized in that Each of the particles corresponds to a plurality of motion state quantity samples, and the plurality of motion state quantity samples are sampled from a normal distribution with a preset motion state quantity as a mean; Calculating the cumulative cost based on all the instantaneous costs during the second preset number of iterations includes: For each motion state quantity sample, the cumulative cost corresponding to the motion state quantity sample is calculated based on all the instantaneous costs in the second preset number of iterations, and for each motion state quantity sample, the motion state quantity at the first moment in the first iteration in the second preset number of iterations is the motion state quantity sample; Determining the fitness according to the accumulated cost includes: The fitness is determined according to the accumulated costs corresponding to each of the plurality of motion state quantity samples.

8. The method according to claim 6, characterized in that The second predetermined condition includes at least one of the following: The number of update iterations of the fitness of the plurality of particles reaches a fourth preset number; In the fifth preset number of consecutive update iterations of the fitness of any particle, the difference between the fitness obtained by any two update iterations is less than the first preset value.

9. The method according to claim 6, characterized in that The first predetermined condition includes at least one of the following: The number of iterations of the optimization process of the controller reaches a sixth preset number; In the seventh consecutive preset number of iterations in the optimization process of the controller, the difference between the actual cumulative costs corresponding to the controller parameters obtained in any two iterations is less than the second preset value, and the actual cumulative cost corresponding to the controller parameters obtained in any one iteration is calculated based on the actual driving data generated by controlling the movement of the movable equipment in the water area using the controller obtained in this iteration and the preset cost function.

10. An electronic device, characterized in that: include: A memory and a processor, wherein The memory is used to store programs; The processor is coupled to the memory, and is used to execute the program stored in the memory to implement the controller optimization method according to any one of claims 1 to 9.

11. A movable device in water area, characterized in that: The electronic device comprising claim 10.

12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a computer, the controller optimization method according to any one of claims 1 to 9 can be implemented.