Parameter determination model training, parameter determination method, device, equipment and medium

CN117422126BActive Publication Date: 2026-09-22CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311417577.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2026-09-22
Estimated Expiration
2043-10-27

AI Technical Summary

Technical Problem

波浪能转换器提供了一种将海洋中的波浪能转换为电能的方法,但在波浪能转换器的使用过程中,会存在来波相位与浮子运动相位不匹配的问题,导致波浪能吸收效率降低,而通过取力器(Power Take Off,PTO)装置对浮子的运动状态进行主动干预,则可以提高波浪能转换器的能量吸收效率

Benefits of technology

[0026]本公开实施例提供的技术方案与现有技术相比具有如下优点:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117422126B_ABST
    Figure CN117422126B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a parameter determination model training, a parameter determination method, an apparatus, a device and a medium. The method comprises: obtaining training sample data; based on the training sample data, obtaining a second winch force gradient parameter at a current time through a parameter output network in a parameter determination model; based on the training sample data and the second winch force gradient parameter at the current time, obtaining an evaluation parameter through an evaluation network in the parameter determination model; calculating a first loss of the parameter output network; calculating a second loss of the evaluation network; and adjusting model parameters of the parameter determination model based on the first loss and the second loss, so that the first loss and the second loss converge. By training the parameter determination model, the present disclosure can fully consider environmental factors when determining the winch force gradient parameter, improve the accuracy of parameter determination, and improve energy absorption efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a parameter determination model training, parameter determination method, apparatus, device and medium. Background Technology

[0002] Ocean wave energy is an abundant and renewable clean energy source, with currently proven reserves estimated at between 0.2 TW and 10 TW. This means that even the lowest proven level of 0.2 TW could meet a significant portion of global energy demand. Wave energy converters offer a method to convert ocean wave energy into electricity. However, during their operation, a mismatch between the incoming wave phase and the buoy's motion phase can occur, leading to reduced wave energy absorption efficiency. Actively intervening in the buoy's motion using a power take-off (PTO) device can improve the energy absorption efficiency of the wave energy converter. However, existing PTO control methods often rely on pre-modeling to determine a PTO force gradient parameter, which is then used to control the buoy's motion. But discrepancies frequently exist between the actual situation and the pre-established model, resulting in low final energy absorption efficiency. Summary of the Invention

[0003] To address the aforementioned technical problems, this disclosure provides a parameter determination model training, parameter determination method, apparatus, device, and medium.

[0004] A first aspect of this disclosure provides a parameter determination model training method, the method comprising:

[0005] Acquire training sample data, which includes wave characteristic parameters, float motion parameters, and the first power take-off force gradient parameters of the previous moment;

[0006] Based on the training sample data, the parameter output network in the parameter determination model is used to obtain the force gradient parameters of the second force take-off device at the current moment;

[0007] Based on the training sample data and the force gradient parameters of the second power take-off at the current moment, the evaluation network in the model is determined through the parameters to obtain the evaluation parameters;

[0008] Based on the training sample data, the second force take-off force gradient parameter at the current moment, the evaluation parameter, and the preset first loss function, calculate the first loss of the parameter output network;

[0009] Based on the evaluation parameters and the preset second loss function, the second loss of the evaluation network is calculated;

[0010] The model parameters of the parameter determination model are adjusted based on the first loss and the second loss to make the first loss and the second loss converge.

[0011] A second aspect of this disclosure provides a parameter determination method, the method comprising:

[0012] Acquire the target wave characteristic parameters, the target buoy motion parameters, and the first target power take-off force gradient parameters from the previous moment;

[0013] The target wave characteristic parameters, the target float motion parameters, and the first target power take-off force gradient parameters of the previous moment are input into the parameter determination model to obtain the second target power take-off force gradient parameters of the current moment output by the parameter determination model. The parameter determination model is trained by the training method described in the first aspect above.

[0014] A third aspect of this disclosure provides a parameter determination model training apparatus, the apparatus comprising:

[0015] The data acquisition module is used to acquire training sample data, which includes wave characteristic parameters, float motion parameters, and the first power take-off force gradient parameters of the previous moment.

[0016] The output module is used to obtain the second force take-off force gradient parameters at the current moment by using the parameter determination network in the model based on the training sample data.

[0017] The evaluation module is used to determine the evaluation network in the model based on the training sample data and the force gradient parameters of the second force take-off at the current moment, and to obtain the evaluation parameters.

[0018] The first loss determination module is used to calculate the first loss of the parameter output network based on the training sample data, the second force take-off force gradient parameter at the current moment, the evaluation parameter and the preset first loss function.

[0019] The second loss determination module is used to calculate the second loss of the evaluation network based on the evaluation parameters and a preset second loss function.

[0020] An adjustment module is used to adjust the model parameters of the parameter-determined model based on the first loss and the second loss, so that the first loss and the second loss converge.

[0021] A fourth aspect of the present disclosure provides a parameter determining apparatus, the apparatus comprising:

[0022] The parameter acquisition module is used to acquire target wave characteristic parameters, target buoy motion parameters, and the first target power take-off force gradient parameters of the previous moment.

[0023] The parameter determination module is used to input the target wave characteristic parameters, the target float motion parameters, and the first target power take-off force gradient parameters of the previous moment into the parameter determination model to obtain the second target power take-off force gradient parameters of the current moment output by the parameter determination model. The parameter determination model is trained by the training method described in the first aspect above.

[0024] A fifth aspect of this disclosure provides a computer device including a memory and a processor, and a computer program, wherein the memory stores the computer program, and when the computer program is executed by the processor, it implements the parameter determination model training method of the first aspect or the parameter determination method of the second aspect described above.

[0025] A sixth aspect of this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the parameter determination model training method of the first aspect or the parameter determination method of the second aspect described above.

[0026] The technical solution provided in this disclosure has the following advantages compared with the prior art:

[0027] In this embodiment, training sample data is acquired, including wave feature parameters, buoy motion parameters, and the first power take-off force gradient parameters from the previous moment. Based on the training sample data, the second power take-off force gradient parameters for the current moment are obtained through the parameter output network in the parameter determination model. Based on the training sample data and the second power take-off force gradient parameters for the current moment, the evaluation network in the parameter determination model is used to obtain evaluation parameters. Based on the training sample data, the second power take-off force gradient parameters for the current moment, the evaluation parameters, and a preset first loss function, the first loss of the parameter output network is calculated. Based on the evaluation parameters and the preset second loss function, the second loss of the evaluation network is calculated. Based on the first loss and the second loss, the model parameters of the parameter determination model are adjusted to make the first loss and the second loss converge. This allows the second power take-off force gradient parameters for the current moment to be determined based on environmental factors such as wave feature parameters and buoy motion parameters, combined with the first power take-off force gradient parameters from the previous moment, when constructing the parameter determination model, so that the finally determined power take-off force gradient parameters are adapted to the actual situation. At the same time, the evaluation network is introduced during the model training process to improve the accuracy of the model. Attached Figure Description

[0028] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0029] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a flowchart of a parameter determination model training method provided in an embodiment of this disclosure;

[0031] Figure 2 This is a flowchart of a method for obtaining training sample data provided in an embodiment of this disclosure;

[0032] Figure 3 This is a flowchart of a method for calculating a first loss provided in an embodiment of this disclosure;

[0033] Figure 4 This is a flowchart of a method for calculating a second loss provided in an embodiment of this disclosure;

[0034] Figure 5 This is a flowchart of a method for determining reward evaluation parameters provided in an embodiment of this disclosure;

[0035] Figure 6 This is a flowchart of a parameter determination method provided in an embodiment of this disclosure;

[0036] Figure 7 This is a schematic diagram of the structure of a parameter determination model training device provided in an embodiment of this disclosure;

[0037] Figure 8 This is a schematic diagram of the structure of a parameter determination device provided in an embodiment of this disclosure;

[0038] Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure. Detailed Implementation

[0039] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0040] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0041] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0042] Figure 1 This is a flowchart illustrating a parameter determination model training method provided in an embodiment of this disclosure. This method can be executed by a parameter determination model training device, which can be located in an electronic device. For example, the electronic device can be a portable mobile device such as a laptop computer, a personal digital assistant (PDA), or a wearable device for workers; it can also be a fixed device such as a personal computer, smart home appliance, or server. The server can be a single server, a server cluster, or a distributed or centralized cluster. Figure 1 As shown, the parameter determination model training method provided in this embodiment includes the following steps:

[0043] S101. Obtain training sample data, which includes wave characteristic parameters, float motion parameters, and the first power take-off force gradient parameters of the previous moment.

[0044] The wave characteristic parameters in this disclosure are parameters used to characterize the wave characteristics of ocean states. These parameters may include the wave height of the incoming wave, the amount of change in wave height in the vertical direction, and the rate of change. In some embodiments, the amount of change in wave height in the vertical direction and the rate of change of wave height can be calculated based on the current wave height of the incoming wave.

[0045] In some embodiments of this disclosure, the parameter determination model training device can calculate the rate of change of the wave height in the vertical direction of the incoming wave using the wave height of the incoming wave acquired by wave acquisition and the sampling time interval. Specifically, V can be used. w =(H t -H t-1 ) / t s The velocity V of the wave height in the vertical direction was calculated. w H t H represents the wave height data collected at the current moment. t-1 This represents the wave height data collected at the previous moment. The sampling time interval for each data point is expressed in t. s This indicates the sampling interval between two sets of data.

[0046] The float motion parameters in this embodiment are parameters used to characterize the current motion state of the float in the device. The float motion parameters may include the float's position offset relative to its initial position and its motion velocity parameters. In some embodiments, the float's motion velocity parameters can be calculated based on the float's position offset parameters.

[0047] In some embodiments of this disclosure, the parameter determination model training device can calculate the buoy's motion velocity parameters by collecting the buoy's position offset parameters. Specifically, V can be used. b =(D t -D t-1 ) / t s The velocity parameter V of the float was calculated. b D t D represents the vertical position offset parameter of the float acquired at the current moment. t-1 This represents the vertical position offset parameter of the float acquired at the previous moment. The sampling time interval for each data session is expressed in terms of t. s express.

[0048] In this embodiment of the disclosure, the force gradient parameter of the power take-off can be understood as the rate of change of the control force output by the power take-off. The first force gradient parameter of the power take-off can be understood as the force gradient parameter of the power take-off at the previous moment. The first force gradient parameter of the power take-off at the initial moment is 0. The first force gradient parameter of the power take-off at any moment after the initial moment is the second force gradient parameter of the power take-off output by the parameter output network at the previous moment of that arbitrary moment.

[0049] In this embodiment of the present disclosure, the parameter determination model training device can pre-collect data including wave characteristic parameters, float motion parameters and the first power take-off force gradient parameters of the previous moment, and determine them as training sample data.

[0050] S102. Based on the training sample data, the parameter output network in the parameter determination model is used to obtain the force gradient parameters of the second force take-off at the current moment.

[0051] The parameter determination model in this embodiment can employ various deep reinforcement learning neural networks, such as Deep Q-Networks (DQN) and its derivative Double DQN network, Proximal Policy Optimization (PPO) network, Twin Delayed Deep Deterministic Policy Gradient (TD3) network, and Soft Actor Critic (SAC) network, etc., without limitation.

[0052] The parameter output network in this embodiment can be understood as a neural network included in the parameter determination model, used to determine the second power take-off force gradient parameter at the current moment based on wave characteristic parameters, float motion parameters and the first power take-off force gradient parameter at the previous moment. The parameter output network and the evaluation network together constitute the parameter determination model.

[0053] In this embodiment of the disclosure, the parameter determination model training device can input the training sample data into the parameter output network after obtaining the training sample data, and the parameter output network can process the training sample data to obtain the second force take-off force gradient parameter at the current moment.

[0054] S103. Based on the training sample data and the force gradient parameters of the second force take-off device at the current moment, the evaluation network in the model is determined through the parameters to obtain the evaluation parameters.

[0055] The evaluation network in this embodiment can be understood as a neural network that evaluates the force gradient parameters of the second power take-off at the current moment, taking into account future effects. The evaluation parameters can be understood as the evaluation results. For example, the larger the value of the evaluation parameters, the better the evaluation result.

[0056] In this embodiment of the present disclosure, the parameter determination model training device can, after obtaining the second force take-off force gradient parameters output by the parameter output network, input the training sample data and the second force take-off force gradient parameters together into the evaluation network, and the evaluation network evaluates the training sample data and the second force take-off force gradient parameters to obtain the evaluation parameters.

[0057] S104. Based on the training sample data, the force gradient parameters of the second power take-off at the current moment, the evaluation parameters, and the preset first loss function, calculate the first loss of the parameter output network.

[0058] In this embodiment of the present disclosure, the parameter determination model training device can, after obtaining the second force take-off force gradient parameters output by the parameter output network and the evaluation parameters output by the evaluation network, substitute the training sample data, the second force take-off force gradient parameters and the evaluation parameters into a preset first loss function to calculate the first loss corresponding to the parameter output network.

[0059] S105. Calculate the second loss of the evaluation network based on the evaluation parameters and the preset second loss function.

[0060] In this embodiment of the present disclosure, the parameter determination model training device can, after obtaining the evaluation parameters output by the evaluation network, substitute the evaluation parameters into a preset second loss function to calculate the second loss corresponding to the evaluation network.

[0061] S106. Adjust the model parameters of the parameter-determined model based on the first loss and the second loss to make the first loss and the second loss converge.

[0062] In this embodiment of the disclosure, the parameter determination model training device can adjust the model parameters of the parameter determination model based on the first loss and the second loss after obtaining the first loss and the second loss. Specifically, the model parameters of the parameter output network can be adjusted based on the first loss, and the model parameters of the evaluation network can be adjusted based on the second loss. The gradient descent method is used to update the parameter output network and the evaluation network with the goal of reducing the first loss and the second loss, so that the first loss and the second loss eventually converge. At this time, it can be determined that the parameter determination model training is complete.

[0063] In one exemplary embodiment of this disclosure, the parameter determination model training device can normalize the training sample data after obtaining it, thereby better eliminating training errors that may be caused by data of different scales, accelerating model training, and improving the final training effect. Specifically, the parameter determination model training device can achieve normalization by mapping the training sample data to the [0,1] interval range using the following formula:

[0064]

[0065] in, These are the normalized training sample data, where x is the collected training sample data. min It is the minimum value corresponding to various types of training sample data, x. max This refers to the maximum value corresponding to various preset training sample data. Alternatively, the normalization of training sample data to the [-1, 1] interval can be achieved using the following formula:

[0066]

[0067] After obtaining the normalized data corresponding to each set of training sample data, the parameter determination model training device can use the normalized data of each set of training samples to train the pre-built parameter determination model, obtain the second force take-off force gradient parameters corresponding to the normalized data of each set of training samples, and execute the steps in S103-S106 to realize the training of the parameter determination model.

[0068] This embodiment of the disclosure acquires training sample data, including wave characteristic parameters, buoy motion parameters, and the first power take-off force gradient parameters of the previous moment. Based on the training sample data, a parameter output network in the parameter determination model is used to obtain the second power take-off force gradient parameters of the current moment. Based on the training sample data and the second power take-off force gradient parameters of the current moment, an evaluation network in the parameter determination model is used to obtain evaluation parameters. Based on the training sample data, the second power take-off force gradient parameters of the current moment, the evaluation parameters, and a preset first loss function, a first loss of the parameter output network is calculated. Based on the evaluation parameters and a preset second loss function, a second loss of the evaluation network is calculated. Based on the first loss and the second loss, the model parameters of the parameter determination model are adjusted to make the first loss and the second loss converge. This allows the second power take-off force gradient parameters of the current moment to be determined based on environmental factors such as wave characteristic parameters and buoy motion parameters, combined with the first power take-off force gradient parameters of the previous moment, when constructing the parameter determination model, so that the finally determined power take-off force gradient parameters are adapted to the actual situation. At the same time, the evaluation network is introduced during the model training process to improve the accuracy of the model.

[0069] Figure 2 This is a flowchart of a method for obtaining training sample data provided in an embodiment of this disclosure, such as... Figure 2 As shown, based on the above embodiments, training sample data can be obtained through the following methods.

[0070] S201. Obtain wave characteristic parameters detected by wave data detection equipment and float motion parameters detected by motion state detection equipment.

[0071] In this embodiment of the disclosure, the parameter determination model training device can detect wave feature parameters through wave data detection equipment and detect float motion parameters through motion state detection equipment, and obtain the detected wave feature parameters and float motion parameters.

[0072] S202. Based on the wave characteristic parameters detected by the wave data detection equipment, the wave generation parameters and wave dissipation parameters corresponding to the target sea state are determined by computational fluid dynamics.

[0073] In this embodiment of the disclosure, the target sea state can be understood as any sea state other than the sea state corresponding to the measured wave characteristic parameters that needs to be simulated. Wave-generating parameters can be understood as parameters used to control the wave state simulated in the numerical water tank under a certain sea state, and wave-dissipating parameters can be understood as parameters corresponding to the wave-generating parameters used to control the elimination of waves in the numerical water tank. To ensure the accuracy of data measurement, after simulating sea state waves using wave-generating parameters and measuring relevant data, wave-dissipating parameters are needed to eliminate waves in the numerical water tank before the next wave generation.

[0074] In this embodiment of the disclosure, the parameter determination model training device can construct a sea state wave simulation model based on a large number of measured wave characteristic parameters after obtaining the wave characteristic parameters detected by the wave data detection device, and use computational fluid dynamics (CFD) to determine the wave generation parameters and wave dissipation parameters corresponding to various target sea states through the sea state wave simulation model.

[0075] S203. Obtain wave characteristic parameters and float motion parameters collected from the numerical water tank in the target state, where the target state is the state of the numerical water tank when the wave generation parameters are used to control the numerical water tank.

[0076] In this embodiment of the present disclosure, the parameter determination model training device can, after determining the wave generation parameters and wave dissipation parameters, acquire wave characteristic parameters and float motion parameters collected from the numerical water tank when the numerical water tank is in the target state under the control of the wave generation parameters. Specifically, wave characteristic parameters and float motion parameters can be collected by placing wave data detection equipment and motion state detection equipment in the numerical water tank.

[0077] This embodiment of the disclosure acquires wave characteristic parameters detected by wave data detection equipment and float motion parameters detected by motion state detection equipment. Based on the wave characteristic parameters detected by wave data detection equipment, computational fluid dynamics is used to determine the wave generation parameters and wave dissipation parameters corresponding to the target sea state. Wave characteristic parameters and float motion parameters are collected from the numerical water tank under the target state, which is the state of the numerical water tank when the wave generation parameters are used to control the numerical water tank. This not only allows the acquisition of training sample data through detection equipment, but also allows the acquisition of more data that is difficult to detect in the real marine environment by simulating sea states, thereby expanding the model training set and improving the model training effect.

[0078] Figure 3 This is a flowchart of a method for calculating a first loss provided in an embodiment of this disclosure, as shown below. Figure 3 As shown, based on the above embodiments, the first loss can be calculated by the following method, wherein the evaluation network includes a first evaluation network and a second evaluation network.

[0079] S301. Input the training sample data and the force gradient parameters of the second power take-off at the current moment into the first evaluation network and the second evaluation network respectively to obtain the first evaluation parameters and the second evaluation parameters.

[0080] The first evaluation network and the second evaluation network in this embodiment can be understood as two sets of evaluation networks set up to avoid over-selection of parameters and getting trapped in local optima. Optionally, the first evaluation network and the second evaluation network can be the same or different in structure, but they use different model parameters.

[0081] In this embodiment of the present disclosure, the parameter determination model training device can input the training sample data and the force gradient parameters of the second power take-off device into the first evaluation network and the second evaluation network respectively when obtaining the evaluation parameters through the evaluation network, so as to obtain the first evaluation parameters and the second evaluation parameters output by the first evaluation network and the second evaluation network.

[0082] S302. The evaluation parameter with the smallest value among the first evaluation parameter and the second evaluation parameter is determined as the target evaluation parameter.

[0083] In this embodiment of the present disclosure, the parameter determination model training device can compare the first evaluation parameter and the second evaluation parameter after obtaining them, and select the evaluation parameter with the smallest value among the first evaluation parameter and the second evaluation parameter as the target evaluation parameter.

[0084] S303. Calculate the first loss based on the training sample data, the force gradient parameters of the second power take-off at the current moment, the target evaluation parameters, and the preset first loss function.

[0085] In this embodiment of the disclosure, the parameter determination model training device can, after determining the target evaluation parameters, substitute the training sample data, the force gradient parameters of the second force take-off device, and the target evaluation parameters into a preset first loss function to calculate the first loss. The specific calculation method is as follows:

[0086]

[0087] Among them, Loss a Let S represent the first loss, N represent the number of training sample data or corresponding second force gradient parameters contained in a training batch, and S represent the second loss. i Let a' represent the training sample data. i π represents the force gradient parameter of the second power take-off unit. θ (a′ i |S i ) indicates in S i In the case of a′ i The probability, where α represents the weight, can be a set value. Q(x) minThis represents the target evaluation parameter of the evaluation network based on the training sample data and the corresponding second force take-off force gradient parameter x.

[0088] This embodiment of the disclosure obtains a first evaluation parameter and a second evaluation parameter by inputting training sample data and the force gradient parameter of the second power take-off at the current moment into the first evaluation network and the second evaluation network, respectively. The evaluation parameter with the smaller value between the first evaluation parameter and the second evaluation parameter is determined as the target evaluation parameter. Based on the training sample data, the force gradient parameter of the second power take-off at the current moment, the target evaluation parameter, and the preset first loss function, the first loss is calculated. When calculating the first loss of the parameter output network, the smaller target evaluation parameter between the evaluation parameters output by the first evaluation network and the second evaluation network is used for calculation. Therefore, when training the parameter output network according to the first loss, the over-selection of parameters is avoided, which may lead to getting trapped in local optima.

[0089] In some embodiments of this disclosure, the parameter determination model training device can perform soft updates on the first and second evaluation networks after the evaluation network has been trained and updated for a certain number of rounds based on the second loss. The specific method is as follows:

[0090] Q′ i ←τQ′ i +(1-τ)Q i i = 1, 2

[0091] Among them, Q′ i Let Q represent the i-th final evaluation network. i Let represent the i-th evaluation network, and τ represent the soft update parameter, which can be a set value.

[0092] The initial state of the final evaluation network is the same as that of the evaluation network. During the training of the evaluation network, there are differences between the final evaluation network and the evaluation network after soft updates.

[0093] Figure 4 This is a flowchart illustrating a method for calculating a second loss provided in an embodiment of this disclosure. Figure 4 As shown, based on the above embodiments, the second loss can be calculated using the following method.

[0094] S401. Calculate the reward evaluation parameters based on the training sample data and the force gradient parameters of the second power take-off at the current moment.

[0095] In this embodiment, the reward evaluation parameter can be understood as a parameter used to evaluate the quality of the second power take-off force gradient parameter output by the model at the current moment. It can be calculated based on the power generation data of the wave energy converter, and the power generation can be calculated based on the power take-off force gradient parameter. For example, a larger value for the reward evaluation parameter indicates better performance of the current model.

[0096] In this embodiment of the present disclosure, the parameter determination model training device can calculate the reward evaluation parameters based on the second force take-off force gradient parameters at the current moment and the training sample data after obtaining the second force take-off force gradient parameters at the current moment of the model output.

[0097] S402. For the first evaluation network, calculate the second loss corresponding to the first evaluation network based on the first evaluation parameters, the reward evaluation parameters, and the second loss function.

[0098] In this embodiment of the present disclosure, the parameter determination model training device can, when calculating the second loss corresponding to the first evaluation network, calculate the second loss corresponding to the first evaluation network based on the first evaluation parameters, the reward evaluation parameters, and the second loss function.

[0099] In an exemplary embodiment of this disclosure, the parameter determination model training device can input training sample data and the force gradient parameters of the second force take-off device into the first final evaluation network and the second final evaluation network, respectively, to obtain the first final evaluation parameter and the second final evaluation parameter. The device then selects the final evaluation parameter with the smaller value from the first and second final evaluation parameters and determines it as the target final evaluation parameter. When calculating the second loss, for the first evaluation network, the first evaluation parameter, the target final evaluation parameter, and the reward evaluation parameter are substituted into the second loss function to calculate the second loss corresponding to the first evaluation network. The specific calculation method is as follows:

[0100]

[0101] Among them, Loss c1 S represents the second loss corresponding to the first evaluation network, N represents the number of training sample data or corresponding second force take-off force gradient parameters and reward evaluation parameters contained in a training batch, and S represents the second loss corresponding to the first evaluation network. i Represents the training sample data, a i R represents the force gradient parameter of the second power take-off unit. i Let Q1(x) represent the reward evaluation parameters, Q′ represent the first evaluation parameters of the first evaluation network for the training sample data and the corresponding force gradient parameters x of the second force take-off, ε represent the preset discount factor, and Q′ represent the reward evaluation parameters. min (S t ,a t ) represents the final evaluation parameter of the objective, [R t +εQ′ min (S t ,a t )] 2 This represents the scoring parameter considering the entire control process, based on the training sample data and the corresponding force gradient parameters of the second power take-off. Therefore, Lossc1 This can also be understood as the error between the rating parameters predicted by the first evaluation network and the actual rating parameters.

[0102] S403. For the second evaluation network, calculate the second loss corresponding to the second evaluation network based on the second evaluation parameters, reward evaluation parameters, and the second loss function.

[0103] In this embodiment of the present disclosure, the parameter determination model training device can calculate the second loss corresponding to the second evaluation network based on the second evaluation parameters, the reward evaluation parameters, and the second loss function when calculating the second loss corresponding to the second evaluation network.

[0104] In an exemplary embodiment of this disclosure, the parameter determination model training device can, when calculating the second loss corresponding to the second evaluation network, substitute the second evaluation parameters, the target final evaluation parameters, and the reward evaluation parameters into the second loss function to calculate the second loss corresponding to the second evaluation network. The specific calculation method is similar to S402:

[0105]

[0106] Among them, Loss c2 Let Q2(x) represent the second loss corresponding to the second evaluation network, and let Q2(x) represent the second evaluation parameter of the second evaluation network based on the training sample data and the corresponding force gradient parameter x of the second force take-off. c2 This can also be understood as the error between the rating parameters predicted by the second evaluation network and the actual rating parameters.

[0107] This embodiment calculates reward evaluation parameters based on training sample data and the force gradient parameters of the second power take-off at the current moment. For the first evaluation network, a second loss is calculated based on the first evaluation parameters, the reward evaluation parameters, and the second loss function. Similarly, for the second evaluation network, a second loss is calculated based on the second evaluation parameters, the reward evaluation parameters, and the second loss function. This allows the reward evaluation parameters to be introduced when calculating the second loss of the evaluation network. The second loss is calculated by combining the output results of the current round and multiple output results throughout the training process. The evaluation network is adjusted by the error between the predicted scoring parameters and the actual scoring parameters, making the evaluation network more accurate and thus helping to improve the accuracy of the parameter output network.

[0108] Figure 5 This is a flowchart of a method for determining reward evaluation parameters provided in an embodiment of this disclosure, such as... Figure 5 As shown, based on the above embodiments, the reward evaluation parameters can be determined by the following method.

[0109] S501. Based on the force gradient parameters of the second power take-off at the current moment, calculate the power take-off control parameters at the current moment.

[0110] The power take-off control parameters in this embodiment can be understood as control parameters used to control the power take-off to perform corresponding actions, thereby controlling the motion state of the float.

[0111] In this embodiment of the disclosure, the parameter determination model training device can determine the power take-up control parameters at the current moment based on the conversion relationship between the power take-up force gradient parameters and the power take-up control parameters after obtaining the second power take-up force gradient parameters output by the parameter output network at the current moment. The specific calculation method is as follows:

[0112]

[0113] Among them, f t+1 f represents the power take-off control parameters at the current moment. t α represents the power take-off control parameters at the previous moment. t and t s These represent the second power take-off force gradient parameter and the sampling time interval at the current moment, respectively. Similar to the power take-off force gradient parameter, the power take-off control parameter has a value of 0 at the initial moment.

[0114] S502. Determine whether the float motion parameters and the current power take-off control parameters meet the preset conditions.

[0115] The preset conditions in this embodiment can be understood as pre-set judgment conditions used to determine whether the float motion state and power take-off control parameters have deviated significantly.

[0116] In this embodiment of the present disclosure, the parameter determination model training device can determine whether the float motion parameters and the current power take-off control parameters meet preset conditions. Specifically, the position offset parameter of the float relative to its initial position can be compared with the float offset threshold, the motion velocity parameter of the float can be compared with the float velocity threshold, and the power take-off control parameters can be compared with the power take-off control parameter threshold. If all three conditions are met—position offset parameter less than float offset threshold, motion velocity parameter less than float velocity threshold, and power take-off control parameter less than power take-off control parameter threshold—it can be determined that the float motion parameters and the current power take-off control parameters meet the preset conditions. If at least one of the above three conditions cannot be met, it is determined that the float motion parameters and the current power take-off control parameters do not meet the preset conditions.

[0117] S503. If satisfied, the reward evaluation parameters are calculated based on the float motion parameters, the current power take-off control parameters, and the change in the power take-off force gradient parameters.

[0118] In this embodiment of the present disclosure, the parameter determination model training device can, when determining that the float motion parameters and the current power take-off control parameters meet preset conditions, calculate the change in the power take-off force gradient parameters based on the first power take-off force gradient parameters from the previous moment included in the training sample data and the second power take-off force gradient parameters from the model output at the current moment. Then, based on the float motion parameters, the current power take-off control parameters, and the change in the power take-off force gradient parameters, the reward evaluation parameters are calculated. The specific calculation method is as follows:

[0119] R i =-f PTO d f -τ1|a i -a i-1 |-τ2|f PTO |

[0120] Among them, R i f represents the reward evaluation parameter. PTO d represents the power take-off control parameters at the current moment. f a represents the velocity parameter of the float. i The second power take-off force gradient parameter, a, represents the current moment. i-1 |a| represents the force gradient parameter of the first power take-off unit at the previous moment. i -a i-1 | represents the change in the force gradient parameter of the power take-off unit. Parameters τ1 and τ2 are the sensitivity coefficients for controlling the mutation penalty term and the bias penalty term, respectively, and can be set values.

[0121] Specifically, -f PTO d f This parameter directly evaluates the improvement in wave energy absorption efficiency of the wave energy converter, characterizing the absorption effect of the second power take-off (PTO) force gradient parameter on the incoming wave energy at the current moment. A positive value indicates that the PTO force gradient parameter has a positive effect on wave absorption, while a negative value indicates a negative effect, or a less than ideal effect. -τ1|a i -a i-1 | represents the evaluation of the sudden change in the force gradient parameter of the power take-off. It can only be negative. The larger the sudden change, the more negative the evaluation of the force gradient parameter of the second power take-off at the current moment.

[0122] S504. If not satisfied, the value of the reward evaluation parameter is set to a preset value.

[0123] The preset value in this embodiment can be understood as a preset value used to characterize the negative evaluation of the force gradient parameter of the second power take-off. Since the larger the value of the reward evaluation parameter, the better the model performance, the preset value can be set to a small value close to 0, such as 0.01, which is not limited here.

[0124] In this embodiment of the present disclosure, the parameter determination model training device can determine the value of the reward evaluation parameter as a preset value when the float motion parameters and the current power take-off control parameters meet preset conditions, namely, at least one of the following conditions: the position offset parameter is greater than or equal to the float offset threshold, the motion speed parameter is greater than or equal to the float speed threshold, and the power take-off control parameter is greater than or equal to the power take-off control parameter threshold.

[0125] This embodiment calculates the power take-off control parameters at the current moment based on the second power take-off force gradient parameters at the current moment. It then determines whether the float motion parameters and the power take-off control parameters at the current moment meet preset conditions. If they do, the reward evaluation parameters are calculated based on the float motion parameters, the power take-off control parameters at the current moment, and the change in the power take-off force gradient parameters. If they do not meet the conditions, the reward evaluation parameters are set to preset values. This allows for the determination of accurate reward evaluation parameters, which are then used to train the model and further improve its accuracy.

[0126] In some embodiments of this disclosure, when adjusting the model parameters of the parameter determination model training device, it can further determine the training status of the model based on the reward evaluation parameters, in addition to determining the model training status based on the convergence of the first loss and the second loss. Specifically, it can calculate the sum of multiple reward evaluation parameters corresponding to multiple output results of the model in one round of training, and determine whether the accumulated value of the reward evaluation parameters in the latest round of training has converged to a stable value based on the multiple summation results corresponding to multiple rounds of training. If the accumulated value of the reward evaluation parameters converges to a stable value, it is determined that the parameter determination model training is complete.

[0127] Figure 6 This is a flowchart of a parameter determination method provided in an embodiment of this disclosure, which can be executed by a parameter determination device. Figure 6 As shown, the parameter determination method provided in this embodiment includes the following steps:

[0128] S601. Obtain the target wave characteristic parameters, target buoy motion parameters, and the first target power take-off force gradient parameters of the previous moment.

[0129] The target wave characteristic parameters in this embodiment can be understood as wave characteristic parameters acquired in real time, and the target buoy motion parameters are buoy motion parameters acquired in real time.

[0130] In this embodiment of the disclosure, the parameter determination device can acquire the target wave characteristic parameters and target float motion parameters detected in real time, as well as the first target power take-off force gradient parameters used when controlling the power take-off at the previous moment.

[0131] S602. Input the target wave characteristic parameters, target float motion parameters and the first target power take-off force gradient parameters of the previous moment into the parameter determination model to obtain the second target power take-off force gradient parameters of the current moment output by the parameter determination model.

[0132] In this embodiment of the disclosure, the parameter determination model is trained using the parameter determination model training method described in the above embodiments.

[0133] In this embodiment of the present disclosure, after obtaining the target wave characteristic parameters, the target float motion parameters, and the first target power take-off force gradient parameters of the previous moment, the parameter determination device can input them into the trained parameter determination model, which will then process the input data to obtain output data and determine the output data as the second target power take-off force gradient parameters of the current moment.

[0134] In one exemplary embodiment of the present disclosure, the parameter determination device can, after determining the second target power take-off force gradient parameter at the current moment, calculate the target power take-off control parameter at the current moment based on the second target power take-off force gradient parameter, and then convert the target power take-off control parameter into a control signal and input the control signal into the power take-off so that the power take-off performs the action corresponding to the control signal, controls the motion state of the float, thereby improving the energy absorption efficiency of the wave energy converter.

[0135] This embodiment of the present disclosure obtains the target wave characteristic parameters, the target buoy motion parameters, and the first target power take-off force gradient parameters of the previous moment. The target wave characteristic parameters, the target buoy motion parameters, and the first target power take-off force gradient parameters of the previous moment are input into the parameter determination model to obtain the second target power take-off force gradient parameters of the current moment output by the parameter determination model. This can determine more accurate power take-off force gradient parameters by combining environmental factors, so that the matching degree between the buoy motion phase controlled by the parameter and the incoming wave phase is higher, thereby improving the energy absorption efficiency of wave energy.

[0136] Figure 7 This is a schematic diagram of the structure of a parameter determination model training device provided in an embodiment of this disclosure. Figure 7As shown, the parameter determination model training device 700 includes: a data acquisition module 710, an output module 720, an evaluation module 730, a first loss determination module 740, a second loss determination module 750, and an adjustment module 760. The data acquisition module 710 is used to acquire training sample data, which includes wave characteristic parameters, buoy motion parameters, and the first power take-off force gradient parameters from the previous moment. The output module 720 is used to obtain the second power take-off force gradient parameters at the current moment based on the training sample data and through the parameter output network in the parameter determination model. The evaluation module 730 is used to evaluate the second power take-off force gradient parameters at the current moment based on the training sample data and the current... The second force gradient parameter of the power take-off at the current time is used to determine the evaluation network in the model, thereby obtaining the evaluation parameters; the first loss determination module 740 is used to calculate the first loss of the parameter output network based on the training sample data, the second force gradient parameter of the power take-off at the current time, the evaluation parameters, and a preset first loss function; the second loss determination module 750 is used to calculate the second loss of the evaluation network based on the evaluation parameters and a preset second loss function; the adjustment module 760 is used to adjust the model parameters of the parameter determination model based on the first loss and the second loss, so that the first loss and the second loss converge.

[0137] Optionally, the data acquisition module 710 is specifically used to acquire wave characteristic parameters detected by the wave data detection device and float motion parameters detected by the motion state detection device.

[0138] Optionally, the data acquisition module 710 includes: a first determining unit, used to determine the wave-generating parameters and wave-dissipating parameters corresponding to the target sea state based on the wave characteristic parameters detected by the wave data detection device and by using computational fluid dynamics; and an acquisition unit, used to acquire the wave characteristic parameters and float motion parameters collected from the numerical pool under the target state, wherein the target state is the state of the numerical pool when the wave-generating parameters are used to control the numerical pool.

[0139] Optionally, the evaluation network includes a first evaluation network and a second evaluation network. The evaluation module 730 is specifically used to input the training sample data and the second power take-off force gradient parameter at the current time into the first evaluation network and the second evaluation network respectively to obtain a first evaluation parameter and a second evaluation parameter. The first loss determination module 740 includes: a second determination unit, used to determine the evaluation parameter with the smallest value among the first evaluation parameter and the second evaluation parameter as the target evaluation parameter; and a first calculation unit, used to calculate the first loss based on the training sample data, the second power take-off force gradient parameter at the current time, the target evaluation parameter, and a preset first loss function.

[0140] Optionally, the parameter determination model training device 700 further includes: a reward evaluation parameter calculation module, used to calculate reward evaluation parameters based on the training sample data and the second force take-off force gradient parameters at the current time; and a second loss determination module 750, including: a second calculation unit, used to calculate a second loss corresponding to the first evaluation network based on the first evaluation parameters, the reward evaluation parameters, and the second loss function; and a third calculation unit, used to calculate a second loss corresponding to the second evaluation network based on the second evaluation parameters, the reward evaluation parameters, and the second loss function.

[0141] Optionally, the reward evaluation parameter calculation module includes: a fourth calculation unit, used to calculate the power take-off control parameters at the current moment based on the second power take-off force gradient parameters at the current moment; a judgment unit, used to judge whether the float motion parameters and the power take-off control parameters at the current moment meet preset conditions; a fifth calculation unit, used to calculate the reward evaluation parameters based on the float motion parameters, the power take-off control parameters at the current moment, and the change in the power take-off force gradient parameters if the conditions are met; and a third determination unit, used to determine that the reward evaluation parameters are preset values ​​if the conditions are not met.

[0142] The parameter determination model training device provided in this embodiment can execute the parameter determination model training method described in any of the above embodiments. Its execution method and beneficial effects are similar, and will not be repeated here.

[0143] Figure 8 This is a schematic diagram of the structure of a parameter determination device provided in an embodiment of this disclosure. Figure 8 As shown, the parameter determination device 800 includes: a parameter acquisition module 810 and a parameter determination module 820. The parameter acquisition module 810 is used to acquire target wave characteristic parameters, target float motion parameters, and the first target power take-off force gradient parameters of the previous moment. The parameter determination module 820 is used to input the target wave characteristic parameters, the target float motion parameters, and the first target power take-off force gradient parameters of the previous moment into a parameter determination model to obtain the second target power take-off force gradient parameters of the current moment output by the parameter determination model. The parameter determination model is trained by the parameter determination model training method in the above embodiment.

[0144] The parameter determination device provided in this embodiment can acquire target wave characteristic parameters, target buoy motion parameters, and the first target power take-off force gradient parameters of the previous moment. The target wave characteristic parameters, target buoy motion parameters, and the first target power take-off force gradient parameters of the previous moment are input into the parameter determination model to obtain the second target power take-off force gradient parameters of the current moment output by the parameter determination model. In this way, combined with environmental factors, a more accurate power take-off force gradient parameter is determined, so that the matching degree between the buoy motion phase controlled by the parameter and the incoming wave phase is higher, thereby improving the energy absorption efficiency of wave energy.

[0145] Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure.

[0146] like Figure 9 As shown, the computer device may include a processor 910 and a memory 920 storing computer program instructions.

[0147] Specifically, the processor 910 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0148] Memory 920 may include a large-capacity storage for information or instructions. For example, and not limitingly, memory 920 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 920 may include removable or non-removable (or fixed) media. Where appropriate, memory 920 may be internal or external to the integrated gateway device. In a particular embodiment, memory 920 is a non-volatile solid-state memory. In a particular embodiment, memory 920 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (Electrically Programmable ROM, EPROM), an electrically erasable programmable PROM (EEPROM), an electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0149] The processor 910 reads and executes computer program instructions stored in the memory 920 to perform the steps of the parameter determination model training or parameter determination method provided in the embodiments of this disclosure.

[0150] In one example, the computer device may also include a transceiver 930 and a bus 940. Wherein, as... Figure 9 As shown, the processor 910, memory 920 and transceiver 930 are connected via bus 940 and communicate with each other.

[0151] Bus 940 may include hardware, software, or both. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 940 may include one or more buses. Although specific buses are described and illustrated in the embodiments of this application, this application considers any suitable bus or interconnection.

[0152] This disclosure also provides a computer-readable storage medium that can store a computer program. When the computer program is executed by a processor, the processor enables the processor to implement the parameter determination model training or parameter determination method provided in this disclosure.

[0153] The aforementioned storage medium may include, for example, a memory 920 containing computer program instructions, which can be executed by a processor 910 of a parameter determination model training and parameter determination device to complete the parameter determination model training or parameter determination method provided in the embodiments of this disclosure.

[0154] The computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof. The computer program described above can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on a user computing device, partially on a user device, as a standalone software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0155] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the boxes may occur in a different order than those shown in the accompanying drawings. For example, two consecutively indicated boxes may actually be executed substantially in parallel, or sometimes in reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or operations, or using a combination of dedicated hardware and computer instructions.

[0156] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0157] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for training a parameter-determined model, characterized in that, The method includes: Acquire training sample data, which includes wave characteristic parameters, float motion parameters, and the first power take-off force gradient parameters of the previous moment; Based on the training sample data, the parameter output network in the parameter determination model is used to obtain the force gradient parameters of the second force take-off device at the current moment; Based on the training sample data and the force gradient parameters of the second power take-off at the current moment, the evaluation network in the model is determined through the parameters to obtain the evaluation parameters; Based on the training sample data, the second force take-off force gradient parameter at the current moment, the evaluation parameter, and the preset first loss function, calculate the first loss of the parameter output network; Based on the evaluation parameters and the preset second loss function, the second loss of the evaluation network is calculated; The model parameters of the model are adjusted based on the first loss and the second loss to make the first loss and the second loss converge. The evaluation network includes a first evaluation network and a second evaluation network. The evaluation parameters are obtained by determining the evaluation network in the model based on the training sample data and the force gradient parameters of the second force take-off at the current time, including: The training sample data and the force gradient parameters of the second power take-off at the current moment are respectively input into the first evaluation network and the second evaluation network to obtain the first evaluation parameters and the second evaluation parameters. The calculation of the first loss of the parameter output network based on the training sample data, the force gradient parameter of the second power take-off at the current moment, the evaluation parameters, and the preset first loss function includes: The evaluation parameter with the smallest value among the first evaluation parameter and the second evaluation parameter is determined as the target evaluation parameter; Based on the training sample data, the force gradient parameter of the second power take-off device at the current moment, the target evaluation parameter, and the preset first loss function, the first loss is calculated; After obtaining the force gradient parameters of the second force take-off device at the current moment through the parameter output network in the parameter determination model based on the training sample data, the method further includes: The reward evaluation parameters are calculated based on the training sample data and the force gradient parameters of the second power take-off at the current moment. The step of calculating the second loss of the evaluation network based on the evaluation parameters and a preset second loss function includes: For the first evaluation network, a second loss is calculated based on the first evaluation parameters, the reward evaluation parameters, and the second loss function. For the second evaluation network, based on the second evaluation parameters, the reward evaluation parameters, and the second loss function, the second loss corresponding to the second evaluation network is calculated; The calculation of reward evaluation parameters based on the training sample data and the force gradient parameters of the second power take-off at the current moment includes: Based on the second power take-off force gradient parameters at the current moment, the power take-off control parameters at the current moment are calculated. Determine whether the float motion parameters and the power take-off control parameters at the current moment meet preset conditions; If satisfied, the reward evaluation parameters are calculated based on the float motion parameters, the power take-off control parameters at the current moment, and the change in the power take-off force gradient parameters. If the conditions are not met, then the value of the reward evaluation parameter is determined to be a preset value.

2. The method according to claim 1, characterized in that, The acquisition of training sample data includes: The wave characteristic parameters detected by the wave data detection equipment and the float motion parameters detected by the motion state detection equipment are obtained.

3. The method according to claim 2, characterized in that, The acquisition of training sample data also includes: Based on the wave characteristic parameters detected by the wave data detection equipment, computational fluid dynamics is used to determine the wave generation parameters and wave dissipation parameters corresponding to the target sea state. The wave characteristic parameters and float motion parameters are acquired from the numerical water tank under the target state, wherein the target state is the state of the numerical water tank when the wave generation parameters are used to control the numerical water tank.

4. A method for determining parameters, characterized in that, The method includes: Acquire the target wave characteristic parameters, the target buoy motion parameters, and the first target power take-off force gradient parameters from the previous moment; The target wave characteristic parameters, the target float motion parameters, and the first target power take-off force gradient parameters of the previous moment are input into the parameter determination model to obtain the second target power take-off force gradient parameters of the current moment output by the parameter determination model. The parameter determination model is trained by the training method as described in any one of claims 1-3.

5. A parameter determination model training device, characterized in that, The device includes: The data acquisition module is used to acquire training sample data, which includes wave characteristic parameters, float motion parameters, and the first power take-off force gradient parameters of the previous moment. The output module is used to obtain the second force take-off force gradient parameters at the current moment by using the parameter determination network in the model based on the training sample data. The evaluation module is used to determine the evaluation network in the model based on the training sample data and the force gradient parameters of the second force take-off at the current moment, and to obtain the evaluation parameters. The first loss determination module is used to calculate the first loss of the parameter output network based on the training sample data, the second force take-off force gradient parameter at the current moment, the evaluation parameter and the preset first loss function. The second loss determination module is used to calculate the second loss of the evaluation network based on the evaluation parameters and a preset second loss function. An adjustment module is used to adjust the model parameters of the parameter-determined model based on the first loss and the second loss, so that the first loss and the second loss converge. The evaluation network includes a first evaluation network and a second evaluation network. The evaluation module is used to input the training sample data and the force gradient parameter of the second power take-off device at the current moment into the first evaluation network and the second evaluation network respectively to obtain the first evaluation parameter and the second evaluation parameter. The first loss determination module includes determining the evaluation parameter with the smallest value among the first evaluation parameter and the second evaluation parameter as the target evaluation parameter; Based on the training sample data, the force gradient parameter of the second power take-off device at the current moment, the target evaluation parameter, and the preset first loss function, the first loss is calculated; The reward evaluation parameter calculation module is used to calculate the reward evaluation parameter based on the training sample data and the second power take-off force gradient parameter at the current moment after obtaining the gradient of the second power take-off force gradient parameter at the current moment through the parameter output network in the parameter determination model based on the training sample data. The second loss determination module includes, for the first evaluation network, calculating a second loss corresponding to the first evaluation network based on the first evaluation parameters, the reward evaluation parameters, and the second loss function; For the second evaluation network, based on the second evaluation parameters, the reward evaluation parameters, and the second loss function, the second loss corresponding to the second evaluation network is calculated; The reward evaluation parameter calculation module is used to calculate the power take-off control parameters at the current moment based on the second power take-off force gradient parameters at the current moment. Determine whether the float motion parameters and the power take-off control parameters at the current moment meet preset conditions; If satisfied, the reward evaluation parameters are calculated based on the float motion parameters, the power take-off control parameters at the current moment, and the change in the power take-off force gradient parameters. If the conditions are not met, then the value of the reward evaluation parameter is determined to be a preset value.

6. A parameter determining device, characterized in that, The device includes: The parameter acquisition module is used to acquire target wave characteristic parameters, target buoy motion parameters, and the first target power take-off force gradient parameters of the previous moment. The parameter determination module is used to input the target wave characteristic parameters, the target float motion parameters, and the first target power take-off force gradient parameters of the previous moment into the parameter determination model to obtain the second target power take-off force gradient parameters of the current moment output by the parameter determination model. The parameter determination model is trained by the training method as described in any one of claims 1-3.

7. A computer device, characterized in that, include: Memory; processor; And a computer program; wherein the computer program is stored in the memory and configured to be executed by the processor to implement the parameter determination model training method as described in any one of claims 1-3 or the parameter determination method as described in claim 4.

8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the parameter determination model training method as described in any one of claims 1-3 or the parameter determination method as described in claim 4.

Citation Information

Patent Citations

  • Training method of process parameter adjustment model and process parameter adjustment method and device

    CN112950071A

  • Model parameter determination method and device, equipment and storage medium

    CN113887739A

  • Optimization system and method for parameter configuration of wave energy device

    JP2022136002A