Control methods, systems, electronic equipment, and storage media applicable to wave energy conversion devices
By introducing a prediction-correction mechanism consisting of a strategy network, a state estimation unit, and a control optimization unit into the wave energy conversion device, the problem of balancing energy capture and safety in existing technologies has been solved, and stable operation in complex marine environments has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA THREE GORGES CORPORATION
- Filing Date
- 2026-01-29
- Publication Date
- 2026-06-02
AI Technical Summary
Existing wave energy conversion device control methods are difficult to balance energy capture efficiency and system safety in complex marine environments, and cannot meet the requirements of real-time performance and reliability for long-term operation at sea.
A policy network is used to generate candidate control commands, which are then predicted in multiple steps by a state estimation unit. The control optimization unit corrects the commands when the prediction results indicate potential limits being exceeded, thus constructing an active prediction-correction mechanism to ensure the safety of the control commands.
Without relying on precise physical models, the wave energy conversion device achieves a balance between energy capture efficiency and operational safety in dynamic marine environments, effectively avoiding the risk of exceeding limits and ensuring long-term stable operation.
Smart Images

Figure CN122129382A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of marine energy development and control technology, specifically relating to a control method, system, electronic equipment and storage medium suitable for wave energy conversion devices. Background Technology
[0002] As a key device for utilizing marine renewable energy, wave energy conversion devices are directly related to energy capture efficiency and control performance. By capturing the kinetic and potential energy dispersed in waves, they realize the conversion from marine mechanical energy to electrical energy.
[0003] Existing control methods for wave energy conversion devices mainly fall into two categories: optimization control methods based on physical models and empirical control methods. While model predictive control can explicitly address system constraints, it heavily relies on precise mathematical models. The complex hydrodynamic characteristics of the marine environment make modeling errors unavoidable, resulting in significant nonlinear effects and excessive computational burden, making this method difficult to operate stably under varying sea conditions. Traditional feedback control methods, such as lockout control or damping control, while simple to implement, cannot adapt to drastically changing wave conditions and struggle to balance energy capture efficiency with system safety. They also fail to meet the real-time and reliability requirements for long-term operation at sea. Summary of the Invention
[0004] The purpose of this application is to provide a control method, system, electronic device, and storage medium suitable for wave energy conversion devices, which can solve the problem that existing wave energy conversion device control methods are difficult to meet the real-time and reliability requirements of long-term operation at sea.
[0005] To solve the above-mentioned technical problems, this application is implemented as follows: In a first aspect, embodiments of this application provide a control method suitable for wave energy conversion devices, including: The controller is configured with a strategy network, which takes the real-time motion state and real-time environmental information of the wave energy conversion device as input and outputs candidate control commands for the wave energy conversion device. The state estimation unit is configured with a prediction network. The prediction network performs multi-step predictions of the motion state of the wave energy conversion device at multiple future moments based on the real-time motion state of the wave energy conversion device, the real-time environmental information, and the candidate control commands, and generates prediction results. If the predicted state of motion at any future time exceeds a safety threshold, the control optimization unit modifies the candidate control command to obtain a safe control command, and controls the wave energy conversion device to execute the safe control command.
[0006] Furthermore, the candidate control commands are modified to obtain safety control commands, including: Within the current control cycle, the candidate control commands are adjusted according to a preset direction and a preset step size to obtain the candidate control commands for the current iteration; Based on the real-time motion state of the wave energy conversion device, the real-time environmental information, and the candidate control commands for the current iteration, the prediction result for the current iteration is generated using the prediction network. If the prediction result of the current iteration indicates that the motion state at any future time exceeds the safety threshold, the above steps are repeated until the motion state at all future times in the prediction result of the current iteration does not exceed the safety threshold, then the iteration is stopped and the safety control instruction is obtained.
[0007] Furthermore, the candidate control command includes at least a control force; the prediction result of the current iteration includes at least a predicted displacement; The process of determining the preset direction includes: Preset positive disturbance values and preset negative disturbance values are applied to the candidate control command respectively to obtain positive disturbance command and negative disturbance command; The prediction network is used to predict the positive disturbance command and the negative disturbance command respectively, so as to obtain the predicted displacement of the positive disturbance command and the predicted displacement of the negative disturbance command. If the predicted displacement of the positive disturbance command is greater than the predicted displacement of the negative disturbance command, the preset direction is determined to increase the control force. If the predicted displacement of the positive disturbance command is less than the predicted displacement of the negative disturbance command, the preset direction is determined to reduce the control force.
[0008] Furthermore, the process of determining the preset step size includes: Determine the number of iterations for the candidate control instruction; The preset step size is determined based on the preset basic correction step size and the number of iterations of the candidate control command.
[0009] Furthermore, the real-time motion status of the wave energy conversion device includes: real-time floating body displacement and real-time floating body velocity; the real-time environmental information includes: real-time external wave information; The controller is also configured with a shared feature extraction network, and the method further includes: The shared feature extraction network extracts features from real-time floating body displacement, real-time floating body velocity, and real-time external wave information to generate a shared feature vector. The policy network generates the candidate control command based on the shared feature vector, and the prediction network generates the prediction result based on the shared feature vector.
[0010] Furthermore, the training process for the policy network and the prediction network includes: First training phase: Without enabling the constraint correction process, the policy network aims to maximize energy capture efficiency and learns the dynamic characteristics of the wave energy conversion device under no safety constraints to update the model parameters of the policy network itself; wherein, the safety constraint means that the motion state at all future moments in the prediction results does not exceed the safety threshold. Second training phase: The constraint correction process is activated. Under the condition of satisfying the safety constraints, the policy network and the prediction network update their respective model parameters with the goal of jointly optimizing the energy capture efficiency of the policy network and the prediction network.
[0011] Furthermore, the method also includes: High displacement state samples are sampled to generate high displacement training data; With the goal of collaboratively optimizing the energy capture efficiency of the policy network and the prediction network, the model parameters of the policy network and the prediction network are updated, including: Using the high displacement training data, the model parameters of the policy network and the prediction network are updated with the goal of co-optimizing the energy capture efficiency of the policy network and the prediction network.
[0012] Secondly, embodiments of this application provide a control system suitable for wave energy conversion devices, including: A controller is used to configure a strategy network, which takes the real-time motion state and real-time environmental information of the wave energy conversion device as input and outputs candidate control commands for the wave energy conversion device. A state estimation unit is used to configure a prediction network. The prediction network performs multi-step prediction of the motion state of the wave energy conversion device at multiple future moments based on the real-time motion state of the wave energy conversion device, the real-time environmental information, and the candidate control commands, and generates prediction results. The control optimization unit is used to modify the candidate control command to obtain a safe control command when the predicted result indicates that the motion state at any future time exceeds a safety threshold, and to control the wave energy conversion device to execute the safe control command.
[0013] Thirdly, embodiments of this application provide an electronic device, including: processor; Memory for storing processor-executable instructions; The processor is configured to execute the instructions to implement the steps of the method as described in the first aspect.
[0014] Fourthly, embodiments of this application provide a computer-readable storage medium that, when the instructions in the computer-readable storage medium are executed by a processor of a terminal, enables the terminal to perform the steps of the method described in the first aspect.
[0015] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0016] This application's embodiments generate candidate control commands from real-time motion state and real-time environmental information through a policy network, avoiding direct reliance on complex physical models. By introducing a state estimation unit and a control optimization unit, an active prediction-correction mechanism is constructed: the prediction network can assess the future motion state that candidate control commands may cause; when the prediction results show that the motion state at any future moment will exceed the safety threshold, the control optimization unit will immediately correct the candidate control command. This mechanism of safety verification and dynamic correction before command execution provides more stringent and reliable safety constraints. This enables the wave energy conversion device to achieve a balance between energy capture efficiency and operational safety in a dynamic ocean environment without relying on a precise physical model. Through real-time prediction and online correction, potential risks of exceeding limits are effectively avoided, ensuring the long-term stable operation of the wave energy conversion device. Attached Figure Description
[0017] Figure 1 A flowchart illustrating the steps of a control method for a wave energy conversion device according to an embodiment of this application is shown. Figure 2 An exemplary flowchart of steps for a control method applicable to a wave energy conversion device according to an embodiment of this application is shown; Figure 3 An exemplary flowchart of steps for a control method applicable to a wave energy conversion device, provided in another embodiment of this application, is shown; Figure 4 A flowchart illustrating exemplary training steps for a policy network and a prediction network provided in an embodiment of this application is shown. Figure 5 A structural diagram of a control system for a wave energy conversion device according to an embodiment of this application is shown; Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0020] The control method, system, electronic device, and storage medium applicable to wave energy conversion devices provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.
[0021] Figure 1 A flowchart illustrating the steps of a control method for a wave energy conversion device according to an embodiment of this application is shown. (Refer to...) Figure 1 This application provides a control method suitable for wave energy conversion devices, including: S10. The controller configures a strategy network. The strategy network takes the real-time motion state and real-time environmental information of the wave energy conversion device as input and outputs candidate control commands for the wave energy conversion device.
[0022] Specifically, a wave energy conversion device refers to a device that converts the kinetic or potential energy of ocean waves into usable electrical energy. This includes a floating structure that interacts with the waves to generate relative motion, thereby driving the energy conversion mechanism to generate electricity. The controller is the core component of the wave energy conversion device's control system, responsible for generating and outputting control commands based on received input information to adjust the device's motion state. The strategy network is a neural network-based model that learns and executes control strategies. In the control scenario of the wave energy conversion device, the strategy network takes the device's real-time motion state and environmental data as input and outputs control commands to drive the device. Real-time motion state refers to the dynamic parameters of the wave energy conversion device at the current moment, such as the real-time displacement, velocity, and acceleration of the floating body, indicating the device's immediate operational status. Real-time environmental information refers to the characteristic data of the ocean environment in which the wave energy conversion device is located at the current moment, such as the height, period, and direction of external waves. This real-time environmental information directly affects the wave energy conversion device's motion response. Candidate control commands refer to control signals initially generated by the strategy network based on the real-time motion state and real-time environmental information of the wave energy conversion device. They represent the control actions that the strategy network believes to take at the current moment, but their safety has not yet been verified.
[0023] For example, the controller is configured to include at least one policy network. The policy network receives real-time motion state and real-time environmental information of the wave energy conversion device as input. The real-time motion state can be directly acquired by sensors, such as obtaining the real-time position of the float via a displacement sensor and the real-time velocity of the float via a velocity sensor. The real-time environmental information can be acquired by environmental monitoring equipment, such as obtaining real-time wave height, wave period, and other data via wave buoys or radar. Based on this input information, the policy network generates a candidate control command using a control strategy learned internally. For example, the policy network can be a feedforward neural network whose weights and biases are obtained through offline training and directly perform forward computation at runtime to output the control command.
[0024] In some embodiments, the real-time motion state of the wave energy conversion device includes: real-time floating body displacement and real-time floating body velocity. Real-time environmental information includes: real-time external wave information.
[0025] The controller is also configured with a shared feature extraction network, and the method also includes: The shared feature extraction network extracts features from real-time floating body displacement, real-time floating body velocity, and real-time external wave information to generate a shared feature vector.
[0026] The policy network generates candidate control instructions based on shared feature vectors, and the prediction network generates prediction results based on shared feature vectors.
[0027] Specifically, a shared feature extraction network refers to a network used to uniformly process data from different sensors and information sources. The shared feature extraction network can be a multilayer perceptron (MLP) structure, a convolutional neural network (CNN), or a recurrent neural network (RNN); this application embodiment does not impose specific limitations. The shared feature extraction network transforms raw, heterogeneous input data into a unified, high-dimensional, and representative feature representation for use by subsequent policy and prediction networks. The shared feature extraction network receives real-time floating body displacement, real-time floating body velocity, and real-time external wave information as input and processes them to generate a shared feature vector. The feature extraction process may include data normalization, dimensionality reduction, nonlinear transformation, and other operations, aiming to capture the most useful information for control and prediction tasks from the raw data. This shared feature vector, as a unified intermediate representation, ensures that the policy and prediction networks are based on the same and consistent system perception during decision-making and prediction. By using the shared feature vector as input, the policy network can formulate control decisions based on a comprehensive understanding of the real-time motion state of the wave energy conversion device and real-time environmental information, thereby optimizing energy capture efficiency. Predictive networks can accurately predict the future behavior of devices based on system perception consistent with policy networks, providing a basis for the assessment of security constraints.
[0028] In this embodiment, a strategy network configured by the controller takes the real-time motion state and environmental information of the wave energy conversion device as input and outputs candidate control commands. This replaces traditional feedback methods such as lock-in control and damping control that rely on fixed rules. Through environmental interaction, the network autonomously learns dynamic control strategies, proactively adapting to drastically changing wave conditions (such as random fluctuations in wave height and period). Compared to traditional model predictive control, which requires precise modeling of hydrodynamic characteristics, this significantly reduces the impact of modeling errors, adapts to the nonlinear and highly disturbed characteristics of the marine environment, and achieves lightweight real-time decision-making. By collecting dynamic parameters of the floating body and environmental data in real time through sensors, the strategy network can dynamically adjust the control strategy. For example, when wave height changes abruptly, the strategy network can quickly generate control commands that match the current operating conditions, avoiding the efficiency-safety imbalance problem caused by fixed parameters in traditional feedback control, and achieving a dynamic balance between energy capture efficiency and system safety.
[0029] S20. The state estimation unit is configured with a prediction network. The prediction network performs multi-step predictions of the motion state of the wave energy conversion device at multiple future moments based on the real-time motion state of the wave energy conversion device, real-time environmental information, and candidate control commands, and generates prediction results.
[0030] Specifically, the state estimation unit refers to the unit in the control system of a wave energy conversion device that uses a predictive network to evaluate and predict the future operating state of the device. The predictive network is a neural network-based predictive model that infers the motion state of the wave energy conversion device at multiple future time steps based on current and historical data and candidate control commands. Multi-step prediction refers to the continuous prediction of the state sequence at multiple discrete future time points. The prediction result refers to the state sequence output by the predictive network after multi-step prediction, which includes estimates of key operating indicators such as displacement, velocity, and force of the wave energy conversion device at various future times.
[0031] For example, the state estimation unit is configured to include at least one prediction network. The prediction network receives the real-time motion state of the wave energy conversion device, real-time environmental information, and candidate control commands output by the policy network as inputs, and performs multi-step predictions of the motion state of the wave energy conversion device at multiple future time points based on these inputs, and generates prediction results. For example, the prediction network can be a recurrent neural network or a long short-term memory network, which is capable of processing time-series data and inferring future states based on current inputs and / or historical information.
[0032] In this embodiment, the prediction network takes real-time motion state, environmental information, and candidate control commands as input to perform multi-step predictions of key parameters such as float displacement, velocity, and force at multiple discrete time steps in the future. By capturing the long-term dependencies of time series data, potential safety risks can be identified in advance, realizing a shift from passive response to proactive prediction. The multi-step prediction results transform system constraints (such as maximum allowable displacement and maximum force threshold) into quantifiable safety boundaries. For example, if it is predicted that the displacement will exceed the safety threshold at a certain future moment, a correction process is initiated to avoid the safety risks caused by the delayed response of traditional control methods, thereby improving the reliability of the wave energy conversion device.
[0033] S30. When the predicted state of motion at any future moment exceeds the safety threshold, the control optimization unit corrects the candidate control command to obtain a safe control command and controls the wave energy conversion device to execute the safe control command.
[0034] Specifically, the safety threshold refers to the pre-set upper and / or lower limits used to restrict the operating parameters of the wave energy conversion device. When the motion state or stress condition of the wave energy conversion device exceeds the safety threshold, it indicates that there is a safety risk to the wave energy conversion device. The control optimization unit refers to the function in the control system of the wave energy conversion device that adjusts and corrects the candidate control commands generated by the strategy network when the prediction network predicts a safety risk to ensure that the final executed command meets the requirements for safe operation. The safety control command refers to the control command after being corrected by the control optimization unit. The safety control command maintains or optimizes the energy capture efficiency within the possible range while ensuring the safe operation of the wave energy conversion device.
[0035] For example, if the predicted motion state at any future moment exceeds a safety threshold, the control optimization unit modifies the candidate control commands to obtain safe control commands, and then controls the wave energy conversion device to execute the safe control commands. For instance, if the displacement value at any moment in the future displacement sequence output by the prediction network exceeds a preset maximum allowable displacement threshold, the control optimization unit will initiate a correction process. The correction process may include a simple scaling of the candidate control commands; for example, if the predicted displacement is too large, the amplitude of the control command is reduced proportionally. The corrected command is the safe control command, which is then sent to the actuator of the wave energy conversion device, such as a power output device, to drive the wave energy conversion device to operate according to the safe commands.
[0036] In this embodiment, when the predicted motion state at any future moment exceeds a safety threshold, the control optimization unit modifies the candidate control commands to generate safe control commands. This ensures that the executed control actions avoid exceeding the limit at the source, eliminating the potential for structural damage or operational interruption of the wave energy conversion device. The correction process, while meeting the safety threshold, maintains or optimizes the energy capture efficiency within the possible range, satisfying the reliability and economic requirements for the long-term operation of the wave energy conversion device.
[0037] In some embodiments, S30, modifying the candidate control instructions to obtain security control instructions includes: S31. Within the current control cycle, adjust the candidate control commands according to the preset direction and preset step size to obtain the candidate control commands for the current iteration.
[0038] Specifically, the preset direction refers to the direction of increment or decrement when adjusting candidate control commands. For example, when the candidate control command is a control force, the preset direction could be to increase or decrease the control force. The preset step size refers to the magnitude of each adjustment to the candidate control command. For example, a fixed, small value can be set as the step size to ensure that each adjustment is gradual. Alternatively, the step size can be dynamically adjusted based on the deviation of the current predicted state from the safety threshold; for example, the greater the deviation, the larger the step size can be to accelerate convergence. The candidate control command in the current iteration refers to the candidate control command after one or more adjustments during the correction process.
[0039] In some embodiments, the candidate control command includes at least a control force. The prediction result of the current iteration includes at least a predicted displacement.
[0040] The process of determining the preset direction includes: Preset positive disturbance values and preset negative disturbance values are applied to the candidate control commands to obtain positive disturbance commands and negative disturbance commands.
[0041] By using a prediction network, positive and negative disturbance commands are predicted separately, resulting in the predicted displacements of the positive and negative disturbance commands.
[0042] If the predicted displacement of a positive disturbance command is greater than the predicted displacement of a negative disturbance command, the preset direction will be determined as increasing the control force.
[0043] If the predicted displacement of a positive disturbance command is less than the predicted displacement of a negative disturbance command, the preset direction will be set to reduce the control force.
[0044] Specifically, control force refers to the force applied to the buoy by the power output mechanism of the wave energy conversion device, used to regulate the buoy's motion state to achieve energy capture or safety restraint. Control force can be a continuous value, representing the thrust or damping force output by the power output mechanism, whose magnitude and direction directly affect the buoy's acceleration and velocity; alternatively, control force can be discrete, for example, applied by selecting several preset force levels or modes.
[0045] Predicted displacement refers to the value predicted by the prediction network for the displacement of the floating body of the wave energy conversion device at a future time, based on the current system state and candidate control commands. It is a key indicator for evaluating the operational safety of the device. Predicted displacement can be represented by a sequence of vertical displacements of the floating body over multiple future time steps, output by the prediction network; alternatively, it can be the displacement value of the floating body at a specific critical moment in the future, output by the prediction network. Preset positive and negative disturbance values are manually set quantities used for fine-tuning or testing the impact of control commands.
[0046] Positive and negative disturbance values are two test quantities of equal magnitude but opposite direction, used to detect the impact of control force adjustment direction on the future state of the system. The disturbance value can be a preset fixed value, such as increasing or decreasing the control force command by a fixed amount; or it can be a percentage of the current control force command, such as increasing or decreasing the force command by 5%. Positive and negative disturbance commands refer to two hypothetical control commands formed by adding preset positive and negative disturbance values to the original candidate control command, respectively. They are used to simulate the possible system responses under different control force adjustment directions. A positive disturbance command is a candidate control command plus a preset positive disturbance value, and a negative disturbance command is a candidate control command plus a preset negative disturbance value.
[0047] A prediction network is used to predict both positive and negative disturbance commands, simulating the future motion of the wave energy conversion device under these two hypothetical scenarios. The prediction network receives the current system state, environmental information, and the control command after the disturbance as input, and outputs prediction results for multiple future time points, obtaining the predicted displacements for the positive and negative disturbance commands. If the predicted displacement of the positive disturbance command is greater than the predicted displacement of the negative disturbance command, a preset direction is determined to increase the control force to counteract the trend leading to the large displacement. If the predicted displacement of the positive disturbance command is less than the predicted displacement of the negative disturbance command, a preset direction is determined to decrease the control force, thus determining the correction direction.
[0048] This embodiment applies preset positive and negative disturbance values to candidate control commands, and uses a prediction network to predict the resulting positive and negative disturbance commands, obtaining the impact of different control force adjustment directions on future floating body displacement. Based on the comparison of predicted displacement, it determines whether to increase or decrease the control force, thereby ensuring that each correction is made in a direction that effectively reduces the risk of exceeding limits. This improves the accuracy and convergence speed of the correction process, enabling the wave energy conversion device to adjust control commands more quickly and reliably when facing potential risks of exceeding limits. This ensures the safety of device operation while minimizing energy capture efficiency losses caused by unnecessary or erroneous corrections, further ensuring the safe and stable operation of the wave energy conversion device in complex sea conditions.
[0049] In some embodiments, the process of determining the preset step size includes: Determine the number of iterations for the candidate control instructions.
[0050] The preset step size is determined based on the preset basic correction step size and the number of iterations of the candidate control instructions.
[0051] Specifically, determining the iteration number of the candidate control command refers to the number of repeated operations performed to adjust the candidate control command in order to make the prediction result meet the safety threshold during the process of the control optimization unit correcting the candidate control command. Its function is to provide a quantitative and real-time basis for dynamically adjusting the preset step size, reflecting the progress and complexity of the current correction process.
[0052] The preset step size is determined based on a preset base correction step size and the number of iterations of candidate control commands. This aims to achieve adaptive adjustment of the preset step size, making it no longer a fixed value but dynamically changing according to the actual needs of the correction process. By combining the base correction step size and the number of iterations, the correction process can be more refined in the early stages and converge more quickly in the later stages, thereby improving correction efficiency and reducing computational burden. For example, the preset step size can be determined as the product of the base correction step size and the number of iterations, specifically expressed as: Preset step size = Base correction step size × (1 + Number of iterations × Correction coefficient), where the correction coefficient is a preset scaling factor used to control the rate at which the step size increases with the number of iterations; this embodiment does not impose specific limitations. Furthermore, the preset step size can also be determined using a piecewise function. For example, when the number of iterations is small, the preset step size equals the base correction step size; when the number of iterations increases to a certain range, the preset step size increases to a multiple of the base correction step size; when the number of iterations further increases, the preset step size can be further amplified to accelerate convergence.
[0053] In this embodiment, during the process of revising candidate control commands to obtain safe control commands, the revision step size can be dynamically adjusted according to the actual progress of the revision iteration. This avoids the problems of low revision efficiency or increased computational burden that may occur with a fixed step size. Specifically, in the early stages of revision, a smaller step size can be used to achieve fine adjustment of the control commands, effectively preventing over-revision or oscillation and ensuring the accuracy of the revision. As the number of revision iterations increases, if a safe state is still not reached, the system can adaptively increase the revision step size, thereby accelerating the convergence process, especially when the initial command deviates far from the safe region, enabling the system to find control commands that meet safety constraints more quickly. This dynamic step size adjustment mechanism significantly improves the response speed and revision efficiency of the control optimization unit when dealing with the risk of exceeding limits, reduces the consumption of real-time computing resources, and enables the control system of the wave energy conversion device to operate more efficiently and stably while ensuring operational safety, especially in complex and variable sea conditions, where it can better balance energy capture efficiency and system safety.
[0054] S32. Based on the real-time motion state of the wave energy conversion device, real-time environmental information, and candidate control commands for the current iteration, generate the prediction result for the current iteration using the prediction network.
[0055] Specifically, the prediction result of the current iteration refers to a series of predicted motion state values for future moments generated by the prediction network in response to the candidate control commands of the current iteration during the correction process, such as the displacement and velocity of the floating body at multiple future moments.
[0056] S33. If the prediction result of the current iteration indicates that the motion state at any future moment exceeds the safety threshold, repeat the above steps until the motion state at all future moments in the prediction result of the current iteration does not exceed the safety threshold, stop the iteration, and obtain the safety control instruction.
[0057] Specifically, repeating the above steps means that when a risk of exceeding limits is predicted in the future state, the process returns to adjusting the candidate control command and re-predicts. This closed-loop iterative process ensures the effectiveness of the correction. Stopping iteration means that when the prediction results show that the motion state at all future moments is within the safe threshold range, a safe control command has been determined, and the correction process terminates. The safe control command, after iterative correction, ensures that the wave energy conversion device's motion state will not exceed limits in the future, and this safe control command will be actually executed.
[0058] For example, within each control cycle, the control optimization unit first receives candidate control commands generated by the policy network. If the prediction network indicates that the command may lead to future state exceeding limits, the control optimization unit will make preliminary adjustments to the candidate control command according to a preset direction and preset step size, thereby obtaining a candidate control command for the current iteration. This adjustment is not blind, but directional and measured, aiming to guide the control command towards a safe zone. Subsequently, the control optimization unit inputs this adjusted candidate control command for the current iteration, along with the real-time motion state of the wave energy conversion device and real-time environmental information, into the prediction network. Based on these latest inputs, the prediction network again performs multi-step predictions of the motion state of the wave energy conversion device at multiple future moments, generating new prediction results to ensure that each correction is based on the latest control commands and system state, accurately reflecting the impact of the adjusted commands on future states. The control optimization unit continuously monitors this new prediction result. If the prediction results still show that the motion state at any future moment exceeds the safety threshold, it indicates that the current adjustment is insufficient to eliminate the risk. At this point, the control optimization unit repeats the above adjustment and prediction steps, that is, it further adjusts the candidate control commands of the current iteration according to the preset direction and preset step size, and makes predictions again using the prediction network. This iterative process continues until the prediction results generated by the prediction network show that the motion state at all future moments does not exceed the safety threshold. Once this condition is met, the iterative process stops, and the candidate control commands of the current iteration are determined to be safe control commands and are finally sent to the wave energy conversion device for execution. Through this iterative correction mechanism, this scheme gradually converges the potentially unsafe candidate control commands initially generated by the strategy network into a safe control command that both meets the control objective and strictly adheres to safety constraints, through real-time feedback from the prediction network and iterative adjustments by the control optimization unit, thereby avoiding potential safety hazards. This improves the safety and reliability of the wave energy conversion device while maintaining control efficiency, avoiding the sacrifice of energy capture performance due to over-correction.
[0059] For example, Figure 2 An exemplary flowchart of the control method for a wave energy conversion device provided in an embodiment of this application is shown. Figure 3 An exemplary flowchart of a control method for a wave energy conversion device according to another embodiment of this application is shown. (Refer to...) Figure 2The controller collects real-time motion status (such as float displacement, velocity, and acceleration) and real-time environmental information (such as wave height, period, direction, and other ocean characteristic data) of the wave energy conversion device as input, and uses captured mechanical power as a reward guide (i.e., optimizing energy capture efficiency when learning control strategies) to provide target basis for subsequent control. The strategy network configured by the controller takes real-time motion status and real-time environmental information as input, autonomously learns dynamic control strategies, and outputs actions (i.e., candidate control commands). The state prediction network (i.e., prediction network) receives candidate control commands, real-time status, and environmental information, infers the motion status of the wave energy conversion device at multiple future moments (such as float displacement, velocity, and force), and generates prediction results such as the float predicted displacement sequence. If the prediction results show that the motion status at a future moment exceeds the limit (i.e., displacement or force exceeds the safety threshold), the action adjustment strategy (i.e., correction process) is initiated to adjust the candidate commands and generate safe control commands; if the prediction does not exceed the limit (i.e., the prediction results do not exceed the safety threshold), the candidate commands are directly output as thrust commands (i.e., safe control commands). The predicted and corrected safety control commands are sent to the actuators (such as power regulation and power output modules) of the wave energy conversion device to drive the device to operate.
[0060] Reference Figure 3 The input layer receives state information (i.e., real-time operating data and real-time environmental information of the wave energy conversion device) as the basis for input to all subsequent network modules, used to characterize the current dynamic characteristics of the wave energy conversion device and environmental disturbances. The shared layer of the shared feature extraction network performs nonlinear feature transformation on the input real-time operating data and real-time environmental information of the wave energy conversion device, outputting a shared feature vector, which serves as the unified input to the subsequent policy network, value network, and prediction network, ensuring that different networks make decisions / predictions based on the same perception. The policy network is divided into μ (mean network) and σ (standard deviation network), which jointly output candidate control commands a that follow a normal distribution. N(μ,σ). The value network receives shared feature vectors and outputs state value estimates to evaluate the long-term benefits / costs of the current wave energy conversion device's real-time operating data and real-time environmental information, thus assisting in policy network optimization. The prediction network includes at least two WEC state prediction layers, which output prediction results in two stages to provide a basis for core safety assessment. The first WEC state prediction layer receives shared feature vectors and extracts high-order spatiotemporal features to prepare abstract representations for multi-step prediction. The second WEC state prediction layer outputs prediction results with a dimension of 2 to infer the state variables that represent the floating body displacement or associated safety thresholds at critical future moments. The prediction results are used for safety threshold judgment: if the future state exceeds the limit, the control optimization unit is triggered to iteratively correct candidate instructions.
[0061] This application's embodiments generate candidate control commands from real-time motion state and real-time environmental information through a policy network, avoiding direct reliance on complex physical models. By introducing a state estimation unit and a control optimization unit, an active prediction-correction mechanism is constructed: the prediction network can assess the future motion state that candidate control commands may cause; when the prediction results show that the motion state at any future moment will exceed the safety threshold, the control optimization unit will immediately correct the candidate control command. This mechanism of safety verification and dynamic correction before command execution provides more stringent and reliable safety constraints. This enables the wave energy conversion device to achieve a balance between energy capture efficiency and operational safety in a dynamic ocean environment without relying on a precise physical model. Through real-time prediction and online correction, potential risks of exceeding limits are effectively avoided, ensuring the long-term stable operation of the wave energy conversion device.
[0062] In some embodiments, the training process for the policy network and the prediction network includes: The first training phase involves not enabling the constraint correction process. The policy network, aiming to maximize energy capture efficiency, learns the dynamic characteristics of the wave energy conversion device under no safety constraints to update its own model parameters. The safety constraint refers to ensuring that the motion states at all future moments in the prediction results do not exceed a safety threshold.
[0063] Specifically, in the first training phase, the constraint correction process is not enabled, allowing the policy network to fully explore the environment and learn the inherent dynamic characteristics of the wave energy conversion device without safety restrictions, laying the foundation for safety optimization in subsequent phases. This is achieved by running the policy network in a simulated environment to collect a large amount of state-action-reward data, or by acquiring data through limited, controlled interactions with an actual wave energy conversion device. Not enabling the constraint correction process means that, in specific training phases, the control optimization unit does not intervene to correct the candidate control commands output by the policy network, allowing the policy network to freely explore control strategies, including commands that may lead to exceeding limits, to gain a more comprehensive understanding of the performance boundaries and dynamic response of the wave energy conversion device. In this phase, the policy network aims to maximize energy capture efficiency, ensuring that it can learn efficient energy conversion strategies, thereby improving the power generation performance of the wave energy conversion device.
[0064] Second training phase: The constraint correction process is enabled. Under the condition of satisfying safety constraints, the policy network and the prediction network update their respective model parameters with the goal of jointly optimizing the energy capture efficiency of the policy network and the prediction network.
[0065] Specifically, safety constraints refer to the requirement that the motion states at all future moments in the predicted results do not exceed the safety threshold. This clearly defines the safety conditions that the wave energy conversion device must meet during operation, providing a clear judgment standard for safety corrections in the second training phase and actual operation. This is achieved by setting maximum allowable values for key motion parameters such as float displacement, float velocity, and PTO force; or by defining a safety zone, requiring the wave energy conversion device's trajectory to always remain within that zone. In the second training phase, the system activates a constraint correction process. Building upon the learning in the first phase, safety constraints are introduced to further optimize energy capture efficiency within the safety requirements of the strategy network and enable the prediction network to accurately identify potential risks. By continuing to run the strategy network and prediction network in the simulation environment and activating the constraint correction process, the candidate control commands output by the strategy network are checked and corrected for safety through the intervention of the control optimization unit. This ensures that the control commands generated by the strategy network during training meet safety requirements, preventing overstepping of limits during actual deployment. After constraint correction, the future motion states of the corresponding wave energy conversion devices must be within the safety threshold range, ensuring that the trained strategy network and prediction network can work collaboratively to achieve high performance while ensuring safety. For example, by sharing a reward function that simultaneously considers energy capture efficiency and the satisfaction of safety constraints, and updating the parameters of both the policy network and the prediction network; or by training the policy network and the prediction network alternately, with each training session based on the current state of the other and guided by a collaborative objective.
[0066] In the first training phase of this application, the policy network fully explores the dynamic characteristics of the wave energy conversion device without safety constraints, avoiding insufficient exploration caused by prematurely introducing constraints. This allows the policy network to learn a wider range of more efficient control strategies. Subsequently, in the second training phase, by enabling a constraint correction process, the policy network and the prediction network collaboratively optimize energy capture efficiency while satisfying safety constraints. This collaborative optimization mechanism enables the prediction network to more accurately identify potential limit-crossing risks, while the policy network is guided to generate control commands that maximize energy capture efficiency while strictly meeting safety requirements. Therefore, this application embodiment ensures that the trained controller, in actual operation, not only achieves high energy capture efficiency but also actively avoids potential safety risks, significantly improving the operational reliability and safety of the wave energy conversion device.
[0067] In some embodiments, the method further includes sampling high-displacement state samples to generate high-displacement training data.
[0068] With the goal of jointly optimizing the energy capture efficiency of the policy network and the prediction network, the model parameters of the policy network and the prediction network are updated, including: Using high displacement training data, the model parameters of the policy network and the prediction network are updated with the goal of co-optimizing the energy capture efficiency of the policy network and the prediction network.
[0069] Specifically, sampling high-displacement state samples to generate high-displacement training data refers to collecting system state data when the displacement of the floating body of the wave energy conversion device approaches or reaches a safety threshold during operation. The purpose is to specifically construct a training dataset containing extreme operating condition data to enhance the model's learning ability and accuracy in handling such critical situations. During the interaction between the wave energy conversion device and the environment, the floating body displacement is monitored in real time. When the absolute value of the floating body displacement exceeds a preset displacement threshold (e.g., 80% of the safety threshold), the current motion state, environmental information, and corresponding control commands are recorded as high-displacement state samples.
[0070] The goal of collaboratively optimizing the energy capture efficiency of the policy network and the prediction network involves updating their respective model parameters. This means that during training, the policy network and the prediction network are not trained independently, but rather updated together to maximize energy capture efficiency. This update mechanism aims to ensure functional consistency between the two networks; that is, the control commands generated by the policy network can effectively capture energy, while the prediction network can accurately predict the future states under these commands, thus providing a reliable basis for safety corrections. One implementation involves constructing a joint loss function that includes a reward term related to energy capture (e.g., power output) and a state prediction error term from the prediction network. During training iterations, the weight parameters of both the policy network and the prediction network are updated simultaneously by minimizing this joint loss function. Another implementation involves using an alternating training strategy: within a training cycle, the prediction network is first fixed, and the policy network is optimized to maximize energy capture; then the policy network is fixed again, and the prediction network is optimized to improve state prediction accuracy, especially for high-displacement states, and this process is repeated until convergence.
[0071] Utilizing the high-displacement training data refers to the targeted use of this data, representing extreme operating conditions, during model parameter updates. This aims to improve the model's performance in these key scenarios, compensating for the sparsity of high-displacement samples in regular training data and ensuring the model maintains accuracy and robustness in the face of real-world extreme sea conditions. One approach is to increase the proportion of high-displacement samples in each training batch, giving them a more prominent position in the training data distribution. Another approach is to assign higher weights to the loss term corresponding to high-displacement training data when calculating the loss function, thereby prompting the model to pay more attention to and accurately learn the features of these key samples.
[0072] For example, Figure 4 A flowchart illustrating exemplary training steps for a policy network and a prediction network provided in an embodiment of this application is shown. (Refer to...) Figure 4 In Phase 1 (without the barrier mechanism enabled), the current environment state is processed by the shared network layer, and the controller, which includes two policy networks and one value network, outputs the policy mean and standard deviation to generate alternative actions (there may be unsafe actions that exceed the safety threshold). In this phase, the alternative actions are executed directly without safety checks. After each round, the experience of state, action, reward, etc., is saved to a buffer. When the accuracy of the prediction network reaches the target, the experience is extracted from the buffer and high displacement samples are selected and stored in the high displacement experience pool. In Phase 2 (with the barrier mechanism enabled), safe actions are generated by comparing displacement constraints with multi-step displacement prediction sequences and adjusting thrust with thrust constraints. At the same time, the high displacement experience pool data accumulated in Phase 1 and the data-augmented samples are used to train the policy network to learn to maximize long-term rewards under safety constraints, and the value network to more accurately estimate state value to improve policy stability. Through the experience buffer, high displacement experience pool, and data augmentation, the two phases realize the transition from unconstrained exploration to collaborative optimization under safety constraints, supporting iterative improvements in performance and safety of the control policy and prediction network.
[0073] In this embodiment, during the training of the policy network and prediction network, high-displacement state samples generated by the wave energy conversion device during operation are first specifically sampled to generate a training dataset containing these key extreme cases. This high-displacement training data represents the device's behavior patterns when approaching its safety limits, which is crucial for the prediction network to accurately identify potential exceedance risks. Subsequently, when updating the model parameters of the policy network and prediction network with the goal of collaboratively optimizing energy capture efficiency, this high-displacement training data is utilized selectively. This means that while learning how to maximize energy capture, the model will pay more attention to and accurately learn the system dynamics and predictions under extreme displacement conditions. In this way, the prediction network can more accurately predict the future motion state of the device under extreme sea states, thereby providing a more reliable risk assessment for the control optimization unit. When the prediction network can accurately identify high-displacement risks, the control optimization unit can more timely and effectively correct candidate control commands, ensuring that the wave energy conversion device can still capture as much energy as possible while maintaining safe operation. This mechanism enables the entire control system to balance safety and energy capture efficiency when facing complex and variable marine environments, especially extreme sea states, avoiding performance degradation caused by data sparsity.
[0074] Figure 5 A structural diagram of a control system for a wave energy conversion device according to an embodiment of this application is shown. (Refer to...) Figure 5 This application provides a control system suitable for wave energy conversion devices, including: The controller 10 is used to configure the strategy network. The strategy network takes the real-time motion state and real-time environmental information of the wave energy conversion device as input and outputs candidate control commands for the wave energy conversion device.
[0075] In some embodiments, the real-time motion state of the wave energy conversion device includes: real-time floating body displacement and real-time floating body velocity. Real-time environmental information includes: real-time external wave information.
[0076] The controller 10 is also configured with a shared feature extraction network 11, which extracts features from real-time floating body displacement, real-time floating body velocity and real-time external wave information to generate a shared feature vector.
[0077] The policy network generates candidate control instructions based on shared feature vectors, and the prediction network generates prediction results based on shared feature vectors.
[0078] The state estimation unit 20 is used to configure the prediction network. The prediction network performs multi-step prediction of the motion state of the wave energy conversion device at multiple future moments based on the real-time motion state of the wave energy conversion device, real-time environmental information, and candidate control commands, and generates prediction results.
[0079] The control optimization unit 30 is used to modify the candidate control command when the predicted result indicates that the motion state at any future time exceeds the safety threshold, so as to obtain a safe control command and control the wave energy conversion device to execute the safe control command.
[0080] In some embodiments, the control optimization unit 30 includes: The adjustment subunit 31 is used to adjust the candidate control instructions according to the preset direction and preset step size within the current control cycle to obtain the candidate control instructions for the current iteration.
[0081] In some embodiments, the candidate control command includes at least a control force. The prediction result of the current iteration includes at least a predicted displacement.
[0082] The process of determining the preset direction includes: Preset positive disturbance values and preset negative disturbance values are applied to the candidate control commands to obtain positive disturbance commands and negative disturbance commands.
[0083] By using a prediction network, positive and negative disturbance commands are predicted separately, resulting in the predicted displacements of the positive and negative disturbance commands.
[0084] If the predicted displacement of a positive disturbance command is greater than the predicted displacement of a negative disturbance command, the preset direction will be determined as increasing the control force.
[0085] If the predicted displacement of a positive disturbance command is less than the predicted displacement of a negative disturbance command, the preset direction will be set to reduce the control force.
[0086] The process of determining the preset step size includes: Determine the number of iterations for the candidate control instructions; The preset step size is determined based on the preset basic correction step size and the number of iterations of the candidate control instructions.
[0087] The generation subunit 32 is used to generate the prediction result of the current iteration based on the prediction network, according to the real-time motion state of the wave energy conversion device, real-time environmental information and candidate control commands of the current iteration.
[0088] Stop iteration subunit 33 is used to repeat the above steps when the prediction result of the current iteration represents the motion state at any future time exceeding the safety threshold, until the motion state at all future times in the prediction result of the current iteration does not exceed the safety threshold, then stop the iteration and obtain a safety control instruction.
[0089] In some embodiments, a training unit is further included, the training unit comprising: The first training subunit is used in the first training phase: without enabling the constraint correction process, the policy network learns the dynamic characteristics of the wave energy conversion device to update its own model parameters, aiming to maximize energy capture efficiency under no safety constraints. Here, safety constraints refer to the prediction results ensuring that the motion states at all future moments do not exceed a safety threshold.
[0090] The second training subunit, used in the second training phase, enables the constraint correction process. Under the condition of satisfying safety constraints, the policy network and the prediction network update their respective model parameters with the goal of jointly optimizing the energy capture efficiency of the policy network and the prediction network, including: Using high displacement training data, the model parameters of the policy network and the prediction network are updated with the goal of co-optimizing the energy capture efficiency of the policy network and the prediction network.
[0091] The sampling subunit is used to sample high-displacement state samples to generate high-displacement training data.
[0092] In some embodiments, the method further includes sampling high-displacement state samples to generate high-displacement training data.
[0093] With the goal of jointly optimizing the energy capture efficiency of the policy network and the prediction network, the model parameters of the policy network and the prediction network are updated, including: Using high displacement training data, the model parameters of the policy network and the prediction network are updated with the goal of co-optimizing the energy capture efficiency of the policy network and the prediction network.
[0094] Figure 6 A schematic diagram of the structure of an electronic device according to an embodiment of this application is shown. (Refer to...) Figure 6 This application also provides an electronic device, including: processor.
[0095] Memory is used to store processor-executable instructions.
[0096] The processor is configured to execute instructions to implement any control method applicable to the wave energy conversion device.
[0097] In this embodiment, the computer device includes a processor, memory, and network interface connected via a system bus. The computer device's processor provides computing and control capabilities. Its memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs in the non-volatile storage media. The computer device's database stores data samples. Its network interface is used for communication with external terminals via a network connection. When the processor executes the computer program, it implements any control method applicable to the wave energy conversion device.
[0098] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0099] This application also provides a computer-readable storage medium, which, when the instructions in the computer-readable storage medium are executed by the processor of a terminal, enables the terminal to execute any control method applicable to a wave energy conversion device.
[0100] The aforementioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0101] Optionally, a readable storage medium can be coupled to the processor, enabling the processor to read information from and write information to the readable storage medium. Alternatively, the readable storage medium can be an integral part of the processor. The processor and the readable storage medium can reside in application-specific integrated circuits (ASICs). Of course, the processor and the readable storage medium can also exist as discrete components in the device.
[0102] This application also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described xxx method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0103] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0104] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0105] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0106] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A control method suitable for wave energy conversion devices, characterized in that, include: The controller is configured with a strategy network, which takes the real-time motion state and real-time environmental information of the wave energy conversion device as input and outputs candidate control commands for the wave energy conversion device. The state estimation unit is configured with a prediction network. The prediction network performs multi-step predictions of the motion state of the wave energy conversion device at multiple future moments based on the real-time motion state of the wave energy conversion device, the real-time environmental information, and the candidate control commands, and generates prediction results. If the predicted state of motion at any future time exceeds a safety threshold, the control optimization unit modifies the candidate control command to obtain a safe control command, and controls the wave energy conversion device to execute the safe control command.
2. The method as described in claim 1, characterized in that, The candidate control commands are modified to obtain safety control commands, including: Within the current control cycle, the candidate control commands are adjusted according to a preset direction and a preset step size to obtain the candidate control commands for the current iteration; Based on the real-time motion state of the wave energy conversion device, the real-time environmental information, and the candidate control commands for the current iteration, the prediction result for the current iteration is generated using the prediction network. If the prediction result of the current iteration indicates that the motion state at any future time exceeds the safety threshold, the above steps are repeated until the motion state at all future times in the prediction result of the current iteration does not exceed the safety threshold, then the iteration is stopped and the safety control instruction is obtained.
3. The method as described in claim 2, characterized in that, The candidate control command includes at least a control force; the prediction result of the current iteration includes at least a predicted displacement. The process of determining the preset direction includes: Preset positive disturbance values and preset negative disturbance values are applied to the candidate control command respectively to obtain positive disturbance command and negative disturbance command; The prediction network is used to predict the positive disturbance command and the negative disturbance command respectively, so as to obtain the predicted displacement of the positive disturbance command and the predicted displacement of the negative disturbance command. If the predicted displacement of the positive disturbance command is greater than the predicted displacement of the negative disturbance command, the preset direction is determined to increase the control force. If the predicted displacement of the positive disturbance command is less than the predicted displacement of the negative disturbance command, the preset direction is determined to reduce the control force.
4. The method as described in claim 2 or 3, characterized in that, The process of determining the preset step size includes: Determine the number of iterations for the candidate control instruction; The preset step size is determined based on the preset basic correction step size and the number of iterations of the candidate control command.
5. The method as described in claim 1, characterized in that, The real-time motion status of the wave energy conversion device includes: real-time floating body displacement and real-time floating body velocity; the real-time environmental information includes: real-time external wave information. The controller is also configured with a shared feature extraction network, and the method further includes: The shared feature extraction network extracts features from real-time floating body displacement, real-time floating body velocity, and real-time external wave information to generate a shared feature vector. The policy network generates the candidate control command based on the shared feature vector, and the prediction network generates the prediction result based on the shared feature vector.
6. The method as described in claim 1, characterized in that, The training process for the policy network and the prediction network includes: First training phase: Without enabling the constraint correction process, the policy network aims to maximize energy capture efficiency and learns the dynamic characteristics of the wave energy conversion device under no safety constraints to update the model parameters of the policy network itself; wherein, the safety constraint means that the motion state at all future moments in the prediction results does not exceed the safety threshold. Second training phase: The constraint correction process is activated. Under the condition of satisfying the safety constraints, the policy network and the prediction network update their respective model parameters with the goal of jointly optimizing the energy capture efficiency of the policy network and the prediction network.
7. The method as described in claim 6, characterized in that, The method further includes: High displacement state samples are sampled to generate high displacement training data; With the goal of collaboratively optimizing the energy capture efficiency of the policy network and the prediction network, the model parameters of the policy network and the prediction network are updated, including: Using the high displacement training data, the model parameters of the policy network and the prediction network are updated with the goal of co-optimizing the energy capture efficiency of the policy network and the prediction network.
8. A control system suitable for wave energy conversion devices, characterized in that, include: A controller is used to configure a strategy network, which takes the real-time motion state and real-time environmental information of the wave energy conversion device as input and outputs candidate control commands for the wave energy conversion device. A state estimation unit is used to configure a prediction network. The prediction network performs multi-step prediction of the motion state of the wave energy conversion device at multiple future moments based on the real-time motion state of the wave energy conversion device, the real-time environmental information, and the candidate control commands, and generates prediction results. The control optimization unit is used to modify the candidate control command to obtain a safe control command when the predicted result indicates that the motion state at any future time exceeds a safety threshold, and to control the wave energy conversion device to execute the safe control command.
9. An electronic device, characterized in that, include: processor; Memory for storing processor-executable instructions; The processor is configured to execute the instructions to implement the control method applicable to wave energy conversion devices as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the terminal, the terminal is able to perform the control method applicable to a wave energy conversion device as described in any one of claims 1-7.