A method and system for coordinated control of ship attitude and hydrodynamic multi-actuator

By constructing a ship intelligent agent and using the control dominant mode triangle and the degree of influence of mapping to train the reward value, the robustness and adaptability of multiple actuators on ships under complex sea conditions are solved, and stable collaborative control of ships is achieved.

CN122131769APending Publication Date: 2026-06-02GUANGDONG OCEAN UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG OCEAN UNIVERSITY
Filing Date
2026-03-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing collaborative control methods for multiple actuators on ships exhibit poor robustness and adaptability when facing complex, nonlinear, and time-varying sea conditions, leading to decreased control stability. Furthermore, the linear weighted summation method cannot effectively detect changes in different control objectives, resulting in ship control instability.

Method used

A ship intelligent agent is constructed. The control dominant mode triangle is initialized by pre-acquired attitude data, hydrodynamic data and adjustment command data. The mapping influence of the mapping triangle and historical mapping points is combined to perform recursive division and generate reward values ​​to train the intelligent agent, thereby realizing the cooperative control of the ship.

Benefits of technology

It improves the control stability of ships in complex sea conditions, ensures smooth and coordinated control of multiple actuators, and enhances control stability during navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122131769A_ABST
    Figure CN122131769A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for collaborative control of ship attitude and hydrodynamic multi-actuator systems, belonging to the field of ship control technology. The invention constructs an agent to be trained and initializes several control-dominant mode triangles. Then, at each interaction moment during the agent's training process, the control results are categorized into the control-dominant mode triangles, determining the inferior control target reward value, the mapping triangle, and the current mapping point formed within the mapping triangle. Next, the mapping triangle is recursively divided based on the degree of mapping influence between the current mapping point and historical mapping points, obtaining historical distribution fluctuation reward values. Then, reward values ​​are formed based on the historical distribution fluctuation reward values ​​and the inferior control target reward values, thereby generating empirical data to train the agent, resulting in a ship collaborative control agent. Finally, the ship collaborative control agent outputs the ship's collaborative control scheme, improving the ship's control stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of ship control technology, and in particular relates to a method and system for coordinated control of ship attitude and hydrodynamic multi-actuator. Background Technology

[0002] When ships navigate at sea, they are subject to complex and unstable wind, waves, and currents, which can cause complex rolling motions in the hull, seriously affecting navigational safety, crew comfort, and the smooth operation of maritime tasks. To ensure the ship's navigational performance, modern ships are typically equipped with various actuators, such as anti-roll fins, steering gear, side thrusters, and main propulsion systems. Through the coordinated control of these actuators, rolling is suppressed, course is maintained, energy consumption is reduced, and the ship's attitude is stabilized in adverse sea conditions.

[0003] Currently, traditional solutions for the coordinated control of multiple actuators mostly use PID control. However, PID control relies on precise mathematical models and has fixed parameters. When facing complex, nonlinear, and time-varying sea conditions, its robustness and adaptability are poor, making it difficult to achieve stable coordinated control. With the development of reinforcement learning technology, intelligent agents have begun to be applied to the coordinated control of ships. Most current intelligent agents set multiple reward values ​​for each control objective, and then combine these multiple control objectives into a total reward value using a linear weighted summation method, which serves as the sole signal driving the agent's learning. However, this linear weighted summation method loses the relationship between different control objectives and fails to perceive changes in different control objectives. This leads to situations where one control objective is over-adjusted while another is under-adjusted, resulting in ship control instability. Moreover, the reward value setting method, which is solely guided by the control objective, can lead to excessively large adjustment ranges in practical applications, causing a decrease in ship control stability when facing wind, waves, and current disturbances that require fine-tuning. Summary of the Invention

[0004] The present invention aims to provide a method and system for coordinated control of ship attitude and hydrodynamic multi-actuator to solve the above-mentioned technical problems, so as to realize the smooth control of multiple actuators of ships under complex sea conditions and improve the control stability of ships during navigation.

[0005] To address the aforementioned technical problems, embodiments of the present invention provide a method for coordinated control of ship attitude and hydrodynamic multi-actuator systems, comprising:

[0006] The ship's intelligent agent is constructed based on the pre-acquired ship attitude data, hydrodynamic data, and adjustment command data, and several control dominant mode triangles of the ship are initialized based on the pre-acquired control objectives of the ship. At each interaction moment, the next state and control result are determined based on the ship's current state and actions; the control result is categorized into the control dominance mode triangle, and the inferior control target reward value, mapping triangle, and current mapping point formed in the mapping triangle are determined; The mapping influence of the current mapping point on the historical mapping points at historical interaction moments within the mapping triangle is obtained; and the mapping triangle is recursively divided based on the mapping influence to obtain the historical distribution fluctuation reward value of the control result. The reward value at the current interaction moment is determined based on the disadvantage control target reward value and the historical distribution fluctuation reward value; empirical data for the current interaction moment is constructed based on the current state, action, next state, and reward value; the agent is trained based on the empirical data to obtain the ship cooperative control agent; The system acquires the ship's real-time attitude data and real-time hydrodynamic data, inputs the real-time attitude data and real-time hydrodynamic data into the ship's collaborative control agent, obtains the ship's collaborative control scheme, and performs collaborative control on the ship based on the collaborative control scheme.

[0007] Understandably, compared to existing technologies, this invention constructs an agent based on pre-acquired ship attitude data, hydrodynamic data, and adjustment command data, and initializes several control dominance mode triangles according to multiple preset control objectives of the ship, thus obtaining the agent to be trained. Then, at each interaction moment during the agent training process, the control result is categorized into the control dominance mode triangle, determining the inferior control objective reward value, the mapping triangle, and the current mapping point formed within the mapping triangle; the inferior control objective reward value characterizes the change of the control result under multiple control objectives. Next, by measuring the degree of mapping influence between the current mapping point in the mapping triangle and the historical mapping points of historical interaction moments, the relationship between the control result at the current interaction moment and the data at historical interaction moments is measured. Thus, the mapping triangle is recursively divided according to the degree of mapping influence to obtain a historical distribution fluctuation reward value that characterizes the ship control relationship between the current interaction moment and historical interaction moments. Then, the reward value for the current interaction moment is formed based on the historical distribution fluctuation reward value and the inferior control objective reward value. Furthermore, the empirical data for the current interaction moment is constructed by combining the current state, action, next state, and reward value. The agent is trained using this empirical data to obtain the ship cooperative control agent. Ultimately, the ship's collaborative control scheme is output by the ship collaborative control agent, which enables the smooth control of the ship's multiple actuators in complex sea conditions and improves the control stability during the ship's navigation.

[0008] As a preferred embodiment, the construction of the ship's intelligent agent based on pre-acquired ship attitude data, hydrodynamic data, and adjustment command data includes: The state space is constructed based on the pre-acquired ship attitude and hydrodynamic data; Construct the action space based on the pre-acquired adjustment instruction data; The ship's intelligent agent is constructed based on the action space and the state space.

[0009] As a preferred embodiment, the pre-acquired control objectives of the ship include: roll reduction effect, ship energy consumption, sailing resistance, and comfort; the initialization of several control dominant mode triangles of the ship based on the pre-acquired control objectives includes: A regular tetrahedron is constructed within a pre-defined unit sphere, and several initial triangles are formed based on the regular tetrahedron. Based on the roll reduction effect, ship energy consumption, sailing resistance and comfort, the extreme value components of the control target of the intelligent agent under each of the control objectives are generated; Each of the control target extreme value components is mapped to each vertex of the regular tetrahedron to initialize the vertex control target of each of the initial triangles; Based on the control target of each vertex of the initial triangle, the control dominance mode of each initial triangle is determined, resulting in several control dominance mode triangles for the ship.

[0010] As a preferred embodiment, at each interaction moment, determining the next state and control result based on the ship's current state and actions; classifying the control result into the control dominance mode triangle, determining the disadvantageous control target reward value of the control result, the mapping triangle, and the current mapping point formed in the mapping triangle, includes: At each interaction moment, the control result component of the ship's next state under each control objective is determined based on the ship's current state and actions, and the control result of the ship is formed based on the control result component; Each component of the control result is normalized to obtain several normalized control result components; The normalized control result components are classified into the control dominant mode triangle to obtain the control dominant mode, mapping triangle and the current mapping point formed in the mapping triangle of the control result; Based on the control dominance mode of the control results, the control targets are screened to determine the inferior control targets of the control results; The disadvantage control objective reward value of the control result is determined based on the disadvantage control objective and the normalized control result components.

[0011] As a preferred embodiment, classifying the normalized control result components to the control dominance mode triangle to obtain the control dominance mode, mapping triangle, and current mapping point formed in the mapping triangle of the control result includes: The normalized control result components are transformed based on a preset exponential function to form a reward weight for each normalized control result component. The vertex coordinates of the regular tetrahedron are weighted and summed based on the reward weights to determine the control target direction vector of the control result within the regular tetrahedron. With the center of the regular tetrahedron as the origin, a control target guide ray is formed within the regular tetrahedron along the control target direction vector based on the origin; The intersection points of the control target guide ray and each of the control dominant mode triangles are calculated based on a preset ray triangle intersection algorithm; Based on the intersection point, the control dominance mode triangle is filtered to determine the first intersection point and the first control dominance mode triangle corresponding to the control result; and the control dominance mode of the control result is determined based on the first control dominance mode triangle. The first control dominant mode triangle is filtered according to a preset triangle walking algorithm to determine the mapping triangle of the control result; and the current mapping point is formed in the mapping triangle based on the first intersection point.

[0012] As a preferred embodiment, the step of obtaining the mapping influence degree of the current mapping point on the historical mapping points at historical interaction moments within the mapping triangle; and recursively dividing the mapping triangle based on the mapping influence degree to obtain the historical distribution fluctuation reward value of the control result includes: Based on the historical mapping points of the historical interaction moments within the mapping triangle, determine the set of historical mapping points of the mapping triangle; The current mapping point set of the mapping triangle is determined based on the current mapping point and the historical mapping points; Calculate the relative entropy between the historical mapping point set and the current mapping point set, and determine the degree of influence of the current mapping point on the historical mapping points at historical interaction moments within the mapping triangle based on the relative entropy; Initialize the number of recursive partitions of the mapped triangle; The mapping triangle is recursively divided based on the degree of influence of the mapping, so as to update the number of recursive divisions; The historical distribution fluctuation reward value of the control result is obtained based on the number of recursive partitions completed for updating.

[0013] As a preferred embodiment, the recursive partitioning of the mapping triangle based on the degree of mapping influence to update the recursive partitioning number includes: If the degree of influence of the mapping does not exceed a preset first threshold, then the number of recursive partitions is set to a preset zero value; If the influence of the mapping exceeds the first threshold, the mapping triangle is divided into four equal parts to form several sub-triangles, and the number of recursive divisions is incremented by one; all mapping points in the current mapping point set are sorted, and the mapping triangle of each mapping point is re-determined based on the mapping point order and several sub-triangles of the mapping triangle, and each mapping point is stored in the mapping triangle as the current mapping point of the corresponding mapping triangle.

[0014] As a preferred embodiment, the step of acquiring the ship's real-time attitude data and real-time hydrodynamic data, inputting the real-time attitude data and real-time hydrodynamic data into the ship's cooperative control agent to obtain the ship's cooperative control scheme, and performing cooperative control of the ship based on the cooperative control scheme includes: The real-time attitude data and real-time hydrodynamic data of the ship are acquired, and the real-time attitude data and real-time hydrodynamic data are input into the ship cooperative control agent to obtain several candidate cooperative control actions. The candidate cooperative control actions are simulated in a preset ship navigation simulation environment to determine the control simulation results of each candidate cooperative control action. Each of the control simulation results is categorized into the control dominant mode triangle, and the simulation mapping triangle for each of the candidate cooperative control actions is determined. Quantum annealing simulation is performed on the candidate cooperative control actions based on the historical reference mapping points within the simulation mapping triangle of each candidate cooperative control action to determine the cooperative control scheme of the ship. The ship is controlled collaboratively based on the aforementioned collaborative control scheme.

[0015] As a preferred embodiment, the step of performing quantum annealing simulation on the candidate cooperative control actions based on the historical reference mapping points within the simulation mapping triangle of each candidate cooperative control action to determine the cooperative control scheme of the ship includes: Based on the historical comparison mapping points within the simulation mapping triangle of each candidate cooperative control action, determine the number of mapping points, average reward value, and centroid of the mapping points for each candidate cooperative control action. The centroid of the mapping point is inversely mapped to the surface of the unit sphere to determine the spherical mapping point for each candidate cooperative control action; The local action capability value is determined by the number of mapping points and the average reward value, and the global action capability value is determined by the spherical angle between every two spherical mapping points. An action energy function is constructed based on the local action capability value and the global action capability value, and the action energy function is evolved based on a preset quantum annealing simulation algorithm to determine the energy value of each candidate cooperative control action; Based on the energy value, the candidate cooperative control actions are screened to obtain the cooperative control scheme for the ship.

[0016] Accordingly, embodiments of the present invention provide a ship attitude and hydrodynamic multi-actuator collaborative control system, including: an initialization module, a control result classification module, a mapping effect module, an agent training module, and a ship collaborative control module; The initialization module is used to construct the ship's intelligent agent based on the pre-acquired ship attitude data, hydrodynamic data, and adjustment command data, and to initialize several control dominant mode triangles of the ship based on the pre-acquired control objectives of the ship. The control result classification module is used to determine the next state and control result based on the current state and actions of the ship at each interaction moment; classify the control result into the control dominance mode triangle, and determine the inferior control target reward value, mapping triangle, and the current mapping point formed in the mapping triangle of the control result; The mapping influence module is used to obtain the degree of mapping influence of the current mapping point on the historical mapping points at historical interaction moments within the mapping triangle; and to recursively divide the mapping triangle based on the degree of mapping influence to obtain the historical distribution fluctuation reward value of the control result. The agent training module is used to determine the reward value at the current interaction moment based on the disadvantage control target reward value and the historical distribution fluctuation reward value; to construct empirical data for the current interaction moment based on the current state, action, next state, and reward value; and to train the agent based on the empirical data to obtain a ship cooperative control agent. The ship collaborative control module is used to acquire the ship's real-time attitude data and real-time hydrodynamic data, input the real-time attitude data and real-time hydrodynamic data into the ship collaborative control agent to obtain the ship's collaborative control scheme, and perform collaborative control on the ship based on the collaborative control scheme.

[0017] Understandably, compared to existing technologies, this system constructs an agent based on pre-acquired ship attitude data, hydrodynamic data, and adjustment command data. It initializes several control-dominant mode triangles based on multiple preset control objectives of the ship, resulting in an agent to be trained. Then, at each interaction moment during the agent's training process, the control result is categorized into the control-dominant mode triangle, determining the inferior control objective reward value, the mapping triangle, and the current mapping point formed within the mapping triangle. The inferior control objective reward value characterizes the change in the control result under multiple control objectives. Next, by measuring the degree of mapping influence between the current mapping point in the mapping triangle and the historical mapping points of historical interaction moments, the relationship between the control result at the current interaction moment and the data at historical interaction moments is measured. This allows for recursive division of the mapping triangle based on the degree of mapping influence, yielding a historical distribution fluctuation reward value that characterizes the ship control relationship between the current interaction moment and historical interaction moments. Then, the reward value for the current interaction moment is formed based on the historical distribution fluctuation reward value and the inferior control objective reward value. Finally, by combining the current state, action, next state, and reward value, empirical data for the current interaction moment is constructed. This empirical data is used to train the agent, resulting in a ship cooperative control agent. Ultimately, the ship's collaborative control scheme is output by the ship collaborative control agent, which enables the smooth control of the ship's multiple actuators in complex sea conditions and improves the control stability during the ship's navigation. Attached Figure Description

[0018] Figure 1 A flowchart illustrating the steps of a method for coordinated control of ship attitude and hydrodynamic multi-actuator according to an embodiment of the present invention; Figure 2 This is a schematic diagram of a ship attitude and hydrodynamic multi-actuator collaborative control system provided in an embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Example 1 Please refer to Figure 1 , Figure 1 The flowchart of a method for coordinated control of ship attitude and hydrodynamic multi-actuator provided in an embodiment of the present invention includes steps S101 to S105.

[0021] Step S101: Construct the ship's intelligent agent based on the pre-acquired ship attitude data, hydrodynamic data, and adjustment command data, and initialize several control dominant mode triangles of the ship based on the pre-acquired control objectives of the ship.

[0022] In this embodiment, constructing the ship's intelligent agent based on pre-acquired ship attitude data, hydrodynamic data, and adjustment command data includes: The state space is constructed based on the pre-acquired ship attitude and hydrodynamic data; Construct the action space based on the pre-acquired adjustment instruction data; The ship's intelligent agent is constructed based on the action space and the state space.

[0023] In an optional embodiment, attitude data and hydrodynamic data are obtained by sensors installed on the ship; wherein, attitude data includes: roll angle. Pitch angle Roll angular velocity and pitch angular velocity Hydrodynamic data includes: water flow velocity Water flow direction and the water pressure exerted on the ship by the water flow Therefore, the constructed state space is as follows: The pre-acquired adjustment command data refers to the adjustment commands of each actuator in the ship. The adjustment command data includes: the fin angle of the anti-roll fin. Main thruster thrust servo motor rudder angle Power distribution ratio of side thrusters Therefore, the action space constructed using adjustment instruction data is represented as... Next, after constructing the action space and state space, a deep Q-network (DQN) is used as the basic architecture of the agent to form the ship's intelligent agent.

[0024] This embodiment constructs a state space based on pre-acquired ship attitude and hydrodynamic data, and an action space based on pre-acquired adjustment command data, providing the agent with a complete foundation for environmental perception and action execution. During the construction of the state and action spaces, by incorporating attitude and hydrodynamic data into the state space, the state space acquires comprehensive environmental information reflecting the ship's own state and external disturbances. Then, the action space is constructed based on adjustment command data from multiple actuators, allowing the agent to obtain control options for the ship's actuators. Finally, a reinforcement learning agent is formed based on the state and action spaces, enabling the subsequent collaborative control scheme based on the agent to achieve stable control of the ship's multiple actuators in complex sea conditions, thus improving the control stability during ship navigation.

[0025] In this embodiment, the pre-acquired control objectives of the ship include: roll reduction effect, ship energy consumption, sailing resistance, and comfort; the initialization of several control dominant mode triangles of the ship based on the pre-acquired control objectives includes: A regular tetrahedron is constructed within a pre-defined unit sphere, and several initial triangles are formed based on the regular tetrahedron. Based on the roll reduction effect, ship energy consumption, sailing resistance and comfort, the extreme value components of the control target of the intelligent agent under each of the control objectives are generated; Each of the control target extreme value components is mapped to each vertex of the regular tetrahedron to initialize the vertex control target of each of the initial triangles; Based on the control target of each vertex of the initial triangle, the control dominance mode of each initial triangle is determined, resulting in several control dominance mode triangles for the ship.

[0026] In an optional embodiment, the control objectives pre-acquired by the ship include: roll reduction effect. Ship energy consumption Navigation resistance With comfort Define the current interaction moment as The previous interaction time was ; This represents the time interval between two adjacent interaction moments. The roll reduction effect is measured by the ship's roll angle. and pitch angle The calculated anti-slip effect Ship energy consumption can be defined and measured through ship operation monitoring platforms or by deploying sensors in relevant power facilities. Therefore, ship energy consumption... Represented as Navigation resistance Comfort can be calculated using classic empirical formulas such as the Zvankov formula and the Elco method, or using theoretical models such as the Froude classification method. The ship's roll rate can be used as a reference. and pitch angular velocity Calculated; In particular, regarding the specific calculation method for the control target described in this embodiment, those skilled in the art can choose other methods to obtain it according to actual needs.

[0027] In an optional embodiment, an inscribed regular tetrahedron is constructed within a unit sphere (radius 1). The tetrahedron has four vertices, which form four initial triangles. Next, the extreme control target component is defined as the component where only one control target occupies the full weight, while the weight of the remaining control targets is 0. This control target mechanism component reflects the extreme control situation where, during the coordinated control of multiple actuators, the ship focuses on only one control target while completely ignoring the others. Then, each extreme control target component is sequentially weighted and summed with each vertex of the tetrahedron, thus mapping the extreme control target component to each vertex of the tetrahedron. Through this mapping, each vertex of the tetrahedron can represent each extreme control target. Three vertices form an initial triangle, whose vertex control target naturally lacks one control target. Therefore, the resulting control-dominant mode triangle can represent the control target that focuses on the three vertices while ignoring the remaining control target. This allows the representation of the weaker control target in subsequent control result classification.

[0028] For example, construct an inscribed regular tetrahedron within a unit sphere, with its four vertices defined as follows: Therefore, the control target extreme value component of the roll reduction effect is set as follows: The extreme value component of the control target for ship energy consumption is set as follows: The extreme component of the control target for navigation resistance is set as follows: The extreme component of the comfort control target is set as follows: Therefore, the weighted sum of the extreme components of the control target for the anti-slip effect and the four vertices is expressed as follows: ; its and This vertex corresponds to, therefore the roll reduction effect is mapped to... This vertex; the extreme value component of the ship's energy consumption control target is represented by a weighted sum of the four vertices. , and This vertex corresponds to, therefore the ship's energy consumption is mapped to This vertex; the extreme value component of the control target for navigation resistance is represented by a weighted sum of the four vertices. , and This vertex corresponds to, therefore the ship's energy consumption is mapped to This vertex; the extreme value component of the comfort control objective is represented by a weighted sum of the four vertices. , and This vertex corresponds to, therefore comfort is mapped to This vertex.

[0029] Therefore, by The initial triangle formed by these factors represents the dominant control modes of ship energy consumption, sailing resistance, and comfort, with roll reduction being the corresponding disadvantageous control objective. The initial triangle formed by these factors represents the dominant control modes of ship energy consumption, sailing resistance, and roll reduction, with comfort being the corresponding disadvantageous control objective. The initial triangle formed by these three control modes prioritizes roll reduction, sailing resistance, and comfort, while the corresponding disadvantageous control objective is ship energy consumption. The initial triangle formed by these factors represents the dominant control modes of ship energy consumption, roll reduction effect, and comfort, while the corresponding disadvantageous control objective is sailing resistance.

[0030] This embodiment constructs a regular tetrahedron within a preset unit sphere and generates extreme components of the control objectives based on four control goals: roll reduction effect, ship energy consumption, navigation resistance, and comfort. These components are then mapped to the vertices of the tetrahedron to initialize several triangles with clearly defined control dominance modes. By generating extreme components that characterize the control objectives under extreme conditions and mapping them to different vertices of the regular tetrahedron, multiple abstract control objectives with different dimensions are visualized as vertices of a regular tetrahedron. This allows the trade-offs between multiple control objectives to be intuitively represented in geometric positions, achieving independent representation of control objectives in geometric space. Subsequently, the control dominance mode of the initial triangle is determined by the control objectives at each vertex of the initial triangle. This provides an accurate basis for classifying control results during subsequent agent interaction, offering a structured mathematical foundation for achieving multi-objective dynamic balance control. Furthermore, it enables the analysis of relationships between different control objectives in the control results, avoiding the flattening problem of control objective relationships caused by existing linear weighted summation. This allows the subsequent agent-based collaborative control scheme to achieve stable control of multiple actuators of the ship in complex sea conditions, improving the control stability during ship navigation.

[0031] Step S102: At each interaction moment, determine the next state and control result based on the current state and actions of the ship; classify the control result into the control dominant mode triangle, and determine the inferior control target reward value, mapping triangle and the current mapping point formed in the mapping triangle.

[0032] In this embodiment, at each interaction moment, determining the next state and control result based on the ship's current state and actions; classifying the control result into the control dominance mode triangle, determining the disadvantageous control target reward value, mapping triangle, and the current mapping point formed in the mapping triangle, includes: At each interaction moment, the control result component of the ship's next state under each control objective is determined based on the ship's current state and actions, and the control result of the ship is formed based on the control result component; Each component of the control result is normalized to obtain several normalized control result components; The normalized control result components are classified into the control dominant mode triangle to obtain the control dominant mode, mapping triangle and the current mapping point formed in the mapping triangle of the control result; Based on the control dominance mode of the control results, the control targets are screened to determine the inferior control targets of the control results; The disadvantage control objective reward value of the control result is determined based on the disadvantage control objective and the normalized control result components.

[0033] In this embodiment, after determining the next state based on the current state and action at each interaction moment, the control result components of the ship under each control objective are determined according to the current state and the next state. These control result components are then normalized to obtain normalized control result components that eliminate the influence of different control objective dimensions. Subsequently, the normalized control result components are categorized into the control dominance mode triangle to obtain the control dominance mode, mapping triangle, and current mapping point of the control result at the current interaction moment. Based on the control dominance mode, inferior control objectives are selected, and the inferior control objective reward value representing the worst-performing control objective in the current control result can be obtained. This allows the reward value generated based on the inferior control objective reward value to guide the agent to focus on the shortcomings in the current multi-actuator control, thereby improving the corresponding control of the inferior control objective. This achieves global dynamic balance in collaborative control, preventing the ship's control from becoming unstable due to the long-term neglect of a certain control objective, and realizing stable control of the ship's multiple actuators in complex sea conditions, thus improving the control stability during ship navigation.

[0034] In this embodiment, classifying the normalized control result components to the control dominance mode triangle to obtain the control dominance mode, mapping triangle, and current mapping point formed in the mapping triangle of the control result includes: The normalized control result components are transformed based on a preset exponential function to form a reward weight for each normalized control result component. The vertex coordinates of the regular tetrahedron are weighted and summed based on the reward weights to determine the control target direction vector of the control result within the regular tetrahedron. With the center of the regular tetrahedron as the origin, a control target guide ray is formed within the regular tetrahedron along the control target direction vector based on the origin; The intersection points of the control target guide ray and each of the control dominant mode triangles are calculated based on a preset ray triangle intersection algorithm; Based on the intersection point, the control dominance mode triangle is filtered to determine the first intersection point and the first control dominance mode triangle corresponding to the control result; and the control dominance mode of the control result is determined based on the first control dominance mode triangle. The first control dominant mode triangle is filtered according to a preset triangle walking algorithm to determine the mapping triangle of the control result; and the current mapping point is formed in the mapping triangle based on the first intersection point.

[0035] This embodiment, when classifying the normalized control result components into the control dominance mode triangle, transforms the normalized control result components using a preset exponential function to obtain reward weights. The nonlinear transformation of the exponential function maps the normalized control result components to reward weights of different proportions, allowing subtle differences between control targets to be nonlinearly amplified. Then, the reward weights are used to perform a weighted summation of the vertex coordinates of a regular tetrahedron, yielding a control target direction vector representing the overall direction of the control targets within the tetrahedron. Finally, a ray triangle intersection algorithm and a triangle walking algorithm are used to accurately map the control target direction vector onto the control dominance mode triangle, thereby accurately determining the control dominance. The system employs a control-dominant mode, mapping triangles, and the formation of the current mapping point. It identifies the worst-performing control objective in the current control outcome, guiding the agent to focus on the weaknesses in the current multi-actuator control, thereby improving the control of the inferior target and achieving a global dynamic balance in collaborative control. The mapping triangle and the current mapping point characterize the relationship between the current control outcome and the control outcome at historical interaction moments, providing an accurate data foundation for generating subsequent historical distribution fluctuation reward values. This ensures accurate reward value generation and accurate agent training, enabling stable control of the ship's multi-actuator systems under complex sea conditions and improving control stability during ship navigation.

[0036] In an optional embodiment, after obtaining the control result components of the ship under each control objective, a normalization algorithm (e.g., Min-Max Scaling) is used to normalize the control result components to... The range is used to obtain the normalized control result components; Then set the exponential function as Function, use The function converts the normalized control result components into weights, thereby obtaining the reward weights, and through... The function obtains a sum of reward weights of 1. The magnitude of the reward weights is similar to the extreme component of the control objective described earlier; a larger reward weight indicates a larger proportion of the control objective in the overall control result, meaning the agent is currently focusing on that control objective. Next, the vertex coordinates of the regular tetrahedron are weighted and summed using the reward weights. Since the sum of the reward weights is 1, a control objective direction vector located inside the regular tetrahedron can be obtained through this weighted summation. Then, using the center of the regular tetrahedron as the origin, a control objective guide ray is formed along the direction of the control objective direction vector, representing the control result within the regular tetrahedron. A ray triangle intersection algorithm (e.g., the Möller-Trumbore algorithm) is used to calculate the intersection point between the control objective guide ray and the control dominant mode triangle. Since this embodiment uses a regular tetrahedron, the control target guide ray will only form valid intersections with two of the control dominant mode triangles; the other two control dominant mode triangles will form invalid intersections. For the valid intersections, this embodiment uses the control target's reward weight to perform a weighted sum of the four vertices. Therefore, the resulting control target direction vector will be closer to the vertex with the larger reward weight. Consequently, the corresponding control target guide ray will first intersect with the control dominant mode triangles with larger reward weights, which is represented as the intersection parameter in the Möller-Trumbore algorithm. The smaller the value, the better; therefore, the better the selection of intersection parameters. The smallest intersection point forms the first intersection point, and its corresponding control dominance pattern triangle is the first control dominance pattern triangle. Simultaneously, each control dominance pattern triangle has already formed its own control dominance pattern, so determining the first control dominance pattern triangle confirms the control dominance pattern of the control result. At the same time, the inferior control objective of the control result can also be identified, and the negative value of the normalized control result component corresponding to the inferior control objective is then used as the inferior control objective reward value. Through this operation, compared to the classification methods of more common clustering algorithms in existing technologies, which only focus on the overall value and lose the proportional or magnitude relationship between different control objectives during classification, this example, using spatial coordinates, can fully maintain and consider the proportional relationship between different control objectives, thus accurately reflecting the current agent's attention to different control objectives.

[0037] Since the first control dominant mode triangle may have already been recursively divided during previous interaction moments (see the description below for details), meaning that the first control dominant mode triangle may actually be divided into multiple sub-triangles, it is still necessary to find the minimum enclosing triangle of the first intersection point. In this embodiment, the Triangle Walking Algorithm is used to filter several sub-triangles (including the first control dominant mode triangle itself) within the first control dominant mode triangle to obtain the minimum enclosing triangle of the first intersection point. This embodiment defines the minimum enclosing triangle of the first intersection point as the mapping triangle of the control result. Furthermore, after determining the minimum enclosing triangle of the first intersection point, the first intersection point is directly used as the current mapping point in the mapping triangle.

[0038] For example, four control objectives are defined (roll reduction effect). Ship energy consumption Navigation resistance With comfort The normalized control result components are as follows: ; then use The function converts the normalized control result components into weights, thus obtaining the reward weights in the following order: Then, based on the reward weights, the vertex coordinates of the regular tetrahedron are weighted and summed to obtain the control target direction vector: ; Then, the first control dominant mode triangle was obtained by solving the ray triangle algorithm. Therefore, the dominant control mode of its control results is roll reduction effect, sailing resistance and comfort; the disadvantageous control objective is ship energy consumption, and the normalized control result component corresponding to ship energy consumption ( Taking a negative value represents the disadvantage of the control outcome, which is the reward value for the control objective. ; Specifically, the reward weights are as follows: The smallest value is 0.221, which indicates that the corresponding ship has the worst energy consumption performance. It corresponds to the control dominance mode of the first control dominance mode triangle.

[0039] Step S103: Obtain the degree of influence of the current mapping point on the historical mapping points at historical interaction moments within the mapping triangle; and recursively divide the mapping triangle based on the degree of influence to obtain the historical distribution fluctuation reward value of the control result.

[0040] In this embodiment, the step of obtaining the mapping influence of the current mapping point on the historical mapping points at historical interaction moments within the mapping triangle, and recursively dividing the mapping triangle based on the mapping influence to obtain the historical distribution fluctuation reward value of the control result, includes: Based on the historical mapping points of the historical interaction moments within the mapping triangle, determine the set of historical mapping points of the mapping triangle; The current mapping point set of the mapping triangle is determined based on the current mapping point and the historical mapping points; Calculate the relative entropy between the historical mapping point set and the current mapping point set, and determine the degree of influence of the current mapping point on the historical mapping points at historical interaction moments within the mapping triangle based on the relative entropy; Initialize the number of recursive partitions of the mapped triangle; The mapping triangle is recursively divided based on the degree of influence of the mapping, so as to update the number of recursive divisions; The historical distribution fluctuation reward value of the control result is obtained based on the number of recursive partitions completed for updating.

[0041] In this embodiment, after obtaining the current mapping point, a set of historical mapping points is determined based on the historical mapping points at historical interaction moments within the mapping triangle. This set is then combined with the current mapping point to determine the current mapping point set. By calculating the relative entropy between the historical and current mapping point sets, the degree of mapping influence, representing the deviation of the current control result from the distribution of historical control results, can be obtained. Subsequently, the mapping triangle is recursively divided according to the degree of mapping influence to update the number of recursive divisions. This allows the acquisition of the historical distribution fluctuation reward value, representing the disturbance caused by the deviation between the current and historical control results. This enables subsequent reward values ​​to guide the agent to penalize actions that cause severe fluctuations in the historical distribution, thereby effectively suppressing the problem of excessive adjustment amplitude and guiding the agent to make smooth and precise adjustments. Ultimately, this enables the stable control of multiple actuators of the ship in complex sea conditions, improving the control stability during ship navigation.

[0042] In this embodiment, the step of recursively partitioning the mapping triangle based on the degree of mapping influence to update the number of recursive partitions includes: If the degree of influence of the mapping does not exceed a preset first threshold, then the number of recursive partitions is set to a preset zero value; If the influence of the mapping exceeds the first threshold, the mapping triangle is divided into four equal parts to form several sub-triangles, and the number of recursive divisions is incremented by one; all mapping points in the current mapping point set are sorted, and the mapping triangle of each mapping point is re-determined based on the mapping point order and several sub-triangles of the mapping triangle, and each mapping point is stored in the mapping triangle as the current mapping point of the corresponding mapping triangle.

[0043] In this embodiment, when recursively dividing the mapping triangle based on the degree of mapping influence, if the degree of mapping influence does not exceed a preset first threshold, it indicates that the control result at the current interaction moment is consistent with the control result at the historical interaction moment in the overall direction, so the number of recursive divisions is set to zero. If it exceeds the threshold, the mapping triangle is divided into four equal parts and the number of recursive divisions is incremented once. Then, during the recursive division process, all mapping points in the current mapping point set are sorted and the mapping triangle to which each mapping point belongs is re-determined to obtain the updated number of recursive divisions. Finally, the historical distribution fluctuation is determined based on the updated number of recursive divisions. The reward value, expressed as a recursive partition of the mapping triangle and the calculation of the number of recursive partitions, can accurately reflect the degree of continuous interference between the control result at the current interaction moment and the control result at the historical interaction moment. Thus, the calculation of the number of recursive partitions can reflect the deviation between the control result at the current interaction moment and the historical interaction moment. This allows the subsequent reward value to guide the agent to impose penalties on actions that cause drastic fluctuations in the historical distribution, thereby effectively suppressing the problem of excessive adjustment amplitude and guiding the agent to make smooth and fine adjustments. In turn, it can achieve stable control of multiple actuators of the ship in complex sea conditions and improve the control stability during the ship's navigation.

[0044] In an optional embodiment, step S102 above categorizes the control results, uniquely mapping the control results at different interaction times to the control dominance mode triangle. The mapping triangle is a sub-triangle of the control dominance mode triangle; that is, the mapping points in the control dominance mode triangle share the same disadvantaged control objective. After mapping through numerous interaction times, although the mapping points in the control dominance mode triangle share the same disadvantaged control objective and the same control dominance mode, the control dominance mode is composed of three control objectives. The mapping points will be closer to the vertex corresponding to the control objective with the highest reward weight. Therefore, the distribution of these mapping points in the control dominance mode triangle is actually oriented towards the vertex corresponding to the control objective with the highest reward weight. Therefore, this embodiment proposes a recursive triangle partitioning method to further classify the mapping points in the control dominance mode triangle, ensuring that the mapping points within the mapping triangle have the most similar reward weights among the three control objectives.

[0045] Furthermore, to further characterize the impact of adding a mapping point to its mapping triangle on the distribution of historical mapping points at historical interaction moments within the mapping triangle, this embodiment uses relative entropy. First, the historical mapping point set of the mapping triangle is determined using the historical mapping points at historical interaction moments within the mapping triangle. Then, the current mapping point set of the mapping triangle is formed using the current mapping point and the historical mapping points. After that, the relative entropy between the two historical mapping point sets and the current mapping point set is calculated. A first threshold is set to 1. If the relative entropy is less than 1, it indicates that the addition of the current mapping point has a small impact on the distribution of historical mapping points within the mapping triangle, and also indicates that the current mapping point and the historical mapping point are relatively similar in various reward weights. This indicates that the actual control target's attention level is similarity, thus indicating that the action corresponding to the current mapping point has not deviated excessively from the historical action, so recursive partitioning is not required, and the number of recursive partitioning times is set to 0.

[0046] If the relative entropy is greater than 1, it indicates that the addition of the current mapping point has a significant impact on the distribution of historical mapping points within the mapping triangle, and the action corresponding to the current mapping point deviates significantly from the historical action. In this case, it is necessary to reclassify the mapping points within the mapping triangle. As previously defined, the mapping triangle is the minimum enclosing triangle of the mapping points. Therefore, to achieve reclassification, the mapping triangle needs to be divided into four equal parts to form new sub-triangles, and the number of recursive divisions should be incremented by one. Next, the mapping points are sorted in chronological order according to the interaction time, forming a mapping point sequence. Since the mapping point of the control result at each interaction time is fixed, there is no need to recalculate the mapping point. It is only necessary to redetermine the minimum enclosing triangle of the mapping point from the four sub-triangles of the mapping triangle. This allows for the re-determination of the mapping triangle for that mapping point, and then the mapping point is stored as the current mapping point of the corresponding mapping triangle (for distinction, this embodiment will refer to it as the first mapping triangle). Similarly, the recursive partitioning of the first mapping triangle may still be triggered again (determined by the relative entropy and the first threshold). If the recursive partitioning of the first mapping triangle is triggered again, all mapping points within the first mapping triangle are recursively partitioned according to the above description, and the recursive partitioning count is incremented by one each time the mapping triangle is divided into four equal parts. Finally, the number of recursive partitioning iterations is used as the historical distribution fluctuation reward value of the control result.

[0047] It should be noted that relative entropy, also known as Kullback-Leibler divergence or information divergence, is an asymmetric index in information theory used to measure the difference between two probability distributions.

[0048] Step S104: Determine the reward value at the current interaction moment based on the disadvantage control target reward value and the historical distribution fluctuation reward value; construct empirical data for the current interaction moment based on the current state, action, next state, and reward value; train the agent based on the empirical data to obtain the ship cooperative control agent.

[0049] In one optional embodiment, the reward value of the disadvantage control target and the reward value of the historical distribution fluctuation are multiplied to obtain the reward value at the current interaction moment; then, the experience data of the current interaction moment is constructed with the current state, action, next state and reward value and stored in the experience pool; then, the proximal policy optimization (PPO) algorithm is used to train the agent with the experience data in the experience pool to obtain the ship cooperative control agent.

[0050] Step S105: Obtain the real-time attitude data and real-time hydrodynamic data of the ship, input the real-time attitude data and real-time hydrodynamic data into the ship cooperative control agent to obtain the ship's cooperative control scheme, and perform cooperative control on the ship based on the cooperative control scheme.

[0051] In this embodiment, the steps of acquiring the ship's real-time attitude data and real-time hydrodynamic data, inputting the real-time attitude data and real-time hydrodynamic data into the ship's cooperative control agent to obtain the ship's cooperative control scheme, and performing cooperative control of the ship based on the cooperative control scheme include: The real-time attitude data and real-time hydrodynamic data of the ship are acquired, and the real-time attitude data and real-time hydrodynamic data are input into the ship cooperative control agent to obtain several candidate cooperative control actions. The candidate cooperative control actions are simulated in a preset ship navigation simulation environment to determine the control simulation results of each candidate cooperative control action. Each of the control simulation results is categorized into the control dominant mode triangle, and the simulation mapping triangle for each of the candidate cooperative control actions is determined. Quantum annealing simulation is performed on the candidate cooperative control actions based on the historical reference mapping points within the simulation mapping triangle of each candidate cooperative control action to determine the cooperative control scheme of the ship. The ship is controlled collaboratively based on the aforementioned collaborative control scheme.

[0052] This embodiment obtains real-time attitude and hydrodynamic data of the ship, inputs them into a trained ship cooperative control agent to obtain several candidate cooperative control actions, and simulates the candidate actions in a preset ship navigation simulation environment to obtain the control simulation results of each candidate action. Then, each control simulation result is classified into the control dominant mode triangle to determine the simulation mapping triangle, which can obtain the simulation mapping triangle corresponding to each candidate action and its historical reference mapping points. Then, quantum annealing simulation is performed on the candidate actions based on the historical reference mapping points, so that the cooperative control scheme does not deviate excessively from historical control experience, and can effectively eliminate dangerous actions that may cause the ship to perform multiple cooperative control instability. This avoids the ship cooperative control risks caused by agent errors or sudden environmental changes, and can achieve stable control of the ship's multiple actuators in complex sea conditions, improving the control stability during ship navigation.

[0053] In an optional embodiment, real-time attitude data and real-time hydrodynamic data of the ship are acquired, and the real-time attitude data and real-time hydrodynamic data are input into the ship cooperative control agent to obtain several candidate cooperative control actions output by the ship cooperative control agent; then, the candidate cooperative control actions are simulated in a ship navigation simulation environment (e.g., finite element simulation) to determine the control simulation result of each candidate cooperative control action. The control simulation result refers to the simulation component of the control result of the ship under each control objective. Similar to step S102 above, each control simulation result is classified into the control dominant mode triangle to determine the simulation mapping triangle of each candidate cooperative control action; however, it is not necessary to put the control simulation result into the control dominant mode triangle, it is only necessary to confirm the mapping triangle as the simulation mapping triangle.

[0054] In this embodiment, the step of performing quantum annealing simulation on the candidate cooperative control actions based on the historical reference mapping points within the simulation mapping triangle of each candidate cooperative control action to determine the cooperative control scheme of the ship includes: Based on the historical comparison mapping points within the simulation mapping triangle of each candidate cooperative control action, determine the number of mapping points, average reward value, and centroid of the mapping points for each candidate cooperative control action. The centroid of the mapping point is inversely mapped to the surface of the unit sphere to determine the spherical mapping point for each candidate cooperative control action; The local action capability value is determined by the number of mapping points and the average reward value, and the global action capability value is determined by the spherical angle between every two spherical mapping points. An action energy function is constructed based on the local action capability value and the global action capability value, and the action energy function is evolved based on a preset quantum annealing simulation algorithm to determine the energy value of each candidate cooperative control action; Based on the energy value, the candidate cooperative control actions are screened to obtain the cooperative control scheme for the ship.

[0055] In this embodiment, when performing quantum annealing simulation on candidate actions based on historical reference mapping points within the simulation mapping triangle, the number of mapping points, average reward value, and centroid of the mapping points are determined according to the historical reference mapping points within the simulation mapping triangle for each candidate cooperative control action. The centroid of the mapping points is then inversely mapped to the surface of a unit sphere to determine the spherical mapping points. Subsequently, the local action capability value is determined by the number of mapping points and the average reward value, and the global action capability value is determined by the spherical angle between every two spherical mapping points. This yields an action energy function that integrates local development and global exploration, thereby modeling candidate cooperative control actions as an energy minimization problem in geometric space. The local action capability value reflects the experience richness of candidate cooperative control actions within a specific region, while the global action capability value reflects the differences between candidate cooperative control actions. Furthermore, the quantum annealing simulation algorithm efficiently searches for the global optimum of energy, thus avoiding the one-sidedness of the cooperative control scheme. This enables the smooth control of multiple actuators of a ship in complex sea conditions and improves the control stability during ship navigation.

[0056] In one optional embodiment, after determining the simulation mapping triangle of the candidate cooperative control action, the total number of historical reference mapping points within the simulation mapping triangle of the candidate cooperative control action is counted as the number of mapping points; simultaneously, the reward values ​​corresponding to these historical reference mapping points are summed and averaged to obtain the average reward value; and the centroid of the mapping point is obtained using the coordinates of these historical reference mapping points. Then, the centroid of the mapping point is inversely mapped to the surface of the unit sphere to determine the spherical mapping point of each candidate cooperative control action; since this inverse mapping process is a conventional geometric solution process, it will not be described in detail again in this embodiment.

[0057] The number of points mapped to the sphere is then defined as follows: , No. The coordinates of the spherical mapping points are: The number of corresponding mapping points is The average reward value is Therefore, the value of local motion capability can be expressed as The local action capability value is constructed using the reciprocal of the average reward value and the number of mapping points. A higher local action capability value indicates a lower average reward value and a larger number of mapping points. This suggests that the corresponding candidate cooperative control action could not achieve good cooperative control results during historical training, and the likelihood of selection by the quantum annealing simulation algorithm should be reduced. The spherical angle between every two spherical mapping points is defined as... Therefore, the global action capability value can be expressed as: ; This refers to a tiny value used to avoid a denominator of 0; this setting of the global action capability value, based on the angle of the spherical interface, prevents two closely spaced spherical mapping points from being simultaneously selected by the quantum annealing simulation algorithm; the constructed action energy function is expressed as... Then, the quantum annealing simulation algorithm is used to evolve the action energy function to determine the energy value of each candidate cooperative control action; then, the candidate cooperative control action with the largest energy value is selected to form the ship's cooperative control scheme.

[0058] This embodiment constructs an agent based on pre-acquired ship attitude data, hydrodynamic data, and adjustment command data. Several control dominance mode triangles are initialized according to multiple preset control objectives of the ship, resulting in an agent to be trained. Then, at each interaction moment during the agent's training process, the control result is categorized into the control dominance mode triangle, determining the inferior control objective reward value, the mapping triangle, and the current mapping point formed within the mapping triangle. The inferior control objective reward value characterizes the change in the control result under multiple control objectives. Next, the relationship between the control result at the current interaction moment and the historical mapping points at historical interaction moments is measured by the degree of mapping influence between the current mapping point and the historical mapping points. This allows for recursive division of the mapping triangle based on the degree of mapping influence, yielding a historical distribution fluctuation reward value that characterizes the ship control relationship between the current and historical interaction moments. The reward value for the current interaction moment is then formed based on the historical distribution fluctuation reward value and the inferior control objective reward value. Finally, empirical data for the current interaction moment is constructed by combining the current state, action, next state, and reward value. This empirical data is used to train the agent, resulting in a ship cooperative control agent. Ultimately, the ship's collaborative control scheme is output by the ship collaborative control agent, which enables the smooth control of the ship's multiple actuators in complex sea conditions and improves the control stability during the ship's navigation.

[0059] Example 2 Please refer to Figure 2 , Figure 2 A schematic diagram of a ship attitude and hydrodynamic multi-actuator cooperative control system provided in an embodiment of the present invention includes: an initialization module 201, a control result classification module 202, a mapping influence module 203, an agent training module 204, and a ship cooperative control module 205; The initialization module 201 is used to construct the ship's intelligent agent based on the pre-acquired ship attitude data, hydrodynamic data and adjustment command data, and to initialize several control dominant mode triangles of the ship based on the pre-acquired control objectives of the ship. The control result classification module 202 is used to determine the next state and control result based on the current state and actions of the ship at each interaction moment; classify the control result into the control dominance mode triangle, and determine the disadvantage control target reward value, mapping triangle and the current mapping point formed in the mapping triangle of the control result; The mapping influence module 203 is used to obtain the degree of mapping influence of the current mapping point on the historical mapping points at historical interaction moments within the mapping triangle; and to recursively divide the mapping triangle based on the degree of mapping influence to obtain the historical distribution fluctuation reward value of the control result. The agent training module 204 is used to determine the reward value at the current interaction moment based on the disadvantage control target reward value and the historical distribution fluctuation reward value; to construct empirical data for the current interaction moment based on the current state, action, next state, and reward value; and to train the agent based on the empirical data to obtain a ship cooperative control agent. The ship cooperative control module 205 is used to acquire the ship's real-time attitude data and real-time hydrodynamic data, input the real-time attitude data and real-time hydrodynamic data into the ship cooperative control agent to obtain the ship's cooperative control scheme, and perform cooperative control on the ship based on the cooperative control scheme.

[0060] In this embodiment, the initialization module 201 includes: an agent initialization unit; The initialization unit is used to construct a state space based on the pre-acquired ship attitude data and hydrodynamic data; Construct the action space based on the pre-acquired adjustment instruction data; The ship's intelligent agent is constructed based on the action space and the state space.

[0061] In this embodiment, the initialization module 201 further includes: a control dominant mode triangle initialization unit; In the control-dominant mode triangle initialization unit, the control objectives pre-acquired by the ship include: roll reduction effect, ship energy consumption, sailing resistance and comfort; The control-dominant mode triangle initialization unit is used to construct a regular tetrahedron within a preset unit sphere, and to form several initial triangles based on the regular tetrahedron. Based on the roll reduction effect, ship energy consumption, sailing resistance and comfort, the extreme value components of the control target of the intelligent agent under each of the control objectives are generated; Each of the control target extreme value components is mapped to each vertex of the regular tetrahedron to initialize the vertex control target of each of the initial triangles; Based on the control target of each vertex of the initial triangle, the control dominance mode of each initial triangle is determined, resulting in several control dominance mode triangles for the ship.

[0062] In this embodiment, the control result classification module 202 includes: a control result classification unit; The control result classification unit is used to determine the control result component of the ship's next state under each control objective based on the ship's current state and actions at each interaction moment, and to form the ship's control result based on the control result component. Each component of the control result is normalized to obtain several normalized control result components; The normalized control result components are classified into the control dominant mode triangle to obtain the control dominant mode, mapping triangle and the current mapping point formed in the mapping triangle of the control result; Based on the control dominance mode of the control results, the control targets are screened to determine the inferior control targets of the control results; The disadvantage control objective reward value of the control result is determined based on the disadvantage control objective and the normalized control result components.

[0063] In this embodiment, the control result classification unit includes: a control result classification subunit; The control result classification subunit is used to transform the normalized control result components based on a preset exponential function to form a reward weight for each normalized control result component. The vertex coordinates of the regular tetrahedron are weighted and summed based on the reward weights to determine the control target direction vector of the control result within the regular tetrahedron. With the center of the regular tetrahedron as the origin, a control target guide ray is formed within the regular tetrahedron along the control target direction vector based on the origin; The intersection points of the control target guide ray and each of the control dominant mode triangles are calculated based on a preset ray triangle intersection algorithm; Based on the intersection point, the control dominance mode triangle is filtered to determine the first intersection point and the first control dominance mode triangle corresponding to the control result; and the control dominance mode of the control result is determined based on the first control dominance mode triangle. The first control dominant mode triangle is filtered according to a preset triangle walking algorithm to determine the mapping triangle of the control result; and the current mapping point is formed in the mapping triangle based on the first intersection point.

[0064] In this embodiment, the mapping influence module 203 includes: a mapping influence unit; The mapping influence unit is used to determine the set of historical mapping points of the mapping triangle based on the historical mapping points of historical interaction moments within the mapping triangle; The current mapping point set of the mapping triangle is determined based on the current mapping point and the historical mapping points; Calculate the relative entropy between the historical mapping point set and the current mapping point set, and determine the degree of influence of the current mapping point on the historical mapping points at historical interaction moments within the mapping triangle based on the relative entropy; Initialize the number of recursive partitions of the mapped triangle; The mapping triangle is recursively divided based on the degree of influence of the mapping, so as to update the number of recursive divisions; The historical distribution fluctuation reward value of the control result is obtained based on the number of recursive partitions completed for updating.

[0065] In this embodiment, the mapping influence unit includes: a recursive partitioning subunit; The recursive partitioning subunit is used to set the number of recursive partitions to a preset zero value if the degree of influence of the mapping does not exceed a preset first threshold. If the influence of the mapping exceeds the first threshold, the mapping triangle is divided into four equal parts to form several sub-triangles, and the number of recursive divisions is incremented by one; all mapping points in the current mapping point set are sorted, and the mapping triangle of each mapping point is re-determined based on the mapping point order and several sub-triangles of the mapping triangle, and each mapping point is stored in the mapping triangle as the current mapping point of the corresponding mapping triangle.

[0066] In this embodiment, the ship cooperative control module 205 includes: a ship cooperative control unit; The ship cooperative control unit is used to acquire the ship's real-time attitude data and real-time hydrodynamic data, and input the real-time attitude data and real-time hydrodynamic data into the ship cooperative control agent to obtain several candidate cooperative control actions. The candidate cooperative control actions are simulated in a preset ship navigation simulation environment to determine the control simulation results of each candidate cooperative control action. Each of the control simulation results is categorized into the control dominant mode triangle, and the simulation mapping triangle for each of the candidate cooperative control actions is determined. Quantum annealing simulation is performed on the candidate cooperative control actions based on the historical reference mapping points within the simulation mapping triangle of each candidate cooperative control action to determine the cooperative control scheme of the ship. The ship is controlled collaboratively based on the aforementioned collaborative control scheme.

[0067] In this embodiment, the ship collaborative control unit includes: a quantum annealing simulation subunit; The quantum annealing simulation subunit is used to determine the number of mapping points, average reward value, and centroid of mapping points for each candidate cooperative control action based on the historical comparison mapping points within the simulation mapping triangle of each candidate cooperative control action. The centroid of the mapping point is inversely mapped to the surface of the unit sphere to determine the spherical mapping point for each candidate cooperative control action; The local action capability value is determined by the number of mapping points and the average reward value, and the global action capability value is determined by the spherical angle between every two spherical mapping points. An action energy function is constructed based on the local action capability value and the global action capability value, and the action energy function is evolved based on a preset quantum annealing simulation algorithm to determine the energy value of each candidate cooperative control action; Based on the energy value, the candidate cooperative control actions are screened to obtain the cooperative control scheme for the ship.

[0068] This embodiment constructs an agent based on pre-acquired ship attitude data, hydrodynamic data, and adjustment command data. Several control dominance mode triangles are initialized according to multiple preset control objectives of the ship, resulting in an agent to be trained. Then, at each interaction moment during the agent's training process, the control result is categorized into the control dominance mode triangle, determining the inferior control objective reward value, the mapping triangle, and the current mapping point formed within the mapping triangle. The inferior control objective reward value characterizes the change in the control result under multiple control objectives. Next, the relationship between the control result at the current interaction moment and the historical mapping points at historical interaction moments is measured by the degree of mapping influence between the current mapping point and the historical mapping points. This allows for recursive division of the mapping triangle based on the degree of mapping influence, yielding a historical distribution fluctuation reward value that characterizes the ship control relationship between the current and historical interaction moments. The reward value for the current interaction moment is then formed based on the historical distribution fluctuation reward value and the inferior control objective reward value. Finally, empirical data for the current interaction moment is constructed by combining the current state, action, next state, and reward value. This empirical data is used to train the agent, resulting in a ship cooperative control agent. Ultimately, the ship's collaborative control scheme is output by the ship collaborative control agent, which enables the smooth control of the ship's multiple actuators in complex sea conditions and improves the control stability during the ship's navigation.

[0069] In summary, this embodiment of the invention constructs an agent based on pre-acquired ship attitude data, hydrodynamic data, and adjustment command data, and initializes several control dominance mode triangles according to multiple preset control objectives of the ship to obtain the agent to be trained. Then, at each interaction moment in the agent training process, the control result is categorized into the control dominance mode triangle, and the inferior control objective reward value, mapping triangle, and current mapping point formed in the mapping triangle are determined. The inferior control objective reward value represents the change of the control result under multiple control objectives. Next, the relationship between the control result at the current interaction moment and the historical interaction moment data is measured by the degree of mapping influence between the current mapping point in the mapping triangle and the historical mapping points at historical interaction moments. Thus, the mapping triangle is recursively divided according to the degree of mapping influence to obtain a historical distribution fluctuation reward value that can represent the ship control relationship between the current interaction moment and historical interaction moments. Then, the reward value for the current interaction moment is formed based on the historical distribution fluctuation reward value and the inferior control objective reward value. Furthermore, empirical data for the current interaction moment is constructed by combining the current state, action, next state, and reward value. The agent is trained using this empirical data to obtain a ship cooperative control agent. Ultimately, the ship's collaborative control scheme is output by the ship collaborative control agent, which enables the smooth control of the ship's multiple actuators in complex sea conditions and improves the control stability during the ship's navigation.

[0070] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A method for coordinated control of ship attitude and hydrodynamic multi-actuator, characterized in that, include: The ship's intelligent agent is constructed based on the pre-acquired ship attitude data, hydrodynamic data, and adjustment command data, and several control dominant mode triangles of the ship are initialized based on the pre-acquired control objectives of the ship. At each interaction moment, the next state and control result are determined based on the ship's current state and actions; the control result is categorized into the control dominance mode triangle, and the inferior control target reward value, mapping triangle, and current mapping point formed in the mapping triangle are determined; Obtain the degree of influence of the current mapping point on the historical mapping points at historical interaction moments within the mapping triangle; Based on the degree of influence of the mapping, the mapping triangle is recursively divided to obtain the historical distribution fluctuation reward value of the control result; The reward value at the current interaction moment is determined based on the disadvantage control target reward value and the historical distribution fluctuation reward value; Based on the current state, action, next state, and reward value, construct empirical data for the current interaction moment; The agent is trained based on the empirical data to obtain a ship cooperative control agent; The system acquires the ship's real-time attitude data and real-time hydrodynamic data, inputs the real-time attitude data and real-time hydrodynamic data into the ship's collaborative control agent, obtains the ship's collaborative control scheme, and performs collaborative control on the ship based on the collaborative control scheme.

2. The method for coordinated control of ship attitude and hydrodynamic multi-actuator as described in claim 1, characterized in that, The construction of the ship's intelligent agent based on pre-acquired ship attitude data, hydrodynamic data, and adjustment command data includes: The state space is constructed based on the pre-acquired ship attitude and hydrodynamic data; Construct the action space based on the pre-acquired adjustment instruction data; The ship's intelligent agent is constructed based on the action space and the state space.

3. A method for coordinated control of ship attitude and hydrodynamic multi-actuator as described in claim 1 or 2, characterized in that, The pre-acquired control objectives for the ship include: roll reduction effect, ship energy consumption, sailing resistance, and comfort; the initialization of several control dominant mode triangles for the ship based on the pre-acquired control objectives includes: A regular tetrahedron is constructed within a pre-defined unit sphere, and several initial triangles are formed based on the regular tetrahedron. Based on the roll reduction effect, ship energy consumption, sailing resistance and comfort, the extreme value components of the control target of the intelligent agent under each of the control objectives are generated; Each of the control target extreme value components is mapped to each vertex of the regular tetrahedron to initialize the vertex control target of each of the initial triangles; Based on the control target of each vertex of the initial triangle, the control dominance mode of each initial triangle is determined, resulting in several control dominance mode triangles for the ship.

4. The method for coordinated control of ship attitude and hydrodynamic multi-actuator as described in claim 3, characterized in that, At each interaction moment, the next state and control result are determined based on the ship's current state and actions; the control result is categorized into the control dominance mode triangle, and the disadvantageous control target reward value, mapping triangle, and current mapping point formed in the mapping triangle are determined, including: At each interaction moment, the control result component of the ship's next state under each control objective is determined based on the ship's current state and actions, and the control result of the ship is formed based on the control result component; Each component of the control result is normalized to obtain several normalized control result components; The normalized control result components are classified into the control dominant mode triangle to obtain the control dominant mode, mapping triangle and the current mapping point formed in the mapping triangle of the control result; Based on the control dominance mode of the control results, the control targets are screened to determine the inferior control targets of the control results; The disadvantage control objective reward value of the control result is determined based on the disadvantage control objective and the normalized control result components.

5. The method for coordinated control of ship attitude and hydrodynamic multi-actuator as described in claim 4, characterized in that, The step of classifying the normalized control result components into the control dominance mode triangle to obtain the control dominance mode, mapping triangle, and current mapping point formed in the mapping triangle of the control result includes: The normalized control result components are transformed based on a preset exponential function to form a reward weight for each normalized control result component. The vertex coordinates of the regular tetrahedron are weighted and summed based on the reward weights to determine the control target direction vector of the control result within the regular tetrahedron. With the center of the regular tetrahedron as the origin, a control target guide ray is formed within the regular tetrahedron along the control target direction vector based on the origin; The intersection points of the control target guide ray and each of the control dominant mode triangles are calculated based on a preset ray triangle intersection algorithm; Based on the intersection point, the control dominance mode triangle is filtered to determine the first intersection point and the first control dominance mode triangle corresponding to the control result; and the control dominance mode of the control result is determined based on the first control dominance mode triangle. The first control dominant mode triangle is filtered according to a preset triangle walking algorithm to determine the mapping triangle of the control result; and the current mapping point is formed in the mapping triangle based on the first intersection point.

6. The method for coordinated control of ship attitude and hydrodynamic multi-actuator as described in claim 5, characterized in that, The degree of influence of the current mapping point on the historical mapping points at historical interaction moments within the mapping triangle is obtained. Based on the degree of influence of the mapping, the mapping triangle is recursively divided to obtain the historical distribution fluctuation reward value of the control result, including: Based on the historical mapping points of the historical interaction moments within the mapping triangle, determine the set of historical mapping points of the mapping triangle; The current mapping point set of the mapping triangle is determined based on the current mapping point and the historical mapping points; Calculate the relative entropy between the historical mapping point set and the current mapping point set, and determine the degree of influence of the current mapping point on the historical mapping points at historical interaction moments within the mapping triangle based on the relative entropy; Initialize the number of recursive partitions of the mapped triangle; The mapping triangle is recursively divided based on the degree of influence of the mapping, so as to update the number of recursive divisions; The historical distribution fluctuation reward value of the control result is obtained based on the number of recursive partitions completed for updating.

7. The method for coordinated control of ship attitude and hydrodynamic multi-actuator as described in claim 6, characterized in that, The recursive partitioning of the mapping triangle based on the degree of mapping influence, to update the number of recursive partitions, includes: If the degree of influence of the mapping does not exceed a preset first threshold, then the number of recursive partitions is set to a preset zero value; If the influence of the mapping exceeds the first threshold, the mapping triangle is divided into four equal parts to form several sub-triangles, and the number of recursive divisions is incremented by one; all mapping points in the current mapping point set are sorted, and the mapping triangle of each mapping point is re-determined based on the mapping point order and several sub-triangles of the mapping triangle, and each mapping point is stored in the mapping triangle as the current mapping point of the corresponding mapping triangle.

8. The method for coordinated control of ship attitude and hydrodynamic multi-actuator as described in claim 3, characterized in that, The process of acquiring real-time attitude data and real-time hydrodynamic data of the vessel, inputting the real-time attitude data and real-time hydrodynamic data into the vessel's collaborative control agent to obtain a collaborative control scheme for the vessel, and performing collaborative control of the vessel based on the collaborative control scheme includes: The real-time attitude data and real-time hydrodynamic data of the ship are acquired, and the real-time attitude data and real-time hydrodynamic data are input into the ship cooperative control agent to obtain several candidate cooperative control actions. The candidate cooperative control actions are simulated in a preset ship navigation simulation environment to determine the control simulation results of each candidate cooperative control action. Each of the control simulation results is categorized into the control dominant mode triangle, and the simulation mapping triangle for each of the candidate cooperative control actions is determined. Quantum annealing simulation is performed on the candidate cooperative control actions based on the historical reference mapping points within the simulation mapping triangle of each candidate cooperative control action to determine the cooperative control scheme of the ship. The ship is controlled collaboratively based on the aforementioned collaborative control scheme.

9. The method for coordinated control of ship attitude and hydrodynamic multi-actuator as described in claim 8, characterized in that, The process of performing quantum annealing simulations on the candidate cooperative control actions based on historical reference mapping points within the simulation mapping triangle of each candidate cooperative control action to determine the cooperative control scheme of the ship includes: Based on the historical comparison mapping points within the simulation mapping triangle of each candidate cooperative control action, determine the number of mapping points, average reward value, and centroid of the mapping points for each candidate cooperative control action. The centroid of the mapping point is inversely mapped to the surface of the unit sphere to determine the spherical mapping point for each candidate cooperative control action; The local action capability value is determined by the number of mapping points and the average reward value, and the global action capability value is determined by the spherical angle between every two spherical mapping points. An action energy function is constructed based on the local action capability value and the global action capability value, and the action energy function is evolved based on a preset quantum annealing simulation algorithm to determine the energy value of each candidate cooperative control action; Based on the energy value, the candidate cooperative control actions are screened to obtain the cooperative control scheme for the ship.

10. A collaborative control system for ship attitude and hydrodynamic multi-actuator operation, characterized in that, include: The system includes an initialization module, a control result classification module, a mapping effect module, an agent training module, and a ship cooperative control module. The initialization module is used to construct the ship's intelligent agent based on the pre-acquired ship attitude data, hydrodynamic data, and adjustment command data, and to initialize several control dominant mode triangles of the ship based on the pre-acquired control objectives of the ship. The control result classification module is used to determine the next state and control result based on the current state and actions of the ship at each interaction moment; classify the control result into the control dominance mode triangle, and determine the inferior control target reward value, mapping triangle, and the current mapping point formed in the mapping triangle of the control result; The mapping influence module is used to obtain the degree of mapping influence of the current mapping point on the historical mapping points at historical interaction moments within the mapping triangle; and to recursively divide the mapping triangle based on the degree of mapping influence to obtain the historical distribution fluctuation reward value of the control result. The agent training module is used to determine the reward value at the current interaction moment based on the disadvantage control target reward value and the historical distribution fluctuation reward value; Based on the current state, action, next state, and reward value, construct empirical data for the current interaction moment; The agent is trained based on the empirical data to obtain a ship cooperative control agent; The ship collaborative control module is used to acquire the ship's real-time attitude data and real-time hydrodynamic data, input the real-time attitude data and real-time hydrodynamic data into the ship collaborative control agent to obtain the ship's collaborative control scheme, and perform collaborative control on the ship based on the collaborative control scheme.