Underwater robot cooperative operation task allocation method
By constructing a motion allocation knowledge base and game model for a dual-robot-manipulator system, and combining it with an improved differential evolution algorithm, the problems of redundant degrees of freedom and limited energy for autonomous underwater robots to work collaboratively in complex marine environments were solved, and an efficient and robust task allocation strategy was achieved.
Patent Information
- Application Number
- CN202511049110.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-29
AI Technical Summary
In complex marine environments, multiple autonomous underwater robots equipped with robotic arms face issues of redundant degrees of freedom and limited energy when working together, resulting in low efficiency in motion allocation, poor dynamic adaptability and robustness, making it difficult to complete tasks efficiently.
A knowledge base for motion allocation in a dual-robot-manipulator system is constructed. Through a multi-level knowledge representation model and a game theory model, a payoff function is designed, and an improved differential evolution algorithm is used to solve the Nash equilibrium to realize the motion allocation strategy.
Efficient and robust task allocation for multi-robot collaborative operations was achieved in complex marine environments, improving task efficiency and the accuracy of action allocation while satisfying the Nash equilibrium condition.
Smart Images

Figure CN120848563A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of marine robot technology, and is a method for allocating tasks in collaborative underwater robot operations. Background Technology
[0002] The development of marine resources places higher demands on the collaborative operation capabilities of autonomous underwater vehicles (AUVs). Therefore, multiple AUVs equipped with robotic arms have gradually become a research hotspot, with their advantages in completing actions with high efficiency and precision becoming increasingly prominent. While robotic arms endow AUVs with strong operational capabilities, they also bring redundant degrees of freedom and complex collaborative mechanisms. During operations, each robot must consider not only the collaboration between its own robotic arm and the hull, but also the collaboration with other robots and their robotic arms. Furthermore, due to the limited energy of AUVs, especially when carrying high-power equipment such as robotic arms, the rational use of limited energy for collaborative operations is a key factor in improving mission efficiency.
[0003] For the motion allocation of a dual-robot-dual-manipulator system in collaborative operations, the robots and manipulators possess redundant degrees of freedom in an aquatic environment, making their coordination crucial during task execution. Depending on the operational environment, the robot's posture, the manipulator's movements, and angles need to be rationally allocated to achieve the desired motion allocation during the operation. Previous research on manipulator motion allocation has largely focused on scenarios with fewer influencing factors. However, this invention addresses the task of multiple underwater robots equipped with dual manipulators handling a target object in a marine environment. This means not only allocating the specific movements of the manipulators in a highly turbulent aquatic environment but also fully considering the interaction between the manipulators and the hull. Therefore, it necessitates further solutions to the motion allocation problem under complex multi-degree-of-freedom conditions. Summary of the Invention
[0004] This invention discloses a method for allocating tasks in collaborative underwater robot operations to solve the problem of motion allocation in a dual-robot-dual-manipulator system.
[0005] This invention provides the following technical solutions: A method for assigning tasks in cooperative underwater robot operations, the method comprising the following steps: Step 1: Construct a knowledge base for motion allocation in a dual-robot-manipulator system; Step 2: Construct a dual-matrix game model based on the knowledge base, design payoff functions for different strategies of both sides of the game, and form a payoff matrix; Step 3: Use the multi-payer matrix weighting method to transform the game problem into an optimization problem for solution, and verify the correctness of the game model; Step 4: Solve the Nash equilibrium of the two-matrix game model by applying the differential evolution algorithm, and improve the basic differential evolution algorithm by applying the velocity control factor, the optimal point set theory and the step-type inertia factor.
[0006] Preferably, step 1 specifically comprises: The multi-level knowledge representation model-based knowledge base for task action allocation of autonomous underwater robots also includes the following six levels: scene layer; object layer; agent layer; task layer; skill layer; and action layer. Based on this, a detailed design was made for the action allocation knowledge base, which was constructed in three categories: scenario, agent, and concept.
[0007] Preferably, step 2 specifically comprises: The collaborative operation of the robot-manipulator system is decomposed into a three-layer game of robot posture, manipulator action, and manipulator angle. A corresponding payoff function is designed for each layer of the game, and a payoff matrix is constructed.
[0008] Preferably, for the robot's posture game, based on different seabed environments and ocean current information, the posture is divided into two different postures: sitting on the seabed and hovering, respectively: (1) in, The speed of ocean currents that underwater robots can withstand. The speed of the maximum ocean current that an underwater robot can withstand. To determine the different robots' preferences for sitting postures, This represents the degree of preference of both sides for the suspension posture.
[0009] Preferably, for the game of robotic arm actions, both sides choose actions based on different task types and the physical characteristics of the robotic arm. The payoff functions for the game strategies of cutting and grasping are as follows: (2) in: Let be the preference coefficients of both sides in the game regarding the shearing action. Let be the preference coefficient between the two players for the grabbing action. This refers to the capability coefficients of both sides in the game regarding different cable materials.
[0010] Preferably, regarding the game of the robotic arm angles, considering the robotic arm's range of motion and task requirements, we will continue with the previous example. The angle payoff function for the robotic arm's grasping and shearing is: (3) in, The initial angle of the cable relative to the base coordinate system when making decisions for the robot. Preferably, step 3 specifically comprises: A weighted evaluation model based on the analytic hierarchy process (AHP) is constructed. Through systematic steps such as establishing a judgment matrix, calculating weight vectors, and performing consistency checks, the payoff relationships between actions at each level are effectively quantified, thus representing the payoffs of all single-objective players in a multi-objective model. and Use the following formula for weighting: (4) In the formula: , By using weighted summation, the multi-level payoff matrix is transformed into a comprehensive payoff matrix for both sides of the game.
[0011] An underwater robot collaborative operation task allocation system, the system comprising: The knowledge base module constructs a knowledge base for motion allocation of the dual-robot-manipulator system. The payoff matrix module constructs a dual-matrix game model based on a knowledge base, designs payoff functions for different strategies of the two players, and forms a payoff matrix. The verification module uses the multi-payer matrix weighting method to transform the game problem into an optimization problem for solution and verifies the correctness of the game model. The solution module applies the differential evolution algorithm to solve the Nash equilibrium of the two-matrix game model. It improves the basic differential evolution algorithm by applying the velocity control factor, the optimal point set theory, and the step-type inertia factor.
[0012] A computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement a method for assigning tasks in a cooperative operation of an underwater robot.
[0013] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement a method for allocating tasks for cooperative operation of underwater robots.
[0014] The present invention has the following beneficial effects: This invention proposes a method for assigning actions in cooperative underwater robot operations. The invention studies the problem of action assignment in cooperative operations of a dual-robot / dual-manipulator system in an underwater environment. Addressing the shortcomings of existing technologies, it implements a task assignment strategy under complex conditions by constructing a dual-matrix game model and combining it with an improved differential evolution algorithm.
[0015] The optimization algorithm of this invention was applied to realize the collaborative action allocation of a dual-robot-dual-manipulator system in an underwater simulation environment. The Nash equilibrium of the action allocation problem of the dual matrix game was obtained, and the optimization results were verified to satisfy the necessary and sufficient conditions of Nash equilibrium. Attached Figure Description
[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 A simplified diagram of the knowledge base for assigning actions; Figure 2 Assign a scene-based knowledge base to the action; Figure 3 Assign an agent-class knowledge base to the action; Figure 4 It is a knowledge base for action allocation concepts; Figure 5 Assign a knowledge base to the movements of the dual-robot-manipulator arm; Figure 6 This is a diagram of a three-layer game structure; Figure 7 To improve the differential evolution algorithm for solving the Nash equilibrium graph; Figure 8 A comparison chart showing the optimization results of the differential evolution algorithm. Detailed Implementation
[0018] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] The present invention will be described in detail below with reference to specific embodiments. Specific Implementation Example 1: according to Figures 1 to 8 As shown, the specific optimized technical solution adopted by the present invention to solve the above-mentioned technical problems is: The present invention relates to a method for allocating tasks for collaborative operation of underwater robots.
[0021] This invention provides a method for assigning tasks in cooperative underwater robot operations, the method comprising the following steps: Step 1: Construct a knowledge base for motion allocation in a dual-robot-manipulator system; Step 2: Construct a dual-matrix game model based on the knowledge base, design payoff functions for different strategies of both sides of the game, and form a payoff matrix; Step 3: Use the multi-payer matrix weighting method to transform the game problem into an optimization problem for solution, and verify the correctness of the game model; Step 4: Solve the Nash equilibrium of the two-matrix game model by applying the differential evolution algorithm, and improve the basic differential evolution algorithm by applying the velocity control factor, the optimal point set theory and the step-type inertia factor.
[0022] This invention is primarily applied to the field of multi-robot collaborative operations in complex marine environments. It aims to address the shortcomings of existing technologies, such as inefficiency, weak dynamic adaptability, and poor robustness. It particularly emphasizes solving the task allocation problem through a game theory model, thereby improving the efficiency and robustness of multiple autonomous underwater vehicles (AUVs) collaboratively completing tasks in complex marine environments. The invention specifically includes the following steps: Step 1: Constructing a knowledge base for motion allocation in a dual-robot-manipulator system; Step 2: Constructing a dual-matrix game model based on the knowledge base, designing payoff functions for different strategies of both sides, and forming a payoff matrix; Step 3: Transforming the game problem into an optimization problem using a multi-payoff matrix weighting method, and verifying the correctness of the game model; Step 4: Applying an improved differential evolution algorithm to solve the problem, improving the basic differential evolution algorithm by applying velocity control factors, optimal point set theory, and step-type inertia factors to solve for the Nash equilibrium of the dual-matrix game model. This invention can be applied to the optimization of tasks performed by multiple AUVs. Specific Implementation Example 2: The only difference between Embodiment 2 and Embodiment 1 of this application is that: Step 1 specifically involves: The multi-level knowledge representation model-based knowledge base for task action allocation of autonomous underwater robots also includes the following six levels: scene layer; object layer; agent layer; task layer; skill layer; and action layer. Based on this, a detailed design was made for the action allocation knowledge base, which was constructed in three categories: scenario, agent, and concept. Specific Implementation Example 3: The only difference between Embodiment 3 and Embodiment 2 of this application is that: Step 2 specifically involves: The collaborative operation of the robot-manipulator system is decomposed into a three-layer game of robot posture, manipulator action, and manipulator angle. A corresponding payoff function is designed for each layer of the game, and a payoff matrix is constructed. Specific Implementation Example 4: The only difference between Embodiment 4 and Embodiment 3 of this application is that: Regarding the robot's posture dynamics, based on different seabed environments and ocean current information, two different postures are chosen: sitting on the seabed and hovering. (1) in, The speed of ocean currents that underwater robots can withstand. The speed of the maximum ocean current that an underwater robot can withstand. To determine the different robots' preferences for sitting postures, This represents the degree of preference of both sides for the suspension posture. Specific Implementation Example 5: The difference between Embodiment 5 and Embodiment 4 of the present invention lies only in: For the game theory of robotic arm actions, considering different task types and the physical characteristics of the robotic arm, the payoff functions for the shearing and grasping strategies are as follows: (2) in: Let be the preference coefficients of both sides in the game regarding the shearing action. Let be the preference coefficient between the two players for the grabbing action. This refers to the capability coefficients of both sides in the game regarding different cable materials. Specific Implementation Example Six: The difference between Embodiment Six and Embodiment Five of the present invention lies only in: Regarding the game of angles for the robotic arm, considering the range of motion of the robotic arm and the task requirements, we will continue with the previous example. The angle payoff function for the robotic arm's grasping and shearing is as follows: (3) in, The initial angle of the cable relative to the base coordinate system when making decisions for the robot. Specific Implementation Example 7: The difference between Embodiment Seven and Embodiment Six of the present invention lies only in: Step 3 specifically involves: A weighted evaluation model based on the analytic hierarchy process (AHP) is constructed. Through systematic steps such as establishing a judgment matrix, calculating weight vectors, and performing consistency checks, the payoff relationships between actions at each level are effectively quantified, thus representing the payoffs of all single-objective players in a multi-objective model. and Use the following formula for weighting: (4) In the formula: , By using weighted summation, the multi-level payoff matrix is transformed into a comprehensive payoff matrix for both sides of the game. Specific Implementation Example 8: The difference between Embodiment 8 and Embodiment 7 of the present invention lies only in: This invention provides an underwater robot collaborative operation task allocation system, the system comprising: The knowledge base module constructs a knowledge base for motion allocation of the dual-robot-manipulator system. The payoff matrix module constructs a dual-matrix game model based on a knowledge base, designs payoff functions for different strategies of the two players, and forms a payoff matrix. The verification module uses the multi-payer matrix weighting method to transform the game problem into an optimization problem for solution and verifies the correctness of the game model. The solution module applies the differential evolution algorithm to solve the Nash equilibrium of the two-matrix game model. It improves the basic differential evolution algorithm by applying the velocity control factor, the optimal point set theory, and the step-type inertia factor. Specific Implementation Example Nine: The difference between Embodiment Nine and Embodiment Eight of the present invention lies only in: The present invention provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement a method for allocating tasks for cooperative operation of underwater robots. Specific Implementation Example 10: The only difference between Embodiment 10 and Embodiment 9 of the present invention is that: The present invention provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement a method for allocating tasks for cooperative operation of underwater robots. Specific Implementation Example Eleven: The only difference between Embodiment Eleven and Embodiment Ten of this invention is that: This invention provides a method for assigning collaborative operation actions to an underwater robot, the method comprising the following steps: Step 1: Construct a knowledge base for motion allocation in a dual-robot-manipulator system; Step 2: Construct a dual-matrix game model based on the knowledge base, design payoff functions for different strategies of both sides of the game, and form a payoff matrix; Step 3: Use the multi-payer matrix weighting method to transform the game problem into an optimization problem for solution, and verify the correctness of the game model; Step 4: Solve the Nash equilibrium of the two-matrix game model by applying the improved differential evolution algorithm. The basic differential evolution algorithm is improved by applying the velocity control factor, the optimal point set theory and the step-type inertia factor.
[0032] In step one, a motion allocation knowledge base for the dual-robot-dual-manipulator system was constructed. The specific process is as follows: A multi-level knowledge base for motion assignment in a dual-robot / dual-manipulator system was constructed. This knowledge base is based on a multi-level knowledge representation model, encompassing the scene layer, object layer, agent layer, task layer, skill layer, and action layer. This knowledge base includes not only static environmental information and dynamic robot state information but also real-time updates to adapt to changes in the dynamic environment. This dynamic update mechanism ensures the accuracy and timeliness of the knowledge base during motion assignment.
[0033] (1) Scene layer: including factors in the environment in which multiple robots are located when handling tasks that affect the allocation of actions; (2) Object layer: includes all physical objects in the robot's autonomous operation process, such as submarine cables, buoys and arrays; (3) Intelligent agent layer: including equipment knowledge related to the execution of tasks in autonomous underwater robots, including various underwater robots, robotic arms, sensors, etc. that perform tasks; (4) Task layer: includes autonomous underwater robot operation tasks, including tasks such as handling cables, handling arrays, handling buoys, etc., which can be further broken down into basic tasks such as tracking, detection, grasping, and cutting. (5) Skills layer: This includes the robot's capabilities during autonomous operation, such as the ability to follow, detect, and cut cables. (6) Action layer: includes action primitives in the robot task execution process, including robot movement, selection of robotic arm, opening and closing of gripper, etc.
[0034] exist Figure 1 Based on this, a detailed design is made for the action allocation knowledge base, which is constructed in three categories: scenario, agent, and concept.
[0035] Scenario-based knowledge bases, such as Figure 2 The core node is the "Scene" node. The work objects in action assignment exist in the work environment, and the corresponding relationship is "Disposal Object → Located in → Scene". Work objects are divided into three categories: cables, floats, and arrays. Therefore, "Disposal Object → Contains → Cables, Floats, Arrays". Each work object has its own attributes that affect the action assignment result. Taking the "Cable" node as an example, "Cable → Attributes → Cable Shape, Cable Diameter, Cable Material, Cable Work Position". Each attribute fundamentally affects the action assignment result. Therefore, each attribute is divided in detail to match the work capabilities of each robot in the intelligent agent class, namely "Cable Diameter → Contains → Operable, Inoperable", "Cable Shape → Contains → Straight Cable, Branch Cable, Loop Cable", etc.
[0036] "Cable shape → includes → straight-through cables, branch cables, loop cables"; "Cable → includes → cable operation location", while "Cable shape and geology → affect → cable operation location", and geology includes various seabed landforms, "Geology → includes → muddy, sandy, silty, clayey, rocky", which can be specifically represented as: (1) (2) In the formula: and The two sets represent the coefficients of the payoff functions of the two players for different geological conditions.
[0037] Cable material is one of the factors affecting the completion of tasks by a robotic arm. The relationship between "cable" and "cable material" is as follows: (3) (4) In the formula: and These represent the shearing capacity coefficients of the robotic arms of the two robots for cables of different materials.
[0038] In the task of handling buoys, the state of the buoy is also one of the factors affecting the allocation of actions. The sequence is: "Buoy → Attribute → State", "State → Attribute → Stillness and Drifting". Based on fuzzy theory, the motion state of the buoy is defined as follows: (5) In the formula: This indicates the critical ocean current velocity required to keep the buoy in a dynamic state.
[0039] In addition to the task object, the task environment is also a crucial factor influencing action allocation in scenario-based knowledge bases. Among these, geology and ocean currents have the greatest impact on underwater tasks. The equation "Geology and Ocean Currents → Location → Scenario" applies. Geology includes various seabed landforms: "Geology → Includes → Muddy, Sandy, Muddy-Sand, Clayey, Rocky". Since robots often choose a bottom-sitting posture when performing cable work, not every geological condition is suitable for bottom-sitting; therefore, "Geology → Influences → Cable Work Position". Ocean currents have two attributes: direction and velocity. The equation "Ocean Currents → Attributes → Direction, Velocity" applies. Velocity is closely related to each robot's anti-interference capabilities; therefore, it is expanded to include whether the velocity is tolerable, matching the operational capabilities of intelligent agents: "Velocity → Attributes → Tolerable, Untolerable". The velocity classification is based on: (6) Flow direction is also a factor affecting operational capability, but it only plays a supporting role. It is roughly divided into categories based on the angle formed with the robot's heading: "Flow direction → Attributes → No interference, downstream, upstream." The definition of flow direction is: (7) In the formula: and These represent the angles of the ocean current and the robot's direction of travel in the world coordinate system, respectively.
[0040] Action allocation intelligent agent knowledge base such as Figure 3 As shown, the "robot" node is the core, "containing" two operational robots. Each robot "contains" two robotic arms and sensing devices such as sonar and cameras, and possesses operational capabilities for different tasks. The operational capabilities of each robot can be categorized into two types of nodes, "executable" and "non-executable," through the relationship "capability," and are "affected" by environmental information in the scene, matching the actions in the concept. Each "robotic arm" has four joints: "Joint 1," "Joint 2," "Joint 3," and "gripper." Joints at different levels have a "located" relationship, such as "gripper" "located" in "Joint 3," and "Joint 3" in turn "located" in "Joint 2." Each joint is matched with the rotation angle of the action layer. For different environmental information in different scenes, the combination of the rotation angles of each joint of the robotic arm results in the operation status of each robot.
[0041] The knowledge base for action assignment concepts consists of task modules and action modules, such as... Figure 4 As shown. The task module includes "Task Classification" and "Task Action" nodes. The "Task Classification" node is basically the same as the knowledge base of the task sequence allocation concept, with three basic task categories: "Contains," "Disposes of Cables," "Disposes of Arrays," and "Disposes of Floats." Each task "contains" multiple sub-tasks, and the completion of the sub-tasks matches the robot's actions and skills. The "Task Action" node is constructed based on a dual-robot-manipulator game model, with three categories: "Contains," "Robot Posture," "Manipulator Action," and "Manipulator Angle." The sub-nodes of each node are the strategies of the game model, such as "Contains," "Hovering," "Sitting on the Bottom," "Grasping," and "90 Degrees." Due to the different operational capabilities of the agents, each robot has different tendencies when selecting actions, and these tendencies are incorporated into the action allocation process. (8) (9) In the formula: and The four elements in the diagram represent the two robots' preferences for hovering, sitting, grasping, and shearing actions, respectively.
[0042] Then, the rotation angle and working direction of the robotic arm are "matched": (10) The motion module represents the basic movements of the robot and robotic arm, including "movement," "rotation," "opening and closing," "suspending," and "sitting." Movement refers to the robot's movement in water, which can be categorized as "forward and backward movement," "left and right movement," and "up and down movement" depending on the propeller's operation. Rotation can be divided into "robot rotation" and "robotic arm rotation." Robot rotation refers to the robot's adjustment of its own posture, including "roll," "pitch," and "bow." Robo arm rotation refers to the robotic arm adjusting the angles of its joints to reach different working positions, including "0 degrees," "90 degrees," "180 degrees," and "-90 degrees," based on commonly used working positions. "Opening and closing" represents the state of the robotic arm's gripper; by adjusting the "open" and "closed" postures, tasks such as grasping and shearing are achieved. "Suspension" and "sitting" are the most intuitive choices for adjusting the robot's working posture, serving as sub-nodes of task motion nodes and key components of the motion layer.
[0043] Then, the three types of knowledge bases are integrated to obtain a dual-robot-manipulator motion allocation knowledge base, such as... Figure 5 As shown, all subsequent action allocation strategies are based on this library.
[0044] In step two, a dual-matrix game model is constructed based on the knowledge base, and payoff functions for different strategies of both sides are designed to form a payoff matrix. The specific process is as follows: In terms of game theory model design, this invention constructs a game theory problem for the cooperative operation of an autonomous underwater robot. . Robots and robotic arm systems representing the two sides in a game of strategy and competition. Let them represent the sets of pure strategies for player one and player two, respectively. This indicates that the player has A pure strategy, This indicates that the two players have... A pure strategy. The sum of the payoffs for both sides of the game is used. To indicate, among which , . and Player 1 adopts a pure strategy Furthermore, player two adopts a pure strategy. The actual gains of both sides in the game.
[0045] The collaborative operation of the robot-manipulator system is decomposed into a three-layer game involving robot posture, manipulator motion, and manipulator angle. A corresponding payoff function is designed for each layer of the game, and a payoff matrix is constructed.
[0046] (1) First-level game The first layer of the game model involves a game of attitude between the two robots, with each robot choosing different attitudes depending on the seabed environment and ocean current information. These attitudes can be categorized into two types: hovering and bottoming. The choice of bottoming attitude is primarily based on the seabed environment, with coefficients for both robots set to different seabed conditions. The profit function for the two robots' sitting posture is: (11) In the formula: The speed of ocean currents that underwater robots can withstand. The speed of the maximum ocean current that an underwater robot can withstand. This represents the degree of preference different robots have for the sitting posture.
[0047] The payoff for a hovering posture is primarily influenced by the magnitude of the ocean current it resists. The payoff trend should decrease as the ocean current increases, and the payoff should be zero when the resisted current exceeds the maximum resistable current. The payoff function for the hovering posture for both players is as follows: (12) In the formula: This represents the degree of preference of both sides for the suspension posture.
[0048] (2) Second-level game The second layer of the game model involves a game of action between two robotic arms, where each side chooses actions based on different task types and the physical characteristics of the arms. Assuming the task is cutting an undersea cable, and both robots and their arms are located within the same longitudinal section, the 3D model is simplified to a 2D model as input. When establishing the payoff function for this game model, only the 2D coordinates within this longitudinal section need to be considered. After simplification, the payoff function for the cutting strategy is: (13) In the formula: Let be the preference coefficients of both sides in the game regarding the shearing action. To determine the capability coefficients of both sides for different cable materials, The maximum diameter of the cable that the robotic arms of both sides in the game can cut. This is the diameter of the cable at this time.
[0049] The payoff function of the grasping strategy can be divided into the robot's grasping reward and the relative distance reward. After the first layer of posture game, the absolute positions of both sides are determined, and the cable environment information observed by both robots can provide the relative position information for both sides. The relative position of the robotic arm at the end of its operation to itself It is also known that the reward function for the grabbing action is defined as: (14) In the formula: As a relative distance reward, To claim the reward.
[0050] (15) (16) In the formula: represents the preference coefficients of both sides in the game regarding the grabbing action.
[0051] (3) Third-level game The third layer of the game theory model involves a game of angles between the two robotic arms. Considering the range of motion of the arms and task requirements, a payoff function is designed to optimize the choice of arm angles. It is assumed that different arm angles affect the arm's load capacity and are also a significant factor in task completion efficiency. The maximum working range of both robotic arms is... The scope of work is divided into Consider the robotic arm angle benefit function in the partitioned areas: (17) Determine the angle benefit of the grasping action based on the load on the robotic arm: (18) Determine the benefits of changing the angle based on the initial angle and the working angle of the robotic arm's end effector: (19) In the formula: The initial angle of the cable relative to the base coordinate system when making decisions for the robot.
[0052] (20) The overall benefit of the robotic arm's grasping angle is: (twenty one) The payoff function for shearing angles can be divided into safety distance payoff and angle change payoff. Based on the coordinates of the cable relative to both players, the safety distance payoff is defined as follows: (twenty two) In the formula: and These represent the safe distances for the two sides in the game. and These represent the positions of the cable operation locations in the body coordinate systems of both parties.
[0053] Assume the attitude matrix of the robotic arm's end effector in body coordinates is as follows: The attitude matrix of the shear point in the body coordinate system is: , arrive The rotation matrix is denoted as The relationship between the three is as follows: (twenty three) Given the end effector posture of the robotic arm and the posture of the shearing operation position, the rotation transformation matrix between them can be obtained: (twenty four) During robotic arm operations, quaternions are used to represent the difference between the end effector posture and the shear point posture. A rotation matrix is defined. and quaternion matrix form The relationship between them: (25) We can obtain: (26) In the formula: The trace of the matrix, Representing the rotation matrix No. Line 1 The elements of the column.
[0054] Trigonometric form of quaternions It can be seen that quaternions can represent the number of times a number is about an axis. Rotation angle At the same time, the vector size becomes the original. The rotation angle can be calculated using the conversion relationships between different forms of quaternions. : (27) In the formula: the rotation matrix is an orthogonal identity matrix. .because , The range of values is .
[0055] The magnitude of this value can represent the difference between the current pose and the target pose, therefore The larger the value, the greater the required rotation angle. Based on this principle, we design the payoff function for changes in the shear angle: (28) Therefore, the profit function for the robotic arm's shearing is: (29) By combining the strategy sets of the two sides in the game above, we can obtain the single-objective payoff matrix of different robots at different game levels under different strategies: (30) In the formula: subscript Representing the two sides in the game, the corner marker These represent different game layers: robot state payoff, robotic arm action payoff, and robotic arm angle payoff. This represents when the player adopts a pure strategy Furthermore, player two adopts a pure strategy. Time-based gamers The benefits.
[0056] In step three, the game problem is transformed into an optimization problem using the multi-payer matrix weighting method for solution, and the correctness of the game model is verified. The specific process is as follows: In game theory, due to the differences in payoff functions at different levels and the heterogeneity of player preferences, a scientific method is needed to quantify the comprehensive payoff of different actions. However, the payoff function in a two-matrix game model is defined for different environmental information, exhibiting strong subjectivity and relatively complex application scenarios. Therefore, this paper constructs a weighted evaluation model based on the Analytic Hierarchy Process (AHP), a method suitable for solving complex subjective decision-making problems requiring quantification. The application method involves systematic steps such as establishing a judgment matrix, calculating weight vectors, and performing consistency checks, which effectively quantifies the payoff relationships between actions at different levels. The weighting steps between actions are as follows: (1) Construct the judgment matrix.
[0057] The specific objects considered are the three levels of the payoff function of the game problem: robot posture payoff, robotic arm action payoff, and robotic arm angle payoff. The overall payoff for each game strategy is obtained by weighted summation of these three objectives. Since this evaluation level only includes these three objectives, dividing their importance into five levels is sufficient to comprehensively assess the relative importance among these objectives, as shown in the table below: Table 1 Importance Ranking
[0058] By comparing the objectives, the corresponding players can be identified. The judgment matrix .
[0059] (31) (2) Check the consistency of the judgment matrix.
[0060] Since there are errors in the construction of the judgment matrix, it is necessary to check the consistency of the judgment matrix to ensure that the obtained judgment matrix is reliable and consistent.
[0061] (1) Calculate the consistency index : (32) In the formula: It is to determine the largest eigenvalue of the matrix. It is the dimension of the matrix.
[0062] (2) Find the corresponding average random consistency index. Table 2 shows... hour The value of .
[0063] Table 2 Average Random Consistency Index
[0064] (3) Calculate the consistency ratio : (33) when In this case, the consistency of the judgment matrix is considered acceptable; otherwise, the judgment matrix should be appropriately corrected.
[0065] Calculate the weight of each objective.
[0066] (34) In the formula: , .
[0067] The max-min variational method is applied to transform each single-objective matrix in the model into a normalized matrix, unifying the magnitude and normalizing the objective. Then, the payoff matrices of the k-th layer for both sides of the game are linearly weighted. (35) (36) In the formula: (37) (38) All single-objective payoffs of the two players in a multi-objective model and All are weighted by the formula: (39) In the formula: , By using weighted summation, the multi-level payoff matrix is transformed into a comprehensive payoff matrix for both sides of the game.
[0068] In step four, the improved differential evolution algorithm is applied for solving the problem. The specific process is as follows: Differential Evolution (DE) algorithm mainly includes four steps: population initialization, mutation, crossover, and selection. The flowchart of the standard DE algorithm is shown below. Figure 3 As shown. The basic idea of the DE algorithm is to use the difference between two random vectors as a perturbation to add to the basis vector to obtain the mutation vector V(t). The mutation vector crosses with the baseline vector to obtain the test vector U;(t). The baseline vector X;:(t) competes with the test vector U:(t) and is selected and retained. Through continuous evolution, the population obtains the optimal solution. In the DE algorithm, mutation and crossover realize the blind search of the algorithm, and selection realizes the purpose of guiding the evolutionary direction.
[0069] This paper proposes an improvement to the existing differential evolution algorithm by changing control parameters and adjusting the mutation strategy. The two improvements are: (1) using a novel adaptive mutation technique (MT) to improve the algorithm's selection of mutation strategies; and (2) proposing an error-related adaptive scaling factor (AF) to adaptively adjust the search step size. After the improvement, MT achieves adaptive switching between random mode and best mode, and simultaneously adjusts the scaling degree AF of the mutated individuals according to the error value at the current iteration number, thus realizing fast global optimization of the algorithm.
[0070] (1) Adaptive mutation operator MT In the process of constructing the game theory model, the following was introduced: and As auxiliary variables for the complementary relaxation conditions in the payoff function, they are included in the solution of the model for calculation. and As the optimal payoff for both sides in the iterative process of the game, it is closely related to the optimal solution. Therefore, in order to improve the comprehensive computing power and solution speed of the optimization algorithm, the optimization algorithm is improved based on the characteristics of the objective function in this paper.
[0071] In evolutionary optimization (DE), the choice of mutation strategy is the core of the algorithm, and its result greatly affects the algorithm's capability. This paper uses the adaptive mutation operator MT in the improved algorithm to achieve autonomous adjustment of the mutation strategy during the evolutionary process. The principle for selecting the mutation strategy is shown in the following equation: (40) The control method combines the DE / rand / 1 / bin strategy with the DE / best / l / bin strategy, and the MT control algorithm selects the mutation strategy.
[0072] Wherein, MT is defined as: (41) MT is related to the current iteration number t, the maximum iteration number T, and the power exponent N. The smaller the exponent N, the earlier the algorithm enters the DE / best / 1 / bin evolution mode, the stronger the development capability, and the faster the convergence speed. The power exponent determines the selection principle of the mutation strategy of the evolutionary algorithm. When N is greater than 1, the algorithm emphasizes exploration, and when N is less than 1, the algorithm emphasizes development. The value of T can be determined by experiment based on the specific problem.
[0073] (2) Adaptive scaling factor AF In solving game theory models, the payoffs of the two robots differ under various environments due to their varying capabilities. Sometimes the payoffs are significantly different, making the model easier to solve. Other times, the payoff matrices are very close, requiring a more powerful search algorithm. To meet the needs of different scenarios without excessive redundancy, the optimization algorithm is improved. In the DE algorithm, the scaling factor AF affects the search step size, determining the scaling degree of the mutated individuals. Increasing AF improves the algorithm's exploration ability but leads to excessive complexity. Conversely, decreasing AF improves the algorithm's exploration ability but causes premature convergence. To achieve fast global optimization, this paper proposes an error-dependent adaptive scaling factor AF. The scaling factor is 1 at the start of the algorithm's generation, providing strong exploration ability. When the error value of the current generation is less than a set value... When using the adaptive scaling factor in the following formula, the algorithm's development capability can be improved.
[0074] (42) The error value between the adaptive scaling factor AF and the current iteration number Related, The value can be adjusted according to the specific engineering problem.
[0075] Based on this improvement, the correctness of the algorithm was verified. To verify the excellent performance of the improved differential evolution algorithm in solving the Nash equilibrium of a high-dimensional game model, three different algorithms were compared: Whale Optimization Algorithm (WOA), Particle Swarm Optimization (PSO), and ordinary differential evolution algorithm (DE). During the solution process under the same scenario, the maximum number of iterations was limited to 100, and convergence was considered when the fitness value was below 0.001. The optimization results are as follows. Figure 8 As shown in the figure, in terms of convergence speed and accuracy, PSO exhibits a faster convergence speed but is prone to getting trapped in local optima. WOA has a relatively slow convergence speed, and the DE algorithm has stronger search capabilities but is not easy to converge. None of these three algorithms can meet the convergence accuracy requirements set in the experiment. The APDE algorithm inherits the search capabilities of DE and introduces... This improves the convergence speed of the algorithm while also offering better real-time performance and accuracy. Regarding stability, all the above algorithms use random initialization methods, and during the optimization process, they all experience a brief period of getting trapped in local optima, leading to a lack of stability in gradient descent. However, by introducing an adaptive penalty factor, APDE increases particle diversity, making the optimization process more stable and enabling it to continuously converge to the global optimum without getting trapped in local optima.
[0076] The above description is merely a preferred embodiment of a method for allocating tasks in collaborative underwater robot operations. The scope of protection for this method is not limited to the above embodiments; all technical solutions falling within this conceptual framework are within the scope of protection of this invention. It should be noted that for those skilled in the art, any improvements and variations made without departing from the principles of this invention should also be considered within the scope of protection of this invention.
Claims
1. A method for allocating tasks in cooperative underwater robot operations, characterized by: The method includes the following steps: Step 1: Construct a knowledge base for motion allocation in a dual-robot-manipulator system; Step 2: Construct a dual-matrix game model based on the knowledge base, design payoff functions for different strategies of both sides of the game, and form a payoff matrix; Step 3: Use the multi-payer matrix weighting method to transform the game problem into an optimization problem for solution, and verify the correctness of the game model; Step 4: Solve the Nash equilibrium of the two-matrix game model by applying the differential evolution algorithm, and improve the basic differential evolution algorithm by applying the velocity control factor, the optimal point set theory and the step-type inertia factor.
2. The method according to claim 1, characterized in that: Step 1 specifically involves: The multi-level knowledge representation model-based knowledge base for task action allocation of autonomous underwater robots also includes the following six levels: scene layer; object layer; agent layer; task layer; skill layer; and action layer. Based on this, a detailed design was made for the action allocation knowledge base, which was constructed in three categories: scenario, agent, and concept.
3. The method according to claim 2, characterized in that: Step 2 specifically involves: The collaborative operation of the robot-manipulator system is decomposed into a three-layer game of robot posture, manipulator action, and manipulator angle. A corresponding payoff function is designed for each layer of the game, and a payoff matrix is constructed.
4. The method according to claim 3, characterized in that: Regarding the robot's posture dynamics, based on different seabed environments and ocean current information, two different postures are chosen: sitting on the seabed and hovering. (1) in, The speed of ocean currents that underwater robots can withstand. The speed of the maximum ocean current that an underwater robot can withstand. To determine the different robots' preferences for sitting postures, This represents the degree of preference of both sides for the suspension posture.
5. The method according to claim 3, characterized in that: For the game theory of robotic arm actions, considering different task types and the physical characteristics of the robotic arm, the payoff functions for the shearing and grasping strategies are as follows: (2) in: Let be the preference coefficients of both sides in the game regarding the shearing action. Let be the preference coefficient between the two players for the grabbing action. This refers to the capability coefficients of both sides in the game regarding different cable materials.
6. The method according to claim 3, characterized in that: Regarding the game of angles for the robotic arm, considering the range of motion of the robotic arm and the task requirements, we will continue with the previous example. The angle payoff function for the robotic arm's grasping and shearing is as follows: (3) in, The initial angle of the cable relative to the base coordinate system when making decisions for the robot.
7. The method according to claim 4, 5 or 6, characterized in that: Step 3 specifically involves: A weighted evaluation model based on the analytic hierarchy process (AHP) is constructed. Through systematic steps such as establishing a judgment matrix, calculating weight vectors, and performing consistency checks, the payoff relationships between actions at each level are effectively quantified, thus representing the payoffs of all single-objective players in a multi-objective model. and Use the following formula for weighting: (4) In the formula: , By weighted summation, the multi-level payoff matrix is transformed into a comprehensive payoff matrix for both sides of the game.
8. A task allocation system for cooperative operation of underwater robots, characterized in that: The system includes: The knowledge base module constructs a knowledge base for motion allocation of the dual-robot-manipulator system. The payoff matrix module constructs a dual-matrix game model based on a knowledge base, designs payoff functions for different strategies of the two players, and forms a payoff matrix. The verification module uses the multi-payer matrix weighting method to transform the game problem into an optimization problem for solution and verifies the correctness of the game model. The solution module applies the differential evolution algorithm to solve the Nash equilibrium of the two-matrix game model. It improves the basic differential evolution algorithm by applying the velocity control factor, the optimal point set theory, and the step-type inertia factor.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method as claimed in claims 1-7.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the method of claims 1-7.
Citation Information
Patent Citations
Multi-AUV dynamic maneuvering decision-making method based on interval information game
CN112306070A
Task-grading double-mechanical-arm collaborative planning and control method and related device
CN115401697A
Task allocation method and system based on hybrid game genetic algorithm
CN120124979A
Unmanned aerial vehicle (UAV) task cooperation method based on overlapping coalition formation (OCF) game
US11567512B1