Human-machine cooperation personalized obstacle avoidance method based on game theory
The integration of game theory and human body feature recognition dynamically adjusts robotic arm paths to ensure personalized safety and efficiency in human-robot collaboration, addressing the inflexibility of existing collision avoidance methods.
Patent Information
- Application Number
- CN202510342267.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-15
AI Technical Summary
The existing dynamic obstacle avoidance methods for robotic arm do not fully consider the operator's body posture characteristics, resulting in a fixed obstacle avoidance space, lack of personalized safety guarantees, and cannot effectively balance the robotic arm task efficiency and operator safety needs.
By obtaining the operator's body shape information and motion trajectory in real time, combining game theory decision model, dynamically adjusting the motion trajectory of the robotic arm, using depth cameras and convolutional neural networks to generate a 3D human body model, extract body posture characteristics, calculate personalized safety distances, and solve the optimal obstacle avoidance strategy through game theory to ensure the balance of safety and efficiency.
It realizes personalized safety distance adjustment, improves operator safety and collaboration system efficiency, is suitable for a variety of industrial collaboration scenarios, and enhances the scope of application of human-computer collaboration systems.
Smart Images

Figure CN120307274A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robots, and particularly relates to a personalized obstacle avoidance method for human-robot collaboration based on game theory. Background Art
[0002] With the rapid development of robot technology, robotic arms have been widely used in human-robot collaboration tasks in the industrial and medical fields. The traditional physical fence protection method requires a large amount of space and has poor flexibility, unable to meet the collaborative needs of human-robot coexistence. Existing dynamic obstacle avoidance methods for robotic arms mainly estimate the operator's motion trajectory in real time through machine learning and prediction algorithms, but do not fully consider body characteristics such as height, body shape, etc., resulting in a fixed obstacle avoidance space and lacking personalized safety protection.
[0003] Game theory, as an important branch of operations research, optimizes the decision-making strategies of both parties by modeling the interaction behaviors among multiple intelligent agents. In recent years, significant progress has been made in the application of game theory in the fields of robot task coordination and path planning. How to introduce game theory into the human-robot collaboration obstacle avoidance strategy, and then effectively balance the task efficiency of the robotic arm and the safety needs of the operator, providing a more flexible and intelligent decision-making method for human-robot collaboration, has become the focus of research. Currently, there is a lack of relevant research in this field. Summary of the Invention
[0004] To this end, the present invention proposes a personalized obstacle avoidance method for human-robot collaboration based on game theory. This method dynamically adjusts the motion trajectory of the robotic arm by real-time obtaining the operator's body shape information and motion trajectory, and combining with the game theory decision-making model, ensuring that a personalized safety distance is maintained between the robotic arm and the operator during task execution, and can dynamically adjust the safety distance to balance safety and production efficiency.
[0005] In order to achieve the above technical features, the object of the present invention is achieved as follows: A personalized obstacle avoidance method for human-robot collaboration based on game theory, comprising the following steps: Step 1: Obtain human motion data and the pose of the robotic arm, real-time monitor the operator's motion information, and predict the operator's motion trajectory in the next period of time by using the Kalman filter method; at the same time, extract the operator's body characteristics, and analyze and normalize the operator's body shape; Step 2: Calculate the risk degree in real time according to the operation information obtained in Step 1; Step 3: According to the predicted operator's motion trajectory and the position of the robotic arm working area, determine whether the operator will overlap or get too close to the robotic arm working space; if not entering the working space, the robotic arm maintains the normal working state and continues the current operation; if entering, enter the robotic arm obstacle avoidance strategy process; Step 4: If the operator described in Step 3 enters the working area of the robotic arm, the risk level is calculated according to the obstacle avoidance strategy process of the robotic arm. If the risk level is greater than 1, it means that the minimum safety distance is greater than the current minimum distance between the human and the machine, and the risk level is high, so the robotic arm stops moving. If the risk level is less than or equal to 1, it means that there is a space for game between the human and the robotic arm, and then game decision-making is carried out. Step 5: According to the predicted movement trajectory of the operator in Step 1, a game model is constructed. The revenue functions of the operator and the robotic arm are calculated respectively using the backward induction method to obtain the optimal path of the operator. Based on this path, the robotic arm combines its own revenue function to solve the game equilibrium and generate the optimal game strategy. Step 6: Move according to the selected strategy, and judge whether the target position is reached. If it is reached, the process ends; otherwise, go back to Step 1 and continue.
[0006] Preferably, Step 1 specifically includes: Step 1.1: Use a depth camera to obtain the human RGB-D image, and combine a convolutional neural network to regress the SMPL parameters of the human body to generate an initial 3D human model. Step 1.2: Extract the correspondence between image pixels and 3D grid points through DensePose mapping, and further optimize the three-dimensional position of the human body grid using depth information to improve the accuracy of body shape estimation. Step 1.3: In the iterative optimization process, eliminate pose interference through the T-Pose alignment loss. Finally, extract the coordinates of human key points from the body shape coefficients of the SMPL model, calculate the height information, and estimate personalized body posture information through the position of the waist grid points. Step 1.4: Obtain the body shape information of the operator through the SoY method for the calculation of the risk level, and during the game process, make personalized adjustments to the surrounding space in combination with the body posture characteristics of the operator.
[0007] Preferably, the body posture characteristics of the operator in Step 1.3 specifically include information such as tall, short, fat, and thin.
[0008] Preferably, the personalized adjustment of the surrounding space in combination with the body posture characteristics of the operator in Step 1.4 includes: reserving or reducing the space above the operator for height factors, and dynamically adjusting the space on both sides for body shape width.
[0009] Preferably, in Step 2, it is judged whether the robotic arm enters the game state according to the magnitude of the risk level. The risk level The specific calculation formula is: ; In the formula: is the risk level; is the velocity component of the operator along the direction of the robotic arm; is the velocity component of the robotic arm along the operator's direction; is the total response time, including the time required for the sensor, logic unit, actuator, and the braking or stopping of the machine itself; is the maximum displacement generated by the robotic arm within the total response time; is the reach distance of the protection device; is the distance compensation coefficient, used to correct uncertainties such as sensor measurement deviations, installation errors, or reflections; is the height adjustment system; is the body fat adjustment coefficient; is the human height normalization coefficient, where is the height of the operator, is the reference standard height; is the human body fat normalization coefficient, where is the waist circumference of the operator, is the reference standard waist circumference; is the shortest distance between the operator and the robotic arm; is the minimum safety distance adjusted considering body posture factors.
[0010] Preferably, in step 5, the process of formulating the game strategy: When the operator and the robotic arm play a game, the strategy set of the operator : Move one step forward; : Stop moving; : Step back one step}; The strategy set of the robotic arm : Move one step forward along the original path; : Move one step forward along the original path with deceleration; : Re-plan the path and move one step forward; : Move one step backward along the original path}; where re-planning the path means the movement path for the robotic arm to avoid obstacles with the operator; moving along the original path means the path for the robotic arm to move towards the target point following the pre-planned path; The revenue function of the robotic arm consists of four parts: safety term, task efficiency term, motion smoothness term, and path stability term. The specific revenue function formula of the robotic arm is as follows: ; In the formula: is the revenue function of the robotic arm; is the minimum safety distance adjusted considering body posture factors; is the straight-line distance between the operator and the robotic arm at the next moment; , is the distance between the robotic arm and the target point at the current and next moments; is the acceleration of the robotic arm at the current moment; is the maximum acceleration of the robotic arm, serving as the normalization benchmark; is the path change frequency; is the maximum path change frequency, serving as the normalization benchmark; is the weight coefficient, and .
[0011] Preferably, in step 5, the process of formulating the game strategy: The operator's payoff function consists of three parts: a safety term, a comfort term, and a path stability term. The specific formula is as follows: ; In the formula: U M is the operator's payoff function; , respectively represent the distances between the operator and the robotic arm at the current moment and the next moment; is the acceleration of the robotic arm; is the maximum acceleration of the robotic arm, serving as the normalization benchmark; is the path change frequency; is the maximum path change frequency, serving as the normalization benchmark; , , is the operator payoff weight coefficient, satisfying .
[0012] Preferably, in the process of formulating the game strategy in step 5, the operator and the robotic arm conduct a non-zero-sum game under the conditions of incomplete information dynamic game. The optimal strategy combination is derived through backward induction to balance the payoff functions of both sides and ensure safe obstacle avoidance; Starting from the end decision moment of the game, assume that at the current moment the operator selects strategy , and the robotic arm selects strategy . When the robotic arm gradually approaches the operator, the safety distance decreases, and the first term in the payoff function decreases, prompting the robotic arm to preferentially choose a strategy to move away from the operator, thereby reducing the collision risk; when the robotic arm approaches the target point, the second term in the payoff function increases, driving the robotic arm to be more inclined to execute a strategy to approach the target point and complete the task goal first; on the premise of ensuring safety, the robotic arm prefers to choose a strategy to optimize the task execution efficiency rather than simply avoiding obstacles, avoiding unnecessary damage to production efficiency, and ensuring the smoothness of the robotic arm's motion trajectory by constraining the acceleration term and reducing the potential negative impact of violent motion on system stability. Using backward induction, from the final decision moment to the current moment, assume that at the future moment t +1, the optimal strategy of the robotic arm is , then t the optimal strategy of the operator at the moment At t moment, the robotic arm selects the optimal strategy that maximizes its own benefit by solving the game equilibrium: ; Through this way of backward induction derivation, the optimal strategy combination of the operator and the robotic arm is finally obtained .
[0013] The present invention has the following beneficial effects: Based on human feature recognition and game theory methods, the invention proposes a personalized human-machine collaboration obstacle avoidance strategy. In terms of safety improvement: the safety distance is dynamically adjusted through the personalized robotic arm obstacle avoidance strategy, fully considering the body shape characteristics of the operator to ensure the safety of the operator. In terms of efficiency optimization: the operation efficiency is maximized on the premise of ensuring safety, the stagnation time caused by obstacle avoidance is reduced, and the working efficiency of the collaboration system is improved. In terms of strong versatility: it is applicable to a variety of industrial collaboration scenarios, and can be adaptively adjusted for operators with different body shapes, enhancing the applicable range of the human-machine collaboration system. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The present invention will be further described below with reference to the drawings and embodiments.
[0015] Figure 1 is the flow chart of the human-machine collaboration obstacle avoidance strategy.
[0016] Figure 2 is the schematic diagram of the optimal strategy solution process of the human-machine game.
[0017] Figure 3 is the schematic diagram of the body shape feature extraction process based on the SoY method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The following will detail the technical solutions in the embodiments of the present invention, where the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Without departing from the design concept of the present invention, various improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope of the present invention.
[0019] Embodiment 1: As Figure 1 shown, the purpose of the present invention is to provide a personalized robotic arm obstacle avoidance strategy based on game theory, human body prediction and the Shape of You (SoY) method. This strategy dynamically adjusts the movement trajectory of the robotic arm by real-time obtaining the body shape information and movement trajectory of the operator, and combines the game theory decision-making model to ensure that a personalized safety distance is maintained between the robotic arm and the operator when performing tasks.
[0020] Provide an idea of a collision avoidance strategy for a robotic arm based on game theory and human body postural characteristics in human-robot collaboration, mainly including the following parts: acquisition and prediction of human motion information, calculation of body type characteristics and personalized safety distances, acquisition of robotic arm motion information and path planning, decision-making model based on game theory, risk assessment and strategy selection, personalized safety adjustment and execution.
[0021] Among them, the SoY method in the strategy method based on game theory and human body postural characteristics uses a depth camera to obtain human RGB-D images, combines a convolutional neural network to regress SMPL parameters to generate a 3D human model. By mapping the correspondence between image pixels and 3D grid points through DensePose, and combining depth information to optimize the grid position, the body type estimation accuracy is improved. The T-Pose alignment loss is used to eliminate pose interference, and the SMPL body type coefficients are extracted to obtain personalized body type information such as height and waist circumference. The body type information obtained by the SoY method is used for risk calculation, and the space around the operator is personalized adjusted during the game process. The upper space is reserved or reduced according to height, and the two sides of the space are dynamically adjusted according to the body type width.
[0022] Among them, Shape of You (SoY) is a method for improving 3D human shape estimation in vision-based clothing recommendation systems, proposed by Rohan Sarkar et al. in 2023. This method aims to solve the problem of insufficient accuracy in existing technologies when dealing with diverse body types, and achieves breakthroughs through the following core innovations: 1) Loss function optimization: Two loss functions that can be embedded in existing 3D human reconstruction frameworks are proposed to enhance the adaptability of the model to different body types and reduce shape estimation errors. 2) Test-time optimization: A post-processing optimization process is introduced to further improve the details and accuracy of the reconstruction results, especially performing better on complex body types (such as obese or muscular types). 3) Performance improvement: On the challenging SSP-3D dataset, the SoY method has a 17.7% improvement in shape estimation accuracy compared to the previous Shapy method, providing a more reliable technical basis for personalized clothing recommendations in the fashion industry.
[0023] Among them, the SMPL model is the Skinned Multi-Person Linear Model. The parameters of the SMPL model are divided into two categories: body type parameters and pose parameters, which are used to generate three-dimensional human models with different shapes and actions. SMPL parameters are widely used in the fields of human pose estimation, 3D reconstruction, and animation generation, supporting the recovery of human actions and shapes from 2D images or videos.
[0024] The danger level in the strategy method based on game theory and human body postural characteristics is the ratio of the minimum safety distance considering postural factors to the minimum distance between the robotic arm and the operator. The minimum safety distance is based on the national standard GB / T 41355-2022 and is dynamically adjusted through human feature recognition technology, and is set considering the operator's postural characteristics. It consists of the human height, the normalization coefficient of body fatness, the adjustment coefficient of height and body fatness, etc. A danger level value greater than 1 indicates that the distance between the robotic arm and the operator is less than the minimum safety distance, which is very dangerous, and the robotic arm should be stopped immediately. If the value is less than or equal to 1, the decision-making method based on game theory is used to make the next decision to obtain the optimal strategy.
[0025] For the revenue function of the strategy method based on game theory and human body postural characteristics, the revenue function of the robotic arm consists of four parts: The safety term reflects the change in the distance between the robotic arm and the operator, that is, the difference between the minimum safety distance between the human and the machine at the current moment and the distance between the robotic arm and the human after moving one step according to the selected strategy divided by the minimum safety distance between the human and the machine at the current moment; The task efficiency term reflects the speed at which the robotic arm approaches the target point, that is, the difference between the distance between the current robotic arm and the target point and the distance between the robotic arm and the target point at the next moment divided by the distance between the current robotic arm and the target point; The motion smoothness term measures the smoothness of the robotic arm during movement, that is, the ratio of the current acceleration of the robotic arm to the maximum acceleration of the robotic arm; The path stability term measures the frequency of path change of the robotic arm, that is, the ratio of the path change frequency to the maximum path change frequency. The revenue function of the operator consists of three parts: The safety term reflects the trend of the robotic arm moving away from the operator, that is, the ratio of the difference between the distance between the operator and the robotic arm at the next moment and the distance between the operator and the robotic arm at the current moment to the distance between the operator and the robotic arm at the current moment; The comfort term measures the impact of the robotic arm movement on the operator's comfort, that is, the ratio of the current acceleration of the robotic arm to the maximum acceleration of the robotic arm; The path stability term.
[0026] Example 2: In human-robot collaboration, a robotic arm obstacle avoidance method based on game theory and human body postural characteristics, the method includes the following steps: Step 1: Obtain human motion data and the pose of the robotic arm through a depth camera, real-time monitor the operator's motion information, and predict the operator's future motion trajectory for a period of time through the Kalman filter algorithm. Extract the operator's postural characteristics, and use the SoY method to analyze and normalize the operator's body shape for subsequent personalized safety distance calculation.
[0027] Step 2: Calculate the danger level in real time according to the motion information obtained in Step 1.
[0028] Step 3: According to the predicted operator's motion trajectory and the position of the robotic arm's working area, determine whether the operator will overlap or get too close to the robotic arm's working space. If the operator will not enter the working space, the robotic arm maintains its normal working state and continues the current operation. If the operator will enter, then enter the robotic arm obstacle avoidance strategy process.
[0029] Step 4: If the operator described in Step 3 enters the working area of the robotic arm, then calculate the risk degree in the robotic arm obstacle avoidance strategy process. If the risk degree is greater than 1, indicating that the minimum safety distance is greater than the current minimum human-machine distance, the robotic arm stops moving. If the risk degree is less than or equal to 1, indicating that there is a playable space between the human and the robotic arm, then make a game decision.
[0030] Step 5: According to the operator's motion trajectory predicted in Step 1, construct a game model, and use the backward induction method to calculate the payoff functions of the operator and the robotic arm respectively to obtain the optimal path of the operator. Based on this path, the robotic arm combines its own payoff function to solve the game equilibrium and generate the optimal obstacle avoidance strategy.
[0031] Step 6: Move according to the selected strategy, judge whether the target position is reached. If it is reached, end the process; otherwise, go back to Step 1 and continue.
[0032] Embodiment 3: A robotic arm obstacle avoidance method based on game theory and human body postural characteristics in human-robot collaboration. According to Figure 1 as shown, its content includes the following steps: Step 1: Obtain human motion data and the pose of the robotic arm through a depth camera, real-time monitor the operator's motion information, and predict the operator's motion trajectory in the next period of time through the Kalman filtering algorithm. Extract the postural characteristics of the operator, and use the Figure 3 SoY method in
[0033] to analyze and normalize the operator's body shape for subsequent personalized safety distance calculation.
[0034] ; Among them, the minimum safety distance adjusted considering postural factors and the current minimum human-machine distance ratio constitutes the risk degree including postural factors. If D is greater than 1, indicating that the current minimum human-machine distance is less than or equal to the shortest distance before collision, there is a greater risk of collision, and the robotic arm should immediately stop working; if D is less than or equal to 1, indicating that the current minimum human-machine distance is greater than the shortest distance before collision, the robotic arm can enter the next decision, and choose to take other collision avoidance measures or continue normal work according to the specific situation.
[0035] In the formula: is the risk level; is the velocity component of the operator along the direction of the robotic arm; is the velocity component of the robotic arm along the direction of the operator; is the total response time, including the time required for the sensor, logic unit, actuator, and the braking or stopping of the machine itself; is the maximum displacement generated by the robotic arm within the total response time; is the reach distance of the protection device; is the distance compensation coefficient, used to correct uncertainties such as sensor measurement deviations, installation errors, or reflections; is the height adjustment system; is the body build adjustment coefficient; is the human height normalization coefficient, where is the height of the operator, is the reference standard height; is the human body build normalization coefficient, where is the waist circumference of the operator, is the reference standard waist circumference; is the shortest distance between the operator and the robotic arm; is the minimum safety distance adjusted considering body posture factors.
[0036] Step 3: According to the predicted movement trajectory of the operator and the position of the robotic arm's working area, determine whether the operator will enter the working space of the robotic arm or get too close to it. If the operator will not enter the working space, the robotic arm will maintain its normal working state and continue with the current task; if the operator may enter the working space, then enter the robotic arm obstacle avoidance strategy process.
[0037] Step 4: If the operator enters the working area of the robotic arm as described in Step 3, then calculate the risk level in the robotic arm obstacle avoidance strategy process. If the risk level is greater than 1, it means the minimum safety distance is greater than the current minimum human - machine distance, and the risk level is high, so the robotic arm stops moving. If the risk level is less than or equal to 1, it means there is a space for negotiation between the human and the robotic arm, and then a negotiation decision - making is carried out.
[0038] Step 5: When entering the negotiation decision - making stage, the process is as Figure 2As shown in the figure. Through game theory modeling, the optimal obstacle avoidance strategy in the process of human-robot collaboration is determined. Assume that the strategy set of the operator is \(M = \{M_1:\) Move forward one step; \(M_2:\) Stop moving; \(M_3:\) Move back one step\(\}\); the strategy set of the robotic arm is \(Z = \{Z_1:\) Move forward one step along the original path; \(Z_2:\) Move forward one step along the original path at a reduced speed; \(Z_3:\) Re-plan the path and move forward one step; \(Z_4:\) Move back one step along the original path\(\}\). First, predict the operator's motion trajectory through Step 1 and construct the game model of both sides. The reverse induction method is used for derivation, and the specific process is as follows: For each moment in the future, the possible strategy set of the operator is , and the possible strategy set of the robotic arm is . Use the operator's motion trajectory prediction model to judge the optimal strategy it may adopt. Starting from the last step of the game, assume that the operator selects strategy at the current moment, and the robotic arm selects strategy . The payoff functions of both sides are: ; ; In the formula: is the payoff function of the robotic arm; is the minimum safety distance after considering the adjustment of body posture factors; is the straight-line distance between the operator and the robotic arm at the next moment; , is the distance between the robotic arm and the target point at the current and next moments; is the acceleration of the robotic arm at the current moment; is the maximum acceleration of the robotic arm, used as the normalization benchmark; is the path change frequency; is the maximum path change frequency, used as the normalization benchmark; is the weight coefficient, and ; U M is the payoff function of the operator; , respectively represent the distances between the operator and the robotic arm at the current moment and the next moment; is the acceleration of the robotic arm; is the maximum acceleration of the robotic arm, used as the normalization benchmark; is the path change frequency; is the maximum path change frequency, used as the normalization benchmark; , , is the operator payoff weight coefficient, satisfying .
[0039] When the robotic arm gradually approaches the operator, the safety distance decreases, and the first term in the reward function increases, prompting the robotic arm to prefer strategies that keep it away from the operator to reduce the risk of collision. When the robotic arm approaches the target point, the second term in the reward function increases accordingly, driving the robotic arm to be more inclined to execute strategies that approach the target point and prioritize completing the operation task. On the premise of ensuring safety, the robotic arm is more inclined to choose strategies that optimize the task execution efficiency rather than simply avoid obstacles to prevent unnecessary losses to production efficiency. By constraining the acceleration term, the smoothness of the robotic arm's motion trajectory is ensured, and the adverse effects of violent motion on the system stability are avoided.
[0040] Using backward induction, starting from the final decision-making moment and deriving to the current moment, assuming that at the future moment t +1, the optimal strategy of the robotic arm is , then t the optimal strategy of the operator at moment , at t moment, the robotic arm selects the optimal strategy that maximizes its own reward by solving the game equilibrium: ; Through this way of backward induction derivation, the optimal strategy combination of the operator and the robotic arm is finally obtained, which balances the reward functions of both sides and ensures safe obstacle avoidance.
[0041] Step 6: Execute the motion according to the selected strategy, and determine whether the target position has been reached. If it has been reached, the task ends; if not, return to Step 1 to continue the subsequent process.
[0042] The above embodiments are only used to illustrate the technical solutions of the present invention and do not limit the present invention; although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that any modification or replacement does not make the essence of the corresponding technical solution deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A personalized obstacle avoidance method for human-machine collaboration based on game theory, characterized in that Including the following steps: Step 1: Obtain human motion data and the pose of the robotic arm, monitor the operator's motion information in real time, and predict the operator's motion trajectory in the next period of time through the Kalman filtering method; at the same time, extract the operator's body posture characteristics, and analyze and normalize the operator's body shape; Step 2: Calculate the risk level in real time according to the operation information obtained in Step 1; Step 3: According to the predicted operator's motion trajectory and the position of the robotic arm working area, judge whether the operator will overlap or get too close to the robotic arm working space; if not entering the working space, the robotic arm maintains the normal working state and continues the current operation; if entering, enter the robotic arm obstacle avoidance strategy process; Step 4: If the operator enters the robotic arm working area in Step 3, calculate the risk level in the robotic arm obstacle avoidance strategy process. If the risk level is greater than 1, it means that the minimum safety distance is greater than the current minimum human-machine distance, and the risk level is high, so the robotic arm stops moving; if the risk level is less than or equal to 1, it means that there is a space for game between the human and the robotic arm, then make a game decision; Step 5: According to the operator's motion trajectory predicted in Step 1, construct a game model, use the backward induction method to calculate the payoff functions of the operator and the robotic arm respectively, obtain the optimal path of the operator, and the robotic arm solves the game equilibrium based on this path and combines its own payoff function to generate the optimal game strategy; Step 6: Move according to the selected strategy, judge whether the target position is reached. If reached, end; otherwise, go back to Step 1 and continue.
2. The personalized obstacle avoidance method for human-machine collaboration based on game theory according to claim 1, wherein Step 1 specifically includes: Step 1.1: Use a depth camera to obtain the human RGB-D image, combine a convolutional neural network to regress the SMPL parameters of the human body, and generate an initial 3D human model; Step 1.2: Extract the correspondence between image pixels and 3D grid points through DensePose mapping, and use depth information to further optimize the three-dimensional position of the human body grid to improve the accuracy of body shape estimation; Step 1.3: In the iterative optimization process, eliminate the pose interference through the T-Pose alignment loss. Finally, extract the coordinates of human key points from the body shape coefficients of the SMPL model, calculate the height information, and estimate the personalized body posture information through the position of the waist grid points; Step 1.4: Obtain the operator's body shape information through the SoY method for risk level calculation, and during the game process, make personalized adjustments to the surrounding space in combination with the operator's body posture characteristics.
3. The personalized obstacle avoidance method for human-machine collaboration based on game theory according to claim 2, wherein The operator's body posture characteristics in Step 1.3 specifically include information such as tall, short, fat, and thin.
4. The personalized obstacle avoidance method for human-machine collaboration based on game theory according to claim 2, characterized in that, The personalized adjustment of the surrounding space in Step 1.4 in combination with the operator's body posture characteristics includes: reserving or reducing the space above the operator for height factors, and dynamically adjusting the space on both sides for body width factors.
5. The personalized obstacle avoidance method for human-machine collaboration based on game theory according to claim 2, characterized in that, In step 2, it is determined whether the robotic arm enters the game state by the magnitude of the risk level. The risk level The specific calculation formula is as follows: ; In the formula: is the risk level; is the velocity component of the operator along the direction of the robotic arm; is the velocity component of the robotic arm along the direction of the operator; is the total response time, including the time required for the sensor, logic unit, actuator, and the braking or stopping of the machine itself; is the maximum displacement generated by the robotic arm within the total response time; is the reach distance of the protection device; is the distance compensation coefficient, used to correct uncertainties such as sensor measurement deviation, installation error, or reflection; is the height adjustment coefficient; is the body build adjustment coefficient; is the human height normalization coefficient, where is the height of the operator, is the reference standard height; is the human body build normalization coefficient, where is the waist circumference of the operator, is the reference standard waist circumference; is the shortest distance between the operator and the robotic arm; is the minimum safety distance adjusted considering postural factors.
6. The personalized obstacle avoidance method for human-machine collaboration based on game theory according to claim 2, characterized in that, The game strategy formulation process in Step 5: When the operator and the robotic arm play a game, the strategy set of the operator : Move one step forward continuously; : Stop moving; : Move one step backward}; The strategy set of the robotic arm : Move one step forward along the original path; : Move one step forward along the original path with deceleration; : Re-plan the path and move one step forward; : Move one step backward along the original path}; where re-planning the path means the movement path for the robotic arm to avoid obstacles with the operator Moving along the original path means that the robotic arm follows the path planned in advance towards the target point. The payoff function of the robotic arm consists of four parts: safety term, task efficiency term, motion smoothness term, and path stability term. The specific payoff function formula of the robotic arm is as follows: ; Wherein: is the revenue function of the robotic arm; is the minimum safety distance after adjustment considering postural factors; is the straight-line distance between the operator and the robotic arm at the next moment; , is the distance between the robotic arm and the target point at the current and next moments; is the acceleration of the robotic arm at the current moment; is the maximum acceleration of the robotic arm, serving as the normalization benchmark; is the path change frequency; is the maximum path change frequency, serving as the normalization benchmark; is the weight coefficient, and .
7. The personalized obstacle avoidance method for human-machine collaboration based on game theory according to claim 6, wherein The game strategy formulation process in Step 5: The operator's payoff function consists of three parts: a safety term, a comfort term, and a path stability term. The specific formula is as follows: ; Wherein: U M is the benefit function of the operator; , respectively represent the distances between the operator and the robotic arm at the current moment and the next moment; is the acceleration of the robotic arm; is the maximum acceleration of the robotic arm, serving as the normalization benchmark; is the path change frequency; is the maximum path change frequency, serving as the normalization benchmark; , , is the benefit weight coefficient of the operator, satisfying .
8. The personalized obstacle avoidance method for human-machine collaboration based on game theory according to claim 7, characterized in that In the process of formulating the game strategy in Step 5, the operator and the robotic arm conduct a non-zero-sum game under the conditions of incomplete information dynamic game. The optimal strategy combination is derived through backward induction to make the payoff functions of both sides reach equilibrium and ensure safe obstacle avoidance. Starting from the end decision-making moment of the game, assume that at the current moment, the operator selects strategy , and the robotic arm selects strategy . When the robotic arm gradually approaches the operator, the safety distance decreases, and the first term in the reward function decreases, prompting the robotic arm to preferentially select a strategy of moving away from the operator, thereby reducing the collision risk; when the robotic arm approaches the target point, the second term of the reward function increases, driving the robotic arm to be more inclined to execute the strategy of approaching the target point and giving priority to completing the task objective. On the premise of ensuring safety, the robotic arm prefers to choose strategies that optimize task execution efficiency rather than simply avoid obstacles, so as to avoid unnecessary damage to production efficiency. By constraining the acceleration term, the smoothness of the robotic arm's motion trajectory is guaranteed, and the potential negative impact of violent motion on system stability is reduced. Using backward induction, it is deduced from the final decision-making moment to the current moment. Assuming that at t time +1, the optimal strategy of the robotic arm is , then t the optimal strategy of the operator at the moment . At t time, the robotic arm selects the optimal strategy that maximizes its own benefit by solving the game equilibrium: ; Through this way of backward induction derivation, the optimal strategy combination of the operator and the robotic arm is finally obtained .