Multi-agent reinforcement learning human-machine collaborative perception decision optimization method and system

CN122851608APending Publication Date: 2026-10-02HEBEI CHEM & PHARMA COLLEGE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611130279.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-10-02

AI Technical Summary

Technical Problem

然而,现有方法多假设机器人处于理想健康状态,忽略了设备老化、磨损等因素导致的性能退化对决策安全性的影响

Benefits of technology

(1)本申请将机器人健康状态向量作为协同决策模型的状态输入,并在奖励函数中引入基于健康状态向量的安全惩罚项,使决策系统能够主动感知性能退化并自适应调整策略,解决了可靠性信息与协同决策过程相互割裂的问题,有效避免因突发故障导致的安全事故和生产中断。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122851608A_ABST
    Figure CN122851608A_ABST
Patent Text Reader

Abstract

The application discloses a multi-agent reinforcement learning man-machine-object collaborative perception decision optimization method and system, relates to the field of man-machine-object collaborative perception decision optimization, and comprises the following steps: acquiring multi-source heterogeneous perception information and fusing the same to obtain collaborative perception features, defining man-machine collaborative interaction rules, collecting robot operation data and generating a health state vector, taking the collaborative perception features and the health state vector as state inputs to construct a Markov decision model, limiting an action boundary according to the interaction rules, defining a reward function to optimize a collaborative strategy, and outputting a man-machine joint action, collecting execution feedback to quantitatively evaluate and generate an evaluation feedback signal, updating the value estimation of the collaborative strategy according to the evaluation feedback signal, and adjusting the punishment intensity parameter of a safety punishment item to realize continuous optimization of the strategy; the collaborative strategy can actively perceive and adaptively respond to robot performance degradation, and realizes continuous optimization of the strategy through performance evaluation feedback, thereby improving the safety and adaptability of a man-machine collaborative system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of human-machine-object collaborative perception and decision optimization technology, and in particular to a multi-agent reinforcement learning method and system for human-machine-object collaborative perception and decision optimization. Background Technology

[0002] With the deepening of intelligent manufacturing and industrial development, human-robot collaborative robot systems are widely used in production, logistics and warehousing, and medical rehabilitation. Unlike traditional industrial robots, human-robot collaborative robots emphasize safe and efficient collaborative work between humans and robots in shared spaces, which places higher demands on the system's perception capabilities, decision-making intelligence, and safety reliability. Existing human-robot collaborative systems mainly rely on preset safety rules and fixed task allocation strategies, making it difficult to adapt to dynamically changing manufacturing environments and robot performance degradation, resulting in a significant contradiction between collaborative efficiency and safety. Furthermore, the effective fusion of multi-source heterogeneous perception information, real-time identification of robot health status, and dynamic optimization of collaborative decision-making remain key technological bottlenecks restricting the performance improvement of human-robot collaborative systems.

[0003] In recent years, multi-agent reinforcement learning technology has demonstrated significant advantages in the field of collaborative decision-making, enabling the autonomous allocation and execution of complex tasks through interactive learning among agents. However, existing methods often assume that the robot is in an ideal health state, neglecting the impact of performance degradation caused by factors such as equipment aging and wear on decision-making safety. Furthermore, current performance evaluation systems primarily focus on task completion efficiency and human-machine ergonomics, lacking quantitative consideration of robot reliability and stability, making it difficult to form a complete closed loop of "perception-decision-evaluation-optimization." Therefore, there is an urgent need for a human-machine-object collaborative perception and decision-making method and system that integrates dynamic constraints on the robot's health state into the collaborative decision-making process and supports continuous optimization. Summary of the Invention

[0004] The purpose of this application is to provide a human-machine-object collaborative perception and decision optimization method and system based on multi-agent reinforcement learning, so as to solve the above-mentioned problems existing in the prior art.

[0005] To achieve the above objectives, this application provides a multi-agent reinforcement learning-based human-machine-object collaborative perception and decision optimization method, comprising the following steps: S1: Acquire multi-source heterogeneous sensing information from personnel operation terminals, robot controllers, equipment sensors and environmental monitoring devices, preprocess it, and fuse the preprocessed multi-source features to obtain collaborative sensing features; S2: Establish the topology of the human-machine system and define the rules for human-machine collaborative interaction; S3: Collect samples of continuous fault-free working time and fault repair time of the robot, calculate the mean fault-free time, mean repair time and their temporal variation trends; determine whether the robot has experienced performance degradation based on the temporal variation trends, calculate the coefficient of variation of working time and fault time, and generate a robot health status vector; the health status vector includes a degradation status identifier, a quantified value of the degradation degree, and the coefficient of variation of working time and fault time. S4: Construct a Markov decision model for human-machine-object collaborative decision-making, using collaborative perception features and health state vectors as state inputs; define the action boundaries of the joint action space according to interaction rules; define a reward function, including task completion reward and safety penalty based on health state vector; optimize the collaborative strategy according to the reward function, and output human-machine joint actions that conform to the interaction rule constraints. S5: Collect state data after performing human-machine collaborative actions as execution feedback; construct a human-machine collaboration efficiency evaluation system, the criteria layer of which includes collaboration efficiency, safety, human-machine ergonomics, and robot reliability and stability; perform quantitative evaluation based on execution feedback and generate evaluation feedback signals; S6: Update the value estimate of the collaborative strategy based on the evaluation feedback signal, and adjust the penalty intensity parameter of the security penalty item to achieve continuous optimization of the strategy.

[0006] Preferably, the preprocessing in step S1 includes time alignment, anomaly removal, and feature extraction of the multi-source heterogeneous sensing information.

[0007] Preferably, the fusion in S1 adopts a weighted fusion strategy, specifically including: The fusion weights are determined based on the data quality, signal-to-noise ratio, or task relevance of each modality. The preprocessed multi-source features are then weighted and summed according to the fusion weights to obtain the collaborative perception features.

[0008] Preferably, the human-computer collaborative interaction rules in step S2 include: Human-robot space safety distance constraint: A minimum safety distance threshold is set based on the robot's working range, the end effector's motion trajectory, and the area where personnel are active; when the actual distance between the human and the robot is less than the threshold, a deceleration or stop command is triggered. Human-robot task timing constraints: Define the order, parallel conditions, and switching logic between human-operated tasks and robot-executed tasks to avoid conflicting tasks between humans and robots in the same time and space region; Robot motion range constraints: Based on the current task type and robot health status, limit the maximum displacement, maximum speed, and maximum acceleration of a single robot action.

[0009] Preferably, the generation of the robot health state vector in S3 specifically includes: Collect samples of the robot's continuous trouble-free operating time. and fault repair time sample , i =1,2,…, n ,in n The number of samples required to ensure statistical accuracy; Calculate Mean Time Between Failures and average repair time : ; ; The entire timeline is divided into multiple time periods, and the mean time between failures (MTBF) is calculated for each time period. and average repair time ; monitor and The changing trend, if Decrease over time or If performance degradation occurs over time, the robot is considered to be in a normal state; otherwise, it is considered to be in a normal state. A degradation status identifier is generated based on the degradation determination result: if the degradation determination result is performance degradation, the degradation status identifier is 1; if it is normal, the degradation status identifier is 0. according to The magnitude of the decline and The rate of increase is used to calculate the quantified value of the degree of degradation; Calculate the coefficient of variation of working time and downtime. and : ; ; ; ; The health status vector consists of the degradation status identifier, the degradation degree quantification value, and the coefficient of variation.

[0010] Preferably, the state space of the Markov decision model described in step S4 The joint action space is composed of collaborative perception features and health state vectors. From human action space and robot motion space The robot's action space is constrained by interaction rules; The collaborative strategy is a policy network. Observe the current state at each step. And select joint actions To maximize expected cumulative discount return To optimize the objective, among which As a discount factor, This is the reward function.

[0011] Preferred reward function Represented as: ; in, As a reward for task completion, Let health status vector be... For safety penalty function, The penalty intensity parameter; the safety penalty function is in Indicates a degenerate state and A positive value is taken if the action exceeds the preset conservative action space; otherwise, a zero value is taken.

[0012] Preferably, the collaborative efficiency indicators of the performance evaluation system in step S5 include task completion time, task completion rate, and resource utilization rate; safety indicators include the number of human-machine collisions, the number of safety distance violations, and the number of emergency stop triggers; human-machine ergonomic indicators include operator workload, operational comfort, and cognitive load; and robot reliability and stability indicators include the time-series variance of mean time between failures. Temporal variance of mean repair time The range of fluctuations in the coefficients of variation of working time and downtime.

[0013] Preferably, the quantitative evaluation in step S5 adopts the fuzzy comprehensive evaluation method, specifically including: Determine the set of evaluation indicators and a collection of comments Corresponding scoring criteria ; Constructing a fuzzy comprehensive judgment matrix ,in This represents the membership degree of the i-th evaluation index to the j-th comment level; Calculate the weight vector ; Calculate the comprehensive evaluation vector and overall score .

[0014] A multi-agent reinforcement learning-based human-machine-object collaborative perception and decision optimization system, employing a multi-agent reinforcement learning-based human-machine-object collaborative perception and decision optimization method, includes: The multi-source sensing module is used to acquire multi-source heterogeneous sensing information from human operation terminals, robot controllers, equipment sensors, and environmental monitoring devices. The preprocessing and fusion module is connected to the multi-source sensing module and is used to preprocess and fuse the multi-source heterogeneous sensing information to output collaborative sensing features. The interaction modeling module is used to establish the topology of the human-computer system and define the rules for human-computer collaborative interaction. A health status identification module, connected to the interactive modeling module, is used to generate a health status vector based on robot operation data. The collaborative decision-making module is connected to the preprocessing and fusion module, the health status identification module and the interaction modeling module respectively. It is used to output human-machine joint actions under the constraints of the interaction rules, taking the collaborative perception features and the health status vector as state inputs. The performance evaluation and feedback module, connected to the collaborative decision-making module, is used to evaluate performance based on the execution feedback of the human-machine joint action and generate feedback signals to optimize the decision parameters of the collaborative decision-making module.

[0015] Therefore, the above-mentioned human-machine-object collaborative perception and decision optimization method and system using multi-agent reinforcement learning has the following beneficial effects: (1) This application uses the robot health state vector as the state input of the collaborative decision-making model and introduces a safety penalty term based on the health state vector in the reward function, so that the decision-making system can actively perceive performance degradation and adaptively adjust the strategy, which solves the problem of the separation between reliability information and collaborative decision-making process, and effectively avoids safety accidents and production interruptions caused by sudden failures.

[0016] (2) This application constructs an effectiveness evaluation system that includes collaboration efficiency, safety, human-machine ergonomics and robot reliability and stability. The evaluation feedback signal is used to update the collaboration strategy and adjust the intensity of safety penalty items, forming a closed-loop iterative mechanism of perception-decision-execution-evaluation-optimization. This overcomes the defect that the evaluation results cannot be effectively fed back to the decision model, and realizes the continuous optimization of the collaboration strategy.

[0017] (3) This application limits the action boundaries of the joint action space through interaction rules, and dynamically adjusts the shrinkage range and execution rate of the action space according to the health status vector to achieve a dynamic balance between safety and task efficiency, thus solving the contradiction that safety and efficiency are difficult to balance in the existing human-machine collaboration system.

[0018] (4) This application introduces the coefficient of variation as a reliability and stability index in the health state vector, which describes the robot's health state from two dimensions: mean and dispersion, providing richer equipment status information for collaborative decision-making and improving the decision-making system's perception accuracy and response sensitivity to performance degradation.

[0019] The technical solution of this application will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0020] Figure 1This is a flowchart of a human-machine-object collaborative perception and decision optimization method based on multi-agent reinforcement learning in this application; Figure 2 This is a flowchart illustrating the health status vector generation process in this application embodiment; Figure 3 This is a system framework diagram in an embodiment of this application. Detailed Implementation

[0021] The following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0022] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning as understood by a person of ordinary skill in the art to which this application pertains.

[0023] The terms "comprising" or "including," as used in this application, mean that the element preceding the term encompasses the element listed after it, and do not exclude the possibility of encompassing other elements as well. The terms "inner," "outer," "upper," and "lower," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application. When the absolute position of the described object changes, the relative positional relationship may also change accordingly. In this application, unless otherwise expressly specified and limited, the term "attached," etc., should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral part; it can refer to a direct connection or an indirect connection through an intermediate medium; it can refer to the internal communication of two elements or the interaction relationship between two elements. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0024] Example 1: A multi-agent reinforcement learning-based human-machine-object collaborative perception and decision optimization method, such as... Figure 1 As shown, it includes the following steps: S1: Acquire multi-source heterogeneous sensing information from personnel operation terminals, robot controllers, equipment sensors and environmental monitoring devices, preprocess it, and fuse the preprocessed multi-source features to obtain collaborative sensing features; Preprocessing includes temporal alignment, anomaly removal, and feature extraction of the multi-source heterogeneous sensing information.

[0025] The fusion adopts a weighted fusion strategy, which specifically includes: The fusion weights are determined based on the data quality, signal-to-noise ratio, or task relevance of each modality. The preprocessed multi-source features are then weighted and summed according to the fusion weights to obtain the collaborative perception features.

[0026] Specifically, the multi-source heterogeneous sensing information includes operator posture and command signals collected by the human operating terminal, joint angle and torque data output by the robot controller, force and visual information acquired by equipment sensors, and environmental physical parameters recorded by environmental monitoring devices. Since the above sensing information differs in time reference, data format, and reliability, preprocessing is required. In this embodiment, an interpolation synchronization method is used to align the multi-source data in time. Specifically, using the highest sampling frequency in the modality to be fused as a reference, cubic spline interpolation upsampling is performed on low-frequency data, and sliding window averaging downsampling is performed on high-frequency data. Median filtering and box plot methods are used for anomaly removal. Numerical feature vectors are extracted for different modal data types. For the image modality, a lightweight convolutional neural network is used to extract depth features. For the force modality, a six-dimensional force / torque signal is directly extracted, and its mean and peak values ​​within the sliding window are statistically analyzed. For the pose and joint temporal modality, end-effector position, Euler angles, and joint angles are extracted to form multi-dimensional kinematic features. Z-score standardization is then applied to the features of each modality.

[0027] For the K modal feature vectors obtained after preprocessing, each modal feature is first mapped to the same feature dimension using a projection matrix. Then, the fusion weights are determined based on the data quality, signal-to-noise ratio (SNR), or relevance to the current task for each modality. For example, data quality is comprehensively evaluated by the calibration accuracy and signal integrity of each modal sensor; the SNR is obtained by calculating the ratio of signal power to noise power; and task relevance is set according to the degree of dependence of each modality on the current task type. In industrial human-robot collaboration scenarios, when the robot is in high-speed motion, the joint timing modality contributes significantly to safety decisions, and its fusion weight is correspondingly increased. When the robot is performing precision assembly operations, the weight of the force sensing modality increases; and when the operator approaches the robot's workspace, the weights of the vision and proximity sensors increase. Finally, the modal feature vectors are weighted and summed according to the fusion weights to obtain a unified-dimensional collaborative perception feature vector, which serves as the state input for the subsequent collaborative decision-making model.

[0028] S2: Establish the topology of the human-machine system and define the rules for human-machine collaborative interaction; Human-computer collaborative interaction rules include: Human-robot space safety distance constraint: A minimum safety distance threshold is set based on the robot's working range, the end effector's motion trajectory, and the area where personnel are active; when the actual distance between the human and the robot is less than the threshold, a deceleration or stop command is triggered. Human-robot task timing constraints: Define the order, parallel conditions, and switching logic between human-operated tasks and robot-executed tasks to avoid conflicting tasks between humans and robots in the same time and space region; Robot motion range constraints: Based on the current task type and robot health status, limit the maximum displacement, maximum speed, and maximum acceleration of a single robot action.

[0029] Specifically, the human-machine system topology refers to a structured model constructed to describe the system's components and their interrelationships based on the physical location relationships, information interaction methods, and task collaboration dependencies among personnel, robots, and surrounding environmental elements in an actual production scenario. The construction of the human-machine system topology includes: first, determining the system boundary and incorporating personnel, robots, workpieces, tooling equipment, and environmental monitoring units participating in the collaborative task into a system node set; then, establishing edge connections based on the physical proximity, communication connections, and task dependencies between nodes, where the edge direction represents the direction of information or material flow, and the edge weight represents the interaction strength or distance; finally, storing the above node and edge relationships in a graph data structure as a system topology graph, where the topology graph type includes at least one of serial topology, parallel topology, master-slave topology, and distributed collaborative topology. In a simple collaborative scenario with a single robot and a single personnel, the system topology adopts a serial topology structure, with information transmitted unidirectionally along "personnel → sensor → robot," and the robot performs corresponding avoidance or collaborative actions based on the perceived personnel status information. In complex assembly lines with multiple workers and robots, the system topology adopts a distributed collaborative topology. Robots exchange their position information and task execution status in real time via industrial Ethernet, while personnel broadcast their pose information through wearable devices. The state information of all nodes converges to a central decision-making unit, which performs unified state fusion and task allocation. After the human-machine system topology is established, the information interaction relationships between nodes within the system are determined based on the topology, providing data association for time alignment, anomaly removal, and feature extraction in subsequent multi-source perception information fusion steps. Simultaneously, the topology also limits the state space dimension of the Markov decision model in step S4, including only the perception information of nodes connected by edges in the topology, reducing interference from irrelevant data and lowering the computational complexity of the decision model.

[0030] S3: Collect samples of continuous fault-free working time and fault repair time of the robot, calculate the mean fault-free time, mean repair time and their temporal variation trends; determine whether the robot has experienced performance degradation based on the temporal variation trends, calculate the coefficient of variation of working time and fault time, and generate a robot health status vector; the health status vector includes a degradation status identifier, a quantified value of the degradation degree, and the coefficient of variation of working time and fault time. like Figure 2As shown, this is a sample of the robot's continuous fault-free operating time. and fault repair time sample , i =1,2,…, n ,in n The number of samples required to ensure statistical accuracy; Calculate Mean Time Between Failures and average repair time : ; ; The entire timeline is divided into multiple time periods, and the mean time between failures (MTBF) is calculated for each time period. and average repair time ; monitor and The changing trend, if Decrease over time or If performance degradation occurs over time, the robot is considered to be in a normal state; otherwise, it is considered to be in a normal state. A degradation status identifier is generated based on the degradation determination result: if the degradation determination result is performance degradation, the degradation status identifier is 1; if it is normal, the degradation status identifier is 0. according to The magnitude of the decline and The rate of increase is used to calculate the quantified value of the degree of degradation; Calculate the coefficient of variation of working time and downtime. and : ; ; ; ; The health status vector consists of the degradation status identifier, the degradation degree quantification value, and the coefficient of variation.

[0031] S4: Construct a Markov decision model for human-machine-object collaborative decision-making, using collaborative perception features and health state vectors as state inputs; define the action boundaries of the joint action space according to interaction rules; define a reward function, including task completion reward and safety penalty based on health state vector; optimize the collaborative strategy according to the reward function, and output human-machine joint actions that conform to the interaction rule constraints. State space of Markov decision models The joint action space is composed of collaborative perception features and health state vectors. From human action space and robot motion space The robot's action space is constrained by interaction rules; The collaborative strategy is a policy network. Observe the current state at each step. And select joint actions To maximize expected cumulative discount return To optimize the objective, among which As a discount factor, This is the reward function.

[0032] reward function Represented as: ; in, As a reward for task completion, Let health status vector be used. For safety penalty function, The penalty intensity parameter; the safety penalty function is in Indicates a degenerate state and A positive value is taken if the action exceeds the preset conservative action space; otherwise, a zero value is taken.

[0033] Specifically, the reward function consists of two parts: task completion reward and safety penalty. The task completion reward provides positive incentives based on task progress, operational accuracy, and completion time. For example, a fixed positive reward is given for each percentage increase in task completion, an additional reward is given for completing the task ahead of schedule, and a negative reward is given for exceeding the time limit, thus guiding the policy network to prioritize collaborative tasks. The safety penalty is designed based on the degradation state indicator and degradation degree quantification value in the health state vector. When the health state vector indicates that the robot is in a performance degradation state, and the robot's motion amplitude, speed, or acceleration output by the policy network exceeds the conservative motion space boundary of the degradation state, the safety penalty function takes a positive value and applies a penalty to the total reward; otherwise, it takes zero. The penalty strength parameter is used to balance task execution efficiency and safety. The more severe the degradation, the greater the penalty strength, and the more the policy network tends to output conservative actions, thereby achieving the task objective while ensuring safety. By maximizing the cumulative discounted reward that includes the reward function, the policy network gradually converges to a collaborative decision-making strategy that meets the constraints of interaction rules and human-machine safety requirements during the training process. This enables the robot to maintain efficient operation when it is in a healthy state and to proactively adjust its action strategy to avoid safety risks when its performance degrades.

[0034] S5: Collect state data after performing human-machine collaborative actions as execution feedback; construct a human-machine collaboration efficiency evaluation system, the criteria layer of which includes collaboration efficiency, safety, human-machine ergonomics, and robot reliability and stability; perform quantitative evaluation based on execution feedback and generate evaluation feedback signals; The performance evaluation system includes collaborative efficiency indicators such as task completion time, task completion rate, and resource utilization rate; safety indicators such as the number of human-robot collisions, the number of violations of safe distances, and the number of emergency stop triggers; human-robot ergonomic indicators such as operator workload, operational comfort, and cognitive load; and robot reliability and stability indicators such as the time-series variance of mean time between failures. Temporal variance of mean repair time The range of fluctuations in the coefficients of variation of working time and downtime.

[0035] The quantitative evaluation adopts the fuzzy comprehensive evaluation method, which specifically includes: Determine the set of evaluation indicators and a collection of comments Corresponding scoring criteria ; Constructing a fuzzy comprehensive judgment matrix ,in This represents the membership degree of the i-th evaluation index to the j-th comment level; Calculate the weight vector ; Calculate the comprehensive evaluation vector and overall score .

[0036] S6: Update the value estimate of the collaborative strategy based on the evaluation feedback signal, and adjust the penalty intensity parameter of the security penalty item to achieve continuous optimization of the strategy.

[0037] Specifically, continuous policy optimization is achieved through a dual-track mechanism of value estimation updates and penalty strength parameter adjustments. Value estimation updates use the evaluation feedback signal as a correction term for temporal difference errors to softly update the value network parameters of the policy network. Simultaneously, performance evaluation signals are introduced as additional supervision. When security or reliability stability falls below a threshold, the deviation between the evaluation feedback and the state value is added as an auxiliary loss term to the training objective, enabling value estimation to synchronously reflect the comprehensive evaluation of decision quality by performance assessment.

[0038] The temporal changes in the security score in the evaluation feedback are used as a regulating factor for updating the policy network parameters. When the security score continuously decreases, the update step size is increased, forcing the policy to adjust rapidly to avoid worsening security risks. When the security score stabilizes or increases, the normal update step size is restored to ensure stable policy convergence. The penalty intensity parameter is dynamically adjusted based on the security score and reliability stability score: the lower the score and the worse the stability, the greater the penalty intensity, causing subsequent decisions to impose greater penalties on high-risk actions. After the equipment reliability recovers, the penalty intensity is gradually reduced to restore the pursuit of task efficiency.

[0039] Through this dual-track mechanism, the evaluation feedback signal forms a closed loop with the value estimation and penalty intensity parameters, enabling the continuous optimization of the collaborative strategy to be driven by both the reward function and performance evaluation. As the human-machine collaborative task continues to be executed and the evaluation feedback signal accumulates, the value estimation becomes more accurate, the penalty intensity parameter better matches the environment, and the decision-making performance continuously improves.

[0040] Example 2: A multi-agent reinforcement learning-based human-machine-object collaborative perception and decision-making optimization system, such as Figure 3 As shown, a human-machine-object collaborative perception and decision optimization method using multi-agent reinforcement learning is employed, comprising: The multi-source sensing module is used to acquire multi-source heterogeneous sensing information from human operation terminals, robot controllers, equipment sensors, and environmental monitoring devices. The preprocessing and fusion module is connected to the multi-source sensing module and is used to preprocess and fuse the multi-source heterogeneous sensing information to output collaborative sensing features. The interaction modeling module is used to establish the topology of the human-computer system and define the rules for human-computer collaborative interaction. A health status identification module, connected to the interactive modeling module, is used to generate a health status vector based on robot operation data. The collaborative decision-making module is connected to the preprocessing and fusion module, the health status identification module and the interaction modeling module respectively. It is used to output human-machine joint actions under the constraints of the interaction rules, taking the collaborative perception features and the health status vector as state inputs. The performance evaluation and feedback module, connected to the collaborative decision-making module, is used to evaluate performance based on the execution feedback of the human-machine joint action and generate feedback signals to optimize the decision parameters of the collaborative decision-making module.

[0041] Therefore, this application adopts the above-mentioned multi-agent reinforcement learning human-machine-object collaborative perception and decision optimization method and system. By using the robot's health state vector as the input of the decision model and simultaneously as the basis for the safety penalty term, this application enables the collaborative strategy to actively perceive and adaptively respond to robot performance degradation, solving the problem of the separation between reliability information and collaborative decision-making. At the same time, the strategy is continuously optimized through performance evaluation feedback, forming a complete closed loop of "perception-decision-execution-evaluation-optimization", which improves the safety and adaptability of the human-machine collaborative system.

[0042] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and not to limit them. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of this application, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of this application.

Claims

1. A multi-agent reinforcement learning method for human-machine-object collaborative perception and decision optimization, characterized in that, Includes the following steps: S1: Acquire multi-source heterogeneous sensing information from personnel operation terminals, robot controllers, equipment sensors and environmental monitoring devices, preprocess it, and fuse the preprocessed multi-source features to obtain collaborative sensing features; S2: Establish the topology of the human-machine system and define the rules for human-machine collaborative interaction; S3: Collect samples of continuous fault-free working time and fault repair time of the robot, and calculate the average fault-free time, average repair time and their time-series variation trends. Based on the time-series change trend, determine whether the robot has experienced performance degradation, calculate the coefficient of variation of working time and failure time, and generate a robot health state vector; the health state vector includes a degradation state identifier, a quantified value of the degradation degree, and the coefficients of variation of working time and failure time; S4: Construct a Markov decision model for human-machine-object collaborative decision-making, using collaborative perception features and health state vectors as state inputs; define the action boundaries of the joint action space according to interaction rules; define a reward function, including task completion reward and safety penalty based on health state vector; optimize the collaborative strategy according to the reward function, and output human-machine joint actions that conform to the interaction rule constraints. S5: Collect state data after performing human-machine collaborative actions as execution feedback; construct a human-machine collaboration efficiency evaluation system, the criteria layer of which includes collaboration efficiency, safety, human-machine ergonomics, and robot reliability and stability; Quantitative evaluation is conducted based on execution feedback to generate assessment feedback signals; S6: Update the value estimate of the collaborative strategy based on the evaluation feedback signal, and adjust the penalty intensity parameter of the security penalty item to achieve continuous optimization of the strategy.

2. The human-machine-object collaborative perception and decision optimization method based on multi-agent reinforcement learning as described in claim 1, characterized in that, The preprocessing in step S1 includes time alignment, anomaly removal, and feature extraction of the multi-source heterogeneous sensing information.

3. The human-machine-object collaborative perception and decision optimization method based on multi-agent reinforcement learning as described in claim 1, characterized in that, The fusion strategy in S1 adopts a weighted fusion strategy, specifically including: The fusion weights are determined based on the data quality, signal-to-noise ratio, or task relevance of each modality. The preprocessed multi-source features are then weighted and summed according to the fusion weights to obtain the collaborative perception features.

4. The human-machine-object collaborative perception and decision optimization method based on multi-agent reinforcement learning as described in claim 1, characterized in that, The human-computer collaborative interaction rules mentioned in step S2 include: Human-robot space safety distance constraint: A minimum safety distance threshold is set based on the robot's working range, the end effector's motion trajectory, and the area where personnel are active; when the actual distance between the human and the robot is less than the threshold, a deceleration or stop command is triggered. Human-robot task timing constraints: Define the order, parallel conditions, and switching logic between human-operated tasks and robot-executed tasks to avoid conflicting tasks between humans and robots in the same time and space region; Robot motion range constraints: Based on the current task type and robot health status, limit the maximum displacement, maximum speed, and maximum acceleration of a single robot action.

5. The human-machine-object collaborative perception and decision optimization method based on multi-agent reinforcement learning as described in claim 1, characterized in that, The specific steps involved in generating the robot's health state vector in S3 are: Collect samples of the robot's continuous trouble-free operating time. and fault repair time sample , i =1,2,…, n ,in n The number of samples required to ensure statistical accuracy; Calculate Mean Time Between Failures and average repair time : ; ; The entire timeline is divided into multiple time periods, and the mean time between failures (MTBF) is calculated for each time period. and average repair time ; monitor and The changing trend, if Decrease over time or If performance degradation occurs over time, the robot is considered to be in a normal state; otherwise, it is considered to be in a normal state. A degradation status identifier is generated based on the degradation determination result: if the degradation determination result is performance degradation, the degradation status identifier is 1; if it is normal, the degradation status identifier is 0. according to The magnitude of the decline and The rate of increase is used to calculate the quantified value of the degree of degradation; Calculate the coefficient of variation of working time and downtime. and : ; ; ; ; The health status vector consists of the degradation status identifier, the degradation degree quantification value, and the coefficient of variation.

6. The human-machine-object collaborative perception and decision optimization method based on multi-agent reinforcement learning as described in claim 1, characterized in that, The state space of the Markov decision model described in step S4 The joint action space is composed of collaboratively perceived features and the health state vector. From human action space and robot motion space The robot's action space is constrained by interaction rules; The collaborative strategy is a policy network. Observe the current state at each step. And select joint actions To maximize expected cumulative discount return To optimize the objective, among which As a discount factor, This is the reward function.

7. The human-machine-object collaborative perception and decision optimization method based on multi-agent reinforcement learning as described in claim 6, characterized in that, reward function Represented as: ; in, As a reward for task completion, Let health status vector be used. For safety penalty function, The penalty intensity parameter; the safety penalty function is in Indicates a degenerate state and A positive value is taken if the action exceeds the preset conservative action space; otherwise, a zero value is taken.

8. The human-machine-object collaborative perception and decision optimization method based on multi-agent reinforcement learning as described in claim 1, characterized in that, In step S5, the efficiency evaluation system includes collaborative efficiency indicators such as task completion time, task completion rate, and resource utilization rate; safety indicators such as the number of human-robot collisions, the number of safety distance violations, and the number of emergency stop triggers; human-robot ergonomic indicators such as operator workload, operational comfort, and cognitive load; and robot reliability and stability indicators such as the time-series variance of mean time between failures. Temporal variance of mean repair time The range of fluctuations in the coefficients of variation of working time and downtime.

9. The human-machine-object collaborative perception and decision optimization method based on multi-agent reinforcement learning as described in claim 1, characterized in that, Step S5 uses the fuzzy comprehensive evaluation method for quantitative evaluation, which specifically includes: Determine the set of evaluation indicators and a collection of comments Corresponding scoring criteria ; Constructing a fuzzy comprehensive judgment matrix ,in This represents the membership degree of the i-th evaluation index to the j-th comment level; Calculate the weight vector ; Calculate the comprehensive evaluation vector and overall score .

10. A multi-agent reinforcement learning human-machine-object collaborative perception and decision optimization system, employing the multi-agent reinforcement learning human-machine-object collaborative perception and decision optimization method as described in any one of claims 1-9, characterized in that, include: The multi-source sensing module is used to acquire multi-source heterogeneous sensing information from human operation terminals, robot controllers, equipment sensors, and environmental monitoring devices. The preprocessing and fusion module is connected to the multi-source sensing module and is used to preprocess and fuse the multi-source heterogeneous sensing information to output collaborative sensing features. The interaction modeling module is used to establish the topology of the human-computer system and define the rules for human-computer collaborative interaction. A health status identification module, connected to the interactive modeling module, is used to generate a health status vector based on robot operation data. The collaborative decision-making module is connected to the preprocessing and fusion module, the health status identification module and the interaction modeling module respectively. It is used to output human-machine joint actions under the constraints of the interaction rules, taking the collaborative perception features and the health status vector as state inputs. The performance evaluation and feedback module, connected to the collaborative decision-making module, is used to evaluate performance based on the execution feedback of the human-machine joint action and generate feedback signals to optimize the decision parameters of the collaborative decision-making module.