Man-machine collaborative assembly task allocation and scheduling method integrating operation experience and efficiency
By adopting human-machine collaborative assembly task allocation and scheduling methods during the integrated processing rack assembly process, and using improved genetic algorithms and reinforcement learning to optimize task sequences, the problem of excessive accuracy deviation and repetitive operations in traditional assembly methods is solved, and assembly efficiency and quality are improved.
Patent Information
- Application Number
- CN202510104941.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-27
AI Technical Summary
The traditional assembly method has problems such as accuracy deviation, excessive repetitive operations, and worker fatigue during the assembly process of the comprehensive processing rack, which makes it difficult to ensure assembly quality.
A method of human-machine collaborative assembly task allocation and scheduling that integrates business experience and efficiency is proposed. By considering the human-machine trust relationship, the constraints and mathematical models of task allocation and scheduling are constructed, and the task sequence optimization is achieved using improved genetic algorithms and reinforcement learning.
It improves assembly efficiency and quality, reduces worker fatigue, enhances the smoothness and operation efficiency of human-machine collaboration, and ensures accuracy and consistency during the assembly process.
Smart Images

Figure CN120216162A_ABST
Abstract
Description
Technical Field
[0001] The present invention proposes a method for human - machine collaborative assembly task allocation and scheduling that integrates operation experience and efficiency, belonging to the field of human - machine collaborative assembly, and particularly involves the human - machine collaborative assembly task allocation and scheduling process of an integrated processing rack for an array antenna, and proposes an innovative method considering the human - machine trust relationship. Background Technique
[0002] With the progress of the times, the European Commission has put forward the idea of Industry 5.0, whose core idea is "people - oriented", that is, while improving productivity and production efficiency, the main position of people in the manufacturing industry should be considered to ensure the interests of workers. The integrated processing rack includes not only various rigid components but also flexible components such as cables. Its assembly process shows characteristics such as multi - variety, variable batch, fast update and iteration, and strong demand fluctuations. Traditional assembly methods mainly rely on manual experience operation. The "blind plugging, blind installation, and blind adjustment" assembly method makes it difficult for workers to correct errors in real - time during the assembly process and requires relying on experience to complete highly repetitive fine operations. Moreover, there are a large number of repetitive operations in the assembly process, which is a great test of the patience and energy of operators. This method makes it difficult to guarantee the assembly quality, is prone to precision deviation, increases the possibility of rework, and thus affects the consistency and reliability of products.
[0003] Human - machine collaboration is a new working mode. In the assembly process of the integrated processing rack, the accuracy and stability of the robot are used to share some repetitive, heavy - duty and error - prone tasks. The high - precision automated wire - inserting device at the end of the robot can quickly and accurately complete wire docking according to the preset program, effectively avoiding the occurrence of wrong insertion; the intelligent tightening device at the end of the robot can tighten the screws strictly according to the standard torque to ensure the consistency of assembly quality. Workers can be freed from high - intensity repetitive labor and devote more energy to more creative and flexible work links such as the optimization of assembly processes, quality control, and handling emergencies, thereby improving the efficiency and quality of the overall assembly operation.
[0004] In the process of human-machine collaboration, the allocation and scheduling of assembly tasks are very important links. Reasonable task allocation can give full play to the respective advantages of humans and machines. Humans possess flexibility, creativity, and the ability to perceive and respond to complex situations, while machines excel in precision, speed, and repetitive operations. For some high-precision component docking and fastening tasks, they can be handed over to automated assembly robots, which can ensure the accurate tightening torque of each screw with extremely high precision and can quickly repeat operations, greatly improving the assembly efficiency. For those operations that require flexible adjustment according to the overall layout and have high flexibility requirements, they are completed by human assembly workers with their experience and flexible operation ability. In terms of scheduling, it is necessary to consider the constraints between humans and machines, the states of workers and robots, and combine multiple optimization goals to achieve the task sequence with the shortest overall operation time and the highest human-machine collaboration efficiency.
[0005] The human operation experience is a key factor in the human-machine collaboration process, which is reflected in the physiological and psychological cognitive levels. Physiological feelings include the comfort and fatigue of the body. A comfortable body posture can reduce the discomfort of the operator during work, enabling them to focus more on the task and directly improving the operation efficiency. Reasonably controlling the fatigue level can avoid efficiency losses caused by decreased energy, and at the same time, it also helps to maintain a good operation experience. On the psychological cognitive level, the human-machine trust relationship occupies the dominant factor. The trust relationship originally represents a bond based on reliability cognition and psychological expectation, which is established on the judgment of the other party's ability, intention, and behavior stability. When workers trust the task completion accuracy, operation precision, and safety performance of the robot, they will rest assured to hand over some repetitive tasks to it and achieve close cooperation with the robot. They can then focus more energy on the links that require human unique thinking and creativity, such as the optimization design of complex task processes and flexible decision-making when dealing with emergencies. This trust also enables the robot to perform tasks without excessive interference from workers during the process. A stable human-machine trust relationship can enable operators to cooperate with machines more confidently, reduce psychological pressure, avoid chaos and delays that may be caused by frequent human adjustments, and thus improve the smoothness and operation efficiency of the entire cooperation process. Summary of the Invention
[0006] To solve the problems mentioned above, a human-machine collaborative assembly task allocation and scheduling method that integrates operation experience and efficiency is proposed. The purpose of this method is to generate task allocation and scheduling sequences under multi-objective optimization conditions. First, the process influencing factors of human-machine task allocation and scheduling in the assembly process of the integrated processing rack are studied; secondly, key factors related to operation experience and efficiency are introduced into the influencing factors, and the constraint conditions and mathematical models of task allocation and scheduling are jointly constructed; finally, the algorithm process and framework of human-machine task allocation and scheduling are realized through improved genetic algorithms and reinforcement learning.
[0007] General Overview of the Method: The method studies the human-machine task allocation and scheduling process from two aspects: constraint modeling and allocation scheduling algorithm, including two parts: task allocation and scheduling constraint modeling considering human-machine trust relationship, and allocation and scheduling combining improved genetic algorithm and reinforcement learning. Human-machine task allocation and scheduling need to meet constraints from multiple aspects, which can be classified into mandatory constraints and optimization constraints. Mandatory constraints are those that must be complied with during the allocation and scheduling process, such as the precedence relationship between tasks; job experience and efficiency are incorporated into the optimization constraints to measure the compliance with the constraints through metrics such as scoring, so as to achieve the optimization goal, such as human body posture comfort, human body fatigue degree, and human-machine trust relationship. After constructing the constraint relationship, the mandatory constraints are evaluated through the improved genetic algorithm to generate a preliminary sequence of human-machine task allocation; then a reinforcement learning model is constructed to learn and optimize according to the optimization constraints to achieve the best task allocation and scheduling sequence.
[0008] To illustrate this method, it is necessary to explain in detail the human-machine task allocation and scheduling and the modeling process of mandatory constraints and optimization constraints.
[0009] Furthermore, human-machine task allocation and scheduling include two aspects: task allocation and task scheduling. Task allocation refers to the process of allocating different tasks to designated workers or robots during the human-machine collaboration process, and constraints such as the capabilities and preferences of workers and robots need to be fully considered; task scheduling refers to the process of arranging and coordinating the tasks completed above in terms of time and resource utilization, and constraints such as task execution time and task precedence relationship need to be fully considered.
[0010] Furthermore, mandatory constraints refer to those binding conditions that are inviolable and must be strictly followed throughout the process of carrying out task allocation and scheduling work. In the present invention, they are summarized into three constraints: Mandatory constraint conditions mainly include task priority constraint, task end matching constraint, and resource occupancy constraint. Task priority constraint means that different assembly tasks have a precedence relationship, and a certain task must be executed after a designated task is completed. Therefore, there cannot be a situation where a subsequent task is executed before a preceding task in the task sequence. The task priority constraint can be described by a directed acyclic graph; the task end matching constraint means that for a robot, different tasks require different execution ends. By evaluating the end functions of the robot, it is necessary to ensure that the tasks assigned to the robot can be executed by the robot; the resource occupancy constraint refers to stipulating that within the execution time interval of the task the same worker / robot cannot perform multiple tasks at the same time. The above constraints are mandatory constraints during the task execution process, and the task allocation and scheduling sequence cannot violate the above constraints.
[0011] Furthermore, the optimization constraints refer to specific constraint conditions that can comprehensively weigh and consider factors such as the importance of indicators and relevant scoring criteria during the task assignment and scheduling process, and then optimize and adjust the task assignment and scheduling plan. In the present invention, it is summarized as the total assembly process time constraint related to operation efficiency, the total idle time constraint of workers and robots, as well as the human body posture comfort constraint, the human cognitive fatigue degree constraint, and the human-machine trust relationship constraint related to operation experience. The total task completion time constraint is evaluated by counting the total duration consumed during the assembly task process; the total idle time constraint of workers and robots is evaluated by counting the total idle time of workers and robots. The human body posture comfort constraint refers to evaluating the human body posture during the task execution process and scoring according to the comfort of the posture to form a comfort score constraint; the human cognitive fatigue degree constraint mainly obtains the current fatigue state of the worker by collecting brain wave data to form a fatigue degree score; the human-machine trust relationship constraint mainly refers to the degree of trust of the worker in the robot. To ensure the smooth completion of the task, the worker needs to evaluate the working ability of the robot to assign appropriate tasks to the robot. During the scheduling process, the worker needs to trust the robot to complete the task completely, so as to ensure the seamless execution of the task; in addition, the safety of the robot also needs to be trusted to facilitate the close cooperation between the worker and the robot; combining all the above conditions to evaluate the optimal sequence.
[0012] The present invention proposes the following five important steps for the human-machine collaborative assembly task assignment and scheduling that integrates operation experience and efficiency:
[0013] The first step: Modeling of mandatory constraint conditions. The mandatory constraint conditions represent the constraints that must be observed during the task assignment and scheduling process, so that a large number of assignment and scheduling sequences can be reduced to partial sequences that meet the constraints. It mainly includes modeling of task priority constraints, task end matching constraints, and resource occupancy constraints.
[0014] The second step: Modeling of optimization constraint conditions. The optimization constraint conditions can achieve the optimal evaluation of the sequence according to the scores of each constraint to generate the optimal sequence under multiple objectives. It mainly includes the total task completion time related to operation efficiency, the total idle time of workers and robots, as well as the human body posture comfort constraint, the human cognitive fatigue degree constraint, and the human-machine trust relationship constraint related to operation experience.
[0015] The third step: Modeling of the task sequence objective function. The goal of human-machine task assignment and scheduling is to shorten the task execution time and reduce the idle time of workers and robots to improve efficiency. Therefore, the objective function includes three penalty terms: the total task completion time, the idle time of workers and robots, and the human body posture comfort, the human body fatigue degree, and the degree of trust of the human in the ability of the robot. Based on this objective function, the optimization of the task assignment and scheduling sequence is realized.
[0016] Step 4: Optimization of the forced constraint task sequence based on the improved genetic algorithm. According to the forced constraint conditions, the improved genetic algorithm is used to filter out the task sequences that do not meet the forced constraints, limit the solution set of the task sequences within a reasonable range, and reduce the difficulty of the subsequent sequence optimization process.
[0017] Step 5: Optimization of the optimized constraint task sequence based on reinforcement learning. According to the optimized constraint conditions and the objective function, the reinforcement learning algorithm is used to construct the environment and the learning mechanism to generate the task allocation and scheduling sequences of the optimal objectives under multiple constraints.
[0018] To implement the above steps, the present method includes the following four methods:
[0019] Method 1: Modeling of forced constraint conditions. The forced constraint conditions include task priority constraints, task end matching constraints, and resource occupancy constraints. The task priority constraints are described by a directed acyclic graph, which can be expressed as G=(V, E), where V={v1, v2,..., v n} is a non-empty finite set representing the nodes in the graph, and the actual meaning is subtasks; E={(v i , v j )|v i , v j ∈V, i≠j} is a set composed of ordered pairs, representing the directed edges in the graph from v i to v j , and the actual meaning is the precedence relationship between subtasks, thus constituting the task scheduling constraints. The task end matching constraints are evaluated by the functions of the robot end, and it is necessary to ensure that the tasks assigned to the robot can be executed by the robot. The resource occupancy constraints are restricted by specifying the task execution time interval and the working status of workers / robots. The task node v i has an execution time attribute, and the attribute is expressed as the task start time task completion time where the start time is related to the actual scheduling of the task, and the task completion time is different due to the task being assigned to a worker or a robot. It is necessary to ensure that the same worker / robot cannot perform multiple tasks at the same time, thus constituting the constraints of task allocation and scheduling.
[0020] Method 2: Optimize the modeling of constraint conditions. The optimized constraint conditions include the total task completion time constraint related to operation efficiency, the total idle time constraint of workers and robots, and the human body posture comfort constraint, human cognitive fatigue degree constraint, and human-machine trust relationship constraint related to operation experience. The total task completion time constraint is evaluated by statistically calculating the total duration consumed in the assembly task process; the total idle time constraint of workers and robots is evaluated by statistically calculating the idle time of workers and robots. The human body posture comfort constraint is evaluated by analyzing the human body posture, human joint angles, and the duration of the posture during the task execution process to further optimize the task allocation; the human cognitive fatigue degree constraint is evaluated by monitoring the brain waves of workers and classifying the fatigue degree. The fatigue degree of workers can be divided into four levels: non-fatigued, mildly fatigued, moderately fatigued, and severely fatigued, which can be expressed as FA = {A, B, C, D}. The task allocation is constrained according to the current fatigue degree; the human-machine trust relationship constraint is evaluated by the degree of trust of humans in the capabilities of robots. The capabilities of robots and the trust of workers can form a Bayesian network. The trust situation of workers in robots is the query variable Y, which has multiple child nodes. The nodes are multiple performance and safety indicators of the robot. The performance indicators include task completion accuracy, task completion time, and precision maintenance ability. The safety indicators include safety response mechanism time, historical safety accident frequency, and safety protection range. These indicators can be expressed as the evidence variable set X = {X1, X2,..., X i , …, X n}. For any indicator X i of a robot, its value range is:
[0021]
[0022] Furthermore, the posterior probability P(Y|X) of the query variable Y, that is, the degree of trust of workers in robots, with respect to the evidence variable X can be solved through Bayes' formula:
[0023]
[0024] Furthermore, first clarify the conditional distribution relationship between each variable, and then, according to the value X = {X1 = x1, X2 = x2,..., X n = x n} of the evidence variable X, find the joint distribution P(X, Y) of the X and Y variables, so as to obtain the posterior probability of the query variable Y under the evidence variable. The degree of trust of workers in robots is transformed according to the posterior probability interval. The optimized constraints for task allocation and scheduling are constructed through the above indicators.
[0025] Method 3: Modeling the objective function based on the optimal task sequence. Based on the total task completion time, the total idle time of workers and robots, and three penalty terms of human body posture comfort, human fatigue level, and human-machine trust level, establish the objective function for constructing the optimal task sequence of the task execution sequence:
[0026]
[0027] Among them, T total represents the total task completion time, and ω1 is its weight; T idle represents the total idle time, and ω2 is its weight; and represent the penalty terms of human body posture comfort, human fatigue level, and human-machine trust level respectively. C p represents the human body posture comfort score, C p0 is the human body posture comfort threshold, and λ1 is the weight of the human body posture comfort penalty term; F represents the worker fatigue level, F max is the worker fatigue level threshold, and λ2 is the weight of the human cognitive fatigue level penalty term; P d represents the human-machine trust level, P d0 is the human-machine trust level threshold, and λ3 is the weight of the human-machine trust relationship penalty term;
[0028] Method 4: Generating the forced constraint task assignment scheduling sequence based on the improved genetic algorithm. According to the above-mentioned forced constraint conditions, construct the corresponding improved genetic algorithm model, and output the initial task sequence set that meets the forced constraints. The settings of the genetic algorithm are as follows: First, initialize the task sequence population, which includes the task assignment situation and execution order. Judge whether the current sequence meets the forced constraints, and perform hybridization and mutation behaviors through chromosomes to gradually generate the task sequence set that meets the requirements.
[0029] Method 5: Optimization of the constrained task assignment and scheduling sequence based on reinforcement learning. Optimize the task assignment and scheduling sequence according to the optimization constraints and objective function. First, construct a reinforcement learning environment, abstract the scenarios of task assignment and scheduling into an environmental model with a specific state space, action space, and reward mechanism. The definition of the state space covers the current task queue situation, including the number of tasks, priorities, end requirements of tasks, and the current states of workers / robots, the postural comfort of workers, the cognitive fatigue level of workers, and the trust level of workers in robots. These state variables together form a multi-dimensional state vector; the action space determines the operations that the reinforcement learning agent can take, including the executor of each subtask, the execution order of tasks, and the task execution time; the reward mechanism is set according to the optimization constraints, with the total task completion time and the total idle time as the incentives in the reward mechanism; based on the postural comfort of the human body, the cognitive fatigue level of the human body, and the human-machine trust relationship as the penalties in the reward mechanism.
[0030] The technical solution of the present invention has the following advantages and benefits:
[0031] The present invention provides a task assignment and scheduling method for human-machine collaborative assembly of an integrated processing rack. Aiming at the technological characteristics of the integrated processing rack assembly process and the working characteristics of workers and robots, the present invention constructs comprehensive constraints combining mandatory constraints and optimization constraints, and respectively uses an improved genetic algorithm and reinforcement learning to initially assign and evaluate and optimize the task sequence to achieve the task assignment and planning sequence with the best indicators.
[0032] The present invention provides a human-machine collaborative assembly task assignment and scheduling method that integrates operation experience and efficiency, and extracts key elements related to operation experience in the cooperation process between humans and robots during the human-machine collaboration process, namely the postural comfort of the human body, the fatigue level of the human body, and the human-machine trust relationship. In particular, in terms of the human-machine trust relationship, the performance and safety indicators of the robot are combined. The performance indicators include the task completion accuracy rate, the task completion time, and the accuracy retention ability, and the safety indicators include the safety response mechanism time, the historical safety accident frequency, and the safety protection range. Obtain data through historical data, calibration data, etc., and evaluate each indicator according to the type and attributes of the current task. Through normalization, it can be transformed into a probability model; at the same time, construct a Bayesian network model of the human-machine trust relationship and multiple indicators, and solve the posterior trust probability of the worker in the robot during the execution of the current task through the above probability model data, and use it as a key indicator for task assignment and scheduling, laying a foundation for subsequent task execution and close cooperation.
[0033] The present invention provides a human-machine task allocation and scheduling method combining an improved genetic algorithm and reinforcement learning. First, the improved genetic algorithm is used to perform an initial planning of the task sequence for the mandatory constraints. Under the condition of satisfying the mandatory constraints, a large number of initial task sequences can be generated, thus greatly reducing the number of sequences for task allocation and scheduling. When optimizing the task sequence according to the optimization constraints subsequently, the construction difficulty of the reinforcement learning model can be significantly reduced, its convergence degree can be improved, and a task sequence that meets the constraint requirements and guarantees optimization constraints such as the shortest operation time and the highest operation efficiency can be generated. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is the overall flowchart of this method
[0035] Figure 2 is the flowchart for modeling mandatory constraints.
[0036] Figure 3 is the flowchart for modeling optimization constraints.
[0037] Figure 4 is the flowchart for the allocation and scheduling of the assembly task sequence. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0039] As Figure 1 shown, the overall process of this method includes multiple steps. First, based on the analysis of the assembly process of the integrated processing rack, the assembly tasks are decomposed into tasks at the action level; then, the mandatory constraints proposed above are modeled, and according to this mathematical model, the improved genetic algorithm is used for the allocation and scheduling of the task sequence, so that the output task sequences can all meet the mandatory constraint conditions, simplifying the subsequent optimization process. Secondly, the optimization constraints are modeled, mathematical constraint relationships and evaluation indicators are constructed, an objective function is established, and finally a reinforcement learning model is constructed to realize the evaluation of the sequence in combination with the optimization constraints.
[0040] As Figure 2As shown in the figure, the mandatory constraint modeling includes the following steps. The mandatory constraint includes three aspects of constraints. For the task priority constraint, first construct a directed acyclic graph of task priorities, and construct the description adjacency matrix of the graph, which is used as the input mathematical model; for the task end matching constraint, first classify the executable and non-executable tasks according to the robot end to form a mathematical model of task attributes; for the resource occupancy constraint, it is necessary to construct the execution time interval of the task and determine the working status of workers and robots at the current time node. Finally, these mathematical models are integrated to form a mathematical model of mandatory constraints.
[0041] As Figure 3 shown in the figure, the optimization constraint condition modeling includes the following steps. The optimization constraints are related to the total task completion time constraint and the total idle time constraint of workers and robots related to operation efficiency, as well as the human body posture comfort constraint, the human cognitive fatigue degree constraint, and the human-machine trust relationship constraint related to operation experience. For the total task completion time constraint, it is modeled by statistically calculating the total duration consumed in the assembly task process; for the total idle time constraint of workers and robots, it is modeled by statistically calculating the total idle time of workers and robots; for the human body posture comfort constraint modeling, first evaluate the human body posture comfort and joint angles according to each task situation, and finally form an index of human body posture comfort; for the human cognitive fatigue degree constraint, by constructing a brain wave-fatigue degree model, the brain wave data of workers are collected in real time to construct an index of worker fatigue degree; for the human-machine trust relationship constraint, it is necessary to consider the performance indicators and safety indicators of the robot: in terms of performance indicators, it is necessary to evaluate through parameters such as task completion accuracy, task completion time, and precision maintenance ability; in terms of safety indicators, it is necessary to evaluate through parameters such as safety response time, historical safety accident frequency, and safety protection range; after completing the parameter evaluation, a Bayesian network can be constructed to comprehensively evaluate the human-machine trust degree.
[0042] As Figure 4 shown in the figure, the sequence planning through the algorithm includes the following steps. Improve the construction of algorithm variables and models in the genetic algorithm, perform computational analysis to form a constrained task sequence; in the reinforcement learning algorithm, construct a reinforcement learning environment to form a reinforcement learning agent, and optimize the task sequence through rewards, punishments, etc., and finally form an optimized human-machine task allocation sequence.
[0043] The present invention proposes a human-machine collaborative assembly task allocation and scheduling method that integrates operation experience and efficiency. The core of this method is to construct mandatory constraints for the execution of assembly tasks, and through improving the genetic algorithm, allocate and schedule tasks to form a set of task sequences that meet the requirements; secondly, considering task experience and efficiency, construct an optimization constraint mathematical model, and finally realize the optimization of the task sequence through the reinforcement learning algorithm.
[0044] The above embodiments of the present invention mainly illustrate a human-machine collaborative assembly task allocation and scheduling method that integrates operation experience and efficiency. It is an example given to clearly illustrate the present invention and should not be construed as a limitation of the present invention. The present invention can be extended to other instances of human-machine collaborative task execution. However, those of ordinary skill in the art should understand that the present invention can be implemented in many other forms without departing from its gist and scope. Without departing from the spirit and scope of the present invention as defined by the appended claims, the present invention may cover various modifications and substitutions.
Claims
1. A method for allocating and scheduling human-machine collaborative assembly tasks that integrates work experience and efficiency, characterized in that For the human-machine collaborative assembly scenario, mandatory constraint modeling, optimization constraint modeling and task sequence objective function modeling are carried out, and the work experience is integrated into the optimization constraints to improve work efficiency; based on the established constraints and objective functions, the mandatory constraint task sequence optimization based on the improved genetic algorithm is first carried out, that is, the sequence of human-machine task allocation and scheduling is restricted, and the task sequences that do not meet the constraints are eliminated; then, the optimization constraint task sequence optimization based on reinforcement learning is further realized, that is, the corresponding reinforcement learning model algorithm is constructed to achieve the optimal solution of the task sequence, and finally the optimal task allocation and scheduling plan is output.
2. The method according to claim 1, characterized in that: The steps include: Step 1: Modeling mandatory constraints. The mandatory constraints refer to the constraints that must be followed in the task allocation and scheduling process. Task allocation and scheduling sequences that do not meet these constraints cannot be executed correctly. The mandatory constraints mainly include task priority constraints, task end matching constraints, and resource occupancy constraints. Step 2: Optimization constraint modeling. The optimization constraint refers to the condition that exists in the optimization situation during task allocation and scheduling. It is necessary to consider multiple factors related to work experience and efficiency. According to the value range and score of each constraint, the optimal evaluation of the sequence is achieved, and finally the optimal sequence generation under multiple constraints is achieved, which mainly includes the total task completion time related to work efficiency, the total idle time of the worker robot, and the human posture comfort constraint, human cognitive fatigue constraint and human-machine trust relationship constraint related to work experience. Step 3: Modeling the objective function of the task sequence. The goal of human-machine task allocation and scheduling is to improve work efficiency and enhance work experience. Improving work efficiency includes shortening the task completion time and reducing the total idle time of workers and robots. Enhancing work experience includes improving human posture comfort and human-machine trust relationship and reducing the negative impact of human cognitive fatigue. Therefore, the objective function includes three penalty items: total task completion time, idle time of workers and robots, human posture comfort, human fatigue and human trust in the robot's ability. Based on this objective function, the optimization of task allocation and scheduling sequence is achieved; Step 4: Optimize the mandatory constraint task sequence based on the improved genetic algorithm. According to the mandatory constraint conditions, the improved genetic algorithm is used to filter the task sequence that does not meet the mandatory constraint, limit the solution set of the task sequence to a reasonable range, and reduce the difficulty of executing the subsequent sequence optimization process; Step 5: Optimization of task sequence with optimization constraints based on reinforcement learning. According to the optimization constraints and objective function, the reinforcement learning algorithm is used to construct the reinforcement learning agent, environment and learning mechanism to generate the task allocation and scheduling sequence with the optimal goal under multiple constraints.
3. The method according to claim 2, characterized in that: The following methods are included: Method 1: Modeling of mandatory constraints. Mandatory constraints include task priority constraints, task end matching constraints, and resource occupancy constraints. Task priority constraints are described by a directed acyclic graph, which can be expressed as G = (V, E), where V = {v1, v2, ···, v n } is a non-empty finite set representing the nodes in the graph, which actually means subtasks; E = {(v i ,v j )|v i ,v j ∈V,i≠j} is a set of ordered pairs, representing the graph consisting of v i Point to v j The directed edge of the task actually means the priority relationship between subtasks, thus constituting the task scheduling constraint; the task end matching constraint evaluates the robot end function to ensure that the robot can execute the task assigned to the robot; the resource occupancy constraint is specified within the execution time interval of the task. The same worker / robot cannot perform multiple tasks at the same time; Method 2: Optimization constraint modeling. The optimization constraints include the total task completion time constraint and the total idle time constraint of workers and robots related to work efficiency, as well as the human posture comfort constraint, human cognitive fatigue constraint, and human-machine trust relationship constraint related to work experience. The total task completion time constraint is evaluated by counting the total time consumed by the assembly task process; the total idle time constraint of workers and robots is evaluated by counting the total idle time of workers and robots; the human posture comfort constraint is evaluated by evaluating the human posture, human joint angle, and posture duration during the task execution process to further optimize the task allocation. The human cognitive fatigue level constraint is achieved by monitoring the brain waves of workers and classifying their fatigue levels. The worker's fatigue level can be divided into four levels: no fatigue, mild fatigue, moderate fatigue, and severe fatigue, which can be expressed as FA = {A, B, C, D}. The task allocation is constrained according to the current fatigue level. The human-machine trust relationship constraint is evaluated by the degree of trust people have in the robot's capabilities. The robot's capabilities and the worker's trust can form a Bayesian network. The worker's trust in the robot is the query variable Y, which has multiple child nodes. The nodes are multiple performance and safety indicators of the robot. The performance indicators include task completion accuracy, task completion time, and accuracy retention ability. The safety indicators include safety response mechanism time, historical safety accident frequency, and safety protection range. These indicators can be expressed as the evidence variable set X = {X1, X2, ···, X i ,···,X n }, for any robot indicator X i , its value range is: Furthermore, the posterior probability P(Y|X) of the query variable Y, i.e., the worker's trust in the robot, relative to the evidence variable X can be solved using the Bayesian formula: Furthermore, we first clarify the conditional distribution relationship between the variables, and then according to the value of the evidence variable X, X = {X1 = x1, X2 = x 2, …,X n =x n } Obtain the joint distribution P(X,Y) of variables X and Y, and then obtain the posterior probability of the query variable Y under the evidence variable. According to the posterior probability interval, it is converted into the worker's trust in the robot. The optimization constraints of task allocation and scheduling are constructed through the above indicators. Method 3: Based on the optimal objective function modeling of the task sequence, the execution sequence of the task is established based on the total task completion time, the total idle time of the worker robot, and the three penalty items of human posture comfort, human fatigue level, and human-machine trust level to construct the optimal objective function of the task sequence: Among them, T total represents the total time to complete the task, ω1 is its weight; T idle represents the total idle time, ω2 is its weight; and They represent the penalty items of human posture comfort, human fatigue and human-machine trust, respectively. p Represents the human posture comfort score, C p0 is the threshold of human posture comfort, λ1 is the weight of the human posture comfort penalty term; F represents the worker’s fatigue level, F max is the worker fatigue threshold, λ2 is the weight of the human cognitive fatigue penalty term; P d Represents human-machine trust, P d0 is the human-machine trust threshold, λ3 is the weight of the human-machine trust relationship penalty term; Method 4: Generation of mandatory constraint task allocation scheduling sequence based on improved genetic algorithm. According to the mandatory constraint conditions mentioned above, the corresponding improved genetic algorithm model is constructed to output the initial task sequence set that meets the mandatory constraint. The genetic algorithm is set as follows: first, the task sequence population is initialized. The task sequence includes the task allocation and execution order. It is judged whether the current sequence meets the mandatory constraint. The chromosomes are hybridized and mutated to gradually generate a set of task sequences that meet the requirements. Method 5: Optimization of task allocation and scheduling sequence with optimization constraints based on reinforcement learning. According to the optimization constraints and objective functions, the task allocation and scheduling sequence are optimized. First, a reinforcement learning environment is constructed, and the scenario of task allocation and scheduling is abstracted into an environmental model with a specific state space, action space and reward mechanism. The definition of the state space covers the current situation of the task queue, including the number of tasks, priority, end-of-task requirements, and the current state of the worker / robot, the worker's posture comfort, the worker's fatigue level and the worker's trust in the robot. These state variables together constitute a multidimensional state vector; the action space determines the operations that the reinforcement learning agent can take, including the executor of each subtask, the execution order of tasks and the task execution time; the reward mechanism is set according to the optimization constraints, and the total task completion time and the total idle time are used as incentives in the reward mechanism; the human posture comfort, human cognitive fatigue level and human-machine trust relationship are used as penalties in the reward mechanism.
4. The method according to claim 3, characterized in that: The following steps are involved: Step 1: Combined with the human-machine collaboration process, mandatory constraints are summarized, including task priority constraints, task end matching constraints, and resource occupancy constraints. Under these constraints, task allocation and scheduling sequences can be correctly executed; Step 2: Combining the characteristics of workers and robots, we summarized the constraints of human posture comfort, human fatigue level, and human-machine trust relationship related to the work experience. Under these optimization constraints, we can generate a close and efficient human-machine collaboration sequence that satisfies the comfort of workers and realizes human-machine trust. Step 3: By improving the genetic algorithm and combining it with mandatory constraints to perform initial allocation and scheduling of task sequences, the order of magnitude of the task sequence solution set is greatly reduced, thus reducing the difficulty of subsequent task sequence optimization; Step 4: Through the reinforcement learning algorithm, construct the intelligent agent, environment and reward and punishment mechanism, combine the optimization constraints and objective functions, and realize the optimal task sequence solution under multi-constraint and multi-objective optimization, which can converge quickly and avoid local optimal solutions.