Multi-robot collaborative operation and task distribution system and method

Through the combination of auction algorithm and Q and Sarsa learning algorithms, the problem of inefficient task allocation in multi-robot systems is solved, dynamic optimization and environmental adaptation are achieved, and task completion quality and efficiency are improved.

CN120278433AInactive Publication Date: 2025-07-08SHENZHEN WANZHONG COM TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510329347.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In traditional multi-robot systems, task allocation is inefficient, lacks dynamic adaptability, and cannot cope with task and environmental changes. The Q-learning algorithm is prone to falling into local optimal solutions and lacks exploration.

Method used

The auction algorithm is used to select the robot in combination with the robot's work progress and spatial matching degree, and the Q learning algorithm is used to optimize the task execution order, and determine whether to switch to the Sarsa learning algorithm through correlation and environmental change rate, which improves the flexibility and efficiency of task allocation.

Benefits of technology

The task completion quality and efficiency of multi-robot systems are improved, and the optimization algorithm can be dynamically adjusted to cope with complex and dynamic environments, avoid local optimal solutions, and improve resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278433A_ABST
    Figure CN120278433A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-robot collaborative operation and task allocation system and method, relates to the technical field of robot operation, and is used for solving the problems that in a traditional method, task allocation efficiency is low, resources are wasted, especially in a multi-robot system, tasks cannot be accurately allocated, dynamic adaptability is lacked, task and environment changes cannot be coped, and the task allocation efficiency is low. An optimal decision strategy is obtained step by step; and a single learning algorithm has defects. An auction algorithm is used for solving the task allocation problem of the robot, the work progress and the space matching degree of the robot are considered, a Q learning algorithm is used for determining whether to switch to a Sarsa learning algorithm or not according to the task completion condition of the robot in combination with environment interaction and a continuous optimization strategy, and the task allocation efficiency is improved. The design of the flexible switching strategy makes up for the defects of single reinforcement learning strategy and lack of flexibility and adaptability in the prior art, so that the robot can dynamically adjust the optimization algorithm according to the actual situation, and the task completion quality and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot operation. More specifically, the present invention relates to a multi-robot collaborative operation and task allocation system and method. Background Art

[0002] Multi-robot collaborative operation and task allocation is a highly complex and challenging technology, which involves multiple aspects such as efficient cooperation, task allocation, and communication coordination among multiple robots. With the continuous progress of technology, especially the breakthroughs in the fields of artificial intelligence, robot perception and decision-making, multi-robot systems will be more widely used in various industries and become an important part of future intelligent systems. With the development of technologies such as artificial intelligence, machine learning, and deep learning, future multi-robot systems will be more intelligent, flexible, and efficient. Through adaptive task allocation algorithms, reinforcement learning, and distributed decision-making, robots will be able to better cope with complex and dynamic task environments.

[0003] The existing technology has the following deficiencies: In traditional methods, the task allocation efficiency is low and resources are wasted. Especially in multi-robot systems, tasks cannot be accurately allocated to improve the overall work efficiency. There is a lack of dynamic adaptability in traditional systems and they cannot cope with task and environmental changes to gradually obtain the optimal decision-making strategy, especially in unknown or complex environments. However, the Q-learning algorithm always selects the action with the maximum Q value, which makes it tend to choose the currently known best action while ignoring other potential actions. This strategy that overly relies on known actions easily leads to insufficient exploration. Due to its greedy selection feature, it is easy to fall into local optimal solutions and may not be able to comprehensively explore effectively.

[0004] In view of the above problems, the present invention proposes a solution. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the existing technology, embodiments of the present invention provide a multi-robot collaborative operation and task allocation system and method to solve the problems raised in the above background art.

[0006] To achieve the above object, the present invention provides the following technical solutions: A multi-robot collaborative operation and task allocation method, including the following steps: Step S1: Establish a pre-operation stage, and use the auction algorithm to select the most suitable robot for processing in combination with the work progress and spatial matching degree of the robot to obtain the task sequence matrix of the robot; Step S2: Use the Q-learning algorithm based on the task completion situation of the robot, and continuously update the Q value through interaction with the environment to finally obtain the optimal strategy; Step S3: Determine the relevance between each step based on the task sequence matrix of the robot, obtain the fluctuation of the reward signal and the noise in the environment, and jointly calculate the environmental change rate; Step S4: Comprehensively determine whether to replace it with the Sarsa learning algorithm to optimize the tasks of the robot to obtain the optimal strategy by combining the relevance between steps and the environmental change rate.

[0007] In a preferred embodiment, step S1 includes the following contents: Obtain the time window and starting position of the task to be assigned, and obtain the current task completion progress of the robot and the position where the robot is located when completing the current task; Establish a matching matrix according to the characteristics of the robot and the task. Each element in the matrix represents the "fitness" of a certain robot performing a certain task. The cost of the matching matrix is calculated based on two factors: work progress and spatial matching degree, including: Spatial matching degree: the distance between the robot and the task position; Work progress: the matching degree between the current task completion progress of the robot and the time window of the task to be assigned; Each robot gives an initial offer based on the cost or fitness of the matching matrix. According to the time cost and energy consumption of performing the task, each robot gives a bid based on its own fitness for the task and selects the most suitable task for itself. The robot gives a dynamic bidding offer according to the time progress and resource situation; Allocate it to the corresponding robot according to the bid amount, select the robot with the optimal bid, and mark it as allocated. Other robots will lose the competition qualification for this task. According to the auction result, allocate each task to the most suitable robot, record the execution order of the robot tasks, and update the work progress of the robot. According to the completion time of the task, adjust the idle time of the robot; Generate a task sequence matrix including task allocation, task start and end times, and the work progress of the robot.

[0008] In a preferred embodiment, step S2 includes the following contents: Execute the task sequence matrix of the robot obtained in the pre-operation stage, and use the Q-learning algorithm to dynamically adjust the behavior of the robot to optimize the task execution order and execution method of each robot; The robot estimates the state at each moment through this function and selects the action with the highest expected cumulative reward to execute to achieve the optimal strategy. Look up the Q-table to obtain the action value function of the state, select an action to execute, continuously update the Q value, obtain and enter the next state. The Q-table update method is as follows: , where: s is the current state, a is the current action, r is the immediate reward obtained according to the current action, is the new state transferred to after executing the action, The maximum Q-value action taken in the new state, α is the learning rate that controls the learning speed, and γ is the discount factor that controls the importance of future rewards; The robot interacts with the environment, selects actions, obtains rewards, and updates the Q-value. Eventually, it gradually converges to the optimal strategy. After multiple rounds of Q-learning optimization, the robot will gradually improve the efficiency of task execution and select the optimal task processing strategy.

[0009] In a preferred embodiment, step S3 includes the following: Use correlation analysis cosine similarity to evaluate the relevance between different tasks. By scanning the task sequence matrix, compare the execution order and status of tasks, and identify the precedence relationship and its impact on the final result; Set a function to calculate the reward for each task, and calculate the mean and variance of the reward signal to quantify the reward fluctuation: collect a certain number of reward samples, and calculate their mean μ and variance: : , where: μ represents the mean, is the i-th reward sample, and N is the total number of samples; , represents the variance, is the i-th reward sample, μ is the previously calculated mean, is the square of the deviation of each sample from the mean. The formula for normalizing the variance to [0, 10] using the normalization method is: , where: is the original variance data, are the minimum and maximum values of the variance respectively; Record the noise value of the sound level meter in the environment, use the displayed A-weighted decibels, and assign scores according to common noise levels: 30 - 40dB: assign 1 point, 50 - 60dB: assign 3 points, 70 - 80dB: assign 5 points, 90 - 100dB: assign 7 points, 120dB and above: assign 10 points.

[0010] Use the fluctuation of the reward signal and the noise in the environment to jointly calculate the environmental change rate, , where: Env Change Rate is the environmental change rate, α and β are weight coefficients, indicating the influence degree of the volatility of the reward signal and the noise on the environmental change rate, is the normalized variance of the reward signal, and Snoise is the score of the noise environment.

[0011] In a preferred embodiment, step S4 includes the following: Combined with the relevance between steps and the environmental change rate, comprehensively determine whether to replace it with the Sarsa learning algorithm to optimize the robot's tasks to obtain the optimal strategy: After receiving the relevance between steps and the environmental change rate, define the relevance between steps and the environmental change rate as input variables, and divide them into different fuzzy sets respectively; Define the learning algorithm as the output variable and divide it into a fuzzy set; Formulate fuzzy rules to describe the impact of the definition of the relevance between steps and the environmental change rate on the learning algorithm; Perform fuzzy reasoning according to the fuzzy rules to determine the learning algorithm for optimizing the robot's task.

[0012] A multi-robot cooperative operation and task allocation system, including: a pre-operation module, an algorithm optimization module, a data acquisition module, and an algorithm selection module: Pre-operation module: Establish a pre-operation stage, and use the auction algorithm in combination with the robot's work progress and spatial matching degree to select the most suitable robot for processing, and obtain the task sequence matrix of the robot; Algorithm optimization module: Use the Q-learning algorithm based on the robot's task completion situation, and continuously update the Q value through interaction with the environment to finally obtain the optimal strategy; Data acquisition module: Judge the relevance between each step based on the robot's task sequence matrix, obtain the fluctuation of the reward signal and the noise in the environment, and jointly calculate the environmental change rate; Algorithm selection module: Comprehensively determine whether to switch to the Sarsa learning algorithm to optimize the robot's task and obtain the optimal strategy in combination with the relevance between steps and the environmental change rate.

[0013] Technical effects and advantages of the multi-robot cooperative operation and task allocation system and method of the present invention: Use the auction algorithm to solve the robot task allocation problem. By considering the robot's work progress and spatial matching degree, it can efficiently select the most suitable robot for task processing. Use the Q-learning algorithm to continuously optimize the strategy according to the robot's task completion situation and in combination with environmental interaction. By analyzing the robot's task sequence matrix, evaluating the relevance between steps, and calculating the environmental change rate, it can cope with the noise and uncertainty in the environment. According to the relevance of tasks and the environmental change rate, decide whether to switch to the Sarsa learning algorithm. The design of flexible strategy switching makes up for the deficiencies of single reinforcement learning strategy, lack of flexibility and adaptability in the prior art, enabling the robot to dynamically adjust the optimization algorithm according to the actual situation and improve the quality and efficiency of task completion. Brief Description of the Drawings

[0014] Figure 1 It is a schematic structural diagram of the multi-robot cooperative operation and task allocation method of the present invention.

[0015] Figure 2This is a schematic structural diagram of the multi-robot collaborative operation and task allocation system of the present invention. Specific embodiments

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0017] Embodiment 1 The present invention discloses a multi-robot collaborative operation and task allocation method, as Figure 1 shown, including the steps: Step S1: Establish a pre-operation stage, and use the auction algorithm to select the most suitable robot for processing in combination with the work progress and spatial matching degree of the robot to obtain the task sequence matrix of the robot; Step S2: Use the Q-learning algorithm according to the task completion situation of the robot, and continuously update the Q value through interaction with the environment to finally obtain the optimal strategy; Step S3: Judge the relevance between each step according to the task sequence matrix of the robot, obtain the fluctuation situation of the reward signal and the noise in the environment, and jointly calculate the environmental change rate; Step S4: Comprehensively decide whether to replace it with the Sarsa learning algorithm to optimize the tasks of the robot to obtain the optimal strategy in combination with the relevance between steps and the environmental change rate.

[0018] The specific implementation is as follows: In step S1, a pre-operation stage is established, and the auction algorithm is used to select the most suitable robot for processing in combination with the work progress and spatial matching degree of the robot to obtain the task sequence matrix of the robot. The specific content includes: Step A1: Obtain the time window and starting position of the task to be allocated, and obtain the current task completion progress of the robot and the position where the robot is located when completing the current task.

[0019] Step A2: Establish a matching matrix according to the characteristics of the robot and the task. Each element in the matrix represents the "fitness" of a certain robot executing a certain task. The cost of the matching matrix is calculated based on two factors: work progress and spatial matching degree, including: Spatial matching degree: the distance between the robot and the task position; Work progress: the matching degree between the current task completion progress of the robot and the time window of the task to be allocated.

[0020] Step A3: The auction algorithm simulates an auction process, where the task to be allocated acts as the auction item and the robot acts as the bidder. The auction process includes: Each robot gives an initial offer based on the cost or fitness of the matching matrix, which can be the time cost of task execution, energy consumption, or other evaluation metrics.

[0021] Each robot makes a bid according to its own fitness for the task and selects the most suitable task for itself. The robot gives a dynamic bidding offer based on the time progress, resource situation, etc.

[0022] Step A4: Allocate to the corresponding robot according to the bid amount. If multiple robots bid for the same task, select the robot with the optimal bid and mark it as allocated. Other robots lose the qualification to compete for this task.

[0023] According to the auction results, allocate each task to the most suitable robot, record the execution order of the robot tasks, and update the work progress of the robots. According to the completion time of the tasks, adjust the idle time of the robots to ensure the matching of subsequent tasks.

[0024] Step A5: During the execution process, the auction algorithm can be dynamically optimized in the following ways: Adjust the task priorities according to the real-time situation to ensure that high-priority tasks are processed first. If a certain robot fails to complete the task on time or a new task is inserted, an auction can be carried out again to make the task allocation more flexible. Optimize the task scheduling according to the work progress of the robots and the time windows of the tasks, so that the working time of the robots is more compact and the idle time is reduced.

[0025] Finally, at the end of the pre-operation phase, a task sequence matrix containing task allocation, task start and end times, and the work progress of the robots is generated, providing a clear task order and time arrangement for subsequent task execution.

[0026] In step S2, the Q-learning algorithm is used based on the task completion situation of the robots, and the Q value is continuously updated through interaction with the environment to finally obtain the optimal strategy. The specific content includes: Execute the task sequence matrix of the robots obtained in the pre-operation phase, use the Q-learning algorithm to dynamically adjust the behavior of the robots, optimize the task execution order and execution method of each robot, so as to continuously improve the overall efficiency and task completion rate.

[0027] It should be noted that Q-learning is an algorithm based on reinforcement learning, aiming to learn the optimal behavior strategy through interaction with the environment. By learning the "action value function", that is, the Q value, it helps the agent select the optimal action in a given state to maximize the long-term reward. In Q-learning, the environment can be the state space of task execution, including the position of the robot, the progress of the task, the working status, etc. Each state represents the state of the robot at a certain moment, such as which task is being executed and the progress of task completion.

[0028] The action of the robot selects different task sequences, ways of executing tasks, or adjusts its own execution strategy, etc. The choice of action will affect the state of the robot and the completion of the task.

[0029] The robot estimates the state at each moment through this function and selects the action with the highest expected cumulative reward to execute to achieve the optimal strategy. Q-Learning is a classic algorithm in reinforcement learning. By learning a state-action value function, it is used to represent the expected reward that can be obtained by executing a certain action in any state in the environment. The robot observes the environment at the current moment to obtain the state S, looks up the Q-table to obtain the action value function of the state, selects an action to execute, continuously updates the Q value, obtains and enters the next state. The Q-table update method is as follows: , where: s is the current state, a is the current action, r is the immediate reward obtained according to the current action, is the new state transferred to after executing the action, is the action with the maximum Q value in the new state, α is the learning rate, which controls the learning speed, and γ is the discount factor, which controls the importance of future rewards.

[0030] The robot interacts with the environment, selects actions, obtains rewards and updates the Q value, and finally gradually converges to the optimal strategy to ensure the maximization of task completion efficiency. After multiple rounds of Q-learning optimization, the robot will gradually improve the efficiency of task execution, select the optimal task processing strategy, reduce task execution time, improve resource utilization, etc.

[0031] In step S3, based on the task sequence matrix of the robot, judge the relevance between each step, obtain the fluctuation of the reward signal and the noise in the environment, and jointly calculate the environmental change rate. The specific content includes: Use correlation analysis cosine similarity to evaluate the relevance between different tasks. By scanning the task sequence matrix, compare the execution order and state of tasks, and identify the sequential relationship and its impact on the final result.

[0032] The reward function is designed based on factors such as the task completion efficiency, execution speed, energy consumption, and space matching degree of the robot. Set a function to calculate the reward for each task. For example, if a certain robot can complete the task in the shortest time, the reward value can be relatively high; conversely, if the task execution is not ideal or the delay time is too long, the reward value is negative.

[0033] Calculate the mean and variance of the reward signal to quantify the fluctuation of the reward, which provides a quantitative description of the reward signal. The mean provides the central tendency of the reward, while the variance characterizes the volatility of the reward. The larger the variance, the higher the volatility of the reward, and vice versa. The specific steps are as follows: Collect a certain number of reward samples and calculate their mean μ and variance: : , where: μ represents the mean, is the i-th reward sample, N is the total number of samples; , represents the variance, is the i-th reward sample, μ is the previously calculated mean, is the square of the deviation of each sample from the mean. The formula for normalizing the contrast to [0, 10] using the normalization method is: ,in: is the variance of the original data, are the minimum and maximum values ​​of the variance, respectively.

[0034] Record the noise value of the sound level meter in the environment. Since A-weighted decibels are closer to the auditory characteristics of the human ear, A-weighted decibels are used to display the noise level of the noise environment. According to the measurement results, the noise level of the noise environment is evaluated and points are assigned according to common noise levels: 30-40dB: 1 point, 50-60dB: 3 points, 70-80dB: 5 points, 90-100dB: 7 points, 120dB and above: 10 points.

[0035] The rate of change of the environment is calculated using the fluctuation of the reward signal and the noise in the environment. , where: Env Change Rate is the environment change rate, α and β are weight coefficients, indicating the impact of the volatility and noise of the reward signal on the environment change rate, is the normalized variance of the reward signal, and Snoise is the score for the noisy environment.

[0036] In step S4, the correlation between the steps and the rate of change of the environment are combined to comprehensively decide whether to switch to the Sarsa learning algorithm to optimize the robot's task and obtain the optimal strategy. The specific contents include: In the traditional Q-Learning algorithm, the current Q value is always updated by the maximum Q value of the next state. This selection method based on greedy strategy will lead to insufficient exploration of the robot, and it is seriously biased towards choosing the known optimal action, thus falling into the local optimal strategy. The Sarsa algorithm updates the current Q value by the expected Q value of the next state, which effectively improves the exploration ability of the strategy. It combines the correlation between steps and the rate of change of the environment to comprehensively decide whether to switch to the Sarsa learning algorithm to optimize the robot's task and obtain the optimal strategy.

[0037] After receiving the correlation between the steps and the environmental change rate, the correlation between the steps and the environmental change rate are defined as input variables, and they are divided into different fuzzy sets.

[0038] For example, "Low", "Medium", "High" for relevance, "Low", "Medium", "High" for the environmental change rate.

[0039] Define the learning algorithm as the output variable and divide it into fuzzy sets. For example, "Q-Learning", "Sarsa" for the learning algorithm.

[0040] Formulate a set of fuzzy rules to describe the influence of different input variables on the output variable. The definition of the rules can be based on professional knowledge or obtained through data analysis and experiments. For example: Mark relevance as T, the environmental change rate as B, and the learning algorithm as Algorithm, then the following can be defined Rule 1: IF (T is Low) AND (B is Low) THEN (Algorithm is Q-Learning) Rule 2: IF (T is High) AND (B is High) THEN (Algorithm is Sarsa) Conduct fuzzy reasoning according to the fuzzy rules to determine the solution of the learning algorithm.

[0041] It should be noted that the division of the fuzzy sets can be adjusted according to the actual situation. For example, although this embodiment takes three fuzzy sets as an example, in fact, relevance, the environmental change rate, and the learning algorithm can be divided into more than three sets to facilitate more accurate adjustment according to different temperatures.

[0042] Embodiment 2 A multi-robot collaborative operation and task allocation system, such as Figure 2 shown, includes: a pre-operation module, an algorithm optimization module, a data acquisition module, and an algorithm selection module, and the modules are signal-connected to each other.

[0043] Pre-operation module: Establish a pre-operation stage, use the auction algorithm in combination with the working progress and spatial matching degree of the robots to select the most suitable robot for processing, and obtain the task sequence matrix of the robots; Algorithm optimization module: Use the Q-learning algorithm based on the task completion situation of the robots, and continuously update the Q value through interaction with the environment to finally obtain the optimal strategy; Data acquisition module: Judge the relevance between each step based on the task sequence matrix of the robots, obtain the fluctuation of the reward signal and the noise in the environment, and jointly calculate the environmental change rate; Algorithm selection module: Based on the relevance between steps and the environmental change rate, comprehensively determine whether to replace it with the Sarsa learning algorithm to optimize the robot's tasks and obtain the optimal strategy.

[0044] All the above formulas are dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0045] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.

[0046] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application of the technical solution and the invention constraints. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0047] In addition, in each embodiment of this application, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0048] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0049] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A multi-robot collaborative operation and task allocation method, characterized in that It includes the following steps: Step S1: Establish a pre-operation stage. Use the auction algorithm in combination with the working progress and spatial matching degree of the robot to select the most suitable robot for processing, and obtain the task sequence matrix of the robot; Step S2: Use the Q-learning algorithm based on the task completion situation of the robot, and continuously update the Q value through interaction with the environment to finally obtain the optimal strategy; Step S3: Judge the relevance between each step according to the task sequence matrix of the robot, obtain the fluctuation of the reward signal and the noise in the environment, and jointly calculate the environmental change rate; Step S4: Comprehensively determine whether to replace it with the Sarsa learning algorithm to optimize the tasks of the robot to obtain the optimal strategy by combining the relevance between steps and the environmental change rate.

2. The multi-robot collaborative operation and task allocation method according to claim 1, characterized in that: Obtain the time window and starting position of the task to be allocated, and obtain the current task completion progress of the robot and the position where the robot is located when completing the current task; Establish a matching matrix according to the characteristics of the robot and the task. Each element in the matrix represents the "fitness" of a certain robot to execute a certain task. The cost of the matching matrix is calculated based on two factors: working progress and spatial matching degree, including: spatial matching degree: the distance between the robot and the task position; working progress: the matching degree between the current task completion progress of the robot and the time window of the task to be allocated; Based on the cost or fitness of the matching matrix, each robot gives an initial bid. According to the time cost and energy consumption of executing the task, each robot gives a bid according to its own fitness for the task, and selects the most suitable task for itself. The robot gives a dynamic bidding price according to the time progress and resource situation; Allocate to the corresponding robot according to the bid amount, select the robot with the optimal bid, and mark it as allocated. Other robots will lose the competition qualification for this task. According to the auction result, allocate each task to the most suitable robot, record the execution order of the robot task, and update the working progress of the robot. Adjust the idle time of the robot according to the completion time of the task; Generate a task sequence matrix including task allocation, task start and end times, and the working progress of the robot.

3. The multi-robot collaborative operation and task allocation method according to claim 2, characterized in that: Execute the task sequence matrix of the robot obtained in the pre-operation stage, and use the Q-learning algorithm to dynamically adjust the behavior of the robot to optimize the task execution order and execution method of each robot; The robot estimates the state at each moment through this function and selects the action with the highest expected cumulative reward to execute to achieve the optimal policy. It looks up the Q-table to obtain the action value function of the state, selects an action to execute, continuously updates the Q-value, obtains and enters the next state. The Q-table update method is as follows: , where: s is the current state, a is the current action, r is the immediate reward obtained according to the current action, is the new state transferred to after executing the action, is the action with the maximum Q-value in the new state, α is the learning rate, which controls the learning speed, and γ is the discount factor, which controls the importance of future rewards; The robot interacts with the environment, selects actions, obtains rewards and updates the Q value, and finally gradually converges to the optimal strategy. After multiple rounds of Q-learning optimization, the robot will gradually improve the efficiency of task execution and select the optimal task processing strategy.

4. The multi-robot collaborative operation and task allocation method according to claim 3, wherein ; Use the cosine similarity of correlation analysis to evaluate the relevance between different tasks. By scanning the task sequence matrix, compare the execution order and status of the tasks, and identify the sequence relationship and its impact on the final result; Set a function to calculate the reward for each task, calculate the mean and variance of the reward signal to quantify the fluctuation of the reward: collect a certain number of reward samples and calculate their mean μ and variance: : , where: μ represents the mean, is the i-th reward sample, and N is the total number of samples; , represents the variance, is the i-th reward sample, μ is the previously calculated mean, is the square of the deviation of each sample from the mean; the formula for normalizing the contrast to [0, 10] using the normalization method is: , where: is the original variance data, are the minimum and maximum values of the variance respectively; Record the noise value of the sound level meter in the environment, use the display of A-weighted decibels, and assign scores according to common noise levels: 30 - 40 dB: assign 1 point, 50 - 60 dB: assign 3 points, 70 - 80 dB: assign 5 points, 90 - 100 dB: assign 7 points, 120 dB and above: assign 10 points; Calculate the environmental change rate using the fluctuations of the reward signal and the noise in the environment, , where: Env Change Rate is the environmental change rate, and α and β are weight coefficients representing the influence degrees of the volatility of the reward signal and the noise on the environmental change rate, is the normalized variance of the reward signal, and Snoise is the score of the noisy environment.

5. The multi-robot collaborative operation and task allocation method according to claim 4, wherein: Comprehensively determine whether to replace it with the Sarsa learning algorithm to optimize the tasks of the robot to obtain the optimal strategy by combining the relevance between steps and the environmental change rate; After receiving the relevance between steps and the environmental change rate, define the relevance between steps and the environmental change rate as input variables and divide them into different fuzzy sets respectively; Define the learning algorithm as the output variable and divide it into fuzzy sets; Formulate fuzzy rules to describe the influence of the definition of the relevance between steps and the environmental change rate on the learning algorithm; Perform fuzzy reasoning according to the fuzzy rules to determine the learning algorithm for optimizing the tasks of the robot.

6. A multi-robot collaborative operation and task allocation system for implementing the multi-robot collaborative operation and task allocation method according to any one of claims 1-5, characterized in that, Including: A pre-operation module, an algorithm optimization module, a data acquisition module, and an algorithm selection module: Pre-operation module: Establish a pre-operation stage, use the auction algorithm to select the most suitable robot for processing in combination with the work progress and spatial matching degree of the robot, and obtain the task sequence matrix of the robot; Algorithm optimization module: Use the Q-learning algorithm based on the task completion situation of the robot, and continuously update the Q value through interaction with the environment to finally obtain the optimal strategy; Data acquisition module: Judge the relevance between each step according to the task sequence matrix of the robot, obtain the fluctuation of the reward signal and the noise in the environment, and jointly calculate the environmental change rate; Algorithm selection module: Comprehensively determine whether to replace it with the Sarsa learning algorithm to optimize the tasks of the robot to obtain the optimal strategy by combining the relevance between steps and the environmental change rate.

Citation Information

Cited By

  • Intelligent station unmanned loader obstacle avoidance scheduling method and system

    CN121254860A