Cooperative working method and system of multiple robots and Internet of Things equipment
By generating the capability map of IoT devices and a three-dimensional task priority matrix, combined with the resource matching algorithm of reinforcement learning, the problem of untimely response of IoT devices is solved, and efficient collaborative work between multiple robots and IoT devices is realized, improving the accuracy and real-timeness of task allocation.
Patent Information
- Application Number
- CN202510469139.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-15
AI Technical Summary
In the prior art, IoT devices do not respond to real-time environmental data in a timely manner, dynamic task allocation efficiency is not high, and it is difficult to achieve the coordinated work of multiple robots and IoT devices, especially in complex and changeable environments, which cannot meet the requirements of control accuracy and real-time.
By obtaining the real-time load status of pending tasks and robots, a capability map of IoT device groups is generated, a resource matching algorithm based on dynamic task evaluation models and reinforcement learning is generated, a three-dimensional task priority matrix is generated, the target robot-IoT device combination is allocated, and the resource allocation strategy is dynamically adjusted through an adaptive feedback mechanism.
It realizes efficient collaborative work between multiple robots and IoT devices in multiple scenarios, improves the accuracy and real-time nature of task allocation, and optimizes resource utilization efficiency.
Smart Images

Figure CN120343049A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of warehousing logistics automation, and particularly relates to a collaborative working method and system for multiple robots and Internet of Things devices. Background Art
[0002] With the rapid development of Internet of Things technology, Internet of Things robots have been widely used in multiple fields such as industry, agriculture, medical care, and military. Internet of Things robots collect data of the working site in real time through devices such as sensors, and transmit the data to a remote control center through a network, realizing remote control of the robots.
[0003] However, in the existing collaborative working scenarios of robots and Internet of Things devices, due to factors such as a large number of robots, complex tasks, and changing environments, the Internet of Things devices in the existing distributed system respond to real-time environmental data in a timely manner, and the existing algorithms have low efficiency in dynamic task allocation, making it difficult to ensure the accuracy and real-time nature of control, and unable to meet the collaborative working scenarios of multiple robots and Internet of Things devices in multiple scenarios. Summary of the Invention
[0004] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a collaborative working method for multiple robots and Internet of Things devices, including:
[0005] Obtaining tasks to be processed and the real-time load status of multiple robots;
[0006] Generating a capability map of an Internet of Things device group according to real-time environmental data of multiple Internet of Things devices collected by a distributed sensor network;
[0007] Based on a pre-established dynamic task evaluation model, generating a three-dimensional task priority matrix for the tasks to be processed according to the tasks to be processed, the real-time load status of multiple robots, and the capability map of the Internet of Things device group;
[0008] Adopting a resource matching algorithm based on reinforcement learning, and allocating a target robot-Internet of Things device combination for the tasks to be processed based on the three-dimensional task priority matrix of the tasks to be processed.
[0009] Preferably, the process of establishing the dynamic task evaluation model includes:
[0010] Calculating the topological structure entropy value of the environmental obstacles and the dynamic interference factor based on the capability map of the Internet of Things device group to obtain the real-time environmental complexity of the Internet of Things device group;
[0011] Constructing a multi-dimensional capability vector based on the specification parameters and historical efficiency data of the Internet of Things devices to obtain the capability quantization vector of the Internet of Things devices;
[0012] Based on the real-time environmental complexity and the ability quantization vector, a dynamic task evaluation model is established.
[0013] Preferably, the expression of the three-dimensional task priority matrix is as follows:
[0014] P(t, s, r) = α·T(t) + β·S(s) + γ·R(r)
[0015] Wherein, P(t, s, r) is the three-dimensional task priority matrix, T(t) is the time dimension priority function, S(s) is the space dimension priority function, R(r) is the resource dimension priority function, α is the time dimension priority weight coefficient, β is the space dimension priority weight coefficient, and γ is the resource dimension priority weight coefficient.
[0016] Preferably, the expression of the time dimension priority function T(t) is as follows:
[0017]
[0018] Wherein, d i is the deadline of the task to be processed, t now is the current system time, ε is to prevent zero minimum value, λ is the decay factor, ω i is the remaining workload of the task to be processed;
[0019] The expression of the time dimension priority weight coefficient α is as follows:
[0020]
[0021] Wherein, α0 is the initial time dimension weight coefficient, d i is the deadline of the task to be processed, t now is the current system time, τ critical is the preset critical time threshold;
[0022] The expression of the space dimension priority function S(s) is as follows:
[0023]
[0024] Wherein, C path is the dynamic path cost, (x r , y r ) is the robot position coordinate, (x t , y t ) is the task target position coordinate, σ is the space decay radius; is the sum of all edges (path segments) on the path, v r is the moving speed of the robot, μ is the obstacle penalty coefficient, and obstacle_density(e) is the obstacle density in the area where the path segment e is located. is the time cost;
[0025] The expression of the resource dimension priority function R(r) is as follows:
[0026]
[0027] where n is the total number of resource types, ω i is the weight of the i-th resource type, and η is the overcapacity reward coefficient; is the normalized task requirement value of the i-th resource, is the normalized device capacity value of device j under the i-th resource, is the total margin of the normalized device capacity of device j under the i-th resource and the normalized task requirement value of the i-th resource, and sigmoid maps the total margin to the interval (0,1); d i is the requirement value of the task for the i-th resource, and max(D) is the maximum value in the task requirement vector in; is the requirement value of all resource types, cj i is the capacity value of the Internet of Things device j for the i-th resource, and max(c j ) is the maximum value of the Internet of Things device capacity vector in; is the capacity value of all resource types of device j,
[0028] Preferably, after allocating the target robot-Internet of Things device combination for the work task to be processed, the method further includes:
[0029] During the execution of the work task to be processed by the target robot-Internet of Things device combination, monitor the state parameters of the Internet of Things device group, and dynamically adjust the resource allocation strategy through an adaptive feedback mechanism.
[0030] Preferably, dynamically adjusting the resource allocation strategy through an adaptive feedback mechanism includes:
[0031] Establish a composite evaluation index including task completion timeliness, device reliability, and energy consumption efficiency;
[0032] Based on the composite evaluation index, optimize the control parameters based on the deep deterministic policy gradient algorithm, and adjust the resource allocation strategy through the control parameters.
[0033] Preferably, the expression of the composite evaluation index is as follows:
[0034]
[0035] where Q is the composite evaluation index value, T completeis the task completion timeliness value, E consumed is the energy efficiency value, S reliability is the reliability index of the Internet of Things device, μ1 is the task completion timeliness weight coefficient, μ2 is the energy efficiency weight coefficient, and μ3 is the Internet of Things device reliability weight coefficient;
[0036] The expression of the deep deterministic policy gradient algorithm is as follows:
[0037]
[0038] Among them, π θ (a|s) is the parameterized policy function, θ is the policy network parameter, is the objective function to be maximized through the policy network parameter θ, Ε[Q(s,a)] is the expected value of the action value function Q(s,a), λ is the policy update constraint coefficient, KL(π θ ||π θold ) is the KL divergence between the old and new policies.
[0039] Based on the same inventive concept, the present invention also provides a collaborative working system for multiple robots and Internet of Things devices, including:
[0040] An original data acquisition module for acquiring the task to be processed and the real-time load status of multiple robots;
[0041] A device group status acquisition module for generating a capability map of the Internet of Things device group according to the real-time environment data of multiple Internet of Things devices collected by the distributed sensor network;
[0042] A task priority matrix generation module for generating a three-dimensional task priority matrix of the task to be processed based on a pre-established dynamic task evaluation model according to the task to be processed, the real-time load status of multiple robots, and the capability map of the Internet of Things device group;
[0043] A robot-device allocation module for allocating a target robot-Internet of Things device combination for the task to be processed based on the three-dimensional task priority matrix of the task to be processed by using a resource matching algorithm of reinforcement learning.
[0044] Preferably, the process of establishing the dynamic task evaluation model includes:
[0045] Based on the capability map of the Internet of Things device group, calculating the environmental obstacle topology structure entropy value and the dynamic interference factor to obtain the real-time environmental complexity of the Internet of Things device group;
[0046] Based on the specification parameters and historical efficiency data of the Internet of Things device, constructing a multi-dimensional capability vector to obtain the capability quantization vector of the Internet of Things device;
[0047] Based on the real-time environment complexity and the ability quantization vector, a dynamic task evaluation model is established.
[0048] Preferably, the expression of the three-dimensional task priority matrix is as follows:
[0049] P(t, s, r) = α·T(t) + β·S(s) + γ·R(r)
[0050] Where, P(t, s, r) is the three-dimensional task priority matrix, T(t) is the time dimension priority function, S(s) is the space dimension priority function, R(r) is the resource dimension priority function, α is the time dimension priority weight coefficient, β is the space dimension priority weight coefficient, and γ is the resource dimension priority weight coefficient.
[0051] Preferably, the expression of the time dimension priority function T(t) is as follows:
[0052]
[0053] Where, d i is the deadline of the task to be processed, t now is the current system time, ε is to prevent zero minimum value, λ is the decay factor, ω i is the remaining workload of the task to be processed;
[0054] The expression of the time dimension priority weight coefficient α is as follows:
[0055]
[0056] Where, α0 is the initial time dimension weight coefficient, d i is the deadline of the task to be processed, t now is the current system time, τ critical is the preset critical time threshold;
[0057] The expression of the space dimension priority function S(s) is as follows:
[0058]
[0059] Where, C path is the dynamic path cost, (x r , y r ) is the robot position coordinate, (x t , y t ) is the task target position coordinate, σ is the space decay radius; is the sum of all edges (path segments) on the path, v r is the moving speed of the robot, μ is the obstacle penalty coefficient, and obstacle_density(e) is the obstacle density in the area where the path segment e is located. is the time cost;
[0060] The expression of the resource dimension priority function R(r) is as follows:
[0061]
[0062]
[0063] where n is the total number of resource types, ω i is the weight of the i-th resource type, and η is the over-capacity reward coefficient; is the normalized task demand value of the i-th resource, is the normalized device capacity value of device j under the i-th resource, is the total surplus of the normalized device capacity of device j under the i-th resource and the normalized task demand value of the i-th resource, and sigmoid maps the total surplus to the interval (0, 1); d i is the demand value of the task for the i-th resource, and max(D) is the maximum value in the task demand vector in; is the demand value of all resource types, cj i is the capacity value of the Internet of Things device j for the i-th resource, and max(c j ) is the maximum value of the Internet of Things device capacity vector ; is the capacity value of all resource types of device j,
[0064] Preferably, the system further includes a dynamic resource allocation adjustment module for:
[0065] During the execution of the pending work task by the target robot-Internet of Things device combination, monitoring the state parameters of the Internet of Things device group, and dynamically adjusting the resource allocation strategy through an adaptive feedback mechanism.
[0066] Preferably, the dynamic resource allocation adjustment module is specifically used for:
[0067] Establish a composite evaluation index including task completion timeliness, device reliability, and energy consumption efficiency;
[0068] Based on the composite evaluation index, optimize the control parameters based on the deep deterministic policy gradient algorithm, and adjust the resource allocation strategy through the control parameters.
[0069] Preferably, the expression of the composite evaluation index is as follows:
[0070]
[0071] Among them, Q is the composite evaluation index value, and T complete is the task completion timeliness value, E consumed is the energy efficiency value, S reliability is the reliability index of the Internet of Things device, μ1 is the task completion timeliness weight coefficient, μ2 is the energy efficiency weight coefficient, and μ3 is the Internet of Things device reliability weight coefficient;
[0072] The expression of the deep deterministic policy gradient algorithm is as follows:
[0073]
[0074] Among them, π θ (a|s) is the parameterized policy function, θ is the policy network parameter, is the objective function to be maximized through the policy network parameter θ, Ε[Q(s,a)] is the expected value of the action value function Q(s,a), λ is the policy update constraint coefficient, and KL(π θ ||π θold ) is the KL divergence between the old and new policies.
[0075] Based on the same inventive concept, the present invention also provides an electronic device, including: at least one processor and a memory; the memory and the processor are connected by a bus;
[0076] The memory is used to store one or more programs;
[0077] When the one or more programs are executed by the at least one processor, the collaborative working method of a plurality of robots and Internet of Things devices as described above is implemented.
[0078] Based on the same inventive concept, the present invention also provides a readable storage medium, on which an execution program is stored, and when the execution program is executed, the collaborative working method of a plurality of robots and Internet of Things devices as described above is implemented.
[0079] Compared with the closest prior art, the beneficial effects of the present invention are as follows:
[0080] The present invention provides a collaborative working method for multiple robots and Internet of Things (IoT) devices, including: obtaining a task to be processed and real-time load statuses of multiple robots; generating a capability map of an IoT device group according to real-time environmental data of multiple IoT devices collected by a distributed sensor network; generating a three-dimensional task priority matrix for the task to be processed based on a pre-established dynamic task evaluation model according to the task to be processed, real-time load statuses of multiple robots, and the capability map of the IoT device group; and using a resource matching algorithm based on reinforcement learning to allocate a target robot-IoT device combination for the task to be processed according to the three-dimensional task priority matrix of the task to be processed. By collecting real-time environmental data of IoT devices, generating a capability map of the IoT device group, and inputting the same together with the task to be processed and real-time load statuses of multiple robots into the pre-established dynamic task evaluation model, a three-dimensional task priority matrix is obtained, and then a target robot-IoT device combination is matched for the task to be processed according to the three-dimensional task priority matrix, so as to realize collaborative working of multiple robots and IoT devices in multiple scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 is a schematic flowchart of a collaborative working method for multiple robots and IoT devices provided by the present invention;
[0082] Figure 2 is a structural diagram of a collaborative working system for multiple robots and IoT devices provided by the present invention;
[0083] Figure 3 is a schematic diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0084] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.
[0085] Embodiment 1:
[0086] The present invention provides a collaborative working method for multiple robots and IoT devices. Specifically, Figure 1 is a schematic flowchart of the collaborative working method for multiple robots and IoT devices provided by an embodiment of the present invention. As shown in the figure, the method includes the following steps:
[0087] S1: Obtain a task to be processed and real-time load statuses of multiple robots;
[0088] S2: Generate a capability map of the IoT device group based on the real-time environmental data of multiple IoT devices collected by the distributed sensor network;
[0089] S3: Based on the pre-established dynamic task evaluation model, generate a three-dimensional task priority matrix for the task to be processed according to the task to be processed, the real-time load status of multiple robots, and the capability map of the IoT device group;
[0090] S4: Adopt a resource matching algorithm based on reinforcement learning, and allocate a target robot-IoT device combination for the task to be processed based on the three-dimensional task priority matrix of the task to be processed.
[0091] In the present invention, the real-time environmental data of the IoT devices is collected to generate a capability map of the IoT device group. At the same time, it is input into the pre-established dynamic task evaluation model together with the task to be processed and the real-time load status of multiple robots to obtain a three-dimensional task priority matrix. Furthermore, a target robot-IoT device combination is matched for the task to be processed according to the three-dimensional task priority matrix, realizing the collaborative work of multiple robots and IoT devices in multiple scenarios.
[0092] In the present invention, the task to be processed and the real-time load status of multiple robots are obtained. The task to be processed can be tasks such as handling, inspection, or processing that the robot needs to perform. In some optional embodiments, the task to be processed can be characterized by a task description tuple. Exemplarily, the constructed task description tuple can be T = {TaskID, Type, Priority, ResourceReq, Location, Deadline}. Among them, TaskID is the task authentication information, Type is the task type, which can be tasks such as handling, inspection, or processing, Priority is the task priority, ResourceReq is the task source sequence, Location is the task location, and Deadline is the task completion deadline. It can be understood that the above task description tuple is only an example and does not constitute a limitation to the present invention.
[0093] The real-time load status of the robot can be collected by an embedded monitoring module installed on the robot. The collected real-time load status of the robot includes but is not limited to: power load: battery SOC (State of Charge, battery power), battery SOH (State of Health, battery health status); mechanical load: joint torque sensor data; computing load: processor utilization rate; task queue depth: using a circular buffer to record tasks to be processed.
[0094] After that, real-time environmental data of multiple Internet of Things devices are collected according to the distributed sensor network. Specifically, intelligent sensor nodes with multimodal perception capabilities can be deployed, a data fusion processing channel based on the Kalman filter can be established, and a dynamic sampling frequency adjustment mechanism can be configured to adapt to environmental changes. The real-time environmental data of the collected Internet of Things devices include, but are not limited to: indoor positioning data, environmental monitoring data, and dynamic obstacle detection data. Specifically, indoor positioning data can be obtained through UWB (Ultra-Wideband) + TDOA (Time Difference of Arrival) hybrid positioning; environmental monitoring data such as temperature, humidity, and light intensity can be collected through Modbus RTU (Modbus Remote Terminal Unit, serial communication transmission mode); dynamic obstacle detection data can be obtained through millimeter-wave radar point cloud data in combination with the running trajectories of multiple robots.
[0095] Furthermore, based on the real-time environmental data of multiple collected Internet of Things devices, a capability map of the Internet of Things device group is generated. Specifically, the device capability characteristics can be extracted according to the real-time environmental data, and a five-dimensional capability vector can be constructed. Exemplarily, the constructed five-dimensional capability vector can be C = {Processing, Storage, Communication, Mobility, Sensing}, where Processing is the device computing power, Storage is the device storage space, Communication is the communication bandwidth, Mobility is the device mobility (0 for static devices, 0-1 scale), and Sensing is the sensor accuracy (normalized to the 0-1 interval). It can be understood that the above five-dimensional capability vector is an example of the capability map of the generated Internet of Things device group and does not constitute a limitation to the present invention.
[0096] The above-mentioned task to be processed, the real-time load status of multiple robots, and the capability map of the Internet of Things device group are all initial preparation data. After obtaining the initial preparation data, based on the pre-established dynamic task evaluation model, according to the task to be processed, the real-time load status of multiple robots, and the capability map of the Internet of Things device group, a three-dimensional task priority matrix of the task to be processed is generated.
[0097] Among them, the process of pre-establishing the dynamic task evaluation model includes: calculating the environmental obstacle topology entropy value and the dynamic interference factor based on the capability map of the Internet of Things device group to obtain the real-time environmental complexity of the Internet of Things device group; constructing a multi-dimensional capability vector based on the specification parameters and historical efficiency data of the Internet of Things devices to obtain the capability quantization vector of the Internet of Things devices; and establishing a dynamic task evaluation model based on the real-time environmental complexity and the capability quantization vector.
[0098] The three-dimensional task priority matrix of the task to be processed generated in the above manner includes the time dimension priority, the space dimension priority, the resource dimension priority, and their respective weight coefficients. Among them, the time dimension priority is used to calculate the urgency coefficient between the due time of the task to be processed and the current progress, the space dimension priority is used to evaluate the topological relationship between the location of the task to be processed and the distribution of Internet of Things devices, and the resource dimension priority is used to obtain the matching degree between the requirements of the task to be processed and the capabilities of Internet of Things devices.
[0099] Specifically, the expression of the three-dimensional task priority matrix is as follows:
[0100] P(t, s, r) = α·T(t) + β·S(s) + γ·R(r)
[0101] Among them, P(t, s, r) is the three-dimensional task priority matrix, T(t) is the time dimension priority function, S(s) is the space dimension priority function, R(r) is the resource dimension priority function, α is the time dimension priority weight coefficient, β is the space dimension priority weight coefficient, and γ is the resource dimension priority weight coefficient.
[0102] In the calculation of the time dimension priority, it is necessary to obtain the time dimension priority function and the time dimension priority coefficient. Among them, the expression of the time dimension priority function T(t) is as follows:
[0103]
[0104] Among them, d i is the due time of the task to be processed, t now is the current system time, ε is to prevent zero minimum value, λ is the attenuation factor, ω i is the remaining workload of the task to be processed;
[0105] When the difference between the due time of the task to be processed and the current system time is less than the preset critical time threshold, the time dimension weight coefficient is triggered to increase exponentially. The expression of the time dimension priority weight coefficient α is as follows:
[0106]
[0107] Among them, α0 is the initial time dimension weight coefficient, d i is the due time of the task to be processed, t now is the current system time, τ critical is the preset critical time threshold;
[0108] Dijkstra's algorithm is a classic algorithm for finding the single-source shortest path in a weighted graph. Its core goal is to find the shortest path or one of the shortest paths from the starting point to all other nodes in the graph, and it is applicable to graph structures with non-negative weights.
[0109] The calculation process of the spatial dimension priority is as follows: The improved Dijkstra's algorithm (proposed by the Dutch computer scientist Edsger W. Dijkstra) is used to calculate the dynamic path cost C path , and the expression of the dynamic path cost is as follows:
[0110]
[0111] where v r is the moving speed of the robot, is the sum of all edges (path segments) on the path, ||e|| is the Euclidean length of edge e, μ is the obstacle penalty coefficient, and obstacle_density(e) is the obstacle density in the area where path segment e is located. is the time cost, and μ·obstacle_density(e) is the obstacle density cost.
[0112] The expression of the spatial dimension priority function S(s) is as follows:
[0113]
[0114] where (x r , y r ) is the robot's position coordinates, (x t , y t ) is the task target position coordinates, and σ is the spatial decay radius.
[0115] The calculation process of the resource dimension priority function is as follows:
[0116] The normalization process of the task requirement vector and the device capability vector is as follows:
[0117]
[0118] where the task requirement vector the device capability vector
[0119] The expression of the improved cosine similarity calculation (i.e., the resource dimension priority function) is as follows:
[0120]
[0121] where n is the total number of resource types, ω iis the weight of the i-th resource type, and η is the over-capacity reward coefficient; is the normalized task demand value of the i-th resource, is the normalized device capacity value of device j under the i-th resource, is the total margin between the normalized device capacity of device j under the i-th resource and the normalized task demand value of the i-th resource, and sigmoid maps the total margin to the interval (0, 1); d i is the demand value of the task for the i-th resource, and max(D) is the maximum value of the task demand vector in it, is the demand value of all resource types, cj i is the capacity value of IoT device j for the i-th resource, and max(c j ) is the maximum value of the IoT device capacity vector in it, is the capacity value of device j for all resource types,
[0122] After generating the three-dimensional task priority matrix of the task to be processed in the above manner, an enhanced learning-based resource matching algorithm is used to allocate a target robot-IoT device combination for the task to be processed.
[0123] After obtaining the three-dimensional task priority matrix, matrix fusion and update can be performed on the three-dimensional task priority matrix. Specifically, the three-dimensional tensor of the three-dimensional task priority matrix is expressed as: where T is the priority in the time dimension, S is the priority in the space dimension, and R is the priority in the resource dimension. Update the spatial dimension priority data at a preset interval time, and trigger a local matrix refresh when the task status changes.
[0124] In the architecture process of the enhanced learning-based resource matching algorithm, the input feature is a slice of the three-dimensional priority matrix (the three-dimensional data of the current time slice), extract spatial-resource features through 3D CNN (Convolutional Neural Network), and use LSTM (Long Short-Term Memory) to capture the time dimension dependency relationship.
[0125] After allocating a target robot-IoT device combination for the task to be processed, in order to further improve the accuracy of the allocation, in some alternative embodiments, the method further includes: during the execution of the task to be processed by the target robot-IoT device combination, monitoring the status parameters of the IoT device group, and dynamically adjusting the resource allocation strategy through an adaptive feedback mechanism.
[0126] The establishment of the adaptive feedback mechanism includes: establishing a knowledge graph of the collaborative efficiency of Internet of Things devices; designing a long-term memory model based on the Transformer architecture to achieve online incremental learning optimization of task allocation strategies.
[0127] Specifically, the resource allocation strategy is dynamically adjusted through the adaptive feedback mechanism, including: establishing a composite evaluation index that includes task completion timeliness, device reliability, and energy consumption efficiency; based on the composite evaluation index, optimizing the control parameters based on the deep deterministic policy gradient algorithm, and adjusting the resource allocation strategy through the control parameters.
[0128] The expression of the composite evaluation index is as follows:
[0129]
[0130] Among them, Q is the value of the composite evaluation index, T complete is the value of task completion timeliness, E consumed is the value of energy efficiency, S reliability is the reliability index of Internet of Things devices, μ1 is the weight coefficient of task completion timeliness, μ2 is the weight coefficient of energy efficiency, and μ3 is the weight coefficient of Internet of Things device reliability;
[0131] The expression of the deep deterministic policy gradient algorithm is as follows:
[0132]
[0133] Among them, π θ (a|s) is the parameterized policy function, θ is the policy network parameter, is the objective function to be maximized through the policy network parameter θ, Ε[Q(s,a)] is the expected value of the action value function Q(s,a), λ is the policy update constraint coefficient, KL(π θ ||π θold ) is the KL divergence between the old and new policies.
[0134] The collaborative working method of multiple robots and Internet of Things devices provided by the present invention includes: obtaining the task to be processed and the real-time load status of multiple robots; generating a capability map of the Internet of Things device group according to the real-time environmental data of multiple Internet of Things devices collected by a distributed sensor network; based on a pre-established dynamic task evaluation model, generating a three-dimensional task priority matrix for the task to be processed according to the task to be processed, the real-time load status of multiple robots, and the capability map of the Internet of Things device group; using a resource matching algorithm of reinforcement learning, and based on the three-dimensional task priority matrix of the task to be processed, allocating a target robot-Internet of Things device combination for the task to be processed. By collecting the real-time environmental data of Internet of Things devices, generating a capability map of the Internet of Things device group, and at the same time inputting it together with the task to be processed and the real-time load status of multiple robots into the pre-established dynamic task evaluation model, a three-dimensional task priority matrix is obtained, and then a target robot-Internet of Things device combination is matched for the task to be processed according to the three-dimensional task priority matrix, so as to realize the collaborative work of multiple robots and Internet of Things devices in multiple scenarios.
[0135] Embodiment 2:
[0136] Based on the same inventive concept, the present invention also provides a collaborative working system 200 of multiple robots and Internet of Things devices. The structure of the system is as Figure 2 shown. The system includes:
[0137] An original data acquisition module 201, configured to obtain the task to be processed and the real-time load status of multiple robots;
[0138] A device group status acquisition module 202, configured to generate a capability map of the Internet of Things device group according to the real-time environmental data of multiple Internet of Things devices collected by a distributed sensor network;
[0139] A task priority matrix generation module 203, configured to generate a three-dimensional task priority matrix for the task to be processed based on a pre-established dynamic task evaluation model according to the task to be processed, the real-time load status of multiple robots, and the capability map of the Internet of Things device group;
[0140] A robot-device allocation module 204, configured to use a resource matching algorithm of reinforcement learning and allocate a target robot-Internet of Things device combination for the task to be processed based on the three-dimensional task priority matrix of the task to be processed.
[0141] Preferably, the process of establishing the dynamic task evaluation model includes:
[0142] Based on the capability map of the Internet of Things device group, calculating the topological structure entropy value and dynamic interference factor of the environmental obstacle, and obtaining the real-time environmental complexity of the Internet of Things device group;
[0143] Construct a multi-dimensional capability vector based on the specification parameters and historical performance data of the Internet of Things devices to obtain the capability quantization vector of the Internet of Things devices;
[0144] Based on the real-time environmental complexity and the capability quantization vector, establish a dynamic task evaluation model.
[0145] Preferably, the expression of the three-dimensional task priority matrix is as follows:
[0146] P(t, s, r) = α·T(t) + β·S(s) + γ·R(r)
[0147] Where, P(t, s, r) is the three-dimensional task priority matrix, T(t) is the time dimension priority function, S(s) is the space dimension priority function, R(r) is the resource dimension priority function, α is the time dimension priority weight coefficient, β is the space dimension priority weight coefficient, and γ is the resource dimension priority weight coefficient.
[0148] Preferably, the expression of the time dimension priority function T(t) is as follows:
[0149]
[0150] Where, d i is the deadline of the task to be processed, t now is the current system time, ε is to prevent zero minimum value, λ is the decay factor, ω i is the remaining workload of the task to be processed;
[0151] The expression of the time dimension priority weight coefficient α is as follows:
[0152]
[0153] Where, α0 is the initial time dimension weight coefficient, d i is the deadline of the task to be processed, t now is the current system time, τ critical is the preset critical time threshold;
[0154] The expression of the space dimension priority function S(s) is as follows:
[0155]
[0156] Where, C path is the dynamic path cost, (x r , y r ) is the robot position coordinate, (x t , y t ) is the task target position coordinate, and σ is the space decay radius; is the sum of all edges (path segments) on the path, v r is the moving speed of the robot, μ is the obstacle penalty coefficient, and obstacle_density(e) is the obstacle density in the area where the path segment e is located. is the time cost;
[0157] The expression of the resource dimension priority function R(r) is as follows:
[0158]
[0159] where n is the total number of resource types, ω i is the weight of the i-th resource type, and η is the over-capacity reward coefficient; is the normalized task demand value of the i-th resource, is the normalized device capacity value of device j under the i-th resource, is the total margin of the normalized device capacity of device j under the i-th resource and the normalized task demand value of the i-th resource, and sigmoid maps the total margin to the interval (0, 1); d i is the demand value of the task for the i-th resource, and max(D) is the maximum value in the task demand vector in; is the demand value of all resource types, cj i is the capacity value of the Internet of Things device j for the i-th resource, and max(c j ) is the maximum value of the Internet of Things device capacity vector ; is the capacity value of device j for all resource types,
[0160] Preferably, the system further includes a dynamic resource allocation adjustment module for:
[0161] During the execution of the pending work task by the target robot-Internet of Things device combination, monitor the state parameters of the Internet of Things device group, and dynamically adjust the resource allocation strategy through an adaptive feedback mechanism.
[0162] Preferably, the dynamic resource allocation adjustment module is specifically used for:
[0163] Establish a composite evaluation index including task completion timeliness, device reliability, and energy consumption efficiency;
[0164] Based on the composite evaluation index, optimize the control parameters based on the deep deterministic policy gradient algorithm, and adjust the resource allocation strategy through the control parameters.
[0165] Preferably, the expression of the composite evaluation index is as follows:
[0166]
[0167] Among them, Q is the composite evaluation index value, T complete is the task completion timeliness value, E consumed is the energy efficiency value, S reliability is the reliability index of the Internet of Things device, μ1 is the task completion timeliness weight coefficient, μ2 is the energy efficiency weight coefficient, and μ3 is the Internet of Things device reliability weight coefficient;
[0168] The expression of the deep deterministic policy gradient algorithm is as follows:
[0169]
[0170] Among them, π θ (a|s) is the parameterized policy function, θ is the policy network parameter, is the objective function to be maximized through the policy network parameter θ, Ε[Q(s,a)] is the expected value of the action value function Q(s,a), λ is the policy update constraint coefficient, KL(π θ ||π θold ) is the KL divergence between the old and new policies.
[0171] Embodiment 3:
[0172] Based on the same inventive concept, as Figure 3 shown, the present invention also provides an electronic device, which may be a computer device, a single-chip microcomputer device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, the processor, and the transceiver component are connected by a bus; the memory can be used to store an execution program, and an exemplary execution program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, and the data can be called and / or modified when the instructions are executed.
[0173] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the readable storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of a method for collaborative work of multiple robots and Internet of Things devices in the above embodiments.
[0174] Embodiment 4:
[0175] Based on the same inventive concept, the present invention also provides a readable storage medium, specifically an electronic device-readable storage medium (Memory). The electronic device-readable storage medium is a memory device in the electronic device, used to store programs and data. It can be understood that the readable storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more execution programs (including program codes). It should be noted that the storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the storage medium can be loaded and executed by the processor to implement the steps of a method for collaborative work of multiple robots and Internet of Things devices in the above embodiments.
[0176] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0177] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0178] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0179] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operating steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0180] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit its protection scope. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that after reading the present invention, various changes, modifications, or equivalent replacements can still be made to the specific implementation manners of the application. However, these changes, modifications, or equivalent replacements are all within the protection scope of the pending claims of the application.
Claims
1. A collaborative working method for multiple robots and Internet of Things devices, characterized in that, Including: Obtain the tasks to be processed and the real-time load status of multiple robots; Generate a capability map of the Internet of Things device group based on the real-time environmental data of multiple Internet of Things devices collected by the distributed sensor network; Based on a pre-established dynamic task evaluation model, generate a three-dimensional task priority matrix for the task to be processed according to the task to be processed, the real-time load status of multiple robots, and the capability map of the Internet of Things device group; Adopt a resource matching algorithm based on reinforcement learning, and allocate a target robot-Internet of Things device combination for the task to be processed based on the three-dimensional task priority matrix of the task to be processed.
2. The method according to claim 1, wherein The process of establishing the dynamic task evaluation model includes: Based on the capability map of the Internet of Things device group, calculate the environmental obstacle topology structure entropy value and the dynamic interference factor to obtain the real-time environmental complexity of the Internet of Things device group; Construct a multi-dimensional capability vector based on the specification parameters and historical efficiency data of the Internet of Things devices to obtain the capability quantization vector of the Internet of Things devices; Based on the real-time environmental complexity and the capability quantization vector, establish a dynamic task evaluation model.
3. The method according to claim 2, wherein The expression of the three-dimensional task priority matrix is as follows: P(t, s, r) = α·T(t) + β·S(s) + γ·R(r) Where, P(t, s, r) is the three-dimensional task priority matrix, T(t) is the time dimension priority function, S(s) is the space dimension priority function, R(r) is the resource dimension priority function, α is the time dimension priority weight coefficient, β is the space dimension priority weight coefficient, and γ is the resource dimension priority weight coefficient.
4. The method according to claim 3, wherein The expression of the time dimension priority function T(t) is as follows: Among them, d i is the deadline of the task to be processed, t now is the current system time, ε is to prevent zero minimum value, λ is the attenuation factor, ω i is the remaining workload of the task to be processed; The expression of the time dimension priority weight coefficient α is as follows: Among them, α0 is the initial time dimension weight coefficient, d i is the deadline of the task to be processed, t now is the current system time, τ critical is the preset critical time threshold; The expression of the space dimension priority function S(s) is as follows: Among them, C path is the dynamic path cost, (x r , y r ) is the robot's position coordinates, (x t , y t ) is the task target position coordinates, and σ is the spatial decay radius; is the sum of all edges (path segments) on the path, v r is the moving speed of the robot, μ is the obstacle penalty coefficient, and obstacle_density(e) is the obstacle density in the area where the path segment e is located. is the time cost; The expression of the resource dimension priority function R(r) is as follows: where n is the total number of resource types, ω i is the weight of the i-th resource type, and η is the over-capacity reward coefficient; is the normalized task demand value for the i-th resource, is the normalized device capacity value of device j under the i-th resource, is the total margin of the normalized device capacity and the normalized task demand value of device j under the i-th resource, and sigmoid maps the total margin to the interval (0, 1); d i is the demand value of the task for the i-th resource, and max(D) is the maximum value in the task demand vector ; is the demand value of all resource types, c ji is the capacity value of the IoT device j for the i-th resource, and max(c j ) is the maximum value of the IoT device capacity vector ; is the capacity value of all resource types of device j, 5. The method according to claim 1, wherein After allocating the target robot-Internet of Things device combination for the task to be processed, the method further includes: During the execution of the task to be processed by the target robot-Internet of Things device combination, monitor the status parameters of the Internet of Things device group, and dynamically adjust the resource allocation strategy through an adaptive feedback mechanism.
6. The method according to claim 5, characterized in that, Dynamically adjusting the resource allocation strategy through the adaptive feedback mechanism includes: Establish a composite evaluation index including task completion timeliness, device reliability, and energy consumption efficiency; Based on the composite evaluation index, optimize the control parameters based on the deep deterministic policy gradient algorithm, and adjust the resource allocation strategy through the control parameters.
7. The method according to claim 6, wherein The expression of the composite evaluation index is as follows: Among them, Q is the composite evaluation index value, T complete is the task completion timeliness value, E consumed is the energy efficiency value, S reliability is the reliability index of Internet of Things devices, μ1 is the weight coefficient of task completion timeliness, μ2 is the weight coefficient of energy efficiency, and μ3 is the weight coefficient of Internet of Things device reliability; The expression of the deep deterministic policy gradient algorithm is as follows: where, π θ (a|s) is a parameterized policy function, θ is the policy network parameter, is the objective function to be maximized by the policy network parameter θ, Ε[Q(s,a)] is the expected value of the action-value function Q(s,a), λ is the policy update constraint coefficient, KL(π θ ||π θold ) is the KL divergence between the old and new policies.
8. A collaborative working system for multiple robots and Internet of Things devices, characterized in that, Including: An original data acquisition module for obtaining the tasks to be processed and the real-time load status of multiple robots; A device group status acquisition module for generating a capability map of the Internet of Things device group according to the real-time environmental data of multiple Internet of Things devices collected by the distributed sensor network; A task priority matrix generation module, configured to generate a three-dimensional task priority matrix for the to-be-processed task based on a pre-established dynamic task evaluation model, according to the to-be-processed task, the real-time load status of multiple robots, and the capability map of the Internet of Things device group. A robot-device allocation module, configured to use a resource matching algorithm of reinforcement learning to allocate a target robot-Internet of Things device combination for the to-be-processed task based on the three-dimensional task priority matrix of the to-be-processed task.
9. An electronic device, characterized in that, Comprising: At least one processor and a memory; The memory and the processor are connected by a bus; The memory is configured to store one or more programs; When the one or more programs are executed by the at least one processor, the collaborative working method of multiple robots and Internet of Things devices according to any one of claims 1 to 7 is implemented.
10. A readable storage medium, characterized in that, A program is stored thereon, and when the program is executed, the collaborative working method of multiple robots and Internet of Things devices according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Task allocation method and device, electronic equipment and computer program
CN118796441A
Oil depot tank field inspection robot task allocation method and system based on Internet of Things
CN119417192A
Wireless spectrum intelligent allocation and edge computing cooperation method
CN119562364A
Multi-underwater robot cooperative control system and method based on predictive control
CN119596712A
Multi-task collaborative robot scheduling method and system for intelligent factory
CN119596882A