Robot cluster individual behavior control method for space-time marching task
By training individual behavior control strategy models and dynamically adjusting, the problem that individuals in robot clusters cannot achieve the expected results of the task in space-time travel tasks is solved, automated control is achieved, and task completion efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510319286.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-20
AI Technical Summary
In space-time travel tasks, individuals in robot clusters may not be able to fully achieve the expected results of the task due to factors such as control accuracy error, power limitations and environmental changes, resulting in the failure of space-time coordination of cluster tasks, affecting the task completion efficiency and accuracy.
By obtaining the historical actual action space change data of the robot cluster, training the individual behavior control strategy model, predicting the individual's control strategy for performing tasks, and adjusting based on preset dynamic adjustment strategies, automatically determining the individual behavior control strategy to reduce manual intervention.
It realizes automated individual behavior control, reduces the cost of manual intervention, improves response speed, adapts to the complexity of large-scale cluster collaboration tasks, and improves task completion efficiency and accuracy.
Smart Images

Figure CN120178746A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic control, and particularly to a method for controlling the individual behavior of a robot cluster for spatio-temporal travel tasks. Background Art
[0002] In the context of the development of modern industry and intelligentization, spatio-temporal travel tasks, as a typical group cooperation task, widely exist in application scenarios such as robot clusters, unmanned aerial vehicle groups, autonomous driving vehicle fleets, and collaborative robots in intelligent manufacturing systems. These tasks require each individual in the group to complete high-complexity task objectives through mutual cooperation in a spatio-temporally dynamically changing environment. For example, the dynamic path planning of a robot cluster on an industrial production line, the precise cooperative flight of an unmanned aerial vehicle group in a complex airspace, and even the collaborative operation of automated equipment in a large-scale logistics distribution network all require individuals to respond to task requirements in real time in terms of time and space and make coordinated actions.
[0003] During each task process, each individual in the cluster will perform actions according to the task objectives. However, due to the influence of various factors, such as control precision errors, power limitations, environmental changes, etc., the individuals in the group may not fully achieve the expected task effects. Such deviations may lead to the spatio-temporal coordination failure of the overall cluster task, thereby affecting the task completion efficiency and accuracy of the cluster.
[0004] To make up for these deficiencies, traditional methods rely on expert manual intervention. By analyzing the deviations in task execution, the unqualified individuals are adjusted, such as correcting the path, optimizing task allocation, or re-planning the motion trajectory, so that they can achieve the predetermined task effects. Although this method can ensure the completion of tasks to a certain extent, it has obvious deficiencies, such as high manual intervention costs, slow response speed, and difficulty in adapting to the complexity of large-scale cluster cooperation tasks, and so on. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a method for controlling the individual behavior of a robot cluster for spatio-temporal travel tasks. The technical solution of the present invention is as follows:
[0006] A method for controlling the individual behavior of a robot cluster for spatio-temporal travel tasks, which includes:
[0007] S1, obtaining the state of the i-th individual when performing the j-th task during the execution of the target task by the robot cluster
[0008] S2, inputting into a pre-trained individual behavior control strategy model, and outputting the predicted control strategy for the i-th individual to perform the j-th task by the individual behavior control strategy model
[0009] S3. According to the preset individual dynamic adjustment strategy and Determine the new control strategy for the i-th individual to execute the j-th task
[0010] S4. Control the i-th individual to Execute the j-th task and collect the actual action space change of the i-th individual when executing the j-th task.
[0011] Optionally, before the S2 inputs Into the pre-trained individual behavior control strategy model, it further includes:
[0012] S01. Obtain the actual action space change matrix of the robot cluster when executing the target task in history;
[0013] S02. Construct the spatial weight vector of the i-th individual when executing the target task in history;
[0014] S03. Construct the task number weight vector of the i-th individual when executing the j-th task of the target task in history;
[0015] S04. Construct a spatio-temporal weight matrix centered on the i-th individual according to the spatial weight vector and the task number weight vector;
[0016] S05. Determine the spatio-temporal actual action space change matrix centered on the i-th individual;
[0017] S06. Determine the behavior trajectory of the i-th individual when the robot cluster executes the target task in history according to the spatio-temporal weight matrix and the spatio-temporal actual action space change matrix;
[0018] S07. Train the individual behavior control strategy model according to the behavior trajectory of the i-th individual until the trained individual behavior control strategy model is obtained.
[0019] Optionally, the S01 includes:
[0020] S011. Determine that the robot cluster has n individuals, and each individual executes m tasks when executing the target task in history;
[0021] S012. Determine the actual action space change of the i-th individual when executing the j-th task when executing the target task in history Where Represents the angle change and distance change of individual i in three-dimensional space;
[0022] S013. Construct the actual action space change matrix of the robot cluster according to the actual action space changes of each individual in the robot cluster when completing each task, which is expressed as:
[0023]
[0024] where represents the actual action space change when the i-th individual in the robot cluster executes the j-th task of the target task in history.
[0025] Optionally, the S02 includes:
[0026] S021. Determine all individuals whose distance from the i-th individual is less than the Euclidean distance limit value The number of such individuals is x;
[0027] S022. Calculate the spatial position weight of each of the x individuals for the i-th individual according to the Euclidean distance between each of the x individuals and the i-th individual and the sum of their Euclidean distances from the i-th individual;
[0028] S023. Reorder and number the x individuals according to the spatial position weight of each individual for the i-th individual to obtain the spatial weight vector, which is expressed as: and where the principle of reordering and numbering is: taking the i-th individual as the center, if the l-th individual is compared with the i-th individual and l > i, then the l-th individual is arranged on one side of the i-th individual, and the spatial position weight is larger, the closer it is to the i-th individual; if l < i, then the l-th individual is arranged on the other side of the i-th individual, and the spatial position weight is larger, the closer it is to the i-th individual.
[0029] Optionally, the S03 includes:
[0030] S031. Determine the number of tasks whose difference in the number of tasks between the task executed before the j-th task of the target task executed by the i-th individual in history and the j-th task is less than the task number difference limit value The number of such tasks is y in total;
[0031] S032. Assign weights to the y tasks to obtain the task number weight vector when the i-th individual executes the j-th task of the target task in history, which is expressed as:
[0032] Optionally, the S04 includes:
[0033] S041. Determine that the weight of the row where the i-th individual is located in the spatio-temporal weight matrix is
[0034] S042, according to Determine the weight of the row where the x-th individual related to the i-th individual space is located;
[0035] S403, Combine the weight of the row where the i-th individual is located and the weight of the row where the x-th individual related to the i-th individual space is located to obtain a spatio-temporal weight matrix, expressed as:
[0036]
[0037] Optionally, the S05 includes:
[0038] According to the actual action space change matrix of the robot cluster, the spatio-temporal weight matrix, and the correspondence before and after reordering and renumbering the x individuals related to the i-th individual, determine the spatio-temporal actual action space change matrix, expressed as:
[0039]
[0040] where f(l) = x; l = f -1 (x), and f represents the function of reordering and renumbering.
[0041] Optionally, the S06 includes:
[0042] S061, Take the spatio-temporal weight matrix and the spatio-temporal actual action space change matrix together as the state of the i-th individual when performing the j-th task Or, take the product of the spatio-temporal weight matrix and the spatio-temporal actual action space change matrix as the state of the i-th individual when performing the j-th task;
[0043] S062, Take the spatio-temporal actual action space change of the i-th individual when performing the j-th task in the spatio-temporal actual action space change matrix as the control strategy
[0044] S063, Combine the states and control strategies of the i-th individual when performing tasks multiple times to obtain the behavior trajectory of the i-th individual in the robot cluster when performing the target task in history, expressed as
[0045] Optionally, the S3 includes:
[0046] S31, Obtain and the control strategy of the i-th individual when performing the (j - 1)-th task
[0047] S32, According to and Dynamically adjust to obtain
[0048] Optionally, the S32 includes:
[0049] If equals then control equals
[0050] If is greater than then calculate and the difference between them as the adjustment difference, and control equals the sum of and the adjustment difference;
[0051] If is less than then calculate and the difference between them as the adjustment difference, and control equals the difference between and the adjustment difference.
[0052] All of the above optional technical solutions can be combined arbitrarily, and the present invention does not elaborate on the structures after combination one by one.
[0053] By means of the above solutions, the beneficial effects of the present invention are as follows:
[0054] By training to obtain the individual behavior control strategy model and individual behavior trajectory based on the actual data obtained from the historical execution of the target task by the robot cluster, when the subsequent robot cluster executes the target task, after the i-th individual obtains its state of the j-th task execution based on the individual behavior trajectory, the state is input into the individual behavior control strategy model, and the individual behavior control strategy model outputs a predicted control strategy, and dynamically adjusts the predicted control strategy based on the preset individual dynamic adjustment strategy, providing a way to automatically determine the individual behavior control strategy. By determining the individual behavior control strategy in this way, it is possible to avoid manual intervention, with the advantages of low cost, fast response speed, and the ability to adapt to the complexity of large-scale cluster cooperation tasks.
[0055] The above description is only an overview of the technical solutions of the present invention. In order to be able to understand the technical means of the present invention more clearly and implement them in accordance with the content of the description, the following takes the preferred embodiments of the present invention and combines with the drawings to elaborate in detail as follows. Description of the Drawings
[0056] Figure 1 is a flowchart of the method provided by the embodiment of the present invention.
[0057] Figure 2It is a flowchart for training an individual behavior control strategy model in the method provided by an embodiment of the present invention. Detailed implementation manners
[0058] The following will further describe in detail the specific implementation manners of the present invention in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0059] The individual behavior control method for a robot cluster facing spatio-temporal travel tasks provided by an embodiment of the present invention can be implemented by any electronic device with computing functions, such as a PC, a mobile terminal, or a server, etc. The robot cluster in the embodiment of the present invention can be a group of drones, an autonomous driving fleet, and collaborative robots in an intelligent manufacturing system, etc.
[0060] The implementation background of the embodiment of the present invention is that the target task executed by the robot cluster each time is the same, and when executing the target task each time, the sub-tasks distributed to each individual in the robot cluster are also the same. Assume that the robot cluster has a total of n individuals, and when the robot cluster executes the target task, each individual needs to execute the task m times.
[0061] In each task, each individual in the robot cluster will act according to the task target. However, due to the influence of various factors, such as control precision error, insufficient power, environmental changes, etc., the individuals in the group often fail to fully achieve the task expected effect. When this situation occurs in the historical execution of the target task by the robot cluster, the individual behavior control method adopted is: the individuals that do not achieve the task expected effect are adjusted in a timely manner through expert intervention to make them achieve the task predetermined effect. This individual behavior control method requires manual intervention, and the manual intervention method has problems such as high cost, slow response speed, and difficulty in adapting to the complexity of large-scale cluster cooperation tasks.
[0062] To solve these problems, based on the historical data obtained by manual intervention for the target task, the embodiment of the present invention trains an individual behavior control strategy model, and predicts the predicted control strategy for the i-th individual to execute the j-th task when the current robot cluster executes the target task through the individual behavior control strategy model And after dynamically adjusting the predicted control strategy it is used to control the i-th individual to execute the j-th task, so that the individual behavior control method does not require manual intervention.
[0063] Based on the above content, as Figure 1 shown, the individual behavior control method for a robot cluster facing spatio-temporal travel tasks provided by an embodiment of the present invention includes the following steps S1 to S4:
[0064] S1. Obtain the state of the \(i\)-th individual of the robot cluster when performing the \(j\)-th task during the execution of the target task.
[0065] Specifically, the obtaining method can be obtained from the behavior trajectory of the \(i\)-th individual during the historical execution of the target task by the robot cluster. Regarding the determination method of the behavior trajectory of the \(i\)-th individual, and its specific content will be explained in the following embodiments and will not be described here for the time being.
[0066] S2. Input into the pre-trained individual behavior control strategy model, and the individual behavior control strategy model outputs the predicted control strategy for the \(i\)-th individual to perform the \(j\)-th task.
[0067] Specifically, the input of the individual behavior control strategy model is the state before a certain individual performs a certain task, and the output is the control strategy predicted by the individual behavior control strategy model for a certain individual to perform a certain task.
[0068] It should be noted that before inputting into the pre-trained individual behavior control strategy model, the individual behavior control strategy model needs to be trained first. In the embodiments of the present invention, when training the individual behavior control strategy model, it can be realized through the following steps S01 to S07:
[0069] S01. Obtain the actual action space change matrix of the robot cluster when performing the target task historically.
[0070] In a specific embodiment, S01 can be realized through the following steps S011 to S013:
[0071] S011. Determine that the robot cluster has \(n\) individuals, and each individual performs \(m\) tasks during the historical execution of the target task.
[0072] S012. Determine the actual action space change of the \(i\)-th individual when performing the \(j\)-th task during the historical execution of the target task where represents the angular change and distance change of individual \(i\) in three-dimensional space, expressed as Assume that its angles before and after the task execution are \(r1\) and \(r2\), and the coordinates are \((x1, y1, z1)\) and \((x2, y2, z2)\) respectively, then It can be expressed by the formula:
[0073]
[0074] S013. Construct an actual action space change matrix for the robot swarm according to the actual action space changes of each individual in the robot swarm when completing each task, which is expressed as:
[0075]
[0076] Wherein, represents the actual action space change when the i-th individual in the robot swarm executes the j-th task of the target task in history.
[0077] The success of the group cooperation task depends on the synergy of each individual with the surrounding environment and other individuals. In such tasks, individuals usually influence each other, and the current state of an individual is not only related to the states of surrounding individuals but also has a certain correlation with its previous historical states. Therefore, considering the complexity of group behavior, the embodiments of the present invention model and analyze the behavior of individuals from two perspectives of space and time, which are respectively reflected in the following steps S02 and S03.
[0078] S02. Construct a spatial weight vector for the i-th individual when executing the target task in history.
[0079] Specifically, from the perspective of space, the relationship between an individual and other individuals around it is the basis of cooperation and task execution. To reasonably select which individuals jointly act on a region, the embodiments of the present invention construct a spatial weight vector. The spatial weight vector is based on the spatial distance between individuals and determines how many neighboring individuals each individual cooperates with to complete the task. For the i-th individual executing the j-th task, assuming its position is (x i , y i , z i ), the embodiments of the present invention use the Euclidean distance to calculate the distance between individual i and other individuals. Assuming the position of the l-th individual is (x l , y l , z l ), then the Euclidean distance between individual i and individual l is:
[0080]
[0081] The embodiments of the present invention also pre-set a limit value of the Euclidean distance between individuals according to experience If the Euclidean distance between the l-th individual and the i-th individual is greater than or equal to , it is considered that there is no interaction between the l-th individual and the i-th individual. Therefore, all individuals whose distance from the i-th individual is less than can be obtained. The embodiments of the present invention assume that there are a total of x individuals interacting with the i-th individual.
[0082] Based on the above, S02 is implemented by, but not limited to, the following steps S021 to S023:
[0083] S021, determine all individuals whose distance from the i-th individual is less than the Euclidean distance limit value The number of such individuals is x.
[0084] S022, calculate the spatial position weight of each of the x individuals relative to the i-th individual according to the Euclidean distance between each of the x individuals and the i-th individual and the sum of their Euclidean distances from the i-th individual.
[0085] Specifically, this step is expressed by the formula:
[0086]
[0087] In the formula, represents the Euclidean distance between the l-th individual and the i-th individual among the x individuals screened according to the Euclidean distance limit value in the j-th task; represents the sum of the Euclidean distances between the x individuals screened according to the Euclidean distance limit value and the i-th individual in the j-th task, represents the spatial position weight of the l-th individual relative to the i-th individual in the j-th task.
[0088] S023, reorder and number the x individuals according to the spatial position weight of each individual relative to the i-th individual to obtain a spatial weight vector, expressed as: And represents the spatial position weight of the x-th individual relative to the i-th individual after reordering and numbering the x individuals screened according to the Euclidean distance limit value according to the spatial position weight, represents the spatial weight vector of the i-th individual.
[0089] Among them, the principle of reordering and numbering is: assume there are a total of x + 1 individuals including the x individuals screened according to the Euclidean distance limit value and the i-th individual. For example, the numbers of the x individuals in the robot cluster are [1, 3,..., l, i]. Taking the i-th individual as the center, if the l-th individual and the i-th individual are compared and l > i, then the l-th individual is on one side of the i-th individual, and the larger the spatial position weight , the closer it is to the i-th individual; if l < i, then the l-th individual is on the other side of the i-th individual, and the larger the spatial position weight , the closer it is to the i-th individual.
[0090] For each of the x individuals selected according to the Euclidean distance limit value such sorting is performed, and then the x individuals are re-sorted from the side less than i to the side greater than i, that is, [1, 2, 3, …, x]. The change in sorting can be expressed by the following mathematical relationship:
[0091] f(l) = x; l = f -1 (x);
[0092] where f represents the function of re-sorting and numbering, indicating that the individual originally numbered l in the robot cluster is numbered x after re-sorting and numbering.
[0093] Through the transformation of the function f of re-sorting and numbering, the re-sorting and numbering of the x individuals selected according to the Euclidean distance limit value can be completed, and at the same time, it is also known that after the x individuals selected according to the Euclidean distance limit value are re-sorted, their numbers are sorted in the original entire population, that is, a one-to-one correspondence relationship.
[0094] S03, construct the task number weight vector when the i-th individual executes the j-th task in the history of the target task.
[0095] From the perspective of the number of tasks, the current task of an individual is not only affected by other individuals in the spatial neighborhood, but may also be affected by its historical state (and historical tasks). Therefore, the embodiments of the present invention construct a task number weight vector to describe the influence of previous tasks on the current task. Through the task number weight vector, it is possible to select and analyze how many previous tasks have a significant impact on the current task, so as to better capture the time dependence in the process of individual task execution.
[0096] For the i-th individual and the j-th task, the difference in the number of tasks between the k-th task and the j-th task is j - k times, (k < j). Use to represent the difference in the number of tasks between the j-th task and the k-th task of the i-th individual, and the mathematical formula can be expressed as:
[0097] The embodiments of the present invention set the limit value of the difference in the number of tasks according to experience If the difference in the number of tasks between the k-th task and the j-th task is greater than or equal to then it is considered that the k-th task will not have an impact on the j-th task. Therefore, all tasks with a difference in the number of tasks less than before the j-th task can be obtained. Assuming that there are a total of y tasks less than
[0098] Based on the above, S03 can be implemented through the following steps S031 and S032:
[0099] S031, determine the number of task times between the task executed before the j-th task of the i-th individual's historical execution of the target task and the j-th task, and the difference in the number of task times is less than the task time difference limit value The number of task times is y times; S032, assign weights to the y tasks to obtain the task time weight vector when the i-th individual historically executes the j-th task of the target task, which is expressed as:
[0100] Among them, when assigning weights to the y tasks, methods such as reciprocal weight normalization, linear decay normalization, or exponential decay normalization can be used to reasonably assign weights between different tasks according to requirements. Taking the reciprocal weight normalization method as an example, when assigning weights to the y tasks, the following formula can be used to assign weights to each task:
[0101]
[0102] In the formula, represents the i-th individual, and the weight of the k-th task relative to the j-th task, is the sum of weights relative to the j-th task under the y tasks, and
[0103] S04, construct a spatio-temporal weight matrix centered on the i-th individual according to the spatial weight vector and the task time weight vector.
[0104] Specifically, through the above steps S02 and S03, it is determined which individuals cooperate with the i-th individual to complete the task spatially and how many previous tasks will affect the current task temporally for the i-th individual to execute the j-th task. This step S04 constructs a spatio-temporal weight matrix centered on the i-th individual according to the spatial weight vector and the task time weight vector of the i-th individual executing the j-th task. The spatio-temporal weight matrix is used to express the relative importance of each position.
[0105] In a specific embodiment, S04 can be implemented through the following steps S041 to S043:
[0106] S041, determine that the weight of the row where the i-th individual is located in the spatio-temporal weight matrix is
[0107] S042, according to determine the weight of the row where the x-th individual spatially related to the i-th individual is located;
[0108] S403, Combine the weight sum of the row where the \(i\)-th individual is located with the weight of the row where the \(x\)-th individual related to the \(i\)-th individual's space is located to obtain a spatio-temporal weight matrix, expressed as:
[0109]
[0110] S05, Determine the spatio-temporal actual action space change matrix centered on the \(i\)-th individual.
[0111] Specifically, the spatio-temporal actual action space change matrix is a set of spatio-temporal actual action space changes when the individuals related to the time and space of the \(i\)-th individual's execution of the \(j\)-th task historically executed the target task. In a specific embodiment, when implementing S05, the spatio-temporal actual action space change matrix can be determined according to the actual action space change matrix of the robot cluster, the spatio-temporal weight matrix, and the correspondence relationship before and after reordering and numbering the \(x\) individuals related to the \(i\)-th individual, expressed as:
[0112]
[0113] where \(f(l)=x\); \(l = f\) -1 (x), and \(f\) represents the function of reordering and numbering.
[0114] Based on the fact that the \(i\)-th individual determined by the spatial weight vector and the task number weight vector collaborates with the surrounding \(x\) individuals in space to complete the task, and the previous \(y\) tasks will affect the current task, the spatio-temporal actual action space change matrix centered on the \(i\)-th individual is determined by integrating the actual action space change matrix of the entire robot cluster.
[0115] S06, Determine the behavior trajectory of the \(i\)-th individual when the robot cluster historically executed the target task according to the spatio-temporal weight matrix and the spatio-temporal actual action space change matrix.
[0116] The behavior trajectory of the \(i\)-th individual can characterize the actual action space change of the \(i\)-th individual when executing each task when the robot cluster historically executed the target task. In a specific embodiment, S06 can be implemented through the following steps S061 to S063:
[0117] S061, Take the spatio-temporal weight matrix and the spatio-temporal actual action space change matrix together as the state of the \(i\)-th individual executing the \(j\)-th task Alternatively, take the product of the spatio-temporal weight matrix and the spatio-temporal actual action space change matrix as the state of the \(i\)-th individual executing the \(j\)-th task.
[0118] Specifically, regarding how to determine the state of the \(i\)-th individual executing the \(j\)-th task according to the spatio-temporal weight matrix and the spatio-temporal actual action space change matrix The embodiments of the present invention provide two methods.
[0119] The first method is: taking the spatio-temporal weight matrix and the spatio-temporal actual action space change matrix together as the state of the i-th individual performing the j-th task
[0120] The weight is flexibly variable and has the same dimension as the spatio-temporal actual action space change matrix. Considering retaining more original information, this method takes the spatio-temporal weight matrix and the spatio-temporal actual action space change matrix together as the state of the i-th individual's j-th task, that is This method can enable the model to self-learn and determine the combination method of the spatio-temporal weight matrix and the spatio-temporal actual action space change matrix during the subsequent training process of the individual behavior control strategy model, and may discover some non-intuitive patterns; secondly, this method retains all the original data and weight information and will not cause partial information loss. However, the limitation of this method is that it may lead to an increase in the complexity of model training because the model needs to process more input data, which may increase the training difficulty and computational resource consumption; it may also lead to the need to design a more complex network structure to effectively integrate these two types of information.
[0121] The second method is: taking the product of the spatio-temporal weight matrix and the spatio-temporal actual action space change matrix as the state of the i-th individual performing the j-th task.
[0122] This method synthesizes the spatio-temporal weight matrix and the spatio-temporal actual action space change matrix into a matrix containing the action information after weight adjustment. Taking the matrix containing the action information after weight adjustment as the state of the i-th individual's j-th task, that is Compared with the first method, this method reduces the dimension of the model input and can reduce the complexity of the model; in this way, the weight directly acts on the corresponding action variable, and the adjustment effect of the weight on the action can be intuitively reflected. However, the disadvantage of this method may be that some information is lost because directly multiplying the two matrices may cause some subtle action changes to be amplified or suppressed after applying the weight, thus possibly losing some important information in the original data.
[0123] It should be noted that in specific implementation, calculations other than the product can also be performed on the spatio-temporal weight matrix and the spatio-temporal actual action space change matrix as needed, and the calculation result is taken as the state of the i-th individual performing the j-th task. Regarding the calculation method, the embodiments of the present invention will not be elaborated in detail.
[0124] S062, taking the spatio-temporal actual action space change of the i-th individual performing the j-th task in the spatio-temporal actual action space change matrix as the control strategy
[0125] S063, combine the states and control strategies of the i-th individual performing the task multiple times to obtain the behavior trajectory of the i-th individual when the robot cluster executes the target task in history, denoted as
[0126] S07, train the individual behavior control strategy model according to the behavior trajectory of the i-th individual until the trained individual behavior control strategy model is obtained.
[0127] According to the idea of model learning, in the embodiments of the present invention, each state in the behavior trajectory of the i-th individual is used as a sample in supervised learning, and the corresponding action of the state, that is, the control strategy, is used as a label in supervised learning. The state is regarded as the input of the individual behavior control strategy model, and the output of the individual behavior control strategy model is regarded as the action. Finally, with the expected action, teach the machine to learn the corresponding relationship between the state and the action. Specifically, take as the input of the individual behavior control strategy model, as the output of the individual behavior control strategy model. In this way, the output of the machine is the control strategy that should be output in this state is how much.
[0128] Regarding the specific type of the individual behavior control strategy prediction model, the embodiments of the present invention recommend selecting a neural network architecture that can simultaneously capture temporal dynamic changes and spatial structure characteristics. Suitable networks include ConvLSTM and attention mechanism networks (such as Transformer), which can model long-term and short-term dependencies in the time dimension and extract local and global features in the spatial dimension at the same time. In addition, the embodiments of the present invention can also select a hybrid network architecture, such as combining CNN to process spatial features and LSTM or Transformer to process temporal features, so as to better handle complex spatio-temporal interaction tasks. These networks can effectively learn spatio-temporal features in complex spatio-temporal tasks, thereby improving the adaptability and accuracy of the individual behavior control strategy model.
[0129] Regarding the specific way of training the individual behavior control strategy model, refer to the existing neural network model training methods, and the embodiments of the present invention will not elaborate on this in detail.
[0130] Such as Figure 2 shown, which is the flowchart of training the individual behavior control strategy model in the method provided by the embodiments of the present invention.
[0131] After training the individual behavior control strategy model through the above steps S01 to S07 and obtaining based on the behavior trajectory of the i-th individual when the robot cluster executes the target task in history, then Input the pre-trained individual behavior control strategy model, and the individual behavior control strategy model outputs the predicted control strategy for the i-th individual to execute the j-th task. However, since the current task may be affected by the previous task during the execution of the task by the individual, therefore, in order to minimize the influence of the previous task on the current task, the embodiment of the present invention further performs dynamic adjustment through the following step S3 on it.
[0132] S3. According to the preset individual dynamic adjustment strategy and determine the new control strategy for the i-th individual to execute the j-th task
[0133] In a specific embodiment, the S3 includes: S31. Obtain and the control strategy when the i-th individual executes the (j - 1)-th task S32. According to and dynamically adjust to obtain
[0134] wherein, is the predicted control strategy for the (j - 1)-th task output by the individual behavior control strategy model after inputting , and is the actual action space change when the i-th individual executes the (j - 1)-th task.
[0135] Specifically, the S32 includes:
[0136] If is equal to then control to be equal to
[0137] If is greater than then calculate the difference between and as the adjustment difference, and control to be equal to plus the adjustment difference;
[0138] If is less than then calculate the difference between and as the adjustment difference, and control to be equal to minus the adjustment difference.
[0139] Specifically, it can be expressed as:
[0140] If the predicted control strategy predicted by the individual behavior control strategy model for the i-th individual in the (j - 1)-th task is equal to the actual change in the action space then it is considered that the actual change in the action space of individual i in performing the (j - 1)-th task is in place.
[0141] If the predicted control strategy predicted by the individual behavior control strategy model for the i-th individual in the (j - 1)-th task is greater than the actual change in the action space then it is considered that the actual change in the action space of individual i in performing the (j - 1)-th task is less than the predicted control strategy Therefore, in performing the j-th task, the embodiment of the present invention adds this part with less actual change in the action space to the predicted control strategy of the j-th task on it.
[0142] If the predicted control strategy predicted by the individual behavior control strategy model for the i-th individual in the (j - 1)-th task is less than the actual change in the action space then it is considered that the actual change in the action space of individual i in performing the (j - 1)-th task is more than the predicted control strategy Therefore, in performing the j-th task, the embodiment of the present invention reduces this part with more actual change in the action space to the predicted control strategy of the j-th task on it.
[0143] Through such an individual dynamic adjustment strategy, and are compared to decide how to dynamically adjust according to the previous state information, which is more in line with the adjustment logic of the industrial site.
[0144] It should be noted that when dynamically adjusting through the above individual dynamic adjustment strategy when j = 1, that is, for the first task, directly take equal to
[0145] S4, control the i-th individual to execute the j-th task according to and collect the actual change in the action space of the i-th individual in performing the j-th task.
[0146] As new actual action space changes are continuously collected, embodiments of the present invention utilize the previously mentioned individual behavior control learning method to continuously update and expand the dataset of state-action pairs. The specific operation is to merge the newly collected dataset with the existing dataset and retrain the individual behavior control policy model using the merged data. During this process, the individual behavior control policy model is tested on the validation set to ensure error convergence and achieve continuous optimization of the performance of the individual behavior control policy model. In this way, the database will gradually improve, and the individual behavior control policy prediction model is trained to handle more diverse data, and its prediction and control effects will also be enhanced accordingly. This automatic learning mechanism ensures the gradual evolution of the model over time Continuously improving its accuracy and reliability in practical applications.
[0147] In summary, for the method provided by the embodiments of the present invention, after training the individual behavior control policy prediction model according to historical data, in practical applications, for predicting the i-th individual and the j-th task, according to the method of making the individual behavior trajectory mentioned above, the state of the i-th individual and the j-th task is obtained Input it into the individual behavior control policy model, and the individual behavior control policy model will output the predicted control policy for the i-th individual and the j-th task Then, according to the preset individual dynamic adjustment policy, a new control policy for the i-th individual and the j-th task is obtained based on the previous actual state Send it to the i-th individual through instructions. After waiting for the individual to complete the action, collect the actual action space change of the i-th individual and the j-th task for subsequent production of the individual behavior trajectory.
[0148] Embodiments of the present invention focus on individual behavior control of a robot cluster for spatio-temporal travel tasks, and propose a method for individual behavior control of a robot cluster for spatio-temporal travel tasks. By recording the spatio-temporal behavior trajectories and adjustment methods of individuals in each task, the machine can imitate the process of experts adjusting individual behaviors, gradually learn and master reasonable individual control strategies. On this basis, the predicted control policy learned by the individual behavior control policy model is compared with the actual action space change to optimize the predicted control policy and achieve dynamic optimization of individuals. This method can not only significantly improve the individual response efficiency and task completion accuracy in spatio-temporal travel tasks, but also effectively cope with spatio-temporal task changes in complex dynamic environments, providing theoretical support and practical reference for the behavior control of group cooperation tasks.
[0149] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A robot cluster individual behavior control method for space-time travel tasks, characterized in that: include: S1, obtain the state of the i-th individual when the robot cluster performs the target task when performing the j-th task S2, Input the pre-trained individual behavior control strategy model, and the individual behavior control strategy model outputs the prediction control strategy for the i-th individual to perform the j-th task S3, dynamically adjust strategies and Determine the new control strategy for the i-th individual to perform the j-th task S4, control the i-th individual according to Perform the j-th task and collect the actual action space changes of the i-th individual performing the j-th task.
2. The robot cluster individual behavior control method for space-time travel tasks according to claim 1 is characterized in that: The S2 will Before inputting the pre-trained individual behavior control strategy model, it also includes: S01, obtain the actual action space change matrix of the robot cluster when the robot cluster performs the target task in history; S02, construct the spatial weight vector of the i-th individual when performing the target task in history; S03, construct the task frequency weight vector of the jth task when the i-th individual performs the target task in history; S04, constructing a spatiotemporal weight matrix centered on the i-th individual according to the spatial weight vector and the task frequency weight vector; S05, determining the space-time actual action space change matrix centered on the i-th individual; S06, determining the behavior trajectory of the i-th individual when the robot cluster performs the target task in history according to the spatiotemporal weight matrix and the spatiotemporal actual action space change matrix; S07, training the individual behavior control strategy model according to the behavior trajectory of the i-th individual until a trained individual behavior control strategy model is obtained.
3. The robot cluster individual behavior control method for space-time travel tasks according to claim 2 is characterized in that: The S01 includes: S011, determine that the robot cluster has a total of n individuals, and each individual has performed m tasks when performing the target task in the past; S012, determine the actual action space change of the i-th individual when performing the j-th task in the history of performing the target task in Represents the angle change and distance change of individual i in three-dimensional space; S013, according to the actual action space change of each individual in the robot cluster when completing each task, the actual action space change matrix of the robot cluster is constructed, which is expressed as: in, It represents the actual action space change of the i-th individual in the robot cluster when it performs the target task j times in history.
4. The robot cluster individual behavior control method for space-time travel tasks according to claim 2 or 3 is characterized in that: The S02 includes: S021, determine whether the distance between the i-th individual and the i-th individual is less than the Euclidean distance limit All individuals of , the number of individuals is x; S022, calculate the spatial position weight of each of the x individuals with respect to the i-th individual according to the sum of the Euclidean distance between each of the x individuals and the i-th individual and the Euclidean distance between each of the x individuals and the i-th individual; S023, re - sort and re - number the x individuals according to the spatial position weights of each individual with respect to the i - th individual, obtaining a spatial weight vector, expressed as: and wherein, the principle of re - sorting and re - numbering is: taking the i - th individual as the center, if the l - th individual is compared with the i - th individual and l > i, then the l - th individual is arranged on one side of the i - th individual, and the spatial position weight is larger, the closer it is to the i - th individual; if l < i, then the l - th individual is arranged on the other side of the i - th individual, and the spatial position weight is larger, the closer it is to the i - th individual.
5. The robot cluster individual behavior control method for space-time travel tasks according to claim 4 is characterized in that: The S03 includes: S031, determine that the difference between the number of tasks performed before the jth task of the i-th individual in the history of performing the target task and the j-th task is less than the task number difference limit value The number of tasks is y times in total; S032, assign weights to y tasks, and obtain the task frequency weight vector of the jth task of the i-th individual in history when performing the target task, expressed as:
6. The robot cluster individual behavior control method for space-time travel tasks according to claim 5 is characterized in that: The S04 includes: S041, determine the weight of the row where the i-th individual in the spatiotemporal weight matrix is located S042, according to Determine the weight of the row of the xth individual associated with the i-th individual space; S403, combining the weight of the row where the i-th individual is located and the weight of the row where the x-th individual is located that is spatially related to the i-th individual, to obtain a spatiotemporal weight matrix, which is expressed as:
7. The robot cluster individual behavior control method for space-time travel tasks according to claim 6 is characterized in that: The S05 includes: According to the actual action space change matrix of the robot cluster, the space-time weight matrix and the corresponding relationship before and after the x individuals related to the i-th individual are re-sorted and numbered, the space-time actual action space change matrix is determined, which is expressed as: Where f(l) = x; l = f -1 (x), f represents the reordered and numbered function.
8. The robot cluster individual behavior control method for space-time travel tasks according to claim 7 is characterized in that: The S06 includes: S061, take the spatiotemporal weight matrix and the spatiotemporal actual action space change matrix together as the state of the i-th individual performing the j-th task Alternatively, the product of the spatiotemporal weight matrix and the spatiotemporal actual action space change matrix is taken as the state of the i-th individual performing the j-th task; S062, taking the spatial change of the actual spatial action of the i-th individual performing the j-th task in the spatial change matrix of the actual spatial action as the control strategy S063, combining the states and control strategies of the i-th individual when performing the task multiple times, obtains the behavior trajectory of the i-th individual when the robot cluster performs the target task in history, expressed as 9. The robot cluster individual behavior control method for space-time travel tasks according to claim 1 is characterized in that: The S3 includes: S31, obtain and the control strategy when the i-th individual performs the j-1-th task S32, according to and Dynamic Adjustment get 10. The robot cluster individual behavior control method for space-time travel tasks according to claim 9 is characterized in that: The S32 includes: if equal Then control equal if Greater than Then calculate and The difference is used as the adjustment difference and controls equal and the sum of the adjusted differences; if Less than Then calculate and The difference is used as the adjustment difference and controls equal The difference from the adjusted difference.
Citation Information
Patent Citations
Power station boiler combustion modeling and optimization method facing mass high-dimensional data
CN108038306A
Robot control signal determination method and device and storage medium
CN113134834A
Autonomous learning and intelligent decision-making method for hydraulic support cluster control behaviors
CN119195827A
Empirical game theoretic system and method for adversarial decision analysis
US12223430B1
Adaptive pattern recognition based controller apparatus and method and human-factored interface therefore
US20070061022A1