A manipulator trajectory optimization method and system based on adversarial learning
By discrete the robotic arm motion space into voxels and using adversarial learning methods to generate high-quality robotic arm trajectories, the problem of high computational cost in three-dimensional space is solved, and a more optimized trajectory generation is achieved.
Patent Information
- Application Number
- CN202510554278.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The prior art has studied the optimization of robotic arm trajectory in three-dimensional space, and traditional methods are expensive to calculate under high dimensionality and complex constraints, and the adversarial learning of two-dimensional image training cannot be directly applied to three-dimensional robotic arm trajectory optimization.
The robotic arm motion space is discretized into voxels, and the voxel sequence to be optimized is generated using the RRT algorithm. The voxel sequence labels are calculated through different neighborhood radii and step sizes. The adversarial learning is combined with the discriminator and the generator to optimize the generator to generate high-quality robotic arm trajectory.
Generate shorter and smoother robotic arm trajectories in three-dimensional space, reducing computational costs and improving trajectory optimization.
Smart Images

Figure CN120056139B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robotic arms, and specifically to a method and system for optimizing the trajectory of a robotic arm based on adversarial learning. Background Art
[0002] Robotic arms were initially applied in automobile manufacturing, such as welding and painting. However, they are no longer limited to automobile manufacturing or even the manufacturing industry. In the field of microelectronics manufacturing, robotic arms can complete the assembly of tiny components that are difficult for human hands to reach; in dangerous environments, such as nuclear power plant maintenance or disaster rescue sites, robotic arms can replace humans to perform operations, thus ensuring personnel safety; in the medical field, surgical robots can assist doctors in performing highly precise minimally invasive surgeries, improving the success rate of surgeries and shortening the recovery time of patients. With the continuous development of artificial intelligence, sensor technology, and control methods, robotic arms are evolving towards a more intelligent and autonomous direction, with stronger perception, decision-making, and collaboration capabilities, capable of adapting to more complex and dynamic working environments and performing more delicate and complex tasks. The expansion of application fields also represents an increase in scenarios and complexity. To fully unleash the potential of robotic arms and achieve optimal performance in various application scenarios, a crucial link is trajectory optimization. Traditional trajectory optimization methods, such as methods based on analytical solutions, numerical optimization methods, and sampling-based planning methods, although playing important roles in different application scenarios, also face problems such as high dimensions, complex constraints, and high computational costs.
[0003] In recent years, artificial intelligence technology has advanced by leaps and bounds. As an important branch of artificial intelligence, adversarial learning has also been widely applied in many industries. Adversarial learning mainly utilizes a constructed generator and discriminator and allows them to continuously learn and improve during the process of competing with each other. Many scholars and research institutions have tried to apply adversarial learning to robotic arm trajectory optimization and have shown good performance. However, these studies all train adversarial learning in the form of two-dimensional images and then generate two-dimensional trajectory optimization results. However, in actual production, the trajectories of robotic arms are in three-dimensional space, and there are few studies on optimizing the trajectories of robotic arms in three-dimensional space using adversarial networks. Summary of the Invention
[0004] To solve the above problems, the present invention provides a method for optimizing the trajectory of a robotic arm based on adversarial learning. The method includes the following steps:
[0005] Discretize the motion space of the robotic arm to obtain voxels, obtain a multi-channel image based on whether the voxels contain obstacles, and use the RRT algorithm to obtain a sequence of voxels to be optimized for the robotic arm; calculate voxel sequences under different conditions using different neighborhood radii and step sizes;
[0006] Calculate the labels of the voxel sequence according to the neighborhood radius and step size, extract the features of the multi-channel image through the feature extraction unit of the discriminator, fuse the features of the multi-channel image and the voxel sequence to obtain the output of the discriminator, and pre-train the discriminator based on the labels and the output of the discriminator;
[0007] Use the pre-trained discriminator to train the adversarial network, and input the multi-channel image and the voxel sequence to be optimized into the trained generator to obtain the optimized result of the robotic arm trajectory.
[0008] Preferably, obtaining the multi-channel image based on whether the voxel contains an obstacle specifically includes:
[0009] Layer the voxels according to their positions, and each layer of voxels serves as a channel image of the multi-channel image;
[0010] If the voxel contains an obstacle, the pixel value in the channel image is 1; if the voxel does not contain an obstacle, the pixel value in the channel image is 0, and the pixel values corresponding to the voxels where the starting point and the ending point of the robotic arm are located are 2.
[0011] Preferably, calculating the labels of the voxel sequence according to the neighborhood radius and step size specifically includes:
[0012] Normalize the neighborhood radius and step size respectively;
[0013] Obtain the weights of the neighborhood radius and step size, calculate the exponential function of the product of the normalized neighborhood radius and the weight, and the negative exponential function of the product of the normalized step size and the weight, and normalize the product of the exponential function and the negative exponential function as the label.
[0014] Preferably, using the pre-trained discriminator to train the adversarial network specifically includes:
[0015] Load the pre-trained discriminator model and its weights, and initialize the generator model;
[0016] Obtain a batch of multi-channel images and the corresponding voxel sequences to be optimized from the dataset, input the multi-channel images and the voxel sequences to be optimized into the generator, and obtain the optimized voxel sequence output by the generator;
[0017] Input the input multi-channel images and the optimized voxel sequences output by the generator into the pre-trained discriminator to obtain the output scores of the discriminator;
[0018] Calculate the loss of the generator according to the output scores of the discriminator, and use the calculated loss of the generator to update the parameters of the generator through the optimizer.
[0019] Preferably, training the adversarial network using the pre-trained discriminator further includes:
[0020] For the same multi-channel image, obtain the voxel sequence with the largest label value, and the output of the voxel sequence generated by the generator in the discriminator. Calculate the loss of the discriminator according to the difference between the two voxel sequences, as well as the difference between the label and the output of the generator, and update the discriminator using the loss.
[0021] In the second aspect of the present invention, a robotic arm trajectory optimization system based on adversarial learning is provided. The system includes the following modules:
[0022] A preprocessing module, which is used to discretize the motion space of the robotic arm to obtain voxels, obtain a multi-channel image based on whether the voxels contain obstacles, and use the RRT algorithm to obtain the voxel sequence to be optimized for the robotic arm; calculate voxel sequences under different conditions using different neighborhood radii and step sizes;
[0023] A discriminator pre-training module, which is used to calculate the label of the voxel sequence according to the neighborhood radius and step size, extract the features of the multi-channel image through the feature extraction unit of the discriminator, fuse the features of the multi-channel image and the voxel sequence to obtain the output of the discriminator, and pre-train the discriminator based on the label and the output of the discriminator;
[0024] A trajectory optimization module, which is used to train the adversarial network using the pre-trained discriminator, and input the multi-channel image and the voxel sequence to be optimized into the trained generator to obtain the robotic arm trajectory optimization result.
[0025] Preferably, obtaining the multi-channel image based on whether the voxels contain obstacles specifically includes:
[0026] Stratify the voxels according to their positions, and each layer of voxels serves as a channel image of the multi-channel image;
[0027] If the voxel contains an obstacle, the pixel value in the channel image is 1; if the voxel does not contain an obstacle, the pixel value in the channel image is 0, and the pixel values of the voxels where the starting point and the ending point of the robotic arm are located are 2.
[0028] Preferably, calculating the label of the voxel sequence according to the neighborhood radius and step size specifically includes:
[0029] Normalize the neighborhood radius and step size respectively;
[0030] Obtain the weights of the neighborhood radius and step size, calculate the exponential function of the product of the normalized neighborhood radius and the weight, and the negative exponential function of the product of the normalized step size and the weight. Normalize the product of the exponential function and the negative exponential function, and use the result as the label.
[0031] Preferably, the pre-trained discriminator is used to train the adversarial network, specifically as follows:
[0032] Load the pre-trained discriminator model and its weights, and initialize the generator model;
[0033] Obtain a batch of multi-channel images and corresponding voxel sequences to be optimized from the dataset, and input the multi-channel images and the voxel sequences to be optimized into the generator to obtain the optimized voxel sequences output by the generator;
[0034] Input the input multi-channel images and the optimized voxel sequences output by the generator into the pre-trained discriminator to obtain the output score of the discriminator;
[0035] Calculate the loss of the generator according to the output score of the discriminator, and use the calculated loss of the generator to update the parameters of the generator through the optimizer.
[0036] Preferably, using the pre-trained discriminator to train the adversarial network further includes:
[0037] For the same multi-channel image, obtain the voxel sequence with the largest label value and the output of the voxel sequence generated by the generator in the discriminator. Calculate the loss of the discriminator according to the difference between the two voxel sequences and the difference between the label and the generator output, and use the loss to update the discriminator.
[0038] The present invention generates voxel sequences by using adversarial learning in three-dimensional space, and calculates the labels of the voxel sequences by using the neighborhood radius and step size to realize the training of the discriminator. According to the labels, the discriminator judges the quality of the voxel sequences. At the same time, the generator and the discriminator are continuously trained in the confrontation to improve the quality of the voxel sequences finally generated by the generator. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flowchart of Embodiment 1;
[0040] Figure 2 is a schematic diagram of the relationship between voxels and multi-channel images;
[0041] Figure 3 is a schematic diagram of the discriminator;
[0042] Figure 4 is a schematic diagram of the generator;
[0043] Figure 5 is a comparison chart of the trajectory lengths of the prior art and the present invention;
[0044] Figure 6 is a comparison chart of the trajectory smoothness of the prior art and the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0045] In this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0047] Figure 1 The first embodiment of the present invention is shown, as Figure 1 shown in the robotic arm trajectory optimization method based on adversarial learning, which includes the following steps:
[0048] S1, discretize the robotic arm motion space to obtain voxels, obtain a multi-channel image based on whether the voxels contain obstacles, and use the RRT algorithm to obtain the sequence of voxels to be optimized for the robotic arm; calculate the voxel sequences under different conditions by using different neighborhood radii and step sizes;
[0049] The robotic arm moves in a three-dimensional space. The continuous motion space is divided into many small cubes of the same size, and each small cube is used as a voxel. For each voxel, determine whether it contains an obstacle. If a voxel contains an obstacle, it is marked as occupied; if there is no obstacle, it is marked as free. In an alternative embodiment, the obtaining of the multi-channel image based on whether the voxels contain obstacles is specifically:
[0050] Layer the voxels according to the positions of the voxels, and each layer of voxels serves as a channel image of the multi-channel image;
[0051] If the voxel contains an obstacle, the pixel value in the channel image is 1. If the voxel does not contain an obstacle, the pixel value in the channel image is 0. The pixel values corresponding to the voxels where the starting point and the ending point of the robotic arm are located are 2.
[0052] As shown Figure 2 in the figure, the voxels are divided into multiple layers, and the voxels in each layer serve as a channel in the multi-channel image. The number of pixels in each channel is the same as the number of voxels in each layer and they correspond one by one. If there is an obstacle in the voxel, the corresponding pixel value in the corresponding channel is set to 1; if there is no obstacle, the pixel value is set to 0. The pixel values corresponding to the voxels where the starting point and the ending point of the robotic arm are located are 2. In this way, the multi-channel image not only contains the starting point and ending point information, but also contains information such as obstacles.
[0053] The RRT (Rapidly-exploring Random Tree) algorithm has a relatively fast operation speed, but the path found is generally not the optimal path, and there may be redundant movements or non-direct routes. For example, in an open scene, the path found by RRT may not be a straight line but a curved one. The trajectory generated by the RRT algorithm is not suitable for direct application to the trajectory of the robotic arm movement and needs to be processed. In the present invention, the RRT algorithm is used to obtain the sequence of voxels to be optimized for the robotic arm. Among them, the sequence of voxels to be optimized is a sequence composed of a group of voxels. For example, the sequence of voxels to be optimized is [15, 26, 31, 46, 50], and the trajectory formed by the central points of these voxels is used as the trajectory for the robotic arm to move.
[0054] The RRT* algorithm is an improvement of the RRT algorithm. Compared with the RRT algorithm, as the number of sampling points increases, the path generated by the RRT* algorithm will gradually converge to the optimal path. However, the RRT* algorithm needs to perform additional calculations in each iteration, including finding neighboring nodes and performing reconnection operations, and the calculation process is complex. In the present invention, the RRT* algorithm is used to calculate the sequence of voxels under different neighborhood radii and step sizes. Generally speaking, the larger the neighborhood radius and the smaller the step size, the closer the obtained sequence of voxels is to the optimal solution. The present invention uses the RRT* algorithm to obtain the sequences of voxels under different conditions and trains the discriminator with these sequences.
[0055] S2. Calculate the label of the sequence of voxels according to the neighborhood radius and step size, extract the features of the multi-channel image through the feature extraction unit of the discriminator, perform feature fusion on the features of the multi-channel image and the sequence of voxels to obtain the output of the discriminator, and pre-train the discriminator based on the label and the output of the discriminator;
[0056] As mentioned above, the larger the neighborhood radius and the smaller the step size, the closer the sequence of voxels is to the optimal solution. By running the RRT* algorithm with different neighborhood radii and step sizes, different movement paths of the robotic arm, that is, sequences of voxels, are obtained, and the sequences of voxels are labeled according to the neighborhood radius and step size used when generating these paths. In one embodiment, the calculating the label of the sequence of voxels according to the neighborhood radius and step size specifically includes:
[0057] Normalize the neighborhood radius and step size respectively;
[0058] Obtain the weights of the neighborhood radius and step size, calculate the exponential function of the product of the normalized neighborhood radius and the weight, and the negative exponential function of the product of the normalized step size and the weight, and normalize the product of the exponential function and the negative exponential function as the said label.
[0059] For example, there are two voxel sequences obtained through different neighborhood radii and step sizes. The neighborhood radius of sequence A is 1.0 and the step size is 0.2. The neighborhood radius of sequence B is 0.5 and the step size is 0.4. The weight of the neighborhood radius is 0.6 and the weight of the step size is 0.4. The label of sequence A after normalization is 1, and the label of sequence B is 0. Here there are only two voxel sequences. If there are more voxel sequences, each voxel sequence will have a label. The larger the value of the label, the higher the quality of the voxel sequence, that is, the closer it is to the optimal solution.
[0060] The feature extraction unit in the discriminator converts the input multi-channel image into a feature vector. Preferably, it is implemented through structures such as 3DCNN or CNN. Combine the environmental features extracted from the multi-channel image and the voxel sequence representing the motion path of the robotic arm. There are various fusion methods, such as splicing them together, or using a more complex fusion network to learn the relationship between them. The fused features will be input into the subsequent network layers of the discriminator, and finally the output of the discriminator is obtained. The pre-training process is to use these labeled voxel sequences and multi-channel feature maps to train the discriminator. The training objective is to make the output of the discriminator as close as possible to the true label, that is, to enable the discriminator to distinguish the quality of the voxel sequence. In one embodiment, the loss of the discriminator is obtained using the label and the output of the discriminator, and then the discriminator is updated according to the obtained loss.
[0061] In one embodiment, the discriminator adopts a structure of convolution + MLP. The input of the discriminator is a multi-channel image and a voxel sequence, and the output is a score, as Figure 3 shown. The present invention does not make specific limitations on the structure of the discriminator.
[0062] S3. Use the pre-trained discriminator to train the adversarial network, and input the multi-channel image and the voxel sequence to be optimized into the trained generator to obtain the robotic arm trajectory optimization result.
[0063] After pre-training the discriminator, in one embodiment, fix the parameters of the discriminator, use the generator to generate voxel sequences, and then use the discriminator to judge the quality of the voxel sequences, and further, update the parameters of the generator, that is, train. Specifically, load the pre-trained discriminator model and its weights, and initialize the generator model;
[0064] Obtain a batch of multi-channel images and the corresponding voxel sequences to be optimized from the dataset, input the multi-channel images and the voxel sequences to be optimized into the generator, and obtain the optimized voxel sequences output by the generator;
[0065] Input the input multi-channel images and the optimized voxel sequences output by the generator into the pre-trained discriminator to obtain the output scores of the discriminator;
[0066] Calculate the loss of the generator according to the output scores of the discriminator, and use the calculated loss of the generator to update the parameters of the generator through the optimizer.
[0067] First, load the previously trained discriminator model, including its network structure and the learned parameters. The pre-trained discriminator already has the ability to evaluate the quality of voxel sequences. At the same time, initialize a new generator model, and the initial parameters of the generator are randomly set. Select a batch of data from the training dataset, and each piece of data contains a multi-channel voxel image representing the environment and an initial voxel sequence to be optimized. Input these data into the generator model. The generator will output the optimized voxel sequences according to the multi-channel images and the initial paths.
[0068] In an optional embodiment, the generator model is a 3D convolutional neural network. The 3D CNN first extracts the features of the multi-channel images, and then the RNN or Transformer processes the initial voxel sequences. After fusing the features of the two, the optimized sequences are generated through multiple perceptrons, etc. In one embodiment, the length of the voxel sequences is a unified length, for example, all are 10. If the path is less than 10 voxels, the other voxels are all set to 0. For example, if a voxel sequence is [0, 0, 12, 15, 36, 69, 75, 0, 0, 0], the path corresponding to the voxel sequence is [12, 15, 36, 69, 75]. The input of the generator is the multi-channel images and the voxel sequences to be optimized, and the output is the optimized voxel sequences, as Figure 4 shown. The present invention does not limit the specific structure of the generator.
[0069] The original multi-channel image and the optimized voxel sequence just output by the generator are input into the pre-trained discriminator together. The discriminator will give a score based on the knowledge it has learned before. This score represents the quality of the optimized voxel sequence output by the generator as considered by the discriminator. The higher the score, the closer the generated sequence is considered to be to the desired optimized result by the discriminator. Further, the loss of the generator is calculated based on the output score of the discriminator. The objective of the loss function is that when the score given by the discriminator is higher, the loss of the generator is smaller. Using the calculated loss value, the parameters of the generator model are updated through an optimizer such as Adam, SGD, etc. The above process is iterated continuously until the generator can generate an optimized voxel sequence that is good enough or reaches the preset number of iterations.
[0070] To further train the discriminator in adversarial learning, for the same environmental multi-channel image, find the voxel sequence with the highest label value. The sequence with the highest label value is considered the optimal path. The multi-channel image is input into the generator, and the generator will output an optimized voxel sequence. Then, the optimized voxel sequence generated by the generator and the same multi-channel image are input into the pre-trained discriminator to obtain the evaluation score of the discriminator for this generated sequence.
[0071] In another embodiment, the loss of the discriminator is calculated based on the difference between two voxel sequences and the difference between the label and the output of the generator. Specifically, calculate the difference between the two voxel sequences, such as the average distance and / or smoothness, etc., and then weight the difference between the label and the output of the generator according to the difference of the voxel sequences. Preferably, the greater the difference of the voxel sequences, the greater the weight. Then, use the weight to weight the difference between the label and the output of the generator, and take the weighted difference as the loss to update the parameters of the generator.
[0072] In an alternative embodiment, for the same multi-channel image, obtain the voxel sequence with the largest label value and the output of the voxel sequence generated by the generator in the discriminator, calculate the loss of the discriminator based on the difference between the label and the output of the generator, and use the loss to update the discriminator. Preferably, at this time, the objective function of the discriminator is to maximize the difference between the label and the output of the generator.
[0073] In the second embodiment of the present invention, a robotic arm trajectory optimization system based on adversarial learning is provided. The system includes the following modules:
[0074] A preprocessing module for discretizing the robotic arm motion space to obtain voxels, obtaining a multi-channel image based on whether the voxels contain obstacles, and using the RRT algorithm to obtain the voxel sequence to be optimized for the robotic arm; calculating voxel sequences under different conditions by using different neighborhood radii and step sizes;
[0075] The discriminator pre-training module is used to calculate the labels of the voxel sequence according to the neighborhood radius and the step size, extract the features of the multi-channel image through the feature extraction unit of the discriminator, perform feature fusion on the features of the multi-channel image and the voxel sequence to obtain the output of the discriminator, and pre-train the discriminator based on the labels and the output of the discriminator;
[0076] The trajectory optimization module is used to train the adversarial network with the pre-trained discriminator, and input the multi-channel image and the voxel sequence to be optimized into the trained generator to obtain the optimized result of the robotic arm trajectory.
[0077] Preferably, obtaining the multi-channel image based on whether the voxel contains an obstacle specifically includes:
[0078] Stratify the voxels according to the positions of the voxels, and each layer of voxels serves as a channel image of the multi-channel image;
[0079] If the voxel contains an obstacle, the pixel value in the channel image is 1. If the voxel does not contain an obstacle, the pixel value in the channel image is 0. The pixel values corresponding to the voxels where the starting point and the ending point of the robotic arm are located are 2.
[0080] Preferably, calculating the labels of the voxel sequence according to the neighborhood radius and the step size specifically includes:
[0081] Normalize the neighborhood radius and the step size respectively;
[0082] Obtain the weights of the neighborhood radius and the step size, calculate the exponential function of the product of the normalized neighborhood radius and the weight, and the negative exponential function of the product of the normalized step size and the weight, and normalize the product of the exponential function and the negative exponential function as the label.
[0083] Preferably, training the adversarial network with the pre-trained discriminator specifically includes:
[0084] Load the pre-trained discriminator model and its weights, and initialize the generator model;
[0085] Obtain a batch of multi-channel images and the corresponding voxel sequences to be optimized from the dataset, input the multi-channel images and the voxel sequences to be optimized into the generator, and obtain the optimized voxel sequence output by the generator;
[0086] Input the input multi-channel images and the optimized voxel sequences output by the generator into the pre-trained discriminator to obtain the output scores of the discriminator;
[0087] Calculate the loss of the generator according to the output scores of the discriminator, and use the calculated loss of the generator to update the parameters of the generator through the optimizer.
[0088] Preferably, training the adversarial network using the pre-trained discriminator further includes:
[0089] For the same multi-channel image, obtain the voxel sequence with the largest label value, and the output of the voxel sequence generated by the generator in the discriminator. Calculate the loss of the discriminator according to the difference between the two voxel sequences, as well as the difference between the label and the output of the generator, and update the discriminator using the loss.
[0090] Embodiment 3: The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in Embodiment 1 is implemented.
[0091] Embodiment 4: The present invention also provides a computer device, which at least includes a memory and a processor. A computer program is stored on the memory. When the computer program is executed by the processor, the method described in Embodiment 1 is implemented.
[0092] In the experiment of the present invention, the present invention is compared with RRT, RRT*, and PRM. The experimental results are shown in Table 1.
[0093] Table 1 Experimental results of trajectories of different algorithms
[0094] Trajectory generation method Parameters (neighborhood radius, step size) Trajectory length (centimeters) Trajectory smoothness (radians) RRT - 122 0.41 <![CDATA[RRT* 1 > (0.5, 0.1) 107 0.34 <![CDATA[RRT* 2 > (1.0, 0.1) 105 0.29 <![CDATA[RRT* 3 > (0.5, 0.2) 113 0.31 <![CDATA[RRT* 4 > (1.0, 0.2) 109 0.30 PRM - 116 0.42 The present invention - 103 0.26
[0095] As can be seen from Table 1, the optimized trajectory of the present invention is significantly shorter in length than the trajectories of the RRT, RRT*, and PRM algorithms, and the smoothness of the trajectory is also improved, as Figure 5 、 Figure 6 shown.
[0096] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, can also be implemented by a combination of hardware and software. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it. Other embodiments can also be adopted. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A manipulator trajectory optimization method based on adversarial learning, characterized in that, The method includes the following steps: Discretize the motion space of the robotic arm to obtain voxels, obtain a multi-channel image based on whether the voxels contain obstacles, and use the RRT algorithm to obtain a sequence of voxels to be optimized for the robotic arm; calculate voxel sequences under different conditions by using different neighborhood radii and step sizes; Calculate the labels of the voxel sequence according to the neighborhood radius and step size, extract the features of the multi-channel image through the feature extraction unit of the discriminator, perform feature fusion on the features of the multi-channel image and the voxel sequence to obtain the output of the discriminator, and pre-train the discriminator based on the labels and the output of the discriminator; Use the pre-trained discriminator to train the adversarial network, and input the multi-channel image and the sequence of voxels to be optimized into the trained generator to obtain the optimized result of the robotic arm trajectory; The calculating the labels of the voxel sequence according to the neighborhood radius and step size is specifically: Normalize the neighborhood radius and step size respectively; Obtain the weights of the neighborhood radius and step size, calculate the exponential function of the product of the normalized neighborhood radius and the weight, and the negative exponential function of the product of the normalized step size and the weight, and normalize the product of the exponential function and the negative exponential function as the label; The using the pre-trained discriminator to train the adversarial network is specifically: Load the pre-trained discriminator model and its weights, and initialize the generator model; Obtain a batch of multi-channel images and the corresponding sequences of voxels to be optimized from the dataset, input the multi-channel images and the sequences of voxels to be optimized into the generator, and obtain the optimized voxel sequence output by the generator; Input the input multi-channel image and the optimized voxel sequence output by the generator into the pre-trained discriminator to obtain the output score of the discriminator; Calculate the loss of the generator according to the output score of the discriminator, and use the calculated loss of the generator to update the parameters of the generator through the optimizer; The using the pre-trained discriminator to train the adversarial network further includes: For the same multi-channel image, obtain the voxel sequence with the largest label value, and the output of the voxel sequence generated by the generator in the discriminator, calculate the loss of the discriminator according to the difference between the two voxel sequences, and the difference between the label and the output of the generator, and update the discriminator using the loss.
2. The method according to claim 1, characterized in that The obtaining the multi-channel image based on whether the voxels contain obstacles is specifically: Layer the voxels according to the positions of the voxels, and each layer of voxels serves as a channel image of the multi-channel image; If the voxel contains an obstacle, the pixel value in the channel image is 1, if the voxel does not contain an obstacle, the pixel value in the channel image is 0, and the pixel values corresponding to the voxels where the starting point and the ending point of the robotic arm are located are 2.
3. A robotic arm trajectory optimization system based on adversarial learning, characterized in that, The system includes the following modules: A preprocessing module, which is used to discretize the motion space of the robotic arm to obtain voxels, obtain a multi-channel image based on whether the voxels contain obstacles, use the RRT algorithm to obtain a sequence of voxels to be optimized for the robotic arm; calculate voxel sequences under different conditions by using different neighborhood radii and step sizes; The discriminator pre-training module is used to calculate the labels of the voxel sequence according to the neighborhood radius and the step size, extract the features of the multi-channel image through the feature extraction unit of the discriminator, perform feature fusion on the features of the multi-channel image and the voxel sequence to obtain the output of the discriminator, and pre-train the discriminator based on the labels and the output of the discriminator; The trajectory optimization module is used to train the adversarial network with the pre-trained discriminator, and input the multi-channel image and the voxel sequence to be optimized into the trained generator to obtain the optimized result of the robotic arm trajectory; The calculation of the labels of the voxel sequence according to the neighborhood radius and the step size is specifically as follows: Normalize the neighborhood radius and the step size respectively; Obtain the weights of the neighborhood radius and the step size, calculate the exponential function of the product of the normalized neighborhood radius and the weight, and the negative exponential function of the product of the normalized step size and the weight, and normalize the product of the exponential function and the negative exponential function as the label; The training of the adversarial network with the pre-trained discriminator is specifically as follows: Load the pre-trained discriminator model and its weights, and initialize the generator model; Obtain a batch of multi-channel images and the corresponding voxel sequences to be optimized from the dataset, input the multi-channel images and the voxel sequences to be optimized into the generator, and obtain the optimized voxel sequence output by the generator; Input the input multi-channel images and the optimized voxel sequences output by the generator into the pre-trained discriminator to obtain the output score of the discriminator; Calculate the loss of the generator according to the output score of the discriminator, and use the calculated generator loss to update the parameters of the generator through the optimizer; The training of the adversarial network with the pre-trained discriminator further includes: For the same multi-channel image, obtain the voxel sequence with the largest label value, and the output of the voxel sequence generated by the generator in the discriminator. Calculate the loss of the discriminator according to the difference between the two voxel sequences, and the difference between the label and the generator output, and update the discriminator using the loss.
4. The system according to claim 3, wherein The obtaining of the multi-channel image based on whether the voxel contains an obstacle is specifically as follows: Layer the voxels according to the positions of the voxels, and each layer of voxels serves as a channel image of the multi-channel image; If the voxel contains an obstacle, the pixel value in the channel image is 1. If the voxel does not contain an obstacle, the pixel value in the channel image is 0. The pixel values of the voxels where the starting point and the ending point of the robotic arm are located are 2.
Citation Information
Patent Citations
Vehicle image optimization method and system based on adversarial learning
CN110458060A
Conditional generative adversarial-based three-dimensional model fuzzy texture feature saliency method
CN114119924A