Visual servo method, device and equipment based on diffusion model and medium
Through the visual servo method based on the diffusion model, the current expected image is generated and transmitted using the variational autoencoder and mutual attention mechanism, the problem of low efficiency in obtaining and providing the current expected image in the prior art is solved, the accurate position and orientation adjustment of the robot tool is achieved, and the accuracy and reliability of the operation are improved.
Patent Information
- Application Number
- CN202510144471.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-10
AI Technical Summary
In the prior art, the inefficiency of obtaining and providing currently desired images makes it difficult for the robot vision servo system to accurately adjust the position and orientation of the tool when the desired images are lacking, and it is prone to malfunction.
Using a visual servo method based on the diffusion model, tool images and teaching images are acquired through the camera, feature vectors are fused using variational autoencoder and mutual attention mechanism, diffusion model is trained to generate predicted expected images, and transmitted to the visual servo system to adjust tool position and orientation.
It improves the efficiency of currently expected images, reduces manual intervention, enhances the reliability of image transmission, ensures the accurate position and orientation of the robot tool, and improves the accuracy and reliability of operation.
Smart Images

Figure CN120147587A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of robotics and information technology, and particularly to a visual servo method, device, equipment and medium based on a diffusion model. Background Art
[0002] The current desired image represents the visual state that the visual servo system expects to achieve or the ideal appearance of the target object. The current desired image plays a crucial role in the visual servo system.
[0003] However, the prior art mainly adopts the method of manual acquisition to obtain the current desired image. The manual acquisition method increases the acquisition time of the current desired image and is not conducive to improving the acquisition efficiency of the current desired image. Therefore, how to obtain the current desired image is a technical problem that needs to be solved urgently. In addition, if the current desired image is lacking, the visual servo system of the robot will control the robot to blindly try various actions, attempting to find the correct motion path by means of trial and error. This method is prone to errors. Therefore, how to provide the current desired image to the visual servo system of the robot is also a technical problem that needs to be solved urgently. Summary of the Invention
[0004] The present invention provides a visual servo method, device, computer equipment and storage medium based on a diffusion model to solve the technical problem of how to obtain the current desired image and solve the technical problem of how to provide the current desired image to the visual servo system of the robot.
[0005] In a first aspect, a visual servo method based on a diffusion model is provided, including: Obtain an image captured by a camera of a preset tool, select the image captured by the camera of the preset tool as a preset tool image, and form a training sample by combining the preset tool image, the corresponding preset teaching image of the preset tool image, and the corresponding true desired image of the preset tool image; Train a diffusion model based on the training sample, and fuse the feature vectors of the preset tool image and the preset teaching image to obtain a first fusion vector; Process the first fusion vector through a decoder in the diffusion model to obtain a predicted desired image; Obtain the total loss value between the predicted desired image and the true desired image; When the total loss value is less than a preset value, stop training the diffusion model to obtain a trained diffusion model; Obtain the image captured by the camera for the current tool, select the image captured for the current tool as the current tool image, fuse the feature vector of the current tool image and the feature vector of the current teaching image to obtain a second fused vector, process the second fused vector through the decoder in the trained diffusion model to obtain the current expected image, and transmit the current expected image to the visual servo system of the robot to control the visual servo system to adjust the position and orientation of the current tool according to the current expected image.
[0006] Further, training the diffusion model based on training samples and fusing the feature vector of the preset tool image and the feature vector of the preset teaching image to obtain a first fused vector includes: Train the diffusion model based on training samples, and input the preset tool image and the corresponding preset teaching image in the training samples into the variational autoencoder. Extract the feature vector of the preset tool image through the variational autoencoder, extract the feature vector of the preset teaching image through the variational autoencoder, and fuse the feature vector of the preset tool image and the feature vector of the preset teaching image through the mutual attention mechanism to obtain a first fused vector.
[0007] Further, processing the first fused vector through the decoder in the diffusion model to obtain the predicted expected image includes: Obtain the diffusion model and input the third embedding vector into the decoder in the diffusion model. Process the third embedding vector through the decoder in the diffusion model to obtain the predicted expected image.
[0008] Further, obtaining the total loss value between the predicted expected image and the true expected image includes: Calculate the first loss value between the predicted expected image and the true expected image through the KL divergence loss function, and calculate the second loss value between the predicted expected image and the true expected image through the mean square error loss function. Add the first loss value and the second loss value to obtain the total loss value between the predicted expected image and the true expected image.
[0009] Further, when the total loss value is less than the preset value, stopping training the diffusion model to obtain the trained diffusion model includes: When the total loss value is less than the preset value, stop training the diffusion model and save the model parameters. Load the model parameters into the model structure, and select the model structure loaded with the model parameters as the trained diffusion model.
[0010] Further, the method of obtaining the image captured by the camera for the current tool, selecting the image captured for the current tool as the current tool image, fusing the feature vectors of the current tool image and the current teaching image to obtain a second fusion vector, processing the second fusion vector through the decoder in the trained diffusion model to obtain the current expected image, and transmitting the current expected image to the visual servo system of the robot to control the visual servo system to adjust the position and orientation of the current tool according to the current expected image includes: Obtain the image captured by the camera for the current tool, select the image captured for the current tool as the current tool image, input the current tool image and the corresponding current teaching image of the current tool image into the variational autoencoder, perform feature extraction on the current tool image through the variational autoencoder to obtain the feature vector of the current tool image, and perform feature extraction on the current teaching image through the variational autoencoder to obtain the feature vector of the current teaching image; Through the mutual attention mechanism, fuse the feature vectors of the current tool image and the current teaching image to obtain a second fusion vector, process the second fusion vector through the decoder in the trained diffusion model to obtain the current expected image, obtain a transmission instruction, execute the transmission instruction, transmit the current expected image to the visual servo system of the robot, and control the visual servo system to adjust the position and orientation of the current tool according to the current expected image.
[0011] Further, after the method of obtaining the image captured by the camera for the current tool, selecting the image captured for the current tool as the current tool image, fusing the feature vectors of the current tool image and the current teaching image to obtain a second fusion vector, processing the second fusion vector through the decoder in the trained diffusion model to obtain the current expected image, and transmitting the current expected image to the visual servo system of the robot to control the visual servo system to adjust the position and orientation of the current tool according to the current expected image, the visual servo method includes: Obtain the adjustment result returned by the robot, create a display window, and display the adjustment result through the display window.
[0012] In a second aspect, a visual servo device based on a diffusion model is provided, including: A first acquisition module, configured to obtain the image captured by the camera for the preset tool, select the image captured for the preset tool as the preset tool image, and form a training sample by combining the preset tool image, the corresponding preset teaching image of the preset tool image, and the true expected image corresponding to the preset tool image; A fusion module, configured to train a diffusion model based on the training sample, fuse the feature vectors of the preset tool image and the preset teaching image to obtain a first fusion vector; A processing module, configured to process the first fusion vector through a decoder in a diffusion model to obtain a predicted desired image; A second acquisition module, configured to acquire the total loss value between the predicted desired image and the true desired image; A stop module, configured to stop training the diffusion model to obtain a trained diffusion model when the total loss value is less than a preset value; A control module, configured to acquire an image of the current tool captured by a camera, select the image of the current tool captured as the current tool image, fuse the feature vector of the current tool image and the feature vector of the current teaching image to obtain a second fusion vector, process the second fusion vector through a decoder in the trained diffusion model to obtain a current desired image, and transmit the current desired image to the visual servo system of the robot to control the visual servo system to adjust the position and orientation of the current tool according to the current desired image.
[0013] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above visual servo method are implemented.
[0014] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above visual servo method are implemented.
[0015] The present application provides a visual servo method, device, computer device, and storage medium based on a diffusion model. The beneficial effects are in two aspects. On the one hand, an image of the current tool captured by a camera is acquired, the image of the current tool captured is selected as the current tool image, the feature vector of the current tool image and the feature vector of the current teaching image are fused to obtain a second fusion vector, and the second fusion vector is processed through a decoder in the trained diffusion model to obtain a current desired image, which solves the technical problem of how to obtain the current desired image. Since there is no need to manually obtain the current desired image, the acquisition time of the current desired image is reduced, which is beneficial to improving the acquisition efficiency of the current desired image. On the other hand, the current desired image is transmitted to the visual servo system of the robot to control the visual servo system to adjust the position and orientation of the current tool according to the current desired image, which solves the technical problem of how to provide the current desired image to the visual servo system of the robot. Since the current desired image is automatically transmitted to the visual servo system of the robot and is not affected by manual intervention, it is beneficial to improve the reliability of the current desired image. Description of the Drawings
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments of the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0017] Figure 1 It is a schematic diagram of an application environment of a visual servo method in an embodiment of the present invention; Figure 2 It is a schematic flowchart of a visual servo method provided by an embodiment of the present invention; Figure 3 It is Figure 2 A schematic flowchart of a specific implementation manner of step S23 in Figure 4 It is Figure 2 A schematic flowchart of a specific implementation manner of step S25 in Figure 5 It is Figure 2 A schematic flowchart of a specific implementation manner of step S26 in Figure 6 It is a schematic structural diagram of a visual servo device in an embodiment of the present invention; Figure 7 It is a schematic structural diagram of a computer device in an embodiment of the present invention. Specific Embodiments
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0019] Please refer to Figure 1 , Figure 1 It is a schematic diagram of an application environment of a visual servo method in an embodiment of the present invention. The visual servo method provided by the embodiment of the present invention can be applied in an application environment such as Figure 1 , where the client communicates with the server through a network.
[0020] The server obtains the image captured by the camera of the preset tool through the client, selects the image captured by the camera of the preset tool as the preset tool image, and forms a training sample by combining the preset tool image, the corresponding preset teaching image of the preset tool image, and the corresponding true expected image of the preset tool image; Train a diffusion model based on training samples, fuse the feature vectors of a preset tool image and a preset teaching image to obtain a first fused vector; Process the first fused vector through the decoder in the diffusion model to obtain a predicted desired image; Obtain the total loss value between the predicted desired image and the true desired image; When the total loss value is less than a preset value, stop training the diffusion model to obtain a trained diffusion model; Obtain the image of the current tool captured by the camera, select the image of the current tool captured as the current tool image, fuse the feature vectors of the current tool image and the current teaching image to obtain a second fused vector, process the second fused vector through the decoder in the trained diffusion model to obtain the current desired image, and transmit the current desired image to the visual servo system of the robot to control the visual servo system to adjust the position and orientation of the current tool according to the current desired image.
[0021] In the solution implemented by the above visual servo method, device, equipment and medium, the beneficial effects are in two aspects. On the one hand, obtain the image of the current tool captured by the camera, select the image of the current tool captured as the current tool image, fuse the feature vectors of the current tool image and the current teaching image to obtain a second fused vector, process the second fused vector through the decoder in the trained diffusion model to obtain the current desired image, which solves the technical problem of how to obtain the current desired image. Since there is no need to manually obtain the current desired image, the acquisition time of the current desired image is reduced, which is beneficial to improving the acquisition efficiency of the current desired image. On the other hand, transmit the current desired image to the visual servo system of the robot, control the visual servo system to adjust the position and orientation of the current tool according to the current desired image, which solves the technical problem of how to provide the current desired image to the visual servo system of the robot. Since the current desired image is automatically transmitted to the visual servo system of the robot and is not affected by manual intervention, it is beneficial to improve the reliability of the current desired image.
[0022] Among them, the device running the client is simply referred to as: client device.
[0023] Among them, the device running the server is simply referred to as: server device.
[0024] Among them, the client device includes but is not limited to smart phones, personal computers, vehicle networking terminals, tablet computers and portable wearable devices.
[0025] Among them, the server device can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail through specific embodiments. Please refer toFigure 2 , Figure 2 is a schematic flowchart of a visual servo method provided by an embodiment of the present invention, including the following steps: S21, obtaining an image captured by a camera of a preset tool, selecting the image captured of the preset tool as a preset tool image, and forming a training sample with the preset tool image, a preset teaching image corresponding to the preset tool image, and a true expected image corresponding to the preset tool image; Among them, the preset tool image is a preset tool image.
[0026] Among them, the preset teaching image is a preset teaching image.
[0027] Among them, the true expected image is a true expected image.
[0028] Among them, the predicted expected image is an expected image predicted by a diffusion model during training.
[0029] Among them, the current expected image is an expected image predicted by the diffusion model after training.
[0030] S22, training a diffusion model based on the training sample, and fusing the feature vectors of the preset tool image and the preset teaching image to obtain a first fusion vector; Among them, the training of the diffusion model based on the training sample, and fusing the feature vectors of the preset tool image and the preset teaching image to obtain a first fusion vector includes: training a diffusion model based on the training sample, and inputting the preset tool image and the preset teaching image corresponding to the preset tool image in the training sample into a variational autoencoder; extracting the feature vector of the preset tool image through the variational autoencoder, extracting the feature vector of the preset teaching image through the variational autoencoder, and fusing the feature vectors of the preset tool image and the preset teaching image through a mutual attention mechanism to obtain a first fusion vector.
[0031] Among them, the first fusion vector not only contains the key information of the feature vectors of the preset tool image and the preset teaching image, but also enhances the interaction between the feature vectors of the preset tool image and the preset teaching image through the mutual attention mechanism, enabling the diffusion model to better understand the preset tool image and the preset teaching image.
[0032] Among them, the variational autoencoder can map the high-dimensional data of the preset tool image and the preset teaching image to a low-dimensional latent space, which is beneficial for processing the high-dimensional data of the preset tool image and the preset teaching image.
[0033] Exemplarily, training a diffusion model based on training samples, and inputting a preset tool image and a preset teaching image corresponding to the preset tool image in the training samples into a variational autoencoder, including: Training a diffusion model using training samples, and obtaining the number of training rounds during the training process of the diffusion model; When the number of training rounds is less than a preset number of rounds, input the preset tool image and the preset teaching image corresponding to the preset tool image in the training samples into the variational autoencoder.
[0034] S23, processing the first fusion vector through a decoder in the diffusion model to obtain a predicted expected image; S24, obtaining the total loss value between the predicted expected image and the true expected image; Wherein, obtaining the total loss value between the predicted expected image and the true expected image includes: Calculating a first loss value between the predicted expected image and the true expected image through a KL divergence loss function, and calculating a second loss value between the predicted expected image and the true expected image through a mean square error loss function; Adding the first loss value and the second loss value to obtain the total loss value between the predicted expected image and the true expected image.
[0035] Adding the first loss value and the second loss value to obtain the total loss value between the predicted expected image and the true expected image, including: Adopting a preset total loss value generation model, adding the first loss value and the second loss value to obtain the total loss value between the predicted expected image and the true expected image; Wherein, the total loss value generation model is: ; Wherein, is the total loss value, is the first weight coefficient, is the first loss value, is the second weight coefficient, is the second loss value.
[0036] Wherein, the total loss value generation model is a generation model of the total loss value.
[0037] Wherein, the total loss value between the predicted expected image and the true expected image is used to describe the overall difference degree between the predicted expected image and the true expected image; the larger the total loss value between the predicted expected image and the true expected image, the greater the overall difference degree between the predicted expected image and the true expected image; the smaller the total loss value between the predicted expected image and the true expected image, the smaller the overall difference degree between the predicted expected image and the true expected image.
[0038] S25. When the total loss value is less than the preset value, stop training the diffusion model to obtain the trained diffusion model. Among them, when the total loss value is less than the preset value, it indicates that the performance of the diffusion model during training has reached an acceptable level. At this time, stopping the training of the diffusion model can avoid over-training the diffusion model, which is beneficial to reducing the training time and improving the training efficiency.
[0039] S26. Obtain the image of the current tool captured by the camera, select the image of the current tool captured as the current tool image, fuse the feature vectors of the current tool image and the current teaching image to obtain a second fusion vector, process the second fusion vector through the decoder in the trained diffusion model to obtain the current expected image, and transmit the current expected image to the visual servo system of the robot to control the visual servo system to adjust the position and orientation of the current tool according to the current expected image.
[0040] Among them, the current tool image is the current tool image.
[0041] Among them, the current teaching image is the current teaching image.
[0042] Among them, the second fusion vector not only contains the key information of the feature vectors of the current tool image and the current teaching image, but also enhances the interaction between the feature vectors of the current tool image and the current teaching image through the mutual attention mechanism, enabling the diffusion model to better understand the current tool image and the current teaching image.
[0043] Exemplarily, controlling the visual servo system to adjust the position and orientation of the current tool according to the current expected image includes: Controlling the visual servo system to process the current expected image to obtain the pose information of the current tool, and reading the position coordinates and rotation angles in the pose information; Controlling the end effector of the robot to adjust the position of the current tool according to the position coordinates, and controlling the end effector of the robot to adjust the orientation of the current tool according to the rotation angle.
[0044] Among them, controlling the end effector of the robot to adjust the position of the current tool according to the position coordinates and controlling the end effector of the robot to adjust the orientation of the current tool according to the rotation angle ensure that the current tool can accurately reach the position and perform tasks in the correct orientation, thereby improving the accuracy and reliability of the operation. This precise control not only improves work efficiency and product quality, but also reduces errors and failure rates during operation.
[0045] For the convenience of explanation, taking the current tool as an electric screwdriver as an example, the example is as follows: Control the end effector of the robot to adjust the position of the screwdriver according to the position coordinates, and control the end effector of the robot to adjust the orientation of the screwdriver according to the rotation angle. This precise control ensures that the screwdriver can accurately reach the predetermined position and tighten the screw with the correct orientation, which greatly improves the accuracy and reliability of the operation and avoids problems such as screw loosening or damage caused by position deviation or incorrect orientation.
[0046] For the sake of illustration, taking the current tool as an electric wrench as an example, the following is an example: Control the end effector of the robot to adjust the position of the electric wrench according to the position coordinates, and control the end effector of the robot to adjust the orientation of the electric wrench according to the rotation angle. This precise control ensures that the electric wrench can accurately reach the predetermined position and tighten the bolt with the correct orientation, which greatly improves the accuracy and reliability of the operation and avoids problems such as bolt loosening or damage caused by position deviation or incorrect orientation.
[0047] Among them, after obtaining the image of the current tool captured by the camera, selecting the image of the current tool captured as the current tool image, fusing the feature vectors of the current tool image and the current teaching image to obtain a second fusion vector, processing the second fusion vector by the decoder in the trained diffusion model to obtain the current expected image, and transmitting the current expected image to the visual servo system of the robot to control the visual servo system to adjust the position and orientation of the current tool according to the current expected image, the visual servo method includes: Obtain the adjustment result returned by the robot, create a display window, and display the adjustment result through the display window.
[0048] In the embodiment of the present invention, the beneficial effects are in two aspects. On the one hand, obtaining the image of the current tool captured by the camera, selecting the image of the current tool captured as the current tool image, fusing the feature vectors of the current tool image and the current teaching image to obtain a second fusion vector, and processing the second fusion vector by the decoder in the trained diffusion model to obtain the current expected image solves the technical problem of how to obtain the current expected image. Since there is no need for manual acquisition of the current expected image, the acquisition time of the current expected image is reduced, which is beneficial to improving the acquisition efficiency of the current expected image; on the other hand, transmitting the current expected image to the visual servo system of the robot and controlling the visual servo system to adjust the position and orientation of the current tool according to the current expected image solves the technical problem of how to provide the current expected image to the visual servo system of the robot. Since the current expected image is automatically transmitted to the visual servo system of the robot and is not affected by manual intervention, it is beneficial to improve the reliability of the current expected image.
[0049] Please refer toFigure 3 , Figure 3 is Figure 2 a schematic flowchart of a specific implementation manner of step S23 in S31. Obtain a diffusion model and input the third embedding vector into the decoder in the diffusion model; S32. Process the third embedding vector through the decoder in the diffusion model to obtain a predicted expected image.
[0050] Among them, the predicted expected image is the expected image predicted by the diffusion model during training. The diffusion model during training is in the training stage and is extracting features and learning rules from the training samples. The predicted expected image can be used to evaluate the performance of the diffusion model during the training process.
[0051] In the embodiment of the present invention, the predicted expected image can be used to evaluate the performance of the diffusion model during the training process. By analyzing the predicted expected image, it is possible to understand the degree of understanding of the input data by the diffusion model, and then adjust the training strategy to improve the performance and generalization ability of the diffusion model.
[0052] Please refer to Figure 4 , Figure 4 is Figure 2 a schematic flowchart of a specific implementation manner of step S25 in S41. When the total loss value is less than a preset value, stop training the diffusion model and save the model parameters; When the total loss value is less than the preset value, it indicates that the model parameters have been optimized to the best state, so the model parameters are saved.
[0053] S42. Load the model parameters into the model structure, and select the model structure loaded with the model parameters as the trained diffusion model.
[0054] In the embodiment of the present invention, loading the model parameters into the model structure and selecting the model structure loaded with the model parameters as the trained diffusion model ensure that the trained diffusion model has high accuracy and generalization performance.
[0055] Please refer to Figure 5 , Figure 5 is Figure 2 a schematic flowchart of a specific implementation manner of step S26 in S51. Obtain the image of the current tool captured by the camera, select the image of the current tool captured as the current tool image, input the current tool image and the current teaching image corresponding to the current tool image into the variational autoencoder, extract the feature vector of the current tool image through the variational autoencoder, and extract the feature vector of the current teaching image through the variational autoencoder; Among them, the variational autoencoder can map the high-dimensional data of the current tool image and the current teaching image to a low-dimensional latent space, which is beneficial to processing the high-dimensional data of the current tool image and the current teaching image.
[0056] S52: Through the mutual attention mechanism, fuse the feature vectors of the current tool image and the current teaching image to obtain a second fusion vector. Process the second fusion vector through the decoder in the trained diffusion model to obtain the current expected image, obtain a transmission instruction, execute the transmission instruction, transmit the current expected image to the visual servo system of the robot, and control the visual servo system to adjust the position and orientation of the current tool according to the current expected image.
[0057] Exemplarily, controlling the visual servo system to adjust the position and orientation of the current tool according to the current expected image includes: Controlling the visual servo system to process the current expected image to obtain the pose information of the current tool, and reading the position coordinates and rotation angles in the pose information; Controlling the end effector of the robot to adjust the position of the current tool according to the position coordinates, and controlling the end effector of the robot to adjust the orientation of the current tool according to the rotation angle.
[0058] In the embodiment of the present invention, through the mutual attention mechanism, the feature vectors of the current tool image and the current teaching image are fused to obtain a second fusion vector. The second fusion vector is processed through the decoder in the trained diffusion model to obtain the current expected image. Since there is no need to manually obtain the current expected image, the acquisition time of the current expected image is reduced, which is beneficial to improving the acquisition efficiency of the current expected image.
[0059] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a visual servo device in an embodiment of the present invention. As Figure 6 shown, the visual servo device includes a first acquisition module 101, a fusion module 102, a processing module 103, a second acquisition module 104, a stop module 105, and a control module 106. The detailed description of each functional module is as follows: The first acquisition module 101 is used to acquire the image of the preset tool captured by the camera, select the image of the preset tool captured as the preset tool image, and form a training sample with the preset tool image, the corresponding preset teaching image of the preset tool image, and the true expected image corresponding to the preset tool image; The fusion module 102 is used to train the diffusion model based on the training sample, and fuse the feature vectors of the preset tool image and the preset teaching image to obtain a first fusion vector; The processing module 103 is configured to process the first fusion vector through a decoder in the diffusion model to obtain a predicted desired image; The second acquisition module 104 is configured to acquire the total loss value between the predicted desired image and the true desired image; The stopping module 105 is configured to stop training the diffusion model to obtain a trained diffusion model when the total loss value is less than a preset value; The control module 106 is configured to acquire an image of the current tool captured by a camera, select the image of the current tool captured as the current tool image, fuse the feature vector of the current tool image and the feature vector of the current teaching image to obtain a second fusion vector, process the second fusion vector through a decoder in the trained diffusion model to obtain a current desired image, and transmit the current desired image to the visual servo system of the robot to control the visual servo system to adjust the position and orientation of the current tool according to the current desired image.
[0060] In the embodiments of the present invention, the beneficial effects are in two aspects. On the one hand, an image of the current tool captured by a camera is acquired, the image of the current tool captured is selected as the current tool image, the feature vector of the current tool image and the feature vector of the current teaching image are fused to obtain a second fusion vector, and the second fusion vector is processed through a decoder in the trained diffusion model to obtain a current desired image, which solves the technical problem of how to obtain the current desired image. Since there is no need for manual acquisition of the current desired image, the acquisition time of the current desired image is reduced, which is beneficial to improving the acquisition efficiency of the current desired image. On the other hand, the current desired image is transmitted to the visual servo system of the robot to control the visual servo system to adjust the position and orientation of the current tool according to the current desired image, which solves the technical problem of how to provide the current desired image to the visual servo system of the robot. Since the current desired image is automatically transmitted to the visual servo system of the robot and is not affected by manual intervention, it is beneficial to improve the reliability of the current desired image.
[0061] For the specific limitations of the visual servo device, reference may be made to the limitations on the visual servo method in the above text, which will not be elaborated here.
[0062] Each module in the above visual servo device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0063] Please refer to Figure 7 , Figure 7Another schematic structural diagram of a computer device in an embodiment of the present invention. In one embodiment, a computer device is provided. The computer device is a server device or a client device, and its internal structure diagram can be as shown in Figure 7 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with external devices. When the computer program is executed by the processor, it can implement the functions or steps of a visual servo method based on a diffusion model.
[0064] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor.
[0065] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can achieve, reference can be made to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0066] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU for short), a graphics processing unit (GPU for short), and a network processor (NP for short); it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0067] The above description and the drawings fully illustrate the embodiments of the present disclosure, enabling those skilled in the art to practice them. In this article, each embodiment may focus on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the methods and products disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, the relevant parts can refer to the description of the method part.
[0068] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. In the description corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A visual servoing method based on a diffusion model, characterized in that: include: Obtain an image obtained by shooting a preset tool with a camera, select the image obtained by shooting the preset tool as a preset tool image, and form a training sample by combining the preset tool image, a preset teaching image corresponding to the preset tool image, and a real expected image corresponding to the preset tool image; Based on the training sample, the diffusion model is trained to fuse the feature vector of the preset tool image and the feature vector of the preset teaching image to obtain a first fusion vector; The first fusion vector is processed by a decoder in the diffusion model to obtain a predicted expected image; Get the total loss value between the predicted expected image and the true expected image; When the total loss value is less than the preset value, stop training the diffusion model and obtain the trained diffusion model; The image of the current tool taken by the camera is obtained, and the image of the current tool is selected as the current tool image. The feature vector of the current tool image and the feature vector of the current teaching image are fused to obtain a second fused vector, and the second fused vector is processed by the decoder in the trained diffusion model to obtain the current expected image, and the current expected image is transmitted to the visual servo system of the robot, and the visual servo system is controlled to adjust the position and orientation of the current tool according to the current expected image.
2. The visual servoing method according to claim 1, characterized in that: The method of training the diffusion model based on the training samples, fusing the feature vector of the preset tool image and the feature vector of the preset teaching image to obtain a first fused vector, includes: Based on the training samples, the diffusion model is trained, and the preset tool images in the training samples and the preset teaching images corresponding to the preset tool images are input into the variational autoencoder; Feature extraction is performed on a preset tool image through a variational autoencoder to obtain a feature vector of the preset tool image, feature extraction is performed on a preset teaching image through a variational autoencoder to obtain a feature vector of the preset teaching image, and the feature vector of the preset tool image and the feature vector of the preset teaching image are fused through a mutual attention mechanism to obtain a first fused vector.
3. The visual servoing method according to claim 1, characterized in that: The first fusion vector is processed by a decoder in the diffusion model to obtain a predicted expected image, including: Obtaining a diffusion model, and inputting the third embedding vector into a decoder in the diffusion model; The third embedding vector is processed by the decoder in the diffusion model to obtain the predicted expected image.
4. The visual servoing method according to claim 1, characterized in that: The obtaining of the total loss value between the predicted expected image and the actual expected image includes: The first loss value of the predicted expected image and the true expected image is calculated by the KL divergence loss function, and the second loss value of the predicted expected image and the true expected image is calculated by the mean square error loss function; The first loss value and the second loss value are added to obtain the total loss value between the predicted expected image and the true expected image.
5. The visual servoing method according to claim 1, characterized in that: When the total loss value is less than a preset value, the diffusion model training is stopped to obtain the trained diffusion model, including: When the total loss value is less than the preset value, stop training the diffusion model and save the model parameters; The model parameters are loaded into the model structure, and the model structure loaded with the model parameters is selected as the trained diffusion model.
6. The visual servoing method according to claim 1, characterized in that: The method includes: obtaining an image of the current tool captured by a camera, selecting the image of the current tool captured as the current tool image, fusing a feature vector of the current tool image and a feature vector of the current teaching image to obtain a second fused vector, processing the second fused vector through a decoder in a trained diffusion model to obtain a current expected image, transmitting the current expected image to a visual servo system of the robot, and controlling the visual servo system to adjust the position and orientation of the current tool according to the current expected image, including: Obtain an image of the current tool captured by a camera, select the image captured by the current tool as the current tool image, input the current tool image and the current teaching image corresponding to the current tool image into a variational autoencoder, extract features of the current tool image through the variational autoencoder to obtain a feature vector of the current tool image, and extract features of the current teaching image through the variational autoencoder to obtain a feature vector of the current teaching image; Through the mutual attention mechanism, the feature vector of the current tool image and the feature vector of the current teaching image are fused to obtain a second fused vector. The second fused vector is processed by the decoder in the trained diffusion model to obtain the current expected image, and a transmission instruction is obtained. The transmission instruction is executed to transmit the current expected image to the robot's visual servo system, and the visual servo system is controlled to adjust the position and orientation of the current tool according to the current expected image.
7. The visual servoing method according to claim 1, characterized in that: In the acquisition camera taking the image of the current tool, the image taken by taking the current tool is selected as the current tool image, the feature vector of the current tool image and the feature vector of the current teaching image are fused to obtain a second fused vector, the second fused vector is processed by a decoder in the trained diffusion model to obtain a current expected image, the current expected image is transmitted to the visual servo system of the robot, and the visual servo system is controlled to adjust the position and orientation of the current tool according to the current expected image. The visual servo method includes: The adjustment result returned by the robot is obtained, a display window is created, and the adjustment result is displayed through the display window.
8. A visual servoing device based on a diffusion model, characterized in that: include: A first acquisition module is used to acquire an image obtained by shooting a preset tool with a camera, select the image obtained by shooting the preset tool as a preset tool image, and combine the preset tool image, a preset teaching image corresponding to the preset tool image, and a real expected image corresponding to the preset tool image into a training sample; A fusion module, used for training a diffusion model based on training samples, fusing a feature vector of a preset tool image with a feature vector of a preset teaching image to obtain a first fusion vector; A processing module, used for processing the first fusion vector through a decoder in the diffusion model to obtain a predicted expected image; The second acquisition module is used to obtain the total loss value between the predicted expected image and the actual expected image; A stop module is used to stop training the diffusion model when the total loss value is less than a preset value to obtain a trained diffusion model; A control module is used to obtain an image of the current tool taken by a camera, select the image taken by the current tool as the current tool image, fuse the feature vector of the current tool image and the feature vector of the current teaching image to obtain a second fused vector, process the second fused vector through a decoder in a trained diffusion model to obtain a current expected image, transmit the current expected image to the robot's visual servo system, and control the visual servo system to adjust the position and orientation of the current tool according to the current expected image.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the visual servoing method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the visual servoing method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method and system for visual servo controlling and equipment
CN110000795A
Method for controlling robot to run based on trajectory diffusion network
CN118034047A
Systems and methods for operating robots using visual servoing
US20130041508A1
Feature detection apparatus and methods for training of robotic navigation
US20160096270A1
Cited By
Robot control method, device and system based on diffusion model and medium
CN120828424A