Visual motion control method and device for automobile automatic charging system, storage medium and program product
Through the visual action control method of diffusion strategy, the Transformer structure and noise prediction network are used for end-to-end training, which solves the problem of charging gun insertion and unplugging of the robotic arm in complex environments, and improves the intelligence, stability, adaptability and robustness of the automatic charging system.
Patent Information
- Application Number
- CN202510536440.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-01
AI Technical Summary
In the existing automatic charging system, it is difficult for the robotic arm to accurately identify the charging port position in complex environments, resulting in low plug-in and unplugging success rate, lack of flexibility and independent learning ability, and insufficient intelligence.
The visual action control method using a diffusion strategy is used, and the diffusion model and noise prediction network of the Transformer structure are used to train end-to-end with the visual encoder, optimize the loss function, and improve the intelligence and robustness of the robotic arm through imitation learning and data augmentation.
It improves the stability and adaptability of the plug-and-release task of the robotic arm in complex environments, enhances the robustness to different vehicle models and vehicle positions, ensures the time consistency of the operation and the ability to respond to sudden interference.
Smart Images

Figure CN120395823A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of industrial robots, and particularly to a visual motion control method, a visual motion control device, a storage medium and a program product for an automatic charging system of an electric vehicle. Background Art
[0002] With the rapid growth of the electric vehicle market and the popularization of electric vehicles, the demand for electric vehicle charging infrastructure has also increased. Currently, all charging piles require manual operations for plugging and unplugging the charging gun. However, in this mode, the utilization rate of charging piles is low. Therefore, automatic charging has become a trend. In an automatic charging parking lot, a robotic arm is required to automatically perform the plugging and unplugging operations of the charging gun. These robotic arms usually rely on visual sensors and pre-programmed control algorithms to complete tasks.
[0003] Chinese Patent Application with Publication No. CN108508897A discloses a vision-based automatic charging system. This system uses a camera to sense environmental information, determines the relative position of the charging socket, and controls the robotic arm to align the charging gun with the charging socket and insert it. However, this method has problems of low intelligence, stability, and robustness in practice, and does not consider the changes in complex environments and different vehicle models.
[0004] Most current methods also have many other limitations in practical applications. For example, the intelligence level of the robotic arm is low. When the vehicle parking position is inaccurate, it is difficult for the robotic arm to accurately identify the position of the charging port, resulting in positioning failure and affecting the success rate of plugging and unplugging. In addition, the control system often requires a large amount of manual adjustment, lacking flexibility and the ability of autonomous learning. Summary of the Invention
[0005] To solve the above problems, a method for automatically charging a robotic arm using deep learning technology, especially reinforcement learning, can be used to achieve a more intelligent and flexible charging robotic arm. However, the method based on reinforcement learning has certain limitations in dealing with multi-modal action distributions, high-dimensional action spaces, and training stability. In addition, the utilization rate of samples during the training process is low.
[0006] Diffusion policy is an emerging deep learning method that represents the visual motion policy of a robot as a conditional denoising diffusion process. This method can learn the gradient of the action distribution and optimize it through a series of stochastic Langevin dynamics steps during the inference process. The diffusion policy can not only gracefully handle multi-modal action distributions, but also be applicable to high-dimensional action spaces and exhibit excellent training stability. The diffusion policy provides effective technical support for the automatic charging robotic arm to complete the charging gun plugging and unplugging tasks.
[0007] Technical Problem This invention aims to provide a diffusion-based visual motion control method for an automatic charging robot arm, addressing the existing technical challenges of enabling the robot arm to automatically plug and unplug charging guns in complex environments. Specifically, this invention aims to develop an intelligent control system that enables the robot arm to automatically plug and unplug charging guns in an automatic charging parking lot, thereby improving the stability, adaptability, robustness, and intelligence of the automatic charging robot arm in complex environments.
[0008] Technical Solution According to a first aspect of the present invention, a visual motion control method for an automatic charging system for an automobile is provided, wherein the automatic charging system for an automobile includes two visual sensors, a robotic arm, a controller, and a charging socket. The visual motion control method includes: a diffusion model structure selection step (S100), selecting the structure of a noise prediction network of a diffusion model based on a Transformer structure; a training data collection and processing step (S200), collecting expert demonstration videos of the charging system's robotic arm plugging and unplugging a charging gun, and processing the collected data, wherein the processing includes video segmentation, key frame labeling, and checking the temporal synchronization of visual data and operation motion data; and a diffusion model training and verification step (S300), performing end-to-end training on the noise prediction network selected in the diffusion model structure selection step and the visual encoder used in the diffusion model structure selection step to optimize the loss function.
[0009] As a preferred solution, according to the visual action control method of the present invention, in the diffusion model structure selection step (S100), the noise prediction network of the diffusion model is represented as , and will have Wheel noise action sequence and the number of diffusion iterations The sinusoidal embedding is passed as input to the Transformer decoder, while the number of diffusion iterations The sinusoidal embedding of is the first input value, and the visual observation It serves as the generation condition of the diffusion model to guide the generation of action sequences.
[0010] As a preferred solution, according to the visual action control method of the present invention, the loss function is expressed as follows: ,in, It is the result of k rounds of Gaussian noise superimposed together. The parameter is The noise prediction network, and is the time-synchronized visual sequence information and action sequence information in the training data, where is obtained through visual encoder processing, and is corresponding to correspondingly the mechanical arm insertion and extraction action sequence without added noise at the corresponding moment.
[0011] As a preferred solution, according to the visual action control method of the present invention, wherein, in the visual observation in the diffusion model structure selection step (S100) O t is obtained by mapping the original visual image sequence obtained by the visual sensor by the visual encoder, and wherein, the visual encoder uses the ResNet model.
[0012] As a preferred solution, according to the visual action control method of the present invention, wherein, the visual encoder extracts high-dimensional information from the original visual image and converts it into the input form of the Transformer as the guiding condition for action sequence denoising.
[0013] As a preferred solution, according to the visual action control method of the present invention, wherein, in the diffusion model structure selection step (S100), by using the method based on the denoising diffusion implicit model to decouple the iterative number relationship between noise addition in the training process and denoising in the inference process, so as to accelerate the generation speed of the action sequence.
[0014] As a preferred solution, according to the visual action control method of the present invention, wherein, in the training data collection and processing step (S200), while operating the mechanical arm to complete the insertion and extraction task, two kinds of data records including the recording of the visual image sequence of the visual sensor and the recording of the action instruction sequence of the mechanical arm are carried out, and the two kinds of data records are consistent in time.
[0015] As a preferred solution, according to the visual action control method of the present invention, wherein, in the training data collection and processing step (S200), in the data record, the change of the initial conditions including the relative position of the mechanical arm and the vehicle is considered, and the change of the environmental information including the vehicle type, the position of the charging socket and the illumination condition is considered.
[0016] As a preferred solution, according to the visual action control method of the present invention, wherein, the diffusion model training and verification step (S300) further includes, before jointly training the noise prediction network and the visual encoder end-to-end, performing data augmentation by trimming and transforming the expert demonstration video and the action sequence, and dividing the training set and the verification set.
[0017] According to a second aspect of the present invention, there is provided a visual motion control device for an automatic charging system, the automatic charging system including two visual sensors, a robotic arm, a controller, and a charging socket. The visual motion control device includes: a diffusion model structure selection unit (400) that selects the structure of the noise prediction network of the diffusion model based on the Transformer structure; a training data collection and processing unit (200) that collects expert demonstration videos of the robotic arm of the charging system inserting and removing the charging gun, and processes the collected data, the processing including video segmentation, key frame annotation, and checking the time synchronization of visual data and operation action data; and a diffusion model training and verification unit (300) that performs end-to-end training on the noise prediction network and the visual encoder to optimize the loss function.
[0018] According to a third aspect of the present invention, there is provided a non-transitory storage medium storing a computer program that, when executed by a processor, can implement the visual motion control method for an automotive automatic charging system according to the first aspect of the present invention.
[0019] According to a fourth aspect of the present invention, there is provided a computer program product including computer instructions that, when executed by a processor, can implement the visual motion control method for an automotive automatic charging system according to the first aspect of the present invention.
[0020] Advantageous technical effects Through the visual motion control method and the visual motion control device of the automotive automatic charging system according to the above embodiments, using the diffusion strategy as the visual motion execution strategy of the charging robotic arm, on the one hand, it can make full use of the advantages of data-driven, improve the intelligence level of the charging robotic arm, improve the adaptability to the environment and the robustness against different vehicle models and vehicle positions; on the other hand, it overcomes the problems existing in the learning-based control strategy in practical applications, has a relatively high training stability, and can learn effective strategies from multi-modal training data. In addition, generating action sequences through the diffusion model makes the actions have good temporal consistency and smoothness, and through the method of rolling planning and execution, it can timely respond to sudden disturbances, thus taking into account the continuity of the actions of the robotic arm during the charging gun insertion and removal tasks in terms of time and the responsiveness to disturbances.
[0021] Other features of the present invention will become clear through the following description of exemplary embodiments with reference to the accompanying drawings. Description of the drawings
[0022] Figure 1 is a block diagram showing the hardware configuration of the automotive automatic charging system according to the present invention; Figure 2It is a flowchart showing a visual action control method of an automatic charging system for an automobile according to the present invention; Figure 3 It is a schematic diagram showing an overall overview of a visual action strategy of an automatic charging robotic arm based on a diffusion strategy according to the present invention; Figure 4 It is a schematic diagram showing a diffusion strategy network structure based on a Transformer architecture according to the present invention; Figure 5 It is a functional block diagram of a visual action control device of an automatic charging system according to the present invention. Detailed implementation manners
[0023] The present invention will be described in detail below with reference to the accompanying drawings showing embodiments of the present invention.
[0024] The following describes exemplary embodiments of the present invention in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative configurations of components, numerical representations, and numerical values described in these embodiments do not limit the scope of the present invention.
[0025] The visual action control method of the automatic charging system for an automobile of the present invention can be implemented by a processor in the automatic charging system executing a computer program stored on a memory in the automatic charging system. As an alternative solution, it can also be implemented by the automatic charging system communicating with a server, and a processor in the server executing a computer program stored on the server and feeding back the program execution result to the automatic charging system in real time.
[0026] In the present invention, the term "unit" may refer to a software environment, a hardware environment, or a combination of a software and a hardware environment. In a software environment, the term "unit" refers to a functionality, an application, a software module, a function, a routine, a set of instructions, or a program that can be executed by a programmable processor, such as a microprocessor, a central processing unit (CPU), or a specially designed programmable device, or a controller. The memory contains instructions or programs that, when executed by the CPU, cause the CPU to perform operations corresponding to the unit or function. In a hardware environment, the term "unit" refers to a hardware element, a circuit, a component, a physical structure, a system, a module, or a subsystem. According to a specific embodiment, the term "unit" may include mechanical, optical, or electrical components, or any combination thereof. The term "unit" may include active (e.g., transistors) or passive (e.g., capacitors) components. The term "unit" may include semiconductor devices having a substrate and other material layers having various conductive concentrations. It may include a CPU or a programmable processor that can execute a program stored in the memory to perform a specified function. The term "unit" may include logic elements (e.g., AND, OR) implemented by transistor circuits or any other switching circuits. In a combination of a software and a hardware environment, the term "unit" or "circuit" refers to any combination of the software and hardware environments as described above. In addition, the terms "element", "component", "part", or "device" may also refer to a "circuit" integrated or not integrated with a packaging material.
[0027] [Hardware Configuration of Automatic Charging System] The following will refer to Figure 1 Describe the hardware configuration of the automatic charging system according to an embodiment of the present invention.
[0028] The vehicle automatic charging system 100 of the present invention includes a first vision sensor 110, a second vision sensor 120, a robotic arm 130, a controller 140, and a charging socket 150.
[0029] The first vision sensor 110 and the second vision sensor 120 are responsible for sensing external environment information. The first vision sensor 110 is installed near the end of the robotic arm and is used to capture high-definition images of the charging socket, while the second vision sensor 120 is installed near the robotic arm and is used to capture high-definition images of the overall environment.
[0030] The robotic arm 130 has multiple degrees of freedom and can perform six degrees of freedom of spatial operations. The execution end of the robotic arm 130 is fixedly connected to a charging gun that matches the charging socket. The robotic arm 130 is used to execute the action instructions sent by the controller and insert the charging gun at the end into the charging socket of the electric vehicle.
[0031] The controller 140 may further include a processor 141 (such as a CPU, MPU, etc.), a memory 142, a RAM 143, a communication interface 144, etc. Among them, various application programs (such as a visual motion control program, a visual information processing program, a strategy generation program, etc.) are stored on the memory 142, and the RAM 143 is the working space of the processor 141. Through the communication interface 144, the controller 140 can communicate with other hardware parts of the vehicle automatic charging system 100 (such as the first visual sensor 110, the second visual sensor 120, the robotic arm 130, etc.). Thus, the controller 140 can receive the information sensed by the sensors, send motion instructions to the robotic arm, and perform processes such as visual information processing, generating motions according to visual motion strategies, etc. In addition, through the communication interface 144, the vehicle automatic charging system 100 can also communicate with external devices such as a server (not shown) connected through a network to download application programs from the server, send information to the server, and receive the operation results of the application programs running on the server, etc.
[0032] The charging socket 150 on the electric vehicle, as the target that the end of the robotic arm needs to reach, can start the charging service of the charging pile after inserting the charging gun.
[0033] [Visual Motion Control Method of Automatic Charging System] Next, with reference to Figures 2 to 4 A visual motion control method of a vehicle automatic charging system according to the present invention will be described. This visual motion control method can be implemented by the processor 141 in the vehicle automatic charging system 100 according to the present invention reading the application programs stored on the memory 142, or can be implemented by the vehicle automatic charging system 100 according to the present invention communicating with the server, and the server executing the application programs stored on the server and sending the execution results to the vehicle automatic charging system 100, or can be implemented by the server and the vehicle automatic charging system 100 according to the present invention respectively executing the corresponding application programs and communicating the execution results with each other.
[0034] Figure 2 is a flowchart showing a visual motion control method of a vehicle automatic charging system according to the present invention, Figure 3 is a schematic diagram showing an overall overview of a visual motion strategy of an automatic charging robotic arm based on a diffusion strategy according to the present invention and Figure 4 is a schematic diagram showing a diffusion strategy network structure based on a Transformer architecture according to the present invention.
[0035] First, in step S100, a diffusion model structure is selected. The noise prediction network of the diffusion model in the present invention selects a structure based on Transformer. Specifically, asFigure 4 As shown, with k the action sequence with wheel noise and the number of diffusion iterations k The sine embedding of is fed into the decoder of the Transformer as input. At the same time, the sine embedding of the number of diffusion iterations k is the first input value.
[0036] In addition, the noise prediction network is also related to the visual observation Here, the visual observation is used as the generation condition of the diffusion model to guide the generation of the action sequence. Therefore, the visual observation is transformed into a visual observation embedding sequence through a shared multi-layer perceptron, and then it is also fed into the decoder of the Transformer as input. The output of the decoder of the Transformer represents the predicted noise, that is, the gradient of the action sequence denoising process. So the whole network can be represented by denote.
[0037] The above network structure can reduce the over-smoothing effect in the commonly used convolutional neural network model and improve the performance of the diffusion strategy when facing complex tasks or high-speed action changes.
[0038] The visual observation in step S100 is mapped from the original visual image sequence, and this process is completed by the visual encoder, which is trained end-to-end with the diffusion strategy. Since there are two visual sensors such as cameras in this system, they respectively use independent visual encoders, and the images in each time step are independently encoded, and then they are concatenated to form . Specifically, the visual encoder can use the ResNet model.
[0039] The role of the visual encoder is to extract high-dimensional information from the original image and transform it into the input form of the Transformer to be used as the guiding condition for action sequence denoising. This design will help capture the key features and information in the original image, and through end-to-end training, the visual encoder can better adapt to the needs of the diffusion strategy.
[0040] In the specific implementation process, the control of the manipulator operation requires good real-time performance, and the diffusion strategy is used for the diffusion model of action sequence generation, and each time it needs to go through kThe noise reduction process of the wheel makes the inference speed of the model slow. However, having a fast inference speed is crucial for closed-loop real-time control. Therefore, the present invention draws on the method of the Denoising Diffusion Implicit Model (DDIM), which can decouple the relationship between the number of iterations of adding noise during training and denoising during inference, enabling the algorithm to use fewer iterations for denoising during inference, thereby accelerating the generation speed of the action sequence.
[0041] Then, the process proceeds to step S200. In step S200, the collection and processing of training data are carried out. Expert demonstration videos of the robotic arm of the charging system inserting and removing the charging gun are collected, and the collected data is processed. The processing includes video segmentation, annotation of key frames, and checking the time synchronization of visual data and operation action data.
[0042] Specifically, the diffusion strategy is learned from a large number of expert demonstration videos, which belongs to the category of imitation learning. Therefore, in the task of the robotic arm automatically inserting and removing the charging gun, to train the model, a large number of diverse expert demonstration videos of the robotic arm inserting and removing the charging gun need to be collected.
[0043] In the actual implementation process, the operator controls the robotic arm on the above system to complete the insertion and removal actions of the charging gun through teleoperation. It should be noted that it is necessary to ensure that the operator has received sufficient training, is familiar with the task, and can proficiently complete the operation of the robotic arm to complete the insertion and removal task. During the operation, data recording is required, including the recording of the visual image sequences of two visual sensors such as cameras, and the recording of the action instruction sequences of the robotic arm, and the two parts of data need to ensure temporal consistency. In addition, it should be noted that to increase data diversity, the initial conditions need to be randomly changed, such as the relative position of the robotic arm and the vehicle; the environmental information also needs to be changed, such as the type of vehicle, the position of the charging socket, and the lighting conditions, etc. Specifically, for different initial conditions (such as the relative position of the robotic arm and the vehicle, etc.), two types of data records including the recording of the visual image sequences of the visual sensor and the recording of the action instruction sequences of the robotic arm are carried out, and for different environmental information (such as the type of vehicle, the position of the charging socket, and the lighting conditions, etc.), two types of data records including the recording of the visual image sequences of the visual sensor and the recording of the action instruction sequences of the robotic arm are carried out.
[0044] Next, the collected data is processed. First, video segmentation is performed, that is, the recorded long video is segmented into multiple demonstration segments; then key frames are annotated, that is, key operation points are manually or algorithmically annotated, such as the time point when starting to contact the object; finally, the time synchronization of visual data and operation action data is checked and ensured.
[0045] Next, the process proceeds to step S300, where the model is trained and verified.
[0046] In the diffusion strategy, the noise prediction network and the visual encoder are jointly trained end-to-end to optimize the loss function. According to the above theoretical derivation of the diffusion model, the loss function is expressed as follows: where, is the result of superimposing k rounds of Gaussian noise together, is the noise prediction network with parameter , and are the visually sequential information and action sequential information synchronized in time in the training data, where is obtained by processing through the visual encoder, while is corresponding to and is the robotic arm plugging and unplugging action sequence without added noise at the t moment.
[0047] Before training, data augmentation can be performed by performing certain cropping and transformation on the demonstration video and action sequence, and the training set and validation set can be divided.
[0048] Since there are many parameters in the diffusion strategy, the model parameters need to be adjusted to a certain extent according to the actual training effect. In addition, different parameters, different network structures, and different method strategies can be selected to train some baseline models to compare the performance of the trained visual action model based on the diffusion strategy, so as to reflect the superiority of the diffusion strategy. Finally, the parameters with the best performance are selected and deployed into the diffusion strategy.
[0049] [Software Structure of the Visual Action Control Device for the Automatic Charging System] Next, the software structure of the visual action control device 10 for the automatic charging system according to an embodiment of the present invention will be described with reference to Figure 5 .
[0050] As Figure 5 shown, the visual action control device 10 for the vehicle automatic charging system 100 according to an embodiment of the present invention includes a diffusion model structure selection unit 400, a training data collection and processing unit 200, and a diffusion model training and verification unit 300.
[0051] The diffusion model structure selection unit 400 selects the structure of the noise prediction network of the diffusion model based on the Transformer structure , where the action sequence with rounds of noise and the diffusion iteration times The sine embedding is fed into the decoder of the Transformer as input, along with the number of diffusion iterations The sine embedding serves as the first input value, and the visual observation is used as the generation condition of the diffusion model to guide the generation of the action sequence.
[0052] The training data collection and processing unit 200 collects the expert demonstration videos of the robotic arm plugging and unplugging the charging gun of the charging system, and processes the collected data. The processing includes video segmentation, annotating key frames, and checking the time synchronization of the visual data and the operation action data. The diffusion model training and verification unit 300 performs end-to-end training on the noise prediction network and the visual encoder together. The loss function is expressed as follows: where and are the visually sequential information and the action sequential information that are time-synchronized in the training data. Among them, is obtained by processing through the visual encoder, while is corresponding to the robotic arm plugging and unplugging action sequence without added noise at the
[0053] By using the visual action control method and the visual action control device of the automatic vehicle charging system according to the above embodiments, and using the diffusion strategy as the visual action execution strategy of the charging robotic arm, on the one hand, the advantages of data-driven can be fully utilized to improve the intelligence level of the charging robotic arm, enhance the adaptability to the environment, and the robustness against different vehicle models and vehicle positions; on the other hand, the problems existing in the learning-based control strategy in practical applications are overcome, the training stability is relatively high, and effective strategies can be learned from multi-modal training data. In addition, generating the action sequence through the diffusion model makes the action have good temporal consistency and smoothness, and the method of rolling planning and execution can timely handle sudden disturbances, thus taking into account both the temporal continuity of the action of the robotic arm when performing the charging gun plugging and unplugging task and the responsiveness to disturbances.
[0054] Other embodiments Embodiments of the present invention can also be implemented by a computer of a system or device that reads and executes computer-executable instructions (e.g., one or more programs) recorded on a storage medium (which may also be more fully referred to as a "non-transitory computer-readable storage medium") to perform one or more functions of the above-described embodiments and / or includes one or more circuits (e.g., an application specific integrated circuit (ASIC)) for performing one or more functions of the above-described embodiments. Moreover, embodiments of the present invention can be implemented by a method of, for example, reading and executing the computer-executable instructions from the storage medium by the computer of the system or device to perform one or more functions of the above-described embodiments and / or controlling the one or more circuits to perform one or more functions of the above-described embodiments. The computer may include one or more processors (e.g., a central processing unit (CPU), a microprocessing unit (MPU)), and may include a network of separate computers or separate processors to read and execute the computer-executable instructions. The computer-executable instructions may be provided to the computer, for example, from a network or the storage medium. The storage medium may include, for example, one or more of a hard disk, a random access memory (RAM), a read-only memory (ROM), a memory of a distributed computing system, an optical disc (such as a compact disc (CD), a digital versatile disc (DVD), or a Blu-ray Disc (BD)™), a flash device, and a memory card, etc.
[0055] Embodiments of the present invention can also be implemented by the following method, that is, by providing software (program) that performs the functions of the above-described embodiments to a system or device through a network or various storage media, and the method of reading and executing the program by a computer or a central processing unit (CPU), a microprocessing unit (MPU) of the system or device.
[0056] Although the present invention has been described with reference to exemplary embodiments, it should be understood that the present invention is not limited to the disclosed exemplary embodiments. The scope of the appended claims should be given the broadest interpretation so as to cover all such variations and equivalent structures and functions.
Claims
1. A visual motion control method for an automatic vehicle charging system, characterized in that, The automatic vehicle charging system includes two vision sensors, a robotic arm, a controller, a communication interface, and a charging socket. The vision action control method includes: A diffusion model structure selection step of selecting the structure of the noise prediction network of the diffusion model based on the Transformer structure; A training data collection and processing step of collecting the expert demonstration videos of the robotic arm of the automatic vehicle charging system inserting and removing the charging gun, and processing the collected data, where the processing includes video segmentation, annotating key frames, and checking the time synchronization of visual data and operation action data; and A diffusion model training and verification step of jointly performing end-to-end training on the noise prediction network selected in the diffusion model structure selection step and the vision encoder used in the diffusion model structure selection step to optimize the loss function.
2. The visual action control method according to claim 1, wherein In the diffusion model structure selection step, represent the noise prediction network of the diffusion model as , and use the action sequence with k rounds of noise and the number of diffusion iterations as the sine embedding of the input into the decoder of the Transformer. At the same time, use the sine embedding of the number of diffusion iterations as the first input value, and use the visual observation as the generation condition of the diffusion model to guide the generation of the action sequence.
3. The visual motion control method according to claim 1 or 2, characterized in that, Loss function L is expressed as follows: wherein, is the result of superimposing k rounds of Gaussian noise, is a noise prediction network with a parameter of . and are the visually sequential information and action sequential information that are time-synchronized in the training data. Among them, is the visually sequential information obtained by processing through a visual encoder, while is corresponding to the robot arm plugging and unplugging action sequence without noise at the moment.
4. The vision action control method according to claim 3, wherein Visual Observation in the Diffusion Model Structure Selection Step O t It is obtained by mapping the original visual image sequence obtained by the visual sensor by the visual encoder, and wherein, the visual encoder uses the ResNet model.
5. The vision action control method according to claim 4, wherein The vision encoder extracts high-dimensional information from the original visual image and converts it into the input form of the Transformer as the guiding condition for denoising the action sequence.
6. The vision action control method according to claim 3, wherein In the diffusion model structure selection step, by using the method based on the denoising diffusion implicit model to decouple the iterative number relationship between noise addition in the training process and denoising in the inference process, the generation speed of the action sequence is accelerated.
7. The vision action control method according to claim 3, wherein In the training data collection and processing step, while operating the robotic arm to complete the insertion and removal task, two types of data records are performed, including recording the visual image sequence of the vision sensor and the action instruction sequence of the robotic arm, and the two types of data records are consistent in time.
8. The vision action control method according to claim 7, wherein In the training data collection and processing step, in the data record, the change of the initial condition including the relative position of the robotic arm and the vehicle is increased, and the change of the environmental information including the vehicle type, the position of the charging socket, and the lighting condition is increased.
9. The vision action control method according to claim 3, wherein The diffusion model training and verification step further includes, before jointly performing end-to-end training on the noise prediction network and the vision encoder, performing data augmentation by cropping and transforming the expert demonstration videos and action sequences, and dividing the training set and the verification set.
10. A visual motion control device for an automatic charging system, characterized in that, The automatic charging system includes two vision sensors, a robotic arm, a controller, and a charging socket. The vision action control device includes: A diffusion model structure selection unit that selects the structure of the noise prediction network of the diffusion model based on the Transformer structure; A training data collection and processing unit that collects the expert demonstration videos of the robotic arm of the charging system inserting and removing the charging gun, and processes the collected data, where the processing includes video segmentation, annotating key frames, and checking the time synchronization of visual data and operation action data; and The diffusion model training and validation unit performs end-to-end training on the noise prediction network and the visual encoder together to optimize the loss function.
11. A non-transitory storage medium stores a computer program, characterized in that, When executed by a processor, the computer program is capable of implementing the visual action control method for an automatic charging system according to any one of claims 1-9.
12. A computer program product, comprising computer instructions, characterized in that, When executed by a processor, the computer instructions are capable of implementing the visual action control method for an automatic charging system according to any one of claims 1-9.
Citation Information
Patent Citations
Alignment system and method for automatic charging of robots based on vision
CN108508897A
Charging mechanical arm control method and system
CN112248835A
Visual navigation model training method, system and equipment and readable storage medium
CN119068231A
Robot arm control method based on video diffusion model and related equipment
CN119399676A
Realistic, controllable agent simulation using guided trajectories and diffusion models
US20240160888A1
Cited By
Model training method and device for robots of various configurations and storage medium
CN121340258A