Control method and related equipment

By acquiring and fusion of multimodal data and using reinforcement learning models to control mechanical equipment, the problem of difficulty in achieving high-precision assembly in the environment of task changes is solved, and the success rate and robustness of assembly tasks are improved.

CN120215318APending Publication Date: 2025-06-27HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311832633.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In industrial scenarios, high-precision assembly tasks are automatically performed through mechanical equipment such as robots, especially in the changing environment of tasks, making it difficult to achieve high-reliability plug-in and unplug.

Method used

A control method is adopted to obtain multimodal data (including image modal data, mechanical data and motion data), fuse these data, and use the trained reinforcement learning model to control the mechanical equipment to perform assembly tasks. This method acquires the target pose of the mechanical device in the teaching mode and trains the reinforcement learning model through multiple sets of training data to improve the success rate and robustness of the assembly task.

Benefits of technology

It improves the success rate and robustness of assembly tasks of mechanical equipment in various scenarios, enhances the anti-interference ability when environmental changes, and ensures high reliability of assembly tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215318A_ABST
    Figure CN120215318A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a control method, in which multi-modal data such as images, mechanical data and motion data can be fused to obtain fused data, and multi-dimensional feature information is extracted from the fused data through a trained reinforcement learning model to determine a control strategy for mechanical equipment. Therefore, the mechanical equipment can successfully complete the assembly task according to the multi-dimensional feature information in various scenes, and the success rate and robustness of executing the assembly task by the mechanical equipment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of control technology, and in particular, to a control method and related devices. Background Art

[0002] The rapid development of the Internet and informatization has promoted the development of robotics technology. Robots have gradually emerged in various industries, helping people complete various complex tasks.

[0003] However, currently, in many industrial scenarios, it is still a huge challenge to automatically perform high-precision assembly tasks such as shaft-hole assembly through mechanical devices such as robots.

[0004] Taking the shaft-hole assembly task as an example, in daily life and industrial production, common shaft-hole assembly tasks can include charger plugging and unplugging, key plugging and unplugging, shaft-hole plugging and unplugging, etc. In high-precision industrial-level assembly scenarios, it is still difficult to achieve high-reliability plugging and unplugging in a task-changing environment. Summary of the Invention

[0005] Embodiments of this application provide a control method to solve the problem that it is difficult to complete high-precision assembly tasks such as shaft-hole assembly through mechanical devices such as robots. This application also provides corresponding devices, equipment, computer-readable storage media, computer program products, etc.

[0006] A first aspect of this application provides a control method, which includes: obtaining multi-modal data, where the multi-modal data includes at least two of image modal data, mechanical data, and motion data, and the multi-modal data includes the assembly position indicated by the assembly task and information about the mechanical device to perform the assembly task; fusing the multi-modal data to obtain fused data; and controlling the mechanical device to perform the assembly task according to the fused data through a trained reinforcement learning model.

[0007] In the first aspect, multi-modal data such as images, mechanical data, and motion data can be fused to obtain fused data, and a trained reinforcement learning model can be used to extract multi-dimensional feature information from the fused data to determine the control strategy for the mechanical device, so that the mechanical device can successfully complete the assembly task according to the multi-dimensional feature information in various scenarios, improving the success rate and robustness of the mechanical device in performing the assembly task.

[0008] In a possible implementation of the first aspect, the method further includes: in the teaching mode, obtaining the target pose of the mechanical device, where the target pose indicates that the relative position between the mechanical device and the assembly position indicated by the assembly task meets the specified conditions; according to the target pose, obtaining the assembly path corresponding to each initial pose of the mechanical device, and taking the initial fusion data corresponding to an initial pose and the corresponding assembly path as a set of training data; according to one set or multiple sets of training data, training the reinforcement learning model to obtain the trained reinforcement learning model.

[0009] In this possible implementation, the target pose can be determined through one teaching, so that through the target pose, the repeated positioning of the mechanical device can be realized. It can be seen that the target pose determined in this teaching mode can be used as a virtual expert, and can realize the automatic repeated positioning and correct assembly of the mechanical device at any random initial pose, so as to efficiently obtain multiple sets of training data, and efficiently provide training data for the subsequent training of the reinforcement learning model, without a large amount of manual operation, greatly improving the training efficiency.

[0010] In a possible implementation of the first aspect, the training of the reinforcement learning model includes one or more iteration processes. Among them, the i-th iteration process includes one or more assembly processes. Each assembly process includes at least one action. Each action corresponds to a set of state transition data. Each set of state transition data includes the first fusion data before the corresponding action, the parameter information of the corresponding action, the second fusion data after the corresponding action, and the reward information of the corresponding action. The reward information is determined according to the assembly result of the assembly process to which the corresponding action belongs, and i is a positive integer; in the i-th iteration process, the method includes: updating the reinforcement learning model of the i-th iteration process according to one set or multiple sets of state transition data of the i-th iteration process.

[0011] In a possible implementation of the first aspect, in the j-th assembly process of the i-th iteration process, the method includes: determining at least one action of the j-th assembly process according to the probability of the i-th iteration process and the probability threshold, and according to the initial pose of the i-th iteration process, where j is a positive integer.

[0012] In this possible implementation, the probability can reflect the possibility of successfully completing the assembly process. In different probability intervals (for example, when it is less than the probability threshold and when it is greater than the probability threshold), different methods can be used to determine the actions in the assembly process.

[0013] In a possible implementation of the first aspect, according to the probability of the $i$-th iteration process and the probability threshold, and based on the initial pose of the $i$-th iteration process, at least one action of the $j$-th assembly process is determined, including: if the probability of the $i$-th iteration process is greater than the probability threshold, then based on the first fusion data before the action of the $j$-th assembly process, through the reinforcement learning model of the $i$-th iteration process, the action of the $j$-th assembly process is determined; if the probability of the $i$-th iteration process is not greater than the probability threshold, then based on the pose of the mechanical equipment and the target pose before the action of the $j$-th assembly process, the action of the $j$-th assembly process is determined.

[0014] In this possible implementation, if the probability of the $i$-th iteration process is not greater than the probability threshold, it indicates that the possibility of successfully completing the assembly during the assembly process of the $i$-th iteration is relatively low. At this time, it may be in the early stage of the training process.

[0015] In order to avoid damage to the mechanical equipment and the target position caused by incorrect assembly processes, and to improve the training efficiency and accelerate the convergence speed, when the probability of the $i$-th iteration process is not greater than the probability threshold, the action of the $j$-th assembly process can be determined based on the first fusion data and the target pose before the action of the $j$-th assembly process.

[0016] If the probability of the $i$-th iteration process is greater than the probability threshold, it indicates that the possibility of successfully completing the assembly during the assembly process of the $i$-th iteration is relatively high. At this time, it may be in the later stage of the training process.

[0017] At this time, the first fusion data before the action of the $j$-th assembly process can be processed through the reinforcement learning model of the $i$-th iteration process to obtain the action of the $j$-th assembly process.

[0018] In a possible implementation of the first aspect, the method further includes: updating the probability of the $i$-th iteration process according to the success rate of one or more assembly processes included in the $i$-th iteration process.

[0019] In a possible implementation of the first aspect, controlling the mechanical equipment to perform the assembly task according to the fusion data through the trained reinforcement learning model includes: determining the expected position of the mechanical equipment in the next step according to the fusion data through the trained reinforcement learning model; determining the expected speed of the mechanical equipment in the next step according to the expected position in the next step, as well as the stiffness characteristics, damping characteristics, and inertia characteristics of the mechanical equipment; controlling the mechanical equipment to move according to the expected speed in the next step.

[0020] In this possible implementation, after determining the expected speed for the next step, the mechanical equipment can be controlled according to the expected speed for the next step to achieve flexible motion characteristics, avoiding collisions with objects such as the assembly position during the execution of the assembly task and causing equipment damage, thus ensuring the safety of the mechanical equipment during the execution of the assembly task.

[0021] In one possible implementation of the first aspect, controlling the movement of the mechanical equipment according to the expected speed for the next step includes: controlling the movement of the mechanical equipment according to the expected speed for the next step and the force feedback information of the mechanical equipment.

[0022] In this possible implementation, the force feedback information of the mechanical equipment is used to reflect the external force situation currently received by the mechanical equipment, so as to judge whether the mechanical equipment is currently under impact or other situations. If the force feedback information indicates that the mechanical equipment is not currently subject to external force, the expected speed for the next step is used as the execution speed to control the mechanical equipment to move according to the execution speed; while if the force feedback information indicates that the mechanical equipment is currently subject to external force, the smaller value of the expected speed for the next step and the preset speed can be used as the execution speed to control the mechanical equipment to move according to the execution speed, thereby controlling the mechanical equipment to achieve flexible motion characteristics, avoiding collisions with objects such as the assembly position during the execution of the assembly task and causing equipment damage, and ensuring the safety of the mechanical equipment during the execution of the assembly task.

[0023] In one possible implementation of the first aspect, obtaining multi-modal data includes: obtaining data of different modalities in the multi-modal data through multiple processes respectively.

[0024] In this possible implementation, data of different modalities in the multi-modal data can be obtained through multiple processes respectively, so that parallel acquisition and processing of data of different modalities can be realized, and the acquisition and processing efficiency of data of different modalities can be improved.

[0025] In one possible implementation of the first aspect, the multi-modal data includes image modality data, and also includes mechanical data and / or motion data; fusing the multi-modal data to obtain fused data includes: processing the image modality data through an encoder to obtain a feature representation corresponding to the image modality data; fusing the feature representation with the mechanical data and / or motion data to obtain fused data.

[0026] In this possible implementation, the image modality data can be feature-extracted through an encoder, mapping the image modality data to a low-dimensional feature representation (for example, it can be a one-dimensional feature representation), so that the low-dimensional feature representation can be concatenated with the one-dimensional mechanical data and / or motion data, thereby realizing effective fusion of the multi-modal data and obtaining fused data.

[0027] In a possible implementation of the first aspect, before obtaining the multi-modal data, the method further includes: controlling a mechanical device to perform rough positioning according to the assembly position.

[0028] In this possible implementation, rough positioning means that the distance between the mechanical device and the target position is within a specified range. For example, it is less than a specified distance threshold, so as to facilitate performing a high-precision assembly task with higher precision in the subsequent process. Exemplarily, the mechanical device can be controlled to perform rough positioning according to the depth image information and the assembly position.

[0029] The second aspect of the present application provides a control device, which has the function of implementing the method of the above first aspect or any possible implementation manner of the first aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, such as an acquisition module and a processing module.

[0030] The third aspect of the present application provides an electronic device, which includes at least one processor, a memory, and computer-executable instructions stored in the memory and executable on the processor. When the computer-executable instructions are executed by the processor, the processor executes the method of the above first aspect or any possible implementation manner of the first aspect.

[0031] The fourth aspect of the present application provides a computer-readable storage medium storing one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the processor executes the method of the above first aspect or any possible implementation manner of the first aspect.

[0032] The fifth aspect of the present application provides a computer program product storing one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the processor executes the method of the above first aspect or any possible implementation manner of the first aspect.

[0033] The sixth aspect of the present application provides a chip system, which includes a processor for supporting an electronic device to implement the functions involved in the above first aspect or any possible implementation manner of the first aspect. In a possible design, the chip system may further include a memory, and the memory is used to store necessary program instructions and data of the electronic device. The chip system can be composed of chips or can include chips and other discrete devices.

[0034] Among them, the technical effects brought by the second aspect to the sixth aspect or any possible implementation manner thereof can be referred to the technical effects brought by the first aspect or the related possible implementation manners of the first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is an exemplary schematic diagram of the control system provided by the embodiments of the present application;

[0036] Figure 2 It is an exemplary schematic diagram of the control method provided by the embodiments of the present application;

[0037] Figure 3 It is an exemplary schematic diagram of obtaining multi-modal data provided by the embodiments of the present application;

[0038] Figure 4 It is an exemplary schematic diagram of the training process of the reinforcement learning model provided by the embodiments of the present application;

[0039] Figure 5 It is an exemplary schematic diagram of the control method provided by the embodiments of the present application;

[0040] Figure 6 It is an exemplary schematic diagram of the data processing flow provided by the embodiments of the present application;

[0041] Figure 7 It is a schematic diagram of an embodiment of the control device provided by the embodiments of the present application;

[0042] Figure 8 It is a schematic diagram of the structure of an electronic device provided by the embodiments of the present application. Detailed Embodiments

[0043] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.

[0044] As is known to those of ordinary skill in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0045] In this application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. The terms "first", "second", etc. in the description, claims, and above-mentioned drawings of this application are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that these terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of this application. In addition, the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusion, so that a process, method, system, product, or device including a series of units does not have to be limited to those units, but may include other units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0046] Currently, in many industrial scenarios, it is still a huge challenge to automatically perform high-precision assembly tasks such as shaft-hole assembly through mechanical equipment such as robots.

[0047] Taking the shaft-hole assembly task as an example, in daily life and industrial production and manufacturing, common shaft-hole assembly tasks can include charger plugging and unplugging, key plugging and unplugging, shaft-hole plugging and unplugging, etc. In high-precision industrial-level assembly scenarios, it is still difficult to achieve high-reliability plugging and unplugging in a task-changing environment.

[0048] Specifically, a traditional method for realizing plugging and unplugging is a search-based method. For example, methods based on random search and spiral search. However, the application scenarios of search-based methods have obvious limitations. The success rate depends on the initial error. As the initial error increases, the socket success rate and socket efficiency of this method decrease significantly.

[0049] Another traditional method for realizing plugging and unplugging is a force-feedback-based method. In the force-feedback-based method, a six-axis force sensor or a tactile sensor can be used for perception, and then a socket insertion strategy is executed for closed-loop control. The accuracy of the force-feedback-based method can reach the positioning accuracy of the robotic arm. However, the force-feedback-based method is also not applicable to scenarios with large initial errors and it is difficult to ensure the socket insertion efficiency.

[0050] In addition, another traditional method for realizing plugging and unplugging is the vision feedback-based method. The vision feedback-based method performs perception by utilizing image information. Through vision feedback for servo control, rapid jacking can be achieved even with a large initial error. However, limited by factors such as camera resolution, the existing vision feedback-based methods are difficult to achieve sub-millimeter-level high-precision jacking tasks. In addition, the image information is easily interfered by factors such as illumination and background changes, resulting in jacking failure.

[0051] Based on this, in the embodiments of the present application, a control method is provided, which can improve the accuracy and reliability of mechanical devices such as robots when performing assembly tasks, and can be robust to external interference when the environment changes.

[0052] The control method of the embodiments of the present application will be introduced below.

[0053] The control method of the embodiments of the present application can be applied to a control system. As Figure 1 shown, the control system may include a data acquisition component, a control component, and an execution component.

[0054] In actual application scenarios, the specific spatial positions of different components in the control system can have various situations, which are not limited herein.

[0055] For example, different components in the control system can be different devices respectively, or different components can also be integrated on the same device.

[0056] In some examples, the execution component may include one or more mechanical devices such as a robot, a robotic arm, and an end effector.

[0057] The execution component is used to perform an assembly task. The specific type of the assembly task is not limited herein. For the convenience of description, in the embodiments of the present application, the assembly task is taken as an example of an axis-hole assembly task for introduction.

[0058] The control component in the control system can be an electronic device.

[0059] In some examples, the electronic device can be a component other than the execution component. For example, the electronic device can be a terminal device externally connected to the robot.

[0060] Or, in some other examples, the electronic device can also be integrated with the execution component and / or other components in the control system. For example, the electronic device can be integrated in the robot as a control component in the robot, and the robot can also be integrated with a robotic arm and an end effector, etc.

[0061] The specific type of the electronic device and the included hardware structure are not limited herein.

[0062] Exemplarily, the electronic device may be a terminal device, a single server, or a server cluster, or may also be a virtual machine (VM) or a container, etc.

[0063] When the electronic device is a terminal device, the type of the terminal device is not limited herein.

[0064] Exemplarily, the terminal device may be a mobile phone, a pad, a computer with wireless transceiver function, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a terminal in industrial control, a terminal in self-driving, a terminal in remote medical, a terminal in smart grid, a terminal in transportation safety, a terminal in smart city, a terminal in smart home, a terminal in Internet of things (IoT), a wearable device, a robot, or a combination of one or more of the above.

[0065] In some examples, the software control framework in the control component may include a first control layer and a second control layer. The first control layer and the second control layer may be in different hardware such as different chips, or may be in the same hardware but implemented through different processes. The embodiments of the present application do not limit this.

[0066] The software control framework can collect and fuse multimodal data through two control layers, and can also implement control through the two control layers, which is beneficial to the efficient processing of multimodal data, and can avoid directly controlling mechanical equipment without buffering by one control layer, improving the accuracy and safety of control.

[0067] Optional implementation manners of the software control framework may refer to the relevant examples of collecting multimodal data, fusing multimodal data, and performing assembly tasks in the subsequent embodiments, and will not be elaborated herein.

[0068] The data acquisition component in the control system is used to acquire multimodal data.

[0069] The number and type of components of the data acquisition component may be determined based on the type of multimodal data to be acquired, and the specific structure and location of the data acquisition component in the control system are not limited herein.

[0070] Exemplarily, the data acquisition component may include one or more of a force sensor, an image sensor, a speed sensor, an acceleration sensor, a gyroscope, etc.

[0071] Among them, a force sensor is a device that converts the magnitude of force into relevant electrical signals. For example, a force sensor can detect mechanical quantities such as tension, tensile force, pressure, weight, torque, internal stress, or strain. That is to say, a force sensor can collect mechanical data such as force and torque information.

[0072] In one example, in order to facilitate the acquisition of mechanical data, the force sensor can be arranged on the end effector of the mechanical device.

[0073] An image sensor can be used to collect image modal data such as two-dimensional images and / or depth images. That is to say, the image sensor can include one or more of a two-dimensional image sensor and a depth image sensor.

[0074] In one example, the two-dimensional image sensor can be located at the end of the robotic arm in the mechanical device, while the depth image sensor can be fixed at a position outside components such as the robotic arm, end effector, and jack in the mechanical device, that is, kept fixed in the world coordinate system.

[0075] Motion sensors such as speed sensors, acceleration sensors, gyroscopes, etc. are used to detect motion data.

[0076] Exemplarily, the motion data may include one or more of the speed, acceleration, rotation information, or pose information of movable components such as robotic arms and end effectors in the mechanical device.

[0077] In some examples, the motion sensor can be located on movable components such as robotic arms and end effectors in the mechanical device to detect the motion conditions such as acceleration, speed, and position of the movable components.

[0078] For example, an exemplary control system may include a robotic arm, a speed sensor, a two-dimensional image sensor, a depth image sensor, a wrist force sensor, a universal two-finger gripper, a host computer, a plug shaft, and a jack, etc. Among them, the host computer can be used as the control component in the control system, and the speed sensor, two-dimensional image sensor, depth image sensor, and wrist force sensor can be used as the data acquisition components in the control system. The robotic arm and the universal two-finger gripper can be used as the execution components in the control system, while the plug shaft is the component to be assembled in the assembly task, and the jack is the assembly position in the assembly task.

[0079] In this exemplary control system, the robotic arm is fixed at the origin of the world coordinate system. The speed sensor, wrist joint force sensor, two-finger gripper, and insertion shaft are sequentially fixed to the end flange of the robotic arm. The two-dimensional image sensor is connected to the end flange of the robotic arm through a middleware. The depth image sensor and the jack are fixed in the world coordinate system.

[0080] Based on the above control system, a control method according to an embodiment of the present application can control a mechanical device to accurately perform an assembly task according to multi-modal data through a reinforcement learning model, and can be robust to external interference and have high reliability when the environment changes.

[0081] As Figure 2 shown, this control method involves one or more of the following processing stages:

[0082] Obtaining multi-modal data, fusing multi-modal data, training a reinforcement learning model based on a teaching mode, and performing an assembly task.

[0083] The following will respectively give an exemplary introduction to each processing stage to specifically introduce Figure 2 the relevant steps and data involved in each of the shown processing stages.

[0084] 1. Obtaining multi-modal data.

[0085] In an embodiment of the present application, the multi-modal data includes at least two of image modal data, mechanical data, and motion data.

[0086] Exemplarily, the mechanical data can be a force signal and / or a torque signal detected by a force sensor at the end of the robotic arm; the image modal data can include a depth image and / or a two-dimensional image; the motion data can include one or more of information such as acceleration, speed, or pose at the end of the mechanical device.

[0087] The multi-modal data includes the assembly position indicated by the assembly task and information about the mechanical device to perform the assembly task.

[0088] Among them, the objects involved in different modal data in the multi-modal data can be the same or different.

[0089] Exemplarily, in the multi-modal data, each modal of data includes the assembly position indicated by the assembly task and information about the mechanical device to perform the assembly task.

[0090] Alternatively, in the multi-modal data, different modal of data can involve different objects. For example, if the assembly position remains fixed, the mechanical data can be information about the mechanical device to perform the assembly task without including the information of the assembly position. While in the image modal data, it can include both the information of the assembly position and the information of the mechanical device.

[0091] In some embodiments, obtaining multimodal data includes:

[0092] Obtaining data of different modalities in the multimodal data through multiple processes respectively.

[0093] In the embodiments of the present application, different modalities in the multimodal data can be obtained through different processes; or, any one of the multiple processes can obtain data of one or more modalities, but the data obtained by different processes are different; or, it can also be that a certain process obtains a part of the mechanical data, such as pressure data, etc., while another process obtains another part of the mechanical data (such as torque data, etc.) and motion data. It can be seen that there are various ways to obtain data of different modalities in the multimodal data through multiple processes respectively, which are not limited herein.

[0094] For example, in some examples, image modality data, mechanical data, and motion data can be obtained by different processes respectively.

[0095] Or, in other examples, data of different modalities in the multimodal data are collected through multiple processes respectively, including:

[0096] Obtaining image modality data and motion data through a first process;

[0097] Obtaining mechanical data through a second process.

[0098] Taking the Figure 3 illustrated example as an example for exemplary introduction.

[0099] In Figure 3 the illustrated example, the first process can obtain image modality data and motion data. Among them, the frequency at which the corresponding image sensor sends the image modality data to the first process can be 20 Hz, and the frequency at which the corresponding motion sensor sends the motion data to the first process can be 500 Hz.

[0100] And, Figure 3 in the illustrated example, the second process can obtain mechanical data. Since the force sensor usually collects mechanical data at a relatively high frequency, for example, it can reach 3000 Hz, the force sensor can send the mechanical data to the second process at a relatively high frequency such as 3000 Hz.

[0101] The second process can save the mechanical data within the most recent specified duration through a data queue and periodically send the data in the data queue to the first process.

[0102] The first process can align the mechanical data, image modality data, and motion data in the time dimension according to information such as the timestamps of the obtained mechanical data, image modality data, and motion data, and transfer the aligned mechanical data, image modality data, and motion data to the main process for the main process to perform data fusion, obtain fusion data, and execute subsequent processing based on the fusion data.

[0103] 2. Fuse multi-modal data.

[0104] In the embodiments of the present application, after obtaining multi-modal data, the multi-modal data can be fused to obtain fusion data.

[0105] There can be various ways to fuse the multi-modal data, which are not limited herein.

[0106] Exemplarily, in the time dimension, data of different modalities in the multi-modal data can be aligned, and then, the data of different modalities can be mapped into the same feature space and then fused.

[0107] For example, in one example, the multi-modal data includes image modality data, and also includes mechanical data and / or motion data;

[0108] Fusing multi-modal data to obtain fusion data includes:

[0109] Processing the image modality data through an encoder to obtain a feature representation corresponding to the image modality data;

[0110] Fusing the feature representation with the mechanical data and / or motion data to obtain fusion data.

[0111] In the embodiments of the present application, the image modality data can be feature-extracted through an encoder, and the image modality data can be mapped into a low-dimensional feature representation (for example, it can be a one-dimensional feature representation), so that the low-dimensional feature representation can be concatenated with the one-dimensional mechanical data and / or motion data, thereby realizing effective fusion of multi-modal data and obtaining fusion data.

[0112] The specific type of the encoder is not limited herein.

[0113] Below, taking the encoder as a variational auto encoder (VAE) as an example, the training and encoding methods of the encoder are introduced.

[0114] First, the training process of the variational auto encoder is introduced.

[0115] The variational auto encoder can include an encoder and a decoder p θ (x∣z).

[0116] The training data of the variational autoencoder can be in the image modality.

[0117] During the training process, the training data x can be input into the encoder to encode the image modality data through the encoder q φ (z∣x), where φ are the parameters of the encoder.

[0118] The encoder q φ (z∣x) maps the input data to the latent feature representation z in the latent space. Assume that the prior distribution p(z) of the latent feature representation z is a standard normal distribution, that is:

[0119]

[0120] where I is the identity matrix.

[0121] Then, the decoder p θ (x∣z) maps the latent feature representation z back to the original data space to reconstruct the training data x, where θ are the parameters of the decoder.

[0122] The loss function can include a first loss term and a second loss term. The first loss term is used to indicate the reconstruction error, and the second loss term is used to indicate the KL divergence of the latent feature representation. The variational autoencoder is trained through the loss function to minimize the reconstruction error and the KL divergence of the latent feature representation. Among them, the reconstruction error measures the generation ability of the decoder, and the KL divergence measures the difference between the prior distribution of the latent feature representation and the distribution learned by the encoder.

[0123] Therefore, the loss function can be expressed by the following formula:

[0124]

[0125] where, is used to describe the reconstruction error, and KL(q φ (z∣x)||p(z)) is the KL divergence of the latent feature representation.

[0126] By algorithms such as stochastic gradient descent, minimize the loss function to learn the parameters φ and θ of the encoder and decoder until the corresponding loss value converges or until the number of iterations reaches the number threshold.

[0127] In this way, the variational autoencoder can learn the distribution of the latent feature representation and the low-dimensional feature representation of the training data during the training process.

[0128] In this way, the encoder in the variational autoencoder after training can be used as the encoder in the above embodiments to obtain the feature representation corresponding to the image modality data.

[0129] That is to say, the image modality data in the multimodal data can be input into the encoder in the trained variational autoencoder to obtain the low-dimensional feature representation output by the encoder.

[0130] Then, the feature representation corresponding to the image modality data can be concatenated with the mechanical data and / or motion data to obtain the fused data.

[0131] In some embodiments, different modality data in the multimodal data can be obtained separately through multiple processes, so that parallel acquisition and processing of different modality data can be achieved, and the acquisition and processing efficiency of different modality data can be improved. After different modality data are separately acquired through the multiple processes, the different modality data can be aligned and fused in dimensions such as time and feature space.

[0132] For example, in some embodiments, the image modality data and motion data can be obtained through the first process, and the mechanical data can be obtained through the second process.

[0133] Then, according to the time information, the image modality data and motion data obtained by the first process, and the mechanical data obtained by the second process can be aligned and fused to obtain the fused data.

[0134] Among them, according to time information such as timestamps, the image modality data and motion data obtained by the first process can be aligned and fused.

[0135] For example, Figure 3 In the example shown, the second process can save the mechanical data within the most recent specified duration through a data queue and periodically send the data in the data queue to the first process.

[0136] The first process can align the mechanical data, image modality data, and motion data in the time dimension according to the timestamp and other information of the obtained mechanical data, image modality data, and motion data respectively, and transfer the aligned mechanical data, image modality data, and motion data to the main process, so that the main process can perform data fusion to obtain the fused data and perform subsequent processing according to the fused data.

[0137] Of course, in other examples, the functions of different processes may also vary, and the embodiments of the present application do not limit this.

[0138] In some examples, the collection and subsequent processing of multimodal data can be implemented through at least two control layers in the software control framework.

[0139] The at least two control layers may include a first control layer and a second control layer.

[0140] The second control layer can be regarded as the bottom admittance layer among at least two control layers. The second control layer can include a second process. In this way, as the inner speed control loop in the software control framework, the second control layer can obtain data with a relatively high frequency (for example, a frequency of 3000 Hz) in the multi-modal data, such as obtaining mechanical data. In addition, it can also obtain high-frequency motion data, etc.

[0141] The first control layer can be regarded as the top control layer among at least two control layers. The first control layer can include a first process. In addition, it can also include a main process, etc., to obtain data with a relatively low frequency (for example, a frequency of 20 Hz, etc.) in the multi-modal data, such as obtaining image modal data and relatively low-frequency motion data, etc., and can implement the synchronization and fusion operations of the multi-modal data to obtain fused data for subsequent processing based on the fused data.

[0142] Any of the above embodiments for obtaining multi-modal data and fusing multi-modal data can be used in the subsequent training process of the reinforcement learning model and the process of performing the assembly task. In the subsequent examples, the methods for obtaining multi-modal data and fusing multi-modal data will not be elaborated any further.

[0143] 3. Train the reinforcement learning model based on the teaching mode.

[0144] In some embodiments, the reinforcement learning model can be trained so as to perform the assembly task through the trained reinforcement learning model in the subsequent stage.

[0145] Among them, the structure of the reinforcement learning model can be an existing reinforcement learning model currently or a reinforcement learning model developed in the subsequent stage. The embodiments of the present application do not make any limitations in this regard. Exemplarily, the reinforcement learning model can include one or more layers of multi-layer perceptrons (MLPs), or can also include other structures of deep neutral networks (DNNs).

[0146] The training process of the reinforcement learning model will be introduced exemplarily below.

[0147] Specifically, as Figure 4 shown, in some embodiments, the training process of the reinforcement learning model can include steps 401-403.

[0148] Step 401, in the teaching mode, obtain the target pose of the mechanical device.

[0149] The target pose indicates that the relative position between the mechanical device and the assembly position indicated by the assembly task meets the specified conditions.

[0150] Among them, in the teaching mode, the mechanical equipment is guided to perform specific tasks through manual operations. The operator can control the movement and actions of the mechanical equipment in real time through methods such as buttons, handles, or touch screens to complete demonstrations of complex operation actions.

[0151] In the embodiment of the present application, in the teaching mode, the structures such as the end of the mechanical equipment can be non-contact aligned with the assembly positions (such as jacks, etc.) indicated by the assembly tasks to obtain the target pose of the mechanical equipment.

[0152] The target pose is usually the position where the mechanical equipment can successfully perform the assembly task. For example, for the shaft-hole assembly task, the target pose can indicate that the end of the mechanical equipment is non-contact aligned with the assembly position indicated by the assembly task. At this time, after the mechanical equipment reaches the target pose, it can accurately complete the shaft-hole assembly task through jacking operations.

[0153] Step 402: According to the target pose, obtain the assembly path corresponding to each initial pose of the mechanical equipment, and use the initial multi-modal data corresponding to an initial pose and the corresponding assembly path as a set of training data.

[0154] In the embodiment of the present application, after determining the target pose, in the case of any random initial pose, the mechanical equipment is controlled to complete the assembly task by moving to the target pose.

[0155] Therefore, one or more initial poses can be determined by random means or the like, and the assembly path corresponding to each initial pose of the mechanical equipment can be obtained according to the target pose. Among them, in the assembly path corresponding to any initial pose, the mechanical equipment can be controlled to move to the target pose at this initial pose and complete the assembly task from this target pose (such as completing jacking from this target pose).

[0156] In addition, in the embodiment of the present application, the initial multi-modal data corresponding to the mechanical equipment at any initial pose can also be collected, and the initial multi-modal data is fused to obtain the corresponding initial fusion data at any initial pose.

[0157] In the embodiment of the present application, the initial fusion data corresponding to an initial pose and the corresponding assembly path are used as a set of training data.

[0158] In this way, one or more sets of training data can be obtained according to one or more random initial poses and the target pose.

[0159] Step 403: Train the reinforcement learning model according to one or more sets of training data to obtain the trained reinforcement learning model.

[0160] In the embodiments of the present application, the target pose can be determined through one-time teaching, so that the repetitive positioning of mechanical equipment can be realized through the target pose. It can be seen that the target pose determined in this teaching mode can be used as a virtual expert, and can realize the automatic repetitive positioning and correct assembly of mechanical equipment at any random initial pose, so as to efficiently obtain multiple sets of training data, and efficiently provide training data for the training of the subsequent reinforcement learning model without a large amount of manual operations, greatly improving the training efficiency.

[0161] When training the reinforcement learning model based on this training data, the training process may include at least one iteration process.

[0162] The following gives an exemplary introduction to the i-th iteration process. The i-th iteration process can be any iteration process in the training process. The steps of other iteration processes can be the same as or similar to those of the i-th iteration process, and may include fewer steps than the i-th iteration process, or may include more steps than the i-th iteration process.

[0163] In the i-th iteration process, it may include one or more assembly processes, and each assembly process includes at least one action.

[0164] Among them, there can be various ways to determine the actions of each assembly process.

[0165] For example, in some embodiments, in the j-th assembly process of the i-th iteration process, the method includes:

[0166] Determine at least one action of the j-th assembly process according to the probability of the i-th iteration process and the probability threshold, and according to the initial pose of the i-th iteration process, where j is a positive integer.

[0167] The specific value of this probability threshold is not limited herein. Exemplarily, this probability threshold can be 60%, or 80%, etc.

[0168] In some examples, if the probability of the i-th iteration process is not greater than the probability threshold, then determine the action of the j-th assembly process according to the pose of the mechanical equipment before the action of the j-th assembly process and the target pose.

[0169] In the embodiments of the present application, this probability can reflect the possibility of successfully completing the assembly process.

[0170] It can be seen that if the probability of the i-th iteration process is not greater than the probability threshold, it indicates that the possibility of successfully completing the assembly in the assembly process of the i-th iteration process is relatively low. At this time, it may be in the early stage of the training process.

[0171] To avoid damage to mechanical equipment, the target position, etc. caused by an incorrect assembly process, and to improve training efficiency and accelerate the convergence speed, when the probability in the $i$-th iteration process is not greater than the probability threshold, the action of the $j$-th assembly process can be determined based on the first fusion data and the target pose before the action of the $j$-th assembly process.

[0172] At this time, the action can be determined based on the pose of the mechanical equipment and the target pose before this action.

[0173] Exemplarily, if the initial pose of the $j$-th assembly process is the same as the initial pose corresponding to any set of training data, the path of the mechanical equipment in the $j$-th assembly process can be determined according to the assembly path in this set of training data, that is, each action in the $j$-th assembly process can be determined.

[0174] If the initial pose of the $j$-th assembly process is different from the initial pose corresponding to each set of training data, the pose of the mechanical equipment and the target pose before this action can be compared to determine this action, so that the mechanical equipment approaches or reaches the target pose to correctly execute the assembly task.

[0175] In some other examples, if the probability in the $i$-th iteration process is greater than the probability threshold, the action of the $j$-th assembly process is determined based on the first fusion data before the action of the $j$-th assembly process through the reinforcement learning model of the $i$-th iteration process.

[0176] In this example, if the probability in the $i$-th iteration process is greater than the probability threshold, it indicates that the possibility of successfully completing the assembly is relatively high during the assembly process of the $i$-th iteration. At this time, it may be in the later stage of the training process.

[0177] At this time, the first fusion data before the action of the $j$-th assembly process can be processed through the reinforcement learning model of the $i$-th iteration process to obtain the action of the $j$-th assembly process.

[0178] After determining a certain action in the $j$-th assembly process, a set of state transition data corresponding to this action can be obtained. Each set of state transition data includes the first fusion data before the corresponding action, the parameter information of the corresponding action, the second fusion data after the corresponding action, and the reward information of the corresponding action. The reward information is determined according to the assembly result of the assembly process to which the corresponding action belongs.

[0179] Specifically, the first multi-modal data before the corresponding action can be collected through the data acquisition component in the control system, and the first fusion data can be obtained through data fusion. After executing this action, the second multi-modal data after the corresponding action is collected through the data acquisition component, and the second fusion data is obtained through data fusion.

[0180] The reward information for the corresponding action can be determined according to the reward function.

[0181] Exemplarily, the reward function r can be in the following form:

[0182]

[0183] Wherein, if the assembly process to which the action belongs successfully completes the assembly task (success insertion), the reward information for the actions in the assembly process is all 1. Otherwise, the reward information for the actions in the assembly process is 0.

[0184] Taking the action A as an example, a set of state transition data corresponding to the action A can be (s, a, s′, r). Wherein, s is the first fusion data before the action A, a is the parameter information of the action A, s′ is the second fusion data after the action A, and r is the reward information of the action A.

[0185] In this way, in the i-th iteration process including one or more assembly processes, a set of state transition data corresponding to each action can be obtained.

[0186] After performing one or more assembly processes in the i-th iteration process, the probability of the i-th iteration process can be updated.

[0187] Specifically, in some embodiments, the probability of the i-th iteration process can be updated according to the success rate of one or more assembly processes included in the i-th iteration process.

[0188] Generally speaking, if the success rate of one or more assembly processes included in the i-th iteration process is relatively high, the updated probability of the i-th iteration process is also relatively high; if the success rate of one or more assembly processes included in the i-th iteration process is relatively low, the updated probability of the i-th iteration process is also relatively low.

[0189] For example, the average success rate avg_succ_rate of n assembly processes in the i-th iteration process can be counted, and the probability p of the i-th iteration process can be updated according to the average success rate avg_succ_rate:

[0190]

[0191] In addition, after performing one or more assembly processes in the i-th iteration process, the reinforcement learning model of the i-th iteration process can be updated.

[0192] Specifically, in some embodiments, the reinforcement learning model of the i-th iteration process can be updated according to one or more sets of state transition data of the i-th iteration process.

[0193] For example, if the 100 assembly processes in the i-th iteration process include a total of 600 actions, then 600 sets of state transition data for the i-th iteration process can be obtained.

[0194] In this example, m sets of state transition data can be randomly selected from the 600 sets of state transition data to calculate the gradient of the reinforcement learning model for the i-th iteration process by methods such as gradient descent, so as to update the reinforcement learning model for the i-th iteration process.

[0195] Referring to the above i-th iteration process, the reinforcement learning model is trained through one or more iteration processes until the probability p converges to 1. In addition, the average success rate avg_succ_rate of the assembly process in the iteration process can be made greater than the success rate threshold (e.g., 0.99), or until the number of iterations reaches the specified number, then the training of the reinforcement learning model is ended, and the trained reinforcement learning model is obtained.

[0196] 4. Execute the assembly task.

[0197] After obtaining the trained reinforcement learning model, the assembly task can be executed through the trained reinforcement learning model.

[0198] Specifically, as Figure 5 shown, in some embodiments, a control method includes steps 501-503, and in some embodiments, step 504 may also be included.

[0199] Step 501, obtain multimodal data.

[0200] The multimodal data includes at least two of image modality data, mechanical data, and motion data, and the multimodal data includes the assembly position indicated by the assembly task and information about the mechanical equipment to perform the assembly task.

[0201] In the embodiments of the present application, the manner of obtaining multimodal data can refer to any embodiment of the processing stage of obtaining multimodal data in the above processing stage, which will not be elaborated here.

[0202] Step 502, fuse the multimodal data to obtain fused data.

[0203] In the embodiments of the present application, the manner of fusing multimodal data can refer to any embodiment of the processing stage of obtaining fused multimodal data in the above processing stage, which will not be elaborated here.

[0204] Step 503, according to the fused data, control the mechanical equipment to perform the assembly task through the trained reinforcement learning model.

[0205] In the embodiments of the present application, the fused data can be input into the trained reinforcement learning model, and the next action can be output through the trained reinforcement learning model, so as to control the mechanical equipment to execute the assembly task according to the next action.

[0206] Among them, the execution of the assembly task can include one action or multiple actions, which is not limited herein.

[0207] In one example, when it is necessary to control the mechanical equipment to execute multiple actions to complete the assembly task, and when the assembly task is not completed, after controlling the mechanical equipment to execute one action, return to execute steps 501-502 to obtain new fused data, and through the trained reinforcement learning model, obtain a new next action, and control the mechanical equipment to execute the next action until the assembly task is completed.

[0208] It can be seen that in the embodiments of the present application, multi-modal data such as images, mechanical data, and motion data can be fused to obtain fused data, and multi-dimensional feature information can be extracted from the fused data through the trained reinforcement learning model to determine the control strategy for the mechanical equipment, so that the mechanical equipment can successfully complete the assembly task according to the multi-dimensional feature information in multiple scenarios, improving the success rate and robustness of the mechanical equipment in executing the assembly task.

[0209] In addition, in some embodiments, before executing step 501, step 504 can also be executed.

[0210] Step 504: Control the mechanical equipment to perform rough positioning according to the assembly position.

[0211] Among them, rough positioning means that the distance between the mechanical equipment and the target position is within a specified range, for example, less than a specified distance threshold, so as to facilitate performing a high-precision assembly task with higher precision in the subsequent process.

[0212] In the embodiments of the present application, there can be various ways to instruct the mechanical equipment to perform rough positioning, which is not limited herein.

[0213] Next, taking the control of the mechanical equipment to perform rough positioning according to the depth image information and the assembly position as an example, an exemplary introduction to a rough positioning method will be given.

[0214] Specifically, the rough positioning method can include the following steps:

[0215] Calibrate the depth image sensor to obtain the pose of the depth image sensor in the base coordinate system of the robotic arm of the mechanical equipment

[0216] In the teaching mode, drag the robotic arm so that the end of the robotic arm is non-contact aligned with the assembly position, and record the pose of the end of the robotic arm at this time As the reference pose;

[0217] Through the depth image sensor, depth image acquisition and processing are performed on the reference assembly position to obtain the 3D point cloud data of the reference assembly position. Taking this point cloud as the reference point cloud data and defining the pose of the reference point cloud data to be consistent with the pose of the reference assembly position, the pose of the reference point cloud data in the coordinate system of the depth image sensor can be obtained as

[0218] When an assembly task is to be executed, the pose of the assembly position corresponding to the assembly task is in an unknown state. Through the depth image sensor, depth image acquisition and processing are performed on this assembly position to obtain the target point cloud data of this assembly position and the pose of the target point cloud data of this assembly position in the coordinate system of the depth image sensor

[0219] Based on the iterative closest point (ICP) registration algorithm, the pose of the target point cloud data relative to the reference point cloud data is obtained

[0220]

[0221] Through pose transformation, when rough positioning the assembly position corresponding to the assembly task, the target pose of the end of the robotic arm is obtained

[0222]

[0223] Control the end of the robotic arm to move to the target pose Thus, rough positioning of the assembly position by the mechanical device is realized.

[0224] After rough positioning of the assembly position by the mechanical device is realized, the mechanical device after rough positioning can be used as the mechanical device for the assembly task to be executed, and step 501 and subsequent steps are executed.

[0225] In addition, in some embodiments, in scenarios such as controlling the mechanical device to execute an assembly task and / or controlling the movement of the mechanical device during the process of training a reinforcement learning model, flexible control of the mechanical device can be realized to reduce or even avoid mechanical damage caused by rigid collisions and other situations during the control process of the mechanical device, providing safety guarantee for the control system.

[0226] Below, an exemplary introduction is given by taking the flexible control during the process of executing an assembly task as an example.

[0227] Specifically, in some embodiments, the above step 503 includes:

[0228] Based on the fused data, determine the expected position of the mechanical equipment in the next step through the trained reinforcement learning model;

[0229] Based on the expected position in the next step, and the stiffness characteristics, damping characteristics, and inertia characteristics of the mechanical equipment, determine the expected speed of the mechanical equipment in the next step;

[0230] Control the mechanical equipment to move according to the expected speed in the next step.

[0231] In the embodiment of the present application, by processing the fused data through the trained reinforcement learning model, the expected position of the mechanical equipment in the next step can be output, so that based on the expected position in the next step, and the stiffness characteristics, damping characteristics, and inertia characteristics of the mechanical equipment, the expected speed of the mechanical equipment in the next step can be determined.

[0232] The specific determination method of the expected speed can be obtained based on the dynamic equation.

[0233] Specifically, in some examples, through second-order equation modeling, the following formula is obtained:

[0234]

[0235] where x is the current position of the mechanical equipment, x des is the expected position of the mechanical equipment in the next step, f ext represents the external force received by the mechanical equipment (for example, it can be the external force received by the mechanical equipment due to collision with the assembly position), and K, D, and M respectively represent the stiffness characteristics, damping characteristics, and inertia characteristics of the mechanical equipment.

[0236] Based on the above formula, the expected acceleration can be obtained, and the calculation formula is:

[0237]

[0238] In the embodiment of the present application, by adjusting and testing the three parameters of the stiffness characteristics, damping characteristics, and inertia characteristics of the mechanical equipment, the dynamic attributes of the mechanical equipment can be changed, so that the mechanical equipment can achieve flexible motion characteristics.

[0239] It can be seen that in this example, after determining the current position of the mechanical equipment, the expected position in the next step, and the stiffness characteristics, damping characteristics, and inertia characteristics of the mechanical equipment, based on the calculation formula of the expected acceleration the expected acceleration can be calculated. Then, the expected acceleration can be integrated to obtain the expected speed of the mechanical equipment in the next step The calculation formula of the expected speed is as follows:

[0240]

[0241] In this way, after determining the expected speed for the next step, the flexible motion characteristics of the mechanical equipment can be controlled according to the expected speed for the next step, avoiding collisions with objects such as the assembly position during the execution of the assembly task, resulting in equipment damage, and ensuring the safety of the mechanical equipment during the execution of the assembly task.

[0242] For example, controlling the mechanical equipment to move according to the expected speed for the next step includes:

[0243] Controlling the mechanical equipment to move according to the expected speed for the next step and the force feedback information of the mechanical equipment.

[0244] In the embodiments of the present application, the force feedback information of the mechanical equipment is used to reflect the external force situation currently received by the mechanical equipment, so as to determine whether the mechanical equipment is currently impacted or the like.

[0245] In this way, if the force feedback information indicates that the mechanical equipment is not currently subject to an external force, the expected speed for the next step is used as the execution speed to control the mechanical equipment to move according to the execution speed; if the force feedback information indicates that the mechanical equipment is currently subject to an external force, the smaller value of the expected speed for the next step and the preset speed can be used as the execution speed to control the mechanical equipment to move according to the execution speed, thereby controlling the mechanical equipment to achieve flexible motion characteristics, avoiding collisions with objects such as the assembly position during the execution of the assembly task, resulting in equipment damage, and ensuring the safety of the mechanical equipment during the execution of the assembly task.

[0246] In addition, in some examples, at least two control layers can be combined to achieve flexible control of the mechanical equipment.

[0247] The at least two control layers may include a first control layer and a second control layer.

[0248] As Figure 6 shown in the example, the first control layer, as the top-level control layer, can obtain multi-modal data through the first process and can fuse the multi-modal data to obtain fused data; then, according to the fused data, through the trained reinforcement learning model, determine the expected position of the mechanical equipment in the next step; then, the first control layer determines the expected speed of the mechanical equipment in the next step according to the expected position in the next step, as well as the stiffness characteristics, damping characteristics, and inertia characteristics of the mechanical equipment.

[0249] After the first control layer determines the expected speed for the next step of the mechanical device, the first control layer can transmit the expected speed for the next step to the second control layer. The second control layer can receive mechanical data from a force sensor, and thus obtain force feedback information of the mechanical device based on this mechanical data. In this way, the second control layer can determine the execution speed based on the expected speed for the next step and the force feedback information of the mechanical device, so as to control the mechanical device to move according to the execution speed.

[0250] It can be seen that in this example, the control of the mechanical device can be achieved through two control layers, which is conducive to the efficient processing of multi-modal data, and can avoid directly controlling the mechanical device without buffering according to the expected speed through one control layer, improving the accuracy and safety of control.

[0251] In some embodiments, after controlling the mechanical device to execute the assembly task, the trained reinforcement learning model can also be continuously learned according to the execution situation of the assembly task, so as to further improve the performance of the reinforcement learning model, and thus execute the assembly task more accurately and efficiently.

[0252] Specifically, in some embodiments, one or more actions of the mechanical device during the process of controlling the mechanical device to execute the assembly task can be obtained, and the state transition data corresponding to the one or more actions can be obtained, so as to update the trained reinforcement learning model according to the state transition data corresponding to the one or more actions of the mechanical device during the process of executing the assembly task.

[0253] Among them, the specific content of the state transition data corresponding to one or more actions of the mechanical device during the process of executing the assembly task and the manner of updating the trained reinforcement learning model according to the state transition data can refer to the relevant content in the processing stage of training the reinforcement learning model based on the teaching mode, which will not be elaborated here.

[0254] The above introduces the control method provided by the embodiments of the present application from multiple aspects. Next, in combination with the drawings, the control device provided by the embodiments of the present application will be introduced.

[0255] As Figure 7 shown, the embodiments of the present application provide a control device 70, and the control device 70 includes:

[0256] An acquisition module 701, configured to acquire multi-modal data, where the multi-modal data includes at least two of image modal data, mechanical data, and motion data, and the multi-modal data includes the assembly position indicated by the assembly task and the information of the mechanical device to perform the assembly task;

[0257] A processing module 702, configured to:

[0258] Fuse the multi-modal data to obtain fused data;

[0259] Based on the fusion data, the trained reinforcement learning model is used to control the mechanical equipment to perform the assembly task.

[0260] Optionally, the processing module 702 is further configured to:

[0261] In the teaching mode, obtain the target pose of the mechanical equipment, where the target pose indicates that the relative position between the mechanical equipment and the assembly position indicated by the assembly task meets the specified conditions;

[0262] According to the target pose, obtain the assembly path corresponding to each initial pose of the mechanical equipment, and use the initial fusion data corresponding to an initial pose and the corresponding assembly path as a set of training data;

[0263] Train the reinforcement learning model according to one or more sets of training data to obtain the trained reinforcement learning model.

[0264] Optionally, the training of the reinforcement learning model includes one or more iteration processes. Among them, the i-th iteration process includes one or more assembly processes. Each assembly process includes at least one action. Each action corresponds to a set of state transition data. Each set of state transition data includes the first fusion data before the corresponding action, the parameter information of the corresponding action, the second fusion data after the corresponding action, and the reward information of the corresponding action. The reward information is determined according to the assembly result of the assembly process to which the corresponding action belongs, and i is a positive integer;

[0265] In the i-th iteration process, the processing module 702 is configured to:

[0266] Update the reinforcement learning model of the i-th iteration process according to one or more sets of state transition data of the i-th iteration process.

[0267] Optionally, in the j-th assembly process of the i-th iteration process, the processing module 702 is configured to:

[0268] Determine at least one action of the j-th assembly process according to the probability of the i-th iteration process and the probability threshold, and according to the initial pose of the i-th iteration process, where j is a positive integer.

[0269] Optionally, the processing module 702 is configured to:

[0270] If the probability of the i-th iteration process is greater than the probability threshold, then determine the action of the j-th assembly process through the reinforcement learning model of the i-th iteration process according to the first fusion data before the action of the j-th assembly process;

[0271] If the probability of the i-th iteration process is not greater than the probability threshold, then determine the action of the j-th assembly process according to the pose of the mechanical equipment and the target pose before the action of the j-th assembly process.

[0272] Optionally, the processing module 702 is further configured to:

[0273] Update the probability of the i-th iteration process according to the success rate of one or more assembly processes included in the i-th iteration process.

[0274] Optionally, the processing module 702 is configured to:

[0275] Determine the expected position of the mechanical equipment in the next step according to the fusion data through the trained reinforcement learning model;

[0276] Determine the expected speed of the mechanical equipment in the next step according to the expected position in the next step, and the stiffness characteristics, damping characteristics and inertia characteristics of the mechanical equipment;

[0277] Control the mechanical equipment to move according to the expected speed in the next step.

[0278] Optionally, the processing module 702 is configured to: Control the mechanical equipment to move according to the expected speed in the next step and the force feedback information of the mechanical equipment.

[0279] Optionally, the processing module 702 is configured to:

[0280] Obtain data of different modalities in the multimodal data through multiple processes respectively.

[0281] Optionally, the multimodal data includes image modality data, and also includes mechanical data and / or motion data;

[0282] The processing module 702 is configured to:

[0283] Process the image modality data through an encoder to obtain a feature representation corresponding to the image modality data;

[0284] Fuse the feature representation with the mechanical data and / or motion data to obtain fusion data.

[0285] Optionally, before obtaining the multimodal data, the processing module 702 is further configured to:

[0286] Control the mechanical equipment to perform rough positioning according to the assembly position.

[0287] Figure 8 As shown, it is a possible schematic logical structure diagram of the electronic device 80 provided by the embodiment of the present application. The electronic device 80 is used to implement the functions of the electronic device involved in any of the above embodiments. The electronic device 80 includes: a memory 801, a processor 802, a communication interface 803, and a bus 804. Among them, the memory 801, the processor 802, and the communication interface 803 are communicatively connected to each other through the bus 804.

[0288] The memory 801 may be a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 801 may store a program. When the program stored in the memory 801 is executed by the processor 802, the processor 802 and the communication interface 803 are used to execute one or more steps in the above control method embodiments.

[0289] The processor 802 may be a central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any combination thereof, and is used to execute relevant programs to implement the functions required to be executed by the acquisition module and the processing module in the control device in the above embodiments, or to execute one or more steps in the method embodiments of the present application. The steps of the method disclosed in the embodiments of the present application may be completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read only memory, a programmable read only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 801, and the processor 802 reads the information in the memory 801 and combines its hardware to execute one or more steps in the above control method embodiments.

[0290] The communication interface 803 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the electronic device 80 and other devices or communication networks.

[0291] The bus 804 can implement a path for transmitting information among various components of the electronic device 80 (for example, the memory 801, the processor 802, and the communication interface 803). The bus 804 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 only a thick line is used in Figure 8 to represent it, but it does not mean that there is only one bus or one type of bus.

[0292] In another embodiment of the present application, there is also provided a computer-readable storage medium, in which computer-executable instructions are stored. When the processor of the device executes the computer-executable instructions, the device executes the above-mentioned Figure 8 steps executed by the processor in Figure 8 .

[0293] In another embodiment of the present application, there is also provided a computer program product, which includes computer-executable instructions stored in a computer-readable storage medium; when the processor of the device executes the computer-executable instructions, the device executes the above-mentioned Figure 8 steps executed by the processor in Figure 8 .

[0294] In another embodiment of the present application, there is also provided a chip system, which includes a processor for implementing the steps executed by the above-mentioned Figure 8 processor. In a possible design, the chip system may further include a memory for storing necessary program instructions and data for the device to write data. The chip system can be composed of chips or can include chips and other discrete devices.

[0295] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0296] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0297] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0298] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0299] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs and other various media that can store program codes.

Claims

1. A control method, characterized in that, The method includes: Obtaining multimodal data, where the multimodal data includes at least two of image modality data, mechanical data, and motion data, and the multimodal data includes the assembly position indicated by the assembly task and information on the mechanical equipment to perform the assembly task; Fusing the multimodal data to obtain fused data; Controlling the mechanical equipment to perform the assembly task according to the fused data through a trained reinforcement learning model.

2. The method according to claim 1, wherein The method further includes: In the teaching mode, obtaining the target pose of the mechanical equipment, where the target pose indicates that the relative position between the mechanical equipment and the assembly position indicated by the assembly task meets a specified condition; Obtaining the assembly path corresponding to each initial pose of the mechanical equipment according to the target pose, and taking the initial fused data corresponding to an initial pose and the corresponding assembly path as a set of training data; Training the reinforcement learning model according to one or more sets of the training data to obtain a trained reinforcement learning model.

3. The method according to claim 2, wherein The training of the reinforcement learning model includes one or more iteration processes. Among them, the i-th iteration process includes one or more assembly processes. Each assembly process includes at least one action. Each action corresponds to a set of state transition data. Each set of state transition data includes the first fused data before the corresponding action, the parameter information of the corresponding action, the second fused data after the corresponding action, and the reward information of the corresponding action. The reward information is determined according to the assembly result of the assembly process to which the corresponding action belongs, and i is a positive integer; In the i-th iteration process, the method includes: Updating the reinforcement learning model of the i-th iteration process according to one or more sets of state transition data of the i-th iteration process.

4. The method according to claim 3, wherein In the j-th assembly process of the i-th iteration process, the method includes: Determining at least one action of the j-th assembly process according to the probability of the i-th iteration process and a probability threshold and according to the initial pose of the i-th iteration process, where j is a positive integer.

5. The method according to claim 4, characterized in that, The determining at least one action of the j-th assembly process according to the probability of the i-th iteration process and a probability threshold and according to the initial pose of the i-th iteration process includes: If the probability of the i-th iteration process is greater than the probability threshold, determining the action of the j-th assembly process through the reinforcement learning model of the i-th iteration process according to the first fused data before the action of the j-th assembly process; If the probability of the i-th iteration process is not greater than the probability threshold, determining the action of the j-th assembly process according to the pose of the mechanical equipment before the action of the j-th assembly process and the target pose.

6. The method according to claim 4 or 5, characterized in that, The method further includes: Updating the probability of the i-th iteration process according to the success rate of one or more assembly processes included in the i-th iteration process.

7. The method according to any one of claims 1 to 6, characterized in that, The controlling the mechanical equipment to perform the assembly task according to the fused data through a trained reinforcement learning model includes: Determining the expected position of the mechanical equipment in the next step according to the fused data through a trained reinforcement learning model; Determine the desired speed of the mechanical equipment for the next step according to the desired position of the next step, as well as the stiffness characteristics, damping characteristics, and inertia characteristics of the mechanical equipment; Control the mechanical equipment to move according to the desired speed of the next step.

8. The method according to claim 7, wherein The controlling the mechanical equipment to move according to the desired speed of the next step includes: Controlling the mechanical equipment to move according to the desired speed of the next step and the force feedback information of the mechanical equipment.

9. The method according to any one of claims 1-8, characterized in that The obtaining the multimodal data includes: Obtain data of different modalities in the multimodal data through multiple processes respectively.

10. The method according to any one of claims 1-9, characterized in that, The multimodal data includes image modality data, and also includes the mechanical data and / or the motion data; The fusing the multimodal data to obtain fused data includes: Process the image modality data through an encoder to obtain a feature representation corresponding to the image modality data; Fuse the feature representation with the mechanical data and / or the motion data to obtain the fused data.

11. The method according to any one of claims 1 to 10, characterized in that, Before obtaining the multimodal data, the method further includes: Control the mechanical equipment to perform rough positioning according to the assembly position.

12. A control device, characterized in that, including: An acquisition module, configured to acquire multimodal data, where the multimodal data includes at least two of image modality data, mechanical data, and motion data, and the multimodal data includes the assembly position indicated by the assembly task and information of the mechanical equipment to perform the assembly task; A processing module, configured to: Fuse the multimodal data to obtain fused data; Control the mechanical equipment to perform the assembly task through a trained reinforcement learning model according to the fused data.

13. The apparatus according to claim 12, wherein The processing module is further configured to: In the teaching mode, acquire the target pose of the mechanical equipment, where the target pose indicates that the relative position between the mechanical equipment and the assembly position indicated by the assembly task meets a specified condition; Obtain the assembly path corresponding to each initial pose of the mechanical equipment according to the target pose, and use the initial fused data corresponding to one initial pose and the corresponding assembly path as a set of training data; Train the reinforcement learning model according to one or more sets of the training data to obtain a trained reinforcement learning model.

14. The device according to claim 13, wherein The training of the reinforcement learning model includes one or more iterative processes, where the i-th iterative process includes one or more assembly processes, each assembly process includes at least one action, each action corresponds to a set of state transition data, and each set of state transition data includes the first fused data before the corresponding action, the parameter information of the corresponding action, the second fused data after the corresponding action, and the reward information of the corresponding action, and the reward information is determined according to the assembly result of the assembly process to which the corresponding action belongs, and i is a positive integer; In the i-th iterative process, the processing module is configured to: Update the reinforcement learning model of the i-th iterative process according to one or more sets of state transition data of the i-th iterative process.

15. The apparatus according to claim 14, wherein During the j-th assembly process of the i-th iteration process, the processing module is configured to: Based on the probability of the i-th iteration process and a probability threshold, and based on the initial pose of the i-th iteration process, determine at least one action for the j-th assembly process, where j is a positive integer.

16. The apparatus according to claim 15, wherein: The processing module is configured to: If the probability of the i-th iteration process is greater than the probability threshold, determine the action for the j-th assembly process through the reinforcement learning model of the i-th iteration process based on the first fusion data before the action of the j-th assembly process; If the probability of the i-th iteration process is not greater than the probability threshold, determine the action for the j-th assembly process based on the pose of the mechanical device and the target pose before the action of the j-th assembly process.

17. The apparatus according to any one of claims 12-16, wherein: The processing module is configured to: Based on the fusion data, determine the expected position of the mechanical device in the next step through a trained reinforcement learning model; Based on the expected position in the next step, and the stiffness characteristics, damping characteristics, and inertia characteristics of the mechanical device, determine the expected speed of the mechanical device in the next step; Based on the expected speed in the next step, control the mechanical device to move.

18. The apparatus according to claim 17, wherein: The processing module is configured to: Control the mechanical device to move based on the expected speed in the next step and the force feedback information of the mechanical device.

19. An electronic device, characterized in that, The electronic device includes at least one processor, a memory, and instructions stored on the memory and executable by the at least one processor. The at least one processor executes the instructions to implement the steps of the method according to any one of claims 1-11.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1-11.

Citation Information

Cited By

  • Control method and related apparatus

    WO2025139282A1