Control method and related apparatus

Through the reinforcement learning model of fusion of multimodal data, the reliability and robustness of high-precision assembly tasks are solved, efficient assembly under environmental changes is achieved, and the assembly success rate and safety of mechanical equipment are improved.

WO2025139282A1PCT designated stage expired Publication Date: 2025-07-03HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/126923
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-27
Filing Date
2024-10-24
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high reliability and robustness of shaft hole assembly tasks in high-precision industrial assembly scenarios, especially under environmental changes and external interference, the success rate and efficiency of traditional methods have significantly decreased.

Method used

By fusing images, mechanics and motion data, using reinforcement learning models to control mechanical equipment, combined with teaching mode and multi-layer control layers, the automated repeated positioning and flexible movement of mechanical equipment are realized, and the success rate and robustness of assembly tasks are improved.

Benefits of technology

In various scenarios, the accuracy and reliability of mechanical equipment performing assembly tasks are improved, environmental adaptability is achieved, equipment damage is reduced, and training efficiency and safety is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024126923_03072025_PF_FP_ABST
    Figure CN2024126923_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a control method, a control device (70), an electronic device (80), and a computer-readable storage medium. The control method can fuse multi-modal data, such as images, mechanical data and motion data, to obtain fused data, and uses a trained reinforcement learning model to extract multi-dimensional feature information from the fused data so as to determine a control strategy for a mechanical apparatus, so that the mechanical apparatus can successfully complete an assembly task on the basis of the multi-dimensional feature information in various scenarios, thereby improving the success rate and robustness of the mechanical apparatus executing the assembly task.
Need to check novelty before this filing date? Find Prior Art

Description

A control method and related equipment

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 27, 2023, with Chinese application number 202311832633.4 and invention name “A control method and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of control technology, and in particular to a control method and related equipment. Background Art

[0003] The rapid development of the Internet and informatization has promoted the development of robotics technology. Robots have gradually emerged in all walks of life, helping people complete a variety of complex tasks.

[0004] However, currently, in many industrial scenarios, it is still a huge challenge to automatically perform high-precision assembly tasks such as shaft-hole assembly through mechanical equipment such as robots.

[0005] Taking shaft-and-hole assembly tasks as an example, common tasks in daily life and industrial manufacturing include plugging and unplugging chargers, keys, and shafts. In high-precision industrial assembly scenarios, achieving high-reliability plugging and unplugging in environments with varying tasks remains difficult.

[0006] Summary of the Invention

[0007] The present invention provides a control method to solve the problem that high-precision assembly tasks such as shaft-hole assembly are difficult to complete using mechanical equipment such as robots. The present invention also provides corresponding devices, equipment, computer-readable storage media, and computer program products.

[0008] In a first aspect, the present application provides a control method, which includes: acquiring multimodal data, the multimodal data including at least two of image modal data, mechanical data, and motion data, the multimodal data including information about an assembly position indicated by an assembly task and mechanical equipment to perform the assembly task; fusing the multimodal data to obtain fused data; and controlling the mechanical equipment to perform the assembly task based on the fused data through a trained reinforcement learning model.

[0009] In the first aspect, multimodal data such as images, mechanical data, and motion data can be fused to obtain fused data, and multi-dimensional feature information can be extracted from the fused data through the trained reinforcement learning model to determine the control strategy for the mechanical equipment, so that the mechanical equipment can successfully complete the assembly task based on the multi-dimensional feature information in a variety of scenarios, thereby improving the success rate and robustness of the mechanical equipment in performing assembly tasks.

[0010] In a possible implementation of the first aspect, the method further includes: in a teaching mode, obtaining a target posture of the mechanical equipment, the target posture indicating that the relative position between the mechanical equipment and the assembly position indicated by the assembly task satisfies a specified condition; according to the target posture, obtaining an assembly path corresponding to each initial posture of the mechanical equipment, and taking the initial fusion data corresponding to an initial posture and the corresponding assembly path as a set of training data; training a reinforcement learning model according to one or more sets of training data to obtain a trained reinforcement learning model.

[0011] In this possible implementation, the target pose can be determined through a single teaching demonstration, thereby achieving repeatable positioning of the mechanical equipment. This demonstrates that the target pose determined in this teaching mode can serve as a virtual expert, enabling automated repeatable positioning and correct assembly of mechanical equipment at any random initial pose. This allows for efficient acquisition of multiple sets of training data, effectively providing training data for subsequent reinforcement learning model training without requiring extensive manual effort, significantly improving training efficiency.

[0012] In a possible implementation of the first aspect, the training of the reinforcement learning model includes one or more iterative processes, wherein the i-th iterative process includes one or more assembly processes, each assembly process includes at least one action, each action corresponds to a set of state transition data, each set of state transition data includes first fusion data before the corresponding action, parameter information of the corresponding action, second fusion data after the corresponding action, and reward information of the corresponding action, the reward information is determined based on the assembly result of the assembly process to which the corresponding action belongs, and i is a positive integer; in the i-th iterative process, the method includes: updating the reinforcement learning model of the i-th iterative process based on the one or more sets of state transition data of the i-th iterative process.

[0013] In a possible implementation of the first aspect, in the j-th assembly process of the i-th iterative process, the method includes: determining at least one action of the j-th assembly process according to the probability of the i-th iterative process and the probability threshold and according to the initial posture of the i-th iterative process, where j is a positive integer.

[0014] In this possible implementation, the probability may reflect the likelihood of successful completion of the assembly process. In different probability intervals (eg, when the probability is less than a probability threshold and when the probability is greater than a probability threshold), actions in the assembly process may be determined in different ways.

[0015] In a possible implementation of the first aspect, based on the probability of the i-th iterative process and the probability threshold, at least one action of the j-th assembly process is determined according to the initial posture of the i-th iterative process, including: if the probability of the i-th iterative process is greater than the probability threshold, then based on the first fusion data before the action of the j-th assembly process, the action of the j-th assembly process is determined through the reinforcement learning model of the i-th iterative process; if the probability of the i-th iterative process is not greater than the probability threshold, then based on the posture of the mechanical equipment before the action of the j-th assembly process and the target posture, the action of the j-th assembly process is determined.

[0016] In this possible implementation, if the probability of the i-th iteration is not greater than the probability threshold, it indicates that the probability of successfully completing the assembly during the i-th iteration is low. At this point, the training process may be in the early stages.

[0017] In order to avoid damage to mechanical equipment and target positions caused by incorrect assembly processes, and to improve training efficiency and accelerate convergence, when the probability of the i-th iterative process is not greater than the probability threshold, the action of the j-th assembly process can be determined based on the first fusion data and target posture before the action of the j-th assembly process.

[0018] If the probability of the i-th iteration process is greater than the probability threshold, it indicates that the probability of successfully completing the assembly during the i-th iteration process is high. At this point, the training process may be in the later stage.

[0019] At this time, the first fusion data before the action of the j-th assembly process can be processed by the reinforcement learning model of the i-th iterative process to obtain the action of the j-th assembly process.

[0020] In a possible implementation manner of the first aspect, the method further includes: updating the probability of the i-th iterative process according to the success rate of one or more assembly processes included in the i-th iterative process.

[0021] In a possible implementation of the first aspect, based on the fused data, a trained reinforcement learning model is used to control the mechanical equipment to perform an assembly task, including: based on the fused data, a trained reinforcement learning model is used to determine the expected position of the mechanical equipment in the next step; based on the expected position in the next step, as well as the stiffness characteristics, damping characteristics, and inertia characteristics of the mechanical equipment, the expected speed of the mechanical equipment in the next step is determined; and based on the expected speed in the next step, the mechanical equipment is controlled to move.

[0022] In this possible implementation, after determining the desired speed for the next step, the mechanical equipment can be controlled to achieve flexible motion characteristics based on the desired speed for the next step, thereby avoiding damage to the equipment due to collision with objects such as the assembly position during the execution of the assembly task, thereby ensuring the safety of the mechanical equipment during the execution of the assembly task.

[0023] In a possible implementation of the first aspect, controlling the mechanical device to move according to the expected speed for the next step includes: controlling the mechanical device to move according to the expected speed for the next step and force feedback information of the mechanical device.

[0024] In this possible implementation, the force feedback information of the mechanical device is used to reflect the external force currently applied to the mechanical device, thereby determining whether the mechanical device is currently subjected to an impact, etc. If the force feedback information indicates that the mechanical device is not currently subjected to an external force, the desired speed for the next step is used as the execution speed to control the mechanical device to move according to the execution speed. If the force feedback information indicates that the mechanical device is currently subjected to an external force, the smaller value of the desired speed for the next step or the preset speed can be used as the execution speed to control the mechanical device to move according to the execution speed. This controls the mechanical device to achieve flexible motion characteristics, avoids damage to the device due to collision with objects such as the assembly location during the execution of the assembly task, and ensures the safety of the mechanical device during the execution of the assembly task.

[0025] In a possible implementation of the first aspect, obtaining multimodal data includes: obtaining data of different modalities in the multimodal data respectively through multiple processes.

[0026] In this possible implementation, multiple processes can be used to obtain data of different modes in multimodal data, thereby enabling parallel acquisition and processing of data of different modes, thereby improving the efficiency of data acquisition and processing of data of different modes.

[0027] In a possible implementation of the first aspect, the multimodal data includes image modal data and also includes mechanical data and / or motion data; fusing the multimodal data to obtain fused data includes: processing the image modal data through an encoder to obtain a feature representation corresponding to the image modal data; and fusing the feature representation with the mechanical data and / or motion data to obtain fused data.

[0028] In this possible implementation, the image modal data can be feature extracted through an encoder, and the image modal data can be mapped to a low-dimensional feature representation (for example, a one-dimensional feature representation), so that the low-dimensional feature representation can be spliced ​​with the one-dimensional mechanical data and / or motion data, thereby achieving effective fusion of multimodal data and obtaining fused data.

[0029] In a possible implementation manner of the first aspect, before acquiring the multimodal data, the method further includes: controlling the mechanical equipment to perform coarse positioning according to the assembly position.

[0030] In this possible implementation, coarse positioning refers to ensuring that the distance between the mechanical device and the target location is within a specified range, for example, less than a specified distance threshold, thereby facilitating subsequent high-precision assembly tasks with higher accuracy. For example, the mechanical device can be controlled to perform coarse positioning based on the depth image information and the assembly location.

[0031] A second aspect of the present application provides a control device that has the function of implementing the method of the first aspect or any possible implementation of the first aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-mentioned functions, such as an acquisition module and a processing module.

[0032] The third aspect of the present application provides an electronic device, which includes at least one processor, a memory, and computer-executable instructions stored in the memory and executable by the processor. When the computer-executable instructions are executed by the processor, the processor executes the method as described in the first aspect or any possible implementation of the first aspect.

[0033] The fourth aspect of the present application provides a computer-readable storage medium storing one or more computer-executable instructions. When the computer-executable instructions are executed by a processor, the processor executes the method as described in the first aspect or any possible implementation of the first aspect.

[0034] The fifth aspect of the present application provides a computer program product that stores one or more computer-executable instructions. When the computer-executable instructions are executed by a processor, the processor executes the method as described in the first aspect or any possible implementation of the first aspect.

[0035] A sixth aspect of the present application provides a chip system, which includes a processor for supporting an electronic device in implementing the functions involved in the first aspect or any possible implementation of the first aspect. In one possible design, the chip system may also include a memory for storing program instructions and data necessary for the electronic device. The chip system may be composed of a chip or may include a chip and other discrete devices.

[0036] Among them, the technical effects brought about by the second to sixth aspects or any possible implementation methods thereof can refer to the technical effects brought about by the first aspect or the relevant possible implementation methods of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] FIG1 is an exemplary schematic diagram of a control system provided by an embodiment of the present application;

[0038] FIG2 is an exemplary schematic diagram of a control method provided in an embodiment of the present application;

[0039] FIG3 is an exemplary schematic diagram of obtaining multimodal data provided by an embodiment of the present application;

[0040] FIG4 is an exemplary schematic diagram of the training process of the reinforcement learning model provided in an embodiment of the present application;

[0041] FIG5 is an exemplary schematic diagram of a control method provided in an embodiment of the present application;

[0042] FIG6 is an exemplary schematic diagram of a data processing flow provided in an embodiment of the present application;

[0043] FIG7 is a schematic diagram of an embodiment of a control device provided in an embodiment of the present application;

[0044] FIG8 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0045] The following describes the embodiments of the present application in conjunction with the accompanying drawings. The terms used in the implementation methods of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application.

[0046] Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0047] In this application, "at least one" means one or more, and "more" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable where appropriate. This is merely a way of distinguishing objects with the same properties when describing them in the embodiments of this application. Furthermore, the terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover a non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a list of elements is not necessarily limited to those elements but may include other elements not expressly listed or inherent to such process, method, product, or apparatus.

[0048] Currently, in many industrial scenarios, it remains a huge challenge to automatically perform high-precision assembly tasks such as shaft-hole assembly through mechanical equipment such as robots.

[0049] Taking shaft-and-hole assembly tasks as an example, common tasks in daily life and industrial manufacturing include plugging and unplugging chargers, keys, and shafts. In high-precision industrial assembly scenarios, achieving high-reliability plugging and unplugging in environments with varying tasks remains difficult.

[0050] Specifically, a traditional method for achieving plugging and unplugging is a search-based method, such as random search and spiral search. However, search-based methods have obvious limitations in their application scenarios. Their success rate depends on the initial error. As the initial error increases, the plugging success rate and efficiency of this method decrease significantly.

[0051] Another traditional method for achieving plugging and unplugging is based on force feedback. This method utilizes a six-axis force sensor or tactile sensor for sensing, then executes a plugging strategy for closed-loop control. While the accuracy of this method can rival the positioning accuracy of a robotic arm, it is not suitable for scenarios with large initial errors and struggles to guarantee plugging efficiency.

[0052] Another traditional method for achieving insertion and removal is based on visual feedback. This method utilizes image information for perception. Using visual feedback for servo control enables fast insertion, even with a wide range of initial errors. However, due to limitations such as camera resolution, existing methods based on visual feedback struggle to achieve submillimeter precision insertion. Furthermore, image information is easily affected by factors such as lighting and background variations, leading to insertion failures.

[0053] Based on this, in an embodiment of the present application, a control method is provided that can improve the accuracy and reliability of mechanical equipment such as robots when performing assembly tasks, and can be robust to external interference when the environment changes.

[0054] The control method of the embodiment of the present application is introduced below.

[0055] The control method of the embodiment of the present application can be applied to a control system. As shown in FIG1 , the control system may include a data acquisition component, a control component, and an execution component.

[0056] In actual application scenarios, the specific spatial positions of different components in the control system may be various and are not limited here.

[0057] For example, different components in a control system may be different devices, or different components may be integrated on the same device.

[0058] In some examples, the execution component may include one or more mechanical devices such as a robot, a robotic arm, and an end effector.

[0059] The execution component is used to perform an assembly task. The specific type of the assembly task is not limited here. For ease of description, in the embodiment of the present application, the assembly task is an axis hole assembly task as an example for introduction.

[0060] The control components in the control system may be electronic devices.

[0061] In some examples, the electronic device may be a component other than the execution component. For example, the electronic device may be a terminal device externally connected to the robot.

[0062] Alternatively, in other examples, the electronic device may be integrated with the actuator and / or other components of the control system. For example, the electronic device may be integrated into a robot as a control component of the robot, and the robot may also be integrated with a robotic arm and an end effector.

[0063] The specific type of the electronic device and the included hardware structure are not limited here.

[0064] Exemplarily, the electronic device may be a terminal device, a single server or a server cluster, or a virtual machine (VM) or a container.

[0065] When the electronic device is a terminal device, the type of the terminal device is not limited here.

[0066] Exemplarily, the terminal device may be a mobile phone, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a terminal in industrial control, a terminal in self-driving, a terminal in remote medical, a terminal in a smart grid, a terminal in transportation safety, a terminal in a smart city, a terminal in a smart home, a terminal in the Internet of Things (IoT), a wearable device, a robot, or a combination of one or more thereof.

[0067] In some examples, the software control framework in the control component may include a first control layer and a second control layer. The first control layer and the second control layer may be in different hardware such as different chips, or in the same hardware but implemented through different processes. The embodiments of the present application do not limit this.

[0068] The software control framework can collect and fuse multimodal data through two control layers, and can also realize control through two control layers, which is conducive to the efficient processing of multimodal data and can avoid one control layer directly controlling mechanical equipment without buffering, thereby improving the accuracy and safety of control.

[0069] Optional implementations of the software control framework may refer to the examples of collecting multimodal data, fusing multimodal data, and executing assembly tasks in subsequent embodiments, which will not be described in detail here.

[0070] The data acquisition component in the control system is used to collect multimodal data.

[0071] The number and type of components of the data acquisition component can be determined based on the type of multimodal data to be acquired, and the specific structure and position of the data acquisition component in the control system are not limited here.

[0072] Exemplarily, the data acquisition component may include one or more of a force sensor, an image sensor, a velocity sensor, an acceleration sensor, a gyroscope, and the like.

[0073] A force sensor is a device that converts force measurements into related electrical signals. For example, a force sensor can detect mechanical quantities such as tension, pull, pressure, weight, torque, internal stress, or strain. In other words, a force sensor can collect mechanical data such as force and torque information.

[0074] In one example, in order to facilitate the collection of mechanical data, the force sensor can be set on the end effector of the mechanical equipment.

[0075] The image sensor can be used to collect image modality data such as two-dimensional images and / or depth images. In other words, the image sensor can include one or more of a two-dimensional image sensor and a depth image sensor.

[0076] In one example, the two-dimensional image sensor can be located at the end of a robotic arm in a mechanical device, while the depth image sensor can be fixed at a position outside the robotic arm, end effector, and socket components in the mechanical device, that is, it remains fixed in the world coordinate system.

[0077] Motion sensors such as speed sensors, accelerometers, and gyroscopes are used to detect motion data.

[0078] For example, the motion data may include one or more of speed, acceleration, rotation information, or posture information of a movable component in a mechanical device, such as a robotic arm, an end effector, or the like.

[0079] In some examples, the motion sensor may be located on a movable component of a mechanical device, such as a robotic arm or an end effector, to detect motion conditions such as acceleration, velocity, and position of the movable component.

[0080] For example, an exemplary control system may include a robotic arm, a speed sensor, a two-dimensional image sensor, a depth image sensor, a wrist force sensor, a universal two-finger gripper, a host computer, a plug-in shaft, and a socket. The host computer may serve as the control component of the control system, while the speed sensor, two-dimensional image sensor, depth image sensor, and wrist force sensor may serve as the data acquisition components of the control system. The robotic arm and universal two-finger gripper may serve as the execution components of the control system, while the plug-in shaft is the component to be transferred during the assembly task, and the socket is the assembly location during the assembly task.

[0081] In this exemplary control system, the robotic arm is fixed at the origin of the world coordinate system, the speed sensor, wrist joint force sensor, two-finger gripper, and plug shaft are fixed in sequence on the end flange of the robotic arm, the two-dimensional image sensor is connected to the end flange of the robotic arm through the middleware, and the depth image sensor and the jack are fixed in the world coordinate system.

[0082] Based on the above control system, a control method of an embodiment of the present application can control mechanical equipment to accurately perform assembly tasks based on multimodal data through a reinforcement learning model, and can be robust to external interference when the environment changes, with high reliability.

[0083] As shown in Figure 2, the control method involves one or more of the following processing stages:

[0084] Acquire multimodal data, fuse multimodal data, train reinforcement learning models based on teaching mode, and perform assembly tasks.

[0085] Each processing stage is exemplarily introduced below to specifically introduce the relevant steps and data involved in each processing stage shown in FIG. 2 .

[0086] 1. Acquire multimodal data.

[0087] In an embodiment of the present application, the multimodal data includes at least two of image modality data, mechanical data, and motion data.

[0088] Exemplarily, the mechanical data may be a force signal and / or torque signal detected by a force sensor at the end of the robotic arm; the image modality data may include a depth image and / or a two-dimensional image; the motion data may include one or more information such as acceleration, velocity or posture of the end of the mechanical device.

[0089] The multimodal data includes information about the assembly location indicated by the assembly task and the mechanical equipment to be used to perform the assembly task.

[0090] The objects involved in different modal data in the multimodal data may be the same or different.

[0091] Illustratively, in the multimodal data, the data of each modality includes information about the assembly location indicated by the assembly task and the mechanical equipment to perform the assembly task.

[0092] Alternatively, within this multimodal data, data from different modalities can relate to different objects. For example, if the assembly location remains fixed, the mechanical data can include information about the mechanical equipment to be assembled, without necessarily including information about the assembly location. However, image modality data can include information about both the assembly location and the mechanical equipment.

[0093] In some embodiments, acquiring multimodal data includes:

[0094] Through multiple processes, data of different modes in multimodal data are obtained respectively.

[0095] In the embodiments of the present application, different modes in the multimodal data can be acquired through different processes; alternatively, any one of the multiple processes can acquire data of one or more modes, but the data acquired by different processes may differ; alternatively, one process can acquire a portion of the mechanical data, such as pressure data, while another process can acquire another portion of the mechanical data (such as torque data) and motion data. It can be seen that there are many ways to acquire data of different modes in the multimodal data through multiple processes, which are not limited here.

[0096] For example, in some examples, image modality data, mechanical data, and motion data may be acquired by different processes respectively.

[0097] Alternatively, in other examples, data of different modalities in the multimodal data are collected separately through multiple processes, including:

[0098] Through the first process, image modality data and motion data are acquired;

[0099] Through the second process, mechanical data is obtained.

[0100] An exemplary introduction is given by taking the example shown in FIG3 as an example.

[0101] In the example shown in FIG3 , image modality data and motion data may be acquired by the first process, wherein the frequency at which the corresponding image sensor sends image modality data to the first process may be 20 Hz, and the frequency at which the corresponding motion sensor sends motion data to the first process may be 500 Hz.

[0102] 3, the second process can acquire mechanical data. Since the force sensor typically collects mechanical data at a relatively high frequency, for example, up to 3000 Hz, the force sensor can send the mechanical data to the second process at a relatively high frequency, such as 3000 Hz.

[0103] The second process can save the most recent mechanical data within a specified period of time through a data queue, and periodically send the data in the data queue to the first process.

[0104] The first process can align the mechanical data, image modal data and motion data in the time dimension based on the acquired information such as the timestamps of their respective mechanical data, image modal data and motion data, and pass the aligned mechanical data, image modal data and motion data to the main process, so that the main process performs data fusion, obtains fused data, and performs subsequent processing based on the fused data.

[0105] 2. Fusion of multimodal data.

[0106] In an embodiment of the present application, after acquiring multimodal data, the multimodal data may be fused to obtain fused data.

[0107] There are many ways to fuse the multimodal data, which are not limited here.

[0108] For example, data of different modalities in multimodal data can be aligned in the time dimension, and then the data of different modalities can be mapped to the same feature space and then fused.

[0109] For example, in one example, the multimodal data includes image modality data and also includes mechanical data and / or motion data;

[0110] Fuse multimodal data to obtain fused data, including:

[0111] The image modality data is processed through an encoder to obtain a feature representation corresponding to the image modality data;

[0112] The feature representation is fused with the mechanical data and / or motion data to obtain fused data.

[0113] In an embodiment of the present application, an encoder can be used to extract features from image modal data and map the image modal data to a low-dimensional feature representation (for example, a one-dimensional feature representation), so that the low-dimensional feature representation can be spliced ​​with one-dimensional mechanical data and / or motion data, thereby achieving effective fusion of multimodal data and obtaining fused data.

[0114] The specific type of the encoder is not limited here.

[0115] Next, we take the variational auto encoder (VAE) as an example to introduce the training and encoding methods of the encoder.

[0116] First, we introduce the training process of the variational autoencoder.

[0117] The variational autoencoder can include an encoder and a decoder p θ (x|z).

[0118] The training data of the variational autoencoder can be in the image mode.

[0119] During the training process, the training data x can be input into the encoder to pass through the encoder q φ (z|x) encodes the image modality data, where φ is the parameter of the encoder.

[0120] Encoder q φ (z|x) maps the input data to the potential feature representation z in the latent space. Assume that the prior distribution p(z) of the potential feature representation z is a standard normal distribution, that is:

[0121] where I is the identity matrix.

[0122] Then, the decoder p θ (x|z) maps the latent feature representation z back to the original data space to reconstruct the training data x, where θ is the parameter of the decoder.

[0123] The loss function may include a first loss term and a second loss term, wherein the first loss term is used to indicate the reconstruction error, and the second loss term is used to indicate the KL divergence of the potential feature representation. The variational autoencoder is trained using the loss function to minimize the reconstruction error and the KL divergence of the potential feature representation. The reconstruction error measures the generative ability of the decoder, and the KL divergence measures the difference between the prior distribution of the potential feature representation and the distribution learned by the encoder.

[0124] Therefore, the loss function It can be expressed by the following formula:

[0125] in, Used to describe the reconstruction error, KL(q φ (z|x)||p(z)) is the KL divergence of the latent feature representation.

[0126] The loss function is minimized by algorithms such as stochastic gradient descent, and the parameters φ and θ of the encoder and decoder are learned until the corresponding loss values ​​converge or the number of iterations reaches a threshold.

[0127] In this way, the variational autoencoder can learn the distribution of latent feature representations and the low-dimensional feature representation of the training data during training.

[0128] In this way, the encoder in the variational autoencoder after training can be used as the encoder for obtaining the feature representation corresponding to the image modality data in the above embodiment.

[0129] That is to say, the image modality data in the multimodal data can be input into the encoder in the trained variational autoencoder to obtain a low-dimensional feature representation output by the encoder.

[0130] Then, the feature representation corresponding to the image modality data can be spliced ​​with the mechanical data and / or motion data to obtain fused data.

[0131] In some embodiments, multiple processes can be used to separately acquire data from different modalities within multimodal data, thereby enabling parallel acquisition and processing of data from different modalities, improving the efficiency of data acquisition and processing from different modalities. After acquiring data from different modalities through multiple processes, the data from different modalities can be aligned and fused in dimensions such as time and feature space.

[0132] For example, in some embodiments, image modality data and motion data may be acquired through a first process, and mechanical data may be acquired through a second process.

[0133] Then, the image modality data and motion data acquired in the first process and the mechanical data acquired in the second process can be aligned and fused according to the time information to obtain fused data.

[0134] The image modality data and motion data acquired by the first process may be aligned and fused according to time information such as a timestamp.

[0135] For example, in the example shown in FIG3 , the second process can save the mechanical data within the most recent specified time period through the data queue, and periodically send the data in the data queue to the first process.

[0136] The first process can align the mechanical data, image modal data and motion data in the time dimension based on the acquired information such as the timestamps of their respective mechanical data, image modal data and motion data, and pass the aligned mechanical data, image modal data and motion data to the main process, so that the main process performs data fusion, obtains fused data, and performs subsequent processing based on the fused data.

[0137] Of course, in some other examples, the functions of different processes may also be different, which is not limited in the embodiments of the present application.

[0138] In some examples, the collection and subsequent processing of multimodal data can be achieved through at least two control layers in a software control framework.

[0139] The at least two control layers may include a first control layer and a second control layer.

[0140] The second control layer can be considered the bottom admittance layer of the at least two control layers. The second control layer can include a second process. Thus, the second control layer, as the inner speed control loop in the software control framework, can acquire higher-frequency data (e.g., 3000 Hz) in the multimodal data, such as mechanical data, and can also acquire high-frequency motion data.

[0141] The first control layer can be considered as the top control layer in at least two control layers. The first control layer can include a first process. In addition, it can also include a main process, etc., to obtain data with lower frequencies (for example, a frequency of 20 Hz, etc.) in multimodal data, such as image modality data and lower-frequency motion data, etc., to achieve synchronization and fusion operations of multimodal data, obtain fused data, and perform subsequent processing based on the fused data.

[0142] Any of the above-mentioned embodiments of obtaining multimodal data and fusing multimodal data can be used in the subsequent training process of the reinforcement learning model and the process of performing the assembly task. In the subsequent examples, the method of obtaining multimodal data and fusing multimodal data will not be described in detail.

[0143] 3. Train the reinforcement learning model based on the teaching mode.

[0144] In some embodiments, the reinforcement learning model may be trained so that the assembly task can be subsequently performed using the trained reinforcement learning model.

[0145] The structure of the reinforcement learning model can be a currently existing reinforcement learning model or a subsequently developed reinforcement learning model, and the embodiments of the present application are not limited thereto. For example, the reinforcement learning model can include one or more layers of multilayer perceptrons (MLPs), or can also include other structures of deep neutral networks (DNNs).

[0146] The following is an exemplary introduction to the training process of the reinforcement learning model.

[0147] Specifically, as shown in FIG4 , in some embodiments, the training process of the reinforcement learning model may include steps 401 - 403 .

[0148] Step 401: In the teaching mode, the target posture of the mechanical device is obtained.

[0149] The target pose indicates that the relative position between the mechanical device and the assembly position indicated by the assembly task meets the specified conditions.

[0150] In teaching mode, the machine is guided to perform specific tasks through manual operation. The operator can control the movement and actions of the machine in real time through buttons, handles, or touch screens, completing complex operational demonstrations.

[0151] In the embodiment of the present application, in the teaching mode, the end and other structures of the mechanical equipment can be non-contact aligned with the assembly position (such as a socket, etc.) indicated by the assembly task to obtain the target posture of the mechanical equipment.

[0152] The target pose typically defines the position at which a robot can successfully perform an assembly task. For example, for a shaft-in-hole assembly task, the target pose might indicate that the end effector of the robot should be non-contact aligned with the assembly position specified by the task. In this case, after reaching the target pose, the robot can then accurately complete the shaft-in-hole assembly task by inserting the hole into the shaft.

[0153] In step 402 , an assembly path corresponding to each initial posture of the mechanical device is obtained according to the target posture, and initial multimodal data corresponding to an initial posture and the corresponding assembly path are used as a set of training data.

[0154] In the embodiment of the present application, after the target posture is determined, the mechanical equipment is controlled to complete the assembly task by moving to the target posture in any random initial posture.

[0155] Therefore, one or more initial postures can be determined by random means, and the assembly path corresponding to each initial posture of the mechanical device can be obtained based on the target posture. In the assembly path corresponding to any initial posture, the mechanical device can be controlled to move to the target posture under the initial posture, and the assembly task (for example, completing the jack from the target posture) can be completed from the target posture.

[0156] In addition, in an embodiment of the present application, initial multimodal data corresponding to the mechanical equipment at any initial posture can be collected, and the initial multimodal data can be fused to obtain initial fused data corresponding to any initial posture.

[0157] In an embodiment of the present application, the initial fusion data corresponding to an initial posture and the corresponding assembly path are used as a set of training data.

[0158] In this way, one or more sets of training data can be obtained based on one or more random initial poses and target poses.

[0159] Step 403: Train the reinforcement learning model based on one or more sets of training data to obtain a trained reinforcement learning model.

[0160] In the embodiment of the present application, a target pose can be determined through a single teaching demonstration, thereby achieving repeated positioning of the mechanical equipment through the target pose. It can be seen that the target pose determined in this teaching mode can serve as a virtual expert and can achieve automatic repeated positioning and correct assembly of the mechanical equipment under any random initial pose, thereby efficiently acquiring multiple sets of training data, efficiently providing training data for the subsequent training of the reinforcement learning model without requiring a large amount of manual operation, greatly improving training efficiency.

[0161] When the reinforcement learning model is trained based on the training data, the training process may include at least one iteration process.

[0162] The following is an exemplary introduction to the i-th iteration process. The i-th iteration process can be any iteration process in the training process. The steps of other iteration processes can be the same or similar to the i-th iteration process, and can include fewer steps than the i-th iteration process, or can include more steps than the i-th iteration process.

[0163] During the i-th iteration process, one or more assembly processes may be included, and each assembly process includes at least one action.

[0164] There are many ways to determine the actions of each assembly process.

[0165] For example, in some embodiments, during the j-th assembly process of the i-th iteration, the method includes:

[0166] According to the probability of the i-th iterative process and the probability threshold, and according to the initial posture of the i-th iterative process, at least one action of the j-th assembly process is determined, where j is a positive integer.

[0167] The specific value of the probability threshold is not limited here. For example, the probability threshold may be 60%, or 80%, etc.

[0168] In some examples, if the probability of the i-th iteration process is not greater than the probability threshold, the action of the j-th assembly process is determined based on the posture of the mechanical device before the action of the j-th assembly process and the target posture.

[0169] In the embodiment of the present application, the probability may reflect the possibility of successful completion of the assembly process.

[0170] It can be seen that if the probability of the i-th iteration process is not greater than the probability threshold, it indicates that the probability of successfully completing the assembly during the i-th iteration process is low. At this time, it can be in the early stage of the training process.

[0171] In order to avoid damage to mechanical equipment and target positions caused by incorrect assembly processes, and to improve training efficiency and accelerate convergence, when the probability of the i-th iterative process is not greater than the probability threshold, the action of the j-th assembly process can be determined based on the first fusion data and target posture before the action of the j-th assembly process.

[0172] At this time, the action can be determined based on the posture of the mechanical device before the action and the target posture.

[0173] For example, if the initial posture of the j-th assembly process is the same as the initial posture corresponding to any set of training data, the path of the mechanical equipment in the j-th assembly process can be determined based on the assembly path in the set of training data, that is, each action in the j-th assembly process can be determined.

[0174] If the initial posture of the j-th assembly process is different from the initial posture corresponding to each set of training data, the posture of the mechanical equipment before the action and the target posture can be compared to determine the action so that the mechanical equipment approaches or reaches the target posture to correctly perform the assembly task.

[0175] In other examples, if the probability of the i-th iterative process is greater than the probability threshold, the action of the j-th assembly process is determined based on the first fusion data before the action of the j-th assembly process through the reinforcement learning model of the i-th iterative process.

[0176] In this example, if the probability of the i-th iteration process is greater than the probability threshold, it indicates that the probability of successfully completing the assembly during the i-th iteration process is high. At this point, the training process may be in the later stage.

[0177] At this time, the first fusion data before the action of the j-th assembly process can be processed by the reinforcement learning model of the i-th iterative process to obtain the action of the j-th assembly process.

[0178] After determining an action in the j-th assembly process, a set of state transition data corresponding to the action can be obtained. Each set of state transition data includes the first fusion data before the corresponding action, the parameter information of the corresponding action, the second fusion data after the corresponding action, and the reward information of the corresponding action. The reward information is determined based on the assembly result of the assembly process to which the corresponding action belongs.

[0179] Specifically, the data acquisition component in the control system can collect first multimodal data before the corresponding action, and obtain first fused data through data fusion. After the action is performed, the data acquisition component can collect second multimodal data after the corresponding action, and obtain second fused data through data fusion.

[0180] The reward information for the corresponding action can be determined according to the reward function.

[0181] Exemplarily, the reward function r may be in the following form:

[0182] If the assembly process to which the action belongs successfully completes the assembly task (success insertion), the reward information of the action in the assembly process is 1. Otherwise, the reward information of the action in the assembly process is 0.

[0183] Taking action A as an example, the set of state transition data corresponding to action A can be (s, a, s′, r). Among them, s is the first fused data before action A, a is the parameter information of action A, s′ is the second fused data after action A, and r is the reward information of action A.

[0184] In this way, a set of state transition data corresponding to each action in one or more assembly processes included in the i-th iteration process can be obtained.

[0185] After executing one or more assembly processes of the i-th iteration process, the probability of the i-th iteration process may be updated.

[0186] Specifically, in some embodiments, the probability of the i-th iterative process may be updated according to the success rate of one or more assembly processes included in the i-th iterative process.

[0187] Generally speaking, if the success rate of one or more assembly processes included in the i-th iterative process is higher, the updated probability of the i-th iterative process is also correspondingly higher; if the success rate of one or more assembly processes included in the i-th iterative process is lower, the updated probability of the i-th iterative process is also correspondingly lower.

[0188] For example, the average success rate avg_succ_rate of the n assembly processes in the i-th iteration process can be calculated, and the probability p of the i-th iteration process can be updated according to the average success rate avg_succ_rate:

[0189] In addition, after executing one or more assembly processes of the i-th iteration process, the reinforcement learning model of the i-th iteration process may be updated.

[0190] Specifically, in some embodiments, the reinforcement learning model of the i-th iterative process may be updated according to one or more sets of state transition data of the i-th iterative process.

[0191] For example, the 100 assembly processes of the i-th iteration process include 600 actions in total, and 600 sets of state transition data of the i-th iteration process can be obtained.

[0192] In this example, m groups of state transition data can be randomly selected from the 600 groups of state transition data to calculate the gradient of the reinforcement learning model of the i-th iteration process through methods such as gradient descent, so as to update the reinforcement learning model of the i-th iteration process.

[0193] Referring to the above i-th iterative process, the reinforcement learning model is trained through one or more iterative processes until the probability p converges to 1. In addition, the average success rate avg_succ_rate of the assembly process in the iterative process can be made greater than the success rate threshold (for example, 0.99), or until the number of iterations reaches a specified number, the training of the reinforcement learning model is terminated to obtain the trained reinforcement learning model.

[0194] 4. Perform assembly tasks.

[0195] After obtaining the trained reinforcement learning model, the assembly task can be performed through the trained reinforcement learning model.

[0196] Specifically, as shown in FIG5 , in some embodiments, a control method includes steps 501 - 503 , and, in some embodiments, may further include step 504 .

[0197] Step 501: Acquire multimodal data.

[0198] The multimodal data includes at least two of image modal data, mechanical data, and motion data. The multimodal data includes information about an assembly position indicated by an assembly task and mechanical equipment to be used to perform the assembly task.

[0199] In the embodiment of the present application, the method of obtaining multimodal data can refer to any embodiment of the processing stage of obtaining multimodal data in the above-mentioned processing stage, and will not be repeated here.

[0200] Step 502: Fuse multimodal data to obtain fused data.

[0201] In the embodiment of the present application, the method of fusing multimodal data can refer to any embodiment of the processing stage for obtaining fused multimodal data in the above-mentioned processing stage, and will not be repeated here.

[0202] Step 503: Based on the fused data, the trained reinforcement learning model is used to control the mechanical equipment to perform the assembly task.

[0203] In an embodiment of the present application, the fusion data can be input into a trained reinforcement learning model, and the next action can be output through the trained reinforcement learning model, thereby controlling the mechanical equipment to perform the assembly task according to the next action.

[0204] The execution of the assembly task may include one action or multiple actions, which is not limited here.

[0205] In one example, when it is necessary to control a mechanical device to perform multiple actions to complete the assembly task and the assembly task is not completed, after controlling the mechanical device to perform an action, return to execute steps 501-502 to obtain new fusion data, and obtain a new next action through the trained reinforcement learning model, and control the mechanical device to perform the next action until the assembly task is completed.

[0206] It can be seen that in the embodiments of the present application, multimodal data such as images, mechanical data, motion data, etc. can be fused to obtain fused data, and multi-dimensional feature information can be extracted from the fused data through the trained reinforcement learning model to determine the control strategy for the mechanical equipment, so that the mechanical equipment can successfully complete the assembly task according to the multi-dimensional feature information in a variety of scenarios, thereby improving the success rate and robustness of the mechanical equipment in performing assembly tasks.

[0207] Furthermore, in some embodiments, before executing step 501 , step 504 may also be executed.

[0208] Step 504: Control the mechanical equipment to perform rough positioning according to the assembly position.

[0209] Among them, coarse positioning refers to ensuring that the distance between the mechanical equipment and the target position is within a specified range, for example, less than a specified distance threshold, so as to facilitate the subsequent execution of high-precision assembly tasks with higher accuracy.

[0210] In the embodiment of the present application, there may be multiple ways to instruct the mechanical equipment to perform coarse positioning, which are not limited here.

[0211] Below, a coarse positioning method is exemplarily introduced by taking the example of controlling mechanical equipment to perform coarse positioning based on depth image information and assembly position.

[0212] Specifically, the coarse positioning method may include the following steps:

[0213] Calibrate the depth image sensor to obtain the position and posture of the depth image sensor in the base coordinate system of the mechanical arm of the mechanical equipment

[0214] In the teaching mode, drag the robot arm so that the end of the robot arm is aligned with the assembly position without contact, and record the position of the end of the robot arm at this time As a reference pose;

[0215] The depth image sensor is used to collect and process the depth image of the reference assembly position to obtain the 3D point cloud data of the reference assembly position. The point cloud is used as the reference point cloud data. The posture of the reference point cloud data is defined to be consistent with the posture of the reference assembly position. The position and posture of the reference point cloud data in the coordinate system of the depth image sensor can be obtained as follows:

[0216] When an assembly task is to be performed, the posture of the assembly position corresponding to the assembly task is in an unknown state. The depth image sensor is used to collect and process the depth image of the assembly position to obtain the target point cloud data of the assembly position and the posture of the target point cloud data of the assembly position in the coordinate system of the depth image sensor.

[0217] Based on the iterative closest point (ICP) registration algorithm, the pose of the target point cloud data relative to the reference point cloud data is obtained.

[0218] Through the pose transformation, the target pose of the end of the robot arm is obtained when the assembly position corresponding to the assembly task is roughly positioned.

[0219] Control the end of the robotic arm to move to the target pose This enables the mechanical equipment to roughly position the assembly position.

[0220] After the mechanical device has roughly positioned the assembly position, the roughly positioned mechanical device can be used as the mechanical device to perform the assembly task, and step 501 and subsequent steps can be performed.

[0221] In addition, in some embodiments, flexible control of mechanical equipment can be achieved in scenarios such as controlling mechanical equipment to perform assembly tasks and / or controlling the movement of mechanical equipment during training reinforcement learning models, so as to reduce or even avoid mechanical damage caused by rigid collisions and the like during the control of the mechanical equipment, thereby providing safety protection for the control system.

[0222] The following is an illustrative introduction using the flexible control during the execution of an assembly task as an example.

[0223] Specifically, in some embodiments, the above step 503 includes:

[0224] Based on the fused data, the trained reinforcement learning model is used to determine the desired position of the mechanical equipment in the next step.

[0225] Determine the expected speed of the mechanical device for the next step based on the expected position of the next step and the stiffness, damping and inertia characteristics of the mechanical device;

[0226] Control the mechanical equipment to move according to the desired speed of the next step.

[0227] In an embodiment of the present application, the fusion data is processed by the trained reinforcement learning model to output the expected position of the mechanical device in the next step, thereby determining the expected speed of the mechanical device in the next step based on the expected position in the next step, as well as the stiffness characteristics, damping characteristics and inertia characteristics of the mechanical device.

[0228] The specific method for determining the expected speed can be obtained based on the dynamic equation.

[0229] Specifically, in some examples, the following formula is obtained by modeling through a second-order equation:

[0230] Where x is the current position of the machine, x des is the next desired position of the mechanical equipment, f ext It represents the external force applied to the mechanical device (for example, the external force applied when the mechanical device collides with the assembly position). K, D, and M represent the stiffness, damping, and inertia characteristics of the mechanical device, respectively.

[0231] Based on the above formula, the desired acceleration can be obtained The calculation formula is:

[0232] In the embodiment of the present application, by adjusting and testing the three parameters of the stiffness characteristics, damping characteristics and inertia characteristics of the mechanical device, the dynamic properties of the mechanical device can be changed, so that the mechanical device can achieve flexible motion characteristics.

[0233] It can be seen that in this example, after determining the current position of the mechanical device, the expected position of the next step, and the stiffness, damping and inertia characteristics of the mechanical device, the expected acceleration can be calculated based on the expected acceleration. Calculate the expected acceleration using the calculation formula Then, the desired acceleration can be Integrate to obtain the expected speed of the next step of the mechanical equipment The expected speed The calculation formula is as follows:

[0234] In this way, after determining the expected speed for the next step, the mechanical equipment can be controlled to achieve flexible motion characteristics according to the expected speed for the next step, avoiding damage to the equipment due to collision with objects such as the assembly position during the execution of the assembly task, thereby ensuring the safety of the mechanical equipment during the execution of the assembly task.

[0235] For example, controlling the movement of mechanical equipment based on the desired speed of the next step includes:

[0236] The mechanical device is controlled to move according to the expected speed of the next step and the force feedback information of the mechanical device.

[0237] In the embodiment of the present application, the force feedback information of the mechanical device is used to reflect the external force currently applied to the mechanical device, thereby determining whether the mechanical device is currently subjected to an impact or the like.

[0238] In this way, if the force feedback information indicates that the mechanical device is not currently subjected to external force, the expected speed of the next step will be used as the execution speed to control the mechanical device to move according to the execution speed; if the force feedback information indicates that the mechanical device is currently subjected to external force, the smaller value of the expected speed of the next step and the preset speed can be used as the execution speed to control the mechanical device to move according to the execution speed, thereby controlling the mechanical device to achieve flexible motion characteristics, avoiding collision with objects such as assembly positions during the execution of assembly tasks and causing damage to the equipment, thereby ensuring the safety of the mechanical equipment during the execution of assembly tasks.

[0239] Furthermore, in some examples, at least two control layers may be combined to achieve flexible control of mechanical equipment.

[0240] The at least two control layers may include a first control layer and a second control layer.

[0241] In the example shown in Figure 6, the first control layer, as the top-level control layer, can obtain multimodal data through the first process, and can fuse the multimodal data to obtain fused data; then, based on the fused data, the trained reinforcement learning model is used to determine the expected position of the mechanical equipment in the next step; then, the first control layer determines the expected speed of the mechanical equipment in the next step based on the expected position of the next step, as well as the stiffness characteristics, damping characteristics, and inertia characteristics of the mechanical equipment.

[0242] After the first control layer determines the desired speed for the next step of the mechanical device, it can transmit this desired speed to the second control layer. The second control layer can receive mechanical data from the force sensor and obtain force feedback information about the mechanical device based on this mechanical data. In this way, the second control layer can determine the execution speed based on the desired speed for the next step and the force feedback information from the mechanical device, thereby controlling the mechanical device to move according to the execution speed.

[0243] It can be seen that in this example, the control of the mechanical equipment can be achieved through two control layers, which is conducive to the efficient processing of multimodal data and can avoid directly controlling the mechanical equipment without buffering according to the desired speed through one control layer, thereby improving the accuracy and safety of control.

[0244] In some embodiments, after controlling the mechanical equipment to perform the assembly task, the trained reinforcement learning model can be continuously learned based on the execution status of the assembly task, thereby further improving the performance of the reinforcement learning model and performing the assembly task more accurately and efficiently.

[0245] Specifically, in some embodiments, one or more actions of the mechanical equipment can be obtained in the process of controlling the mechanical equipment to perform an assembly task, and state transition data corresponding to the one or more actions can be obtained, so as to update the trained reinforcement learning model based on the state transition data corresponding to the one or more actions of the mechanical equipment in the process of performing the assembly task.

[0246] Among them, in the process of executing the assembly task, the specific content of the state transition data corresponding to one or more actions of the mechanical equipment and the method of updating the trained reinforcement learning model based on the state transition data can refer to the relevant content in the processing stage of training the reinforcement learning model based on the teaching mode above, and will not be repeated here.

[0247] The control method provided in the embodiments of the present application has been described above from multiple aspects. The control device provided in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0248] As shown in FIG7 , an embodiment of the present application provides a control device 70 , which includes:

[0249] An acquisition module 701 is configured to acquire multimodal data, the multimodal data including at least two of image modality data, mechanical data, and motion data, the multimodal data including information about an assembly location indicated by an assembly task and a mechanical device to be used for the assembly task;

[0250] The processing module 702 is configured to:

[0251] Fuse multimodal data to obtain fused data;

[0252] Based on the fused data, the trained reinforcement learning model is used to control the mechanical equipment to perform assembly tasks.

[0253] Optionally, the processing module 702 is further configured to:

[0254] In the teaching mode, the target pose of the mechanical device is obtained, and the target pose indicates that the relative position between the mechanical device and the assembly position indicated by the assembly task meets the specified conditions;

[0255] According to the target posture, the assembly path corresponding to each initial posture of the mechanical equipment is obtained, and the initial fusion data corresponding to an initial posture and the corresponding assembly path are used as a set of training data;

[0256] The reinforcement learning model is trained according to one or more sets of training data to obtain a trained reinforcement learning model.

[0257] Optionally, the training of the reinforcement learning model includes one or more iterative processes, wherein the i-th iterative process includes one or more assembly processes, each assembly process includes at least one action, each action corresponds to a set of state transition data, each set of state transition data includes first fused data before the corresponding action, parameter information of the corresponding action, second fused data after the corresponding action, and reward information of the corresponding action, the reward information is determined according to the assembly result of the assembly process to which the corresponding action belongs, and i is a positive integer;

[0258] During the i-th iteration, the processing module 702 is used to:

[0259] According to one or more sets of state transition data of the i-th iterative process, the reinforcement learning model of the i-th iterative process is updated.

[0260] Optionally, in the j-th assembly process of the i-th iteration process, the processing module 702 is configured to:

[0261] According to the probability of the i-th iterative process and the probability threshold, and according to the initial posture of the i-th iterative process, at least one action of the j-th assembly process is determined, where j is a positive integer.

[0262] Optionally, the processing module 702 is configured to:

[0263] If the probability of the i-th iteration process is greater than the probability threshold, the action of the j-th assembly process is determined based on the first fusion data before the action of the j-th assembly process through the reinforcement learning model of the i-th iteration process;

[0264] If the probability of the i-th iteration process is not greater than the probability threshold, the action of the j-th assembly process is determined according to the posture of the mechanical equipment before the action of the j-th assembly process and the target posture.

[0265] Optionally, the processing module 702 is further configured to:

[0266] The probability of the i-th iterative process is updated according to the success rate of one or more assembly processes included in the i-th iterative process.

[0267] Optionally, the processing module 702 is configured to:

[0268] Based on the fused data, the trained reinforcement learning model is used to determine the desired position of the mechanical equipment in the next step.

[0269] Determine the expected speed of the mechanical device for the next step based on the expected position of the next step and the stiffness, damping and inertia characteristics of the mechanical device;

[0270] Control the mechanical equipment to move according to the desired speed of the next step.

[0271] Optionally, the processing module 702 is configured to control the mechanical device to move according to the expected speed of the next step and force feedback information of the mechanical device.

[0272] Optionally, the processing module 702 is configured to:

[0273] Through multiple processes, data of different modes in multimodal data are obtained respectively.

[0274] Optionally, the multimodal data includes image modality data and also includes mechanical data and / or motion data;

[0275] The processing module 702 is used to:

[0276] The image modality data is processed through an encoder to obtain a feature representation corresponding to the image modality data;

[0277] The feature representation is fused with the mechanical data and / or motion data to obtain fused data.

[0278] Optionally, before acquiring the multimodal data, the processing module 702 is further configured to:

[0279] According to the assembly position, control the mechanical equipment to perform rough positioning.

[0280] FIG8 is a schematic diagram of a possible logical structure of an electronic device 80 provided in an embodiment of the present application. The electronic device 80 is used to implement the functions of the electronic device involved in any of the above embodiments. The electronic device 80 includes: a memory 801, a processor 802, a communication interface 803, and a bus 804. The memory 801, processor 802, and communication interface 803 are communicatively connected to each other via the bus 804.

[0281] The memory 801 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 801 may store programs. When the program stored in the memory 801 is executed by the processor 802, the processor 802 and the communication interface 803 are used to perform one or more steps in the above-described control method embodiment.

[0282] The processor 802 can be a central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component or any combination thereof, used to execute relevant programs to implement the functions required to be executed by the acquisition module and the processing module in the control device in the above embodiment, or to execute one or more steps in the embodiment of the method of the present application. The steps of the method disclosed in the embodiment of the present application can be executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 801, and the processor 802 reads the information in the memory 801 and executes one or more steps in the above control method embodiment in combination with its hardware.

[0283] The communication interface 803 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the electronic device 80 and other devices or a communication network.

[0284] Bus 804 provides a pathway for transmitting information between the various components of electronic device 80 (e.g., memory 801, processor 802, and communication interface 803). Bus 804 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, control buses, and the like. For ease of illustration, FIG8 shows a single thick line, but this does not imply that there is only one bus or only one type of bus.

[0285] In another embodiment of the present application, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When the processor of the device executes the computer-executable instructions, the device executes the steps executed by the processor in Figure 8 above.

[0286] In another embodiment of the present application, a computer program product is also provided, which includes computer execution instructions stored in a computer-readable storage medium; when the processor of the device executes the computer execution instructions, the device executes the steps performed by the processor in Figure 8 above.

[0287] In another embodiment of the present application, a chip system is provided, comprising a processor configured to implement the steps performed by the processor in FIG. 8 . In one possible design, the chip system may further comprise a memory configured to store program instructions and data necessary for data writing. The chip system may be comprised of a chip alone or may include a chip and other discrete components.

[0288] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0289] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0290] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0291] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0292] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. A control method, characterized in that, The method includes: Obtaining multimodal data, where the multimodal data includes at least two of image modality data, mechanical data, and motion data, and the multimodal data includes the assembly position indicated by the assembly task and information about the mechanical device to perform the assembly task; Fusing the multimodal data to obtain fused data; According to the fused data, controlling the mechanical device to perform the assembly task through a trained reinforcement learning model.

2. The method according to claim 1, characterized in that, The method further includes: In the teaching mode, obtaining the target pose of the mechanical device, where the target pose indicates that the relative position between the mechanical device and the assembly position indicated by the assembly task meets a specified condition; According to the target pose, obtaining the assembly path corresponding to each initial pose of the mechanical device, and taking the initial fused data corresponding to an initial pose and the corresponding assembly path as a set of training data; Training the reinforcement learning model according to one or more sets of the training data to obtain a trained reinforcement learning model.

3. The method according to claim 2, wherein The training of the reinforcement learning model includes one or more iterative processes. Among them, the i-th iterative process includes one or more assembly processes. Each assembly process includes at least one action. Each action corresponds to a set of state transition data. Each set of state transition data includes the first fused data before the corresponding action, the parameter information of the corresponding action, the second fused data after the corresponding action, and the reward information of the corresponding action. The reward information is determined according to the assembly result of the assembly process to which the corresponding action belongs, and i is a positive integer; In the i-th iterative process, the method includes: Updating the reinforcement learning model of the i-th iterative process according to one or more sets of state transition data of the i-th iterative process.

4. The method according to claim 3, wherein In the j-th assembly process of the i-th iterative process, the method includes: Determining at least one action of the j-th assembly process according to the probability of the i-th iterative process and a probability threshold, and according to the initial pose of the i-th iterative process, where j is a positive integer.

5. The method according to claim 4, wherein The determining at least one action of the j-th assembly process according to the probability of the i-th iterative process and a probability threshold, and according to the initial pose of the i-th iterative process includes: If the probability of the i-th iterative process is greater than the probability threshold, then determining the action of the j-th assembly process through the reinforcement learning model of the i-th iterative process according to the first fused data before the action of the j-th assembly process; If the probability of the i-th iterative process is not greater than the probability threshold, then determining the action of the j-th assembly process according to the pose of the mechanical device before the action of the j-th assembly process and the target pose.

6. The method according to claim 4 or 5, characterized in that, The method further includes: Updating the probability of the i-th iterative process according to the success rate of one or more assembly processes included in the i-th iterative process.

7. The method according to any one of claims 1-6, characterized in that The controlling the mechanical device to perform the assembly task according to the fused data through a trained reinforcement learning model includes: Determining the expected position of the mechanical device in the next step according to the fused data through a trained reinforcement learning model; Determine the desired speed of the mechanical equipment for the next step according to the desired position of the next step, as well as the stiffness characteristics, damping characteristics, and inertia characteristics of the mechanical equipment; Control the mechanical equipment to move according to the desired speed of the next step.

8. The method according to claim 7, characterized in that The controlling the mechanical equipment to move according to the desired speed of the next step includes: Controlling the mechanical equipment to move according to the desired speed of the next step and the force feedback information of the mechanical equipment.

9. The method according to any one of claims 1-8, characterized in that, The obtaining the multimodal data includes: Obtaining data of different modalities in the multimodal data through multiple processes respectively.

10. The method according to any one of claims 1-9, characterized in that, The multimodal data includes image modality data, and also includes the mechanical data and / or the motion data; The fusing the multimodal data to obtain fused data includes: Processing the image modality data through an encoder to obtain a feature representation corresponding to the image modality data; Fusing the feature representation with the mechanical data and / or the motion data to obtain the fused data.

11. The method according to any one of claims 1-10, characterized in that, Before obtaining the multimodal data, the method further includes: Controlling the mechanical equipment to perform rough positioning according to the assembly position.

12. A control device, characterized in that, Including: An obtaining module, configured to obtain multimodal data, where the multimodal data includes at least two of image modality data, mechanical data, and motion data, and the multimodal data includes the assembly position indicated by the assembly task and information of the mechanical equipment to perform the assembly task; A processing module, configured to: Fuse the multimodal data to obtain fused data; Control the mechanical equipment to perform the assembly task according to the fused data through a trained reinforcement learning model.

13. The apparatus according to claim 12, wherein The processing module is further configured to: In a teaching mode, obtain a target pose of the mechanical equipment, where the target pose indicates that the relative position between the mechanical equipment and the assembly position indicated by the assembly task meets a specified condition; Obtain an assembly path corresponding to each initial pose of the mechanical equipment according to the target pose, and use the initial fused data corresponding to an initial pose and the corresponding assembly path as a set of training data; Train the reinforcement learning model according to one or more sets of the training data to obtain a trained reinforcement learning model.

14. The device according to claim 13, wherein, The training of the reinforcement learning model includes one or more iterative processes, where the i-th iterative process includes one or more assembly processes, each assembly process includes at least one action, each action corresponds to a set of state transition data, each set of state transition data includes first fused data before the corresponding action, parameter information of the corresponding action, second fused data after the corresponding action, and reward information of the corresponding action, the reward information is determined according to the assembly result of the assembly process to which the corresponding action belongs, and i is a positive integer; In the i-th iterative process, the processing module is configured to: Update the reinforcement learning model of the i-th iterative process according to one or more sets of state transition data of the i-th iterative process.

15. The apparatus according to claim 14, wherein In the j-th assembly process of the i-th iteration process, the processing module is configured to: Determine at least one action of the j-th assembly process according to the probability of the i-th iteration process and a probability threshold, and according to the initial pose of the i-th iteration process, where j is a positive integer.

16. The apparatus according to claim 15, wherein: The processing module is configured to: If the probability of the i-th iteration process is greater than the probability threshold, determine the action of the j-th assembly process through the reinforcement learning model of the i-th iteration process according to the first fusion data before the action of the j-th assembly process; If the probability of the i-th iteration process is not greater than the probability threshold, determine the action of the j-th assembly process according to the pose of the mechanical device and the target pose before the action of the j-th assembly process.

17. The apparatus according to any one of claims 12-16, wherein: The processing module is configured to: Determine the expected position of the mechanical device in the next step according to the fusion data through the trained reinforcement learning model; Determine the expected speed of the mechanical device in the next step according to the expected position in the next step, and the stiffness characteristic, damping characteristic and inertia characteristic of the mechanical device; Control the mechanical device to move according to the expected speed in the next step.

18. The apparatus according to claim 17, wherein: The processing module is configured to: Control the mechanical device to move according to the expected speed in the next step and the force feedback information of the mechanical device.

19. An electronic device, characterized in that, The electronic device includes at least one processor, a memory, and instructions stored on the memory and executable by the at least one processor. The at least one processor executes the instructions to implement the steps of the method according to any one of claims 1-11.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Control method and related equipment

    CN120215318A

  • A flexible assembly system and method based on multi-mode information description

    CN109543823A

  • Reinforcement learning awarding method suitable for movable mechanical arm

    CN111515961A

  • Robot assembly skill learning method and system based on multimode heterogeneous information fusion

    CN112631128A

  • Robot assembly motion planning method based on demonstration trajectory

    CN114800515A