Method and device for training neural network model and method and device for controlling manipulator

By training neural network models in a simulation environment, fusing visual, force and tactile information, the control accuracy problem of the robot in complex insertion tasks is solved, and high-precision robot operation is achieved.

CN120552032APending Publication Date: 2025-08-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410218396.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-27
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The prior art is difficult to control robots with high precision in complex tasks, especially in insertion tasks, and it is impossible to effectively integrate visual, force and tactile information for accurate control.

Method used

By training neural network models in a simulation environment, fuse images, stress feedback signals and tactile feedback signals of the surrounding environment of the robot, use control strategies to predict the network prediction control strategy, and optimize model parameters to achieve precise control of the robot.

Benefits of technology

It realizes high-precision control of the robot in complex insertion tasks, and can complete fine operations such as oblique insertion and rotary insertion. The system is more stable and universal, and there is no need to switch models in stages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120552032A_ABST
    Figure CN120552032A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for training a neural network model, a computer program product and a storage medium. The method comprises the steps that state information of a manipulator in a simulation environment during task execution is obtained, and the state information comprises an image reflecting the surrounding environment of the manipulator, a stress feedback signal of the manipulator and a tactile feedback signal generated when the manipulator makes contact with an object; encoding the image, the stress feedback signal and the tactile feedback signal to obtain an image feature, a force feature and a tactile feature; fusing the image features, the force features and the tactile features to obtain fused features; on the basis of the fusion features, a control strategy for the manipulator is predicted through a control strategy prediction network; calculating a return value corresponding to the control strategy; and predicting parameters of the network based on the return value optimization control strategy. According to the method disclosed by the invention, the trained control strategy prediction network can accurately complete various tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and more specifically, to methods, devices, computer program products, and storage media for training neural network models, as well as methods, devices, computer program products, and storage media for controlling a robotic arm. Background Art

[0002] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0003] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0004] The application of AI in robot control primarily encompasses three areas: perception, planning and decision-making, and control. In terms of perception, AI primarily powers the robot's perception system, including speech perception and machine vision. For example, machine vision enables image-based object detection. In terms of planning and decision-making, AI determines the robot's action plan. For example, neural networks within a reinforcement learning framework are responsible for decision-making, determining the next action based on current environmental input and action results. This approach is more similar to human thinking. In terms of control, currently, end-to-end control based on reinforcement learning or AI-assisted stochastic parameter optimization of traditional controllers are primarily used to improve robot kinematic balance. For example, AI methods are being introduced to map real-world object motion data onto the robot, enabling imitation learning to enable the robot's kinematic capabilities to mimic the performance of the object.

[0005] As the tasks performed by robots become more and more complex, how to use artificial intelligence technology to control robots more accurately is currently one of the key research directions in this field. Summary of the Invention

[0006] In order to use artificial intelligence technology to control a robot to complete complex tasks more accurately, the present disclosure provides a method for training a neural network model, including: for a first task, obtaining first state information of a manipulator in a simulation environment when performing the first task, wherein the first state information of the manipulator includes: an image reflecting the environment surrounding the manipulator, a force feedback signal of the manipulator, and a tactile feedback signal when the manipulator contacts an object; encoding the image, the force feedback signal, and the tactile feedback signal respectively to obtain image features, force features, and tactile features; fusing the image features, the force features, and the tactile features to obtain fused features; based on the fused features, using a control strategy prediction network to predict a first control strategy for the manipulator; calculating a first reward value corresponding to the first control strategy; and optimizing parameters of the control strategy prediction network based on the first reward value.

[0007] An embodiment of the present disclosure also provides a method for controlling a manipulator, comprising: obtaining state information of the manipulator when performing a task, wherein the state information of the manipulator includes: an image reflecting the environment surrounding the manipulator, a force feedback signal of the manipulator, and a tactile feedback signal when the manipulator contacts an object; encoding the image, the force feedback signal, and the tactile feedback signal respectively to obtain image features, force features, and tactile features; fusing the image features, the force features, and the tactile features to obtain fused features; based on the fused features, using a control strategy prediction network to predict a control strategy for the manipulator; and controlling the movement of the manipulator based on the control strategy.

[0008] An embodiment of the present disclosure also provides a device for training a neural network model, comprising: an information acquisition module, configured to: for a first task, obtain first state information of a manipulator in a simulation environment when performing the first task, wherein the first state information of the manipulator includes: an image reflecting the environment surrounding the manipulator, a force feedback signal of the manipulator, and a tactile feedback signal when the manipulator contacts an object; a feature encoding module, configured to: encode the image, the force feedback signal and the tactile feedback signal respectively to obtain image features, force features and tactile features; a feature fusion module, configured to: fuse the image features, the force features and the tactile features to obtain fusion features; a strategy prediction module, configured to: based on the fusion features, use a control strategy prediction network to predict a first control strategy for the manipulator; a calculation module, configured to: calculate a first reward value corresponding to the first control strategy; and a model training module, configured to: optimize the parameters of the control strategy prediction network based on the first reward value.

[0009] An embodiment of the present disclosure also provides a device for controlling a manipulator, comprising: an information acquisition module, configured to acquire state information of the manipulator when performing a task, wherein the state information of the manipulator includes: an image reflecting the environment surrounding the manipulator, a force feedback signal of the manipulator, and a tactile feedback signal when the manipulator contacts an object; a feature encoding module, configured to encode the image, the force feedback signal, and the tactile feedback signal respectively to obtain image features, force features, and tactile features; a feature fusion module, configured to fuse the image features, the force features, and the tactile features to obtain fused features; a strategy prediction module, configured to predict a control strategy for the manipulator based on the fused features using a control strategy prediction network; and a control module, configured to control the movement of the manipulator based on the control strategy.

[0010] An embodiment of the present disclosure further provides a computer program product, which includes computer software code. When the computer software code is executed by a processor, it provides the above method.

[0011] An embodiment of the present disclosure further provides a computer-readable storage medium having computer-executable instructions stored thereon, which provide the above method when executed by a processor.

[0012] The present disclosure fuses the features of vision, force, and touch, and uses a control strategy prediction network to predict the control strategy based on the fused features, which is more accurate and can be applied to oblique insertion tasks, insertion tasks that require rotating the object at a certain angle, insertion tasks for small objects, and other refined insertion tasks. The control strategy prediction network based on the present disclosure can realize a series of control operations from the approach of the object to be inserted and the inserted object to the completion of the insertion, without switching the model at different stages (for example, using a vision-based model before the object to be inserted and the inserted object come into contact, and using a force-based model after the object to be inserted and the inserted object come into contact), and has better universality and system stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some exemplary embodiments of the present disclosure. For those skilled in the art, other drawings can be derived from these drawings without inventive effort.

[0014] Herein, in the accompanying drawings:

[0015] Figure 1A schematic diagram of an information processing process according to an embodiment of the present disclosure is shown.

[0016] Figure 2A-2D A schematic diagram showing an application scenario according to an embodiment of the present disclosure is shown;

[0017] Figure 3A-3C is a schematic diagram illustrating a process of processing a tactile feedback signal according to an embodiment of the present disclosure;

[0018] Figures 4A-4C is a schematic diagram illustrating a process of calibrating a tactile sensor according to an embodiment of the present disclosure;

[0019] Figure 5A is a schematic flow chart illustrating a method for training a neural network model according to an embodiment of the present disclosure;

[0020] Figure 5B is a schematic flow chart illustrating a process of training a neural network model according to an embodiment of the present disclosure;

[0021] Figure 6 is a schematic flow chart illustrating a method for controlling a manipulator according to an embodiment of the present disclosure;

[0022] Figure 7 1 is a schematic diagram showing the composition of an apparatus for training a neural network model according to an embodiment of the present disclosure;

[0023] Figure 8 is a schematic diagram showing the composition of an apparatus for controlling a manipulator according to an embodiment of the present disclosure; and

[0024] Figure 9 2 is a diagram illustrating an architecture of a computing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solutions and advantages of the present disclosure more apparent, the exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.

[0026] Furthermore, in this specification and the drawings, steps and elements having substantially the same or similar features are denoted by the same or similar reference numerals, and repeated descriptions of these steps and elements will be omitted.

[0027] In addition, in this specification and the drawings, elements are described in singular or plural form, depending on the embodiment. However, the singular and plural forms are appropriately selected for the situations presented merely for convenience of explanation and are not intended to limit the present disclosure thereto. Therefore, the singular form may include the plural form, and the plural form may also include the singular form, unless the context clearly indicates otherwise.

[0028] In this specification and the accompanying drawings, substantially the same or similar steps or elements are denoted by the same or similar reference numerals, and repeated descriptions of these steps and elements will be omitted. At the same time, in the description of the present disclosure, the terms "first", "second", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance or ranking.

[0029] To facilitate description of the present disclosure, concepts related to the present disclosure are introduced below.

[0030] Industrial control primarily refers to the use of computer technology, microelectronics, and electrical methods to make factory production and manufacturing processes more automated, efficient, precise, controllable, and visible. Industrial control involves industrial system circuit control (for example, automated equipment control) and industrial resource scheduling (for example, logistics resource scheduling and power grid scheduling). Currently, artificial intelligence-based methods are being used in a growing number of industrial control scenarios. For example, neural networks can be used to analyze the current state of industrial control systems to quickly develop appropriate control strategies for reference, effectively reducing the workload of manual analysis.

[0031] The process of analyzing the current state of an industrial control system based on a neural network to formulate a control strategy can be implemented using a Markov decision process (MDP). A Markov decision process is a mathematical model for sequential decision making, used to simulate the stochastic strategies and rewards achievable by intelligent agents in environments where the system state exhibits Markov properties. The MDP primarily involves the following steps: Defining the state space and action space: Determining the set of states and actions for the decision problem forms the basis for establishing a Markov decision process; Defining the transition probability and reward function: Based on the characteristics of the problem, defining the probability of state transitions and the corresponding reward for each action; Implementing the decision: Based on the current state and the specified strategy, selecting the optimal action, updating the state based on the transition probabilities, and obtaining a reward; Iterative optimization: By continuously iterating the decision process, the strategy is updated based on the reward and state information after each decision to maximize the long-term reward.

[0032] The various neural networks (or neural network models) that can be used in the embodiments of the present disclosure below can all be artificial intelligence models, especially artificial intelligence-based neural network models. Typically, artificial intelligence-based neural network models are implemented as acyclic graphs in which neurons are arranged in different layers. Typically, a neural network model includes an input layer and an output layer, which are separated by at least one hidden layer. The hidden layer transforms the input received by the input layer into a representation that is useful for generating output in the output layer. The network nodes (i.e., neurons) are fully connected to the nodes in the adjacent layers via edges, and there are no edges between the nodes in each layer. The data received at the nodes of the input layer of the neural network are propagated to the nodes of the output layer via any one of the hidden layer, activation layer, pooling layer, convolutional layer, etc. The input and output of the neural network model can take various forms, and the present disclosure is not limited thereto.

[0033] In summary, the present disclosure relates to technologies such as artificial intelligence and industrial control. The following further describes embodiments of the present disclosure in conjunction with the accompanying drawings.

[0034] First refer to Figure 1 The information processing process according to the embodiment of the present disclosure is described.

[0035] In order to avoid damaging the manipulator in the real environment by energizing and simulating the various states of the manipulator, the present disclosure proposes to implement training of a neural network model for controlling the manipulator in a simulation environment. The trained neural network model for controlling the manipulator can be used to control the manipulator in a real environment. The simulation environment of the present disclosure can be built based on simulation software. By modeling and configuring the parameters of the manipulator in the simulation software, the manipulator in the simulation environment of the present disclosure has the same structure, specifications and performance parameters as the manipulator in the real environment, and both are used to complete the same tasks. According to an embodiment of the present disclosure, the manipulator in the simulation environment and the manipulator in the real environment can both have the following characteristics: Figure 1 The structure shown.

[0036] like Figure 1As shown, the manipulator may be equipped with a camera 110 (cameras may be arranged in both the front and rear directions, the number of cameras may not be 1, and only one camera 110 is shown here as an example), a force sensor 120, and tactile sensors 130-1 and 130-2. The force sensor 120 may be installed at the wrist of the manipulator, the elbow of the manipulator, and other joints (optionally, for object grasping tasks, the force sensor 120 may be installed at the wrist closer to the gripper), and the tactile sensors 130-1 and 130-2 are respectively installed on the sides of the gripper for contact with the object. The camera is used to capture an image P1 of the environment surrounding the manipulator, the force sensor is used to capture a force feedback signal F1 when the manipulator grasps an object 140, and the tactile sensor may capture a tactile feedback signal C1 generated when the manipulator contacts the object when grasping the object 140.

[0037] By encoding the collected image P1, force feedback signal F1 and tactile feedback signal C1 respectively, the image feature v can be obtained. vision , force characteristic v force and tactile features v tactile By analyzing the image feature v vision , force characteristic v force and tactile features v tactile By fusing the three, we can get the fusion feature v muti (For example, ). Based on the fusion feature v muti , the control strategy prediction network π can be used to predict the control strategy for the manipulator, and the action of the manipulator can be controlled based on the control strategy to complete tasks such as inserting an object.

[0038] Since the method for controlling a manipulator disclosed in the present invention takes into account both the visual information of the manipulator's surroundings and the force information (including both force feedback signals and tactile feedback signals) when the manipulator inserts an object, the manipulator can be accurately controlled to complete various tasks. For example, the method for controlling a manipulator disclosed in the present invention can be used to achieve Figure 2A The pin-shaped component straight insertion task shown (i.e., the pin-shaped component 20a is grasped by the clamping jaws 10 to be vertically inserted into the hole of the object 30a), Figure 2B The annular member straight insertion task shown (i.e., the annular member 20b is grasped by the clamping jaws 10 for vertical insertion into the column of the object 30b), Figure 2C The illustrated oblique insertion task of the plug-shaped component (i.e., grasping the plug-shaped component 20c with the clamping jaws 10 for oblique insertion into the hole of the object 30c), and Figure 2D The annular component oblique insertion task shown (i.e., the annular component 20d is grasped by the clamping jaws 10 to be obliquely inserted into the column of the object 30d), wherein Figure 2A-2DThe above-mentioned tactile sensor 130 - 1 or 130 - 2 may be installed on the surface of the clamping jaw 10 for contacting the object.

[0039] It should be understood that an image feature encoding network can be used to encode an image of the environment surrounding the manipulator. The image feature encoding network can be a pre-trained image feature extraction network (for example, a convolutional neural network (CNN), a recurrent neural network (RNN), a feature pyramid network, etc.). In addition, a force feature encoding network can be used to perform dimensionality transformation and regularization processing on the force feedback signal, and a tactile feature encoding network can be used to perform dimensionality transformation and regularization processing on the tactile feedback signal. The force feature encoding network and the tactile feature encoding network here include operation matrices for dimensionality transformation and regularization processing. Since the force feature encoding network and the tactile feature encoding network have simple structures and are only used to perform simple calculations, it is not necessary to perform parameter optimization on the force feature encoding network and the tactile feature encoding network during the training process. The neural network model for controlling the manipulator provided in the embodiment of the present disclosure may include: the image feature encoding network, the force feature encoding network, the tactile feature encoding network, and the control strategy prediction network.

[0040] The model training process disclosed in the present invention can be divided into two stages. In the first training stage, the image feature encoding network and the control strategy prediction network are trained using only the images captured by the camera to obtain an initial model (the initial model is used to control the manipulator and includes the image feature encoding network and the control strategy prediction network). The initial model can be used to process simple straight insertion tasks. In the second training stage, based on the initial model, the image feature encoding network and the control strategy prediction network can be trained using the captured images, force feedback signals and tactile feedback signals, so that the control strategy prediction network can handle more complex oblique insertion tasks.

[0041] During the two-stage training process, the parameters of both the image feature encoding network and the control strategy prediction network can be optimized based on the output of the control strategy prediction network. For example, the control strategy prediction network can be used to predict the control strategy for the manipulator; the reward value corresponding to the control strategy can be calculated; and the parameters of the image feature encoding network and the control strategy prediction network can be optimized based on the reward value.

[0042] According to the embodiments of the present disclosure, other neural network models (such as a position relationship prediction network) can also be used to assist in the training of the image feature encoding network. In this case, the image feature encoding network and the position relationship prediction network can together constitute a self-supervised neural network. For example, the image feature encoding network can be used to extract features and stitch features of the two images captured by the camera to obtain the image feature v vision . Then, the image feature v vision The position relationship prediction network is input (the position relationship prediction network may include a multi-layer perceptron (MLP)) to predict the predicted position relationship between the object to be inserted and the object to be inserted (for example, in the x and y directions, if the distance between the object to be inserted and the object to be inserted is positive, the predicted position relationship parameter between the object to be inserted and the object to be inserted is 1, and if the distance between the object to be inserted and the object to be inserted is negative, the predicted position relationship parameter between the object to be inserted and the object to be inserted is 0), the actual position relationship between the object to be inserted and the object to be inserted is read from the simulation environment, and a loss function is constructed based on the difference between the predicted position relationship and the actual position relationship (for example, the loss function F l =|predicted relationship parameter in the x direction−true relationship parameter in the x direction|+|predicted relationship parameter in the y direction−true relationship parameter in the y direction|) to train the self-supervised neural network.

[0043] For the embodiment of using the position relationship prediction network to assist the training of the image feature encoding network, the model training process can also be divided into two stages. Among them, in the first training stage, the self-supervised neural network (including the image feature encoding network and the position relationship prediction network) and the control strategy prediction network are trained using only the images collected by the camera to obtain an initial model (the initial model is used to control the manipulator and includes the image feature encoding network and the control strategy prediction network). The initial model can be used to process simple straight insertion tasks. In the second training stage, the self-supervised neural network and the control strategy prediction network can be trained together based on the initial model using the collected images, force feedback signals and tactile feedback signals, so that the control strategy prediction network can handle more complex oblique insertion tasks. In each of the above two stages, the parameters of the self-supervised neural network and the control strategy prediction network can be optimized separately. For example, the parameters of the control strategy prediction network can be fixed first, and the parameters of the self-supervised neural network can be optimized. Then, the parameters of the self-supervised neural network can be fixed again, and the parameters of the control strategy prediction network can be optimized. The above process is repeated many times until the processing results of the two models meet expectations.

[0044] In the first training phase, each state S is determined based on the currently captured image, and the action A reflects the movement change of the manipulator (e.g., translation change). The process of the manipulator performing the insertion task includes multiple states. In each state S, the image feature v vision Input the control strategy prediction network π to predict the control action A to be executed, and use the simulation software to obtain the next state S' based on the current state S and the predicted control action A, and calculate the reward function value r for executing the action A. Then, for the next state S', the control action A' to be executed is predicted, and the next state S' is determined based on the next state S' and the control action A' to be executed by the simulation software until the direct insertion task is completed. From the initial state of the direct insertion task, the control action to be executed at each step is predicted, and its next state is determined, and then the next state of the next state is determined in sequence for the determined next state until the direct insertion task is completed. In this process, the initial state and each determined next state constitute the multiple states. After the direct insertion task is completed (that is, the object to be inserted is successfully inserted into the object to be inserted, completing the direct insertion task; or the insertion fails because the object to be inserted slips away, or because the force feedback signal and / or the tactile feedback signal is greater than a predetermined threshold, and the direct insertion task is ended), the value of the reward function J(π) can be calculated, and the parameters of the control strategy prediction network π can be optimized based on the calculated value of J(π).

[0045] In the second training phase, each state S is determined based on the currently acquired image, force feedback signal, and tactile feedback signal. The action A reflects the movement change of the manipulator (for example, at least one of the translation change and the torsion change). The process of the manipulator performing the oblique insertion task includes multiple states. In each state S, the fused feature v mutiInput the control strategy prediction network π to predict the control action A to be executed, and use the simulation software to obtain the next state S' based on the current state S and the predicted control action A, and calculate the reward function value r for executing the action A. Then, for the next state S', predict the control action A' to be executed, and determine the next state S' based on the next state S' and the control action A' to be executed through simulation software until the oblique insertion task is completed. Similarly, from the initial state of the start of the oblique insertion task, predict the control action to be executed in each step, and determine its next state, and then continue to determine the next state of the next state for the determined next state in sequence until the oblique insertion task is completed. In this process, the initial state and each determined next state constitute the multiple states. After the oblique insertion task is completed (that is, the object to be inserted is successfully inserted into the object to be inserted, completing the oblique insertion task; or because the object to be inserted slides away, or the force feedback signal and / or the tactile feedback signal is greater than a predetermined threshold, resulting in a failure of insertion, and the oblique insertion task is ended), the value of the reward function J(π) can be calculated, and the parameters of the control strategy prediction network π can be optimized based on the calculated value of J(π).

[0046] For the above training process, the value r of the reward function can be calculated by the following formula:

[0047]

[0048] Where d is the current insertion depth, and D is the expected total insertion depth for the insertion task. If the task ends due to insertion failure, a penalty parameter of -0.2 can be added to r. For example, if the object slips away at the beginning and is not inserted at all, r = -0.2. If the object is inserted at 50% of the total insertion depth D, but the detected force is greater than a preset force threshold (e.g., 10N), resulting in insertion failure, r = 0.5 - 0.2 = 0.3.

[0049] The reward function J(π) can be calculated by the following formula:

[0050]

[0051] Where t is the identification number of each state (or corresponding action), T is the total number of states (or corresponding actions), and the predetermined parameter γ t ∈(0,1],E π It should be understood that since the action executed later has a greater impact on the execution result of the entire insertion task, the γ corresponding to the action executed later can be made t γ corresponding to the action executed first t Bigger.

[0052] The movement of an object is usually reflected by translation (including movement parameters Ex, Ey, Ez in the x, y, and z directions), roll (i.e., the angle Eφ of rotation perpendicular to the x-axis), pitch (i.e., the angle Eρ of rotation perpendicular to the y-axis), and yaw (i.e., the angle Eθ of rotation perpendicular to the z-axis).

[0053] For ease of explanation, this article considers translation and roll, a total of four degrees of freedom, for illustration purposes only. This is not intended to be limiting; the disclosed method can also be applied to scenarios where more degrees of freedom are considered. In the case of four degrees of freedom, in the first training phase, action A is represented by a three-dimensional vector [Δx, Δy, Δz]. In the second training phase, action A is represented by a four-dimensional vector [Δx, Δy, Δz, Δq].

[0054] In the second training phase, the force feature v can be obtained through the following processing force and tactile features v tactile .

[0055] The force sensor installed on the wrist can collect force feedback of multiple degrees of freedom, such as the force feedback signal f of 6 degrees of freedom. raw =[f x, f y ,f z ,τ x ,τ y ,τ z ]. In order to more comprehensively consider the information of the force feedback signal, the most recent readings of the force sensor can be considered to obtain the force characteristics. In order to simplify the processing, the force characteristics of some degrees of freedom can be selected. For example, when considering the forces of the first five degrees of freedom, the five most recent readings of the force sensor can be considered to obtain a 5×5 reading matrix, and the matrix can be transformed and regularized (for example, all readings are processed as values ​​between -1 and 1) to obtain a 25×1 dimensional vector as the force characteristic v force .

[0056] In order to extract the tactile feedback signal related to the task and filter out irrelevant data, the collected tactile feedback signal can be subtracted from the initial tactile feedback signal to obtain the effective tactile feedback signal. In addition, the data read by the tactile sensor can also be reduced in dimension to facilitate processing. For example, Figure 3A As shown, each tactile sensor 130 mounted on the gripper 10 may include a plurality of (eg, 6×12) contact points 330, each of which may collect a touch signal to obtain a reading, such as Figure 3B The reading matrix shown in Figure 1 is shown in Figure 12. Since the effective data obtained by the contacts in the two columns at the edge of the robot when grasping the object is small, in order to reduce the signal dimension, the readings of these two columns can be ignored (i.e., only the readings of the contacts in the two columns at the edge of the robot are considered). Figure 3BIn addition, you can also Figure 3B The data collected from the two adjacent rows of contacts in the readings in the bold box are averaged to obtain Figure 3C The reduced-dimensional matrix shown in FIG. The reduced-dimensional matrix can be transformed and regularized to obtain a 24-dimensional vector (for the tactile sensor on one side). For the tactile sensors on both sides, a 48-dimensional vector is obtained as the tactile feature v tactile .

[0057] In the reasoning process (i.e., the process of controlling the manipulator), the image, force feedback signal, and tactile feedback signal collected in the real environment can be feature encoded to obtain the image feature v vision , force characteristic v force and tactile features v tactile . Then the image feature v vision , force characteristic v force and tactile features v tactile The three are spliced ​​to obtain fusion features

[0058] Then, the trained control strategy prediction network can be used to predict the fusion features v muti To predict the control action to be performed, so as to control the manipulator to complete the insertion task. The control strategy prediction network trained by the method of the present invention can accurately realize a series of complex operations from grasping objects (of different shapes and sizes) (for example, rings, nails, gears, etc.) to inserting objects into holes or columns, without switching the model in different stages, and has better universality and system stability. The control strategy prediction network trained by the method of the present invention can control the manipulator to complete fine and complex operations such as straight insertion, oblique insertion, and insertion tasks that require the insert to be rotated at a specific angle.

[0059] In order to effectively reduce the difference between the simulation environment and the real environment, so that the neural network model for controlling the robot trained in the simulation environment can be better used to control the robot in the real environment, the present disclosure proposes that before encoding the image of the robot's surrounding environment and the tactile feedback signal collected by the tactile sensor, the image of the robot's surrounding environment collected in the simulation environment and the tactile feedback signal collected by the tactile sensor can also be preprocessed.

[0060] Specifically, the present invention randomizes the lighting, color and texture parameters of the images of the robot's surrounding environment captured in the simulation environment, applies Gaussian blur, introduces random shadows, and adds white noise to the rendered image, so that the images of the robot's surrounding environment captured in the simulation environment for model training are closer to the images in the real environment.

[0061] like Figure 4A As shown in FIG, the force measurement value (i.e., tactile feedback force) F0 (unit: N) of the force sensor and the tactile signal (i.e., tactile feedback signal) C0 output by the tactile sensor are usually in correspondence. In order to calibrate the tactile sensor in the simulation environment to make it closer to the tactile sensor in the real environment, the tactile sensor in the simulation environment and the tactile sensor in the real environment can be calibrated separately. Figure 4B The experiment shown.

[0062] like Figure 4B As shown, the contact 440 of the force sensor 120 can be used to lightly press the contact point of the tactile sensor 130. Theoretically, when the contact 440 of the force sensor is used to lightly press the contact point of the tactile sensor, the force between the force sensor and the tactile sensor should be the same because the force is mutual.

[0063] In real and simulated environments, Figure 4B The experiments shown can obtain a set of measurement results (such as Figure 4C ). Figure 4C As can be seen from FIG, there is a linear relationship between the force measurement value F0 of the force sensor and the tactile signal C0 output by the tactile sensor (for example, Figure 4C The fitted curve l1 is shown in FIG.

[0064] That is, for the tactile sensor in the real environment, the tactile signal of the touch point i can be obtained The force f corresponding to the contact point i With the following relationship:

[0065]

[0066] Among them, k i and b i represents the linear parameter.

[0067] Similarly, for the tactile sensor in the simulation environment, the tactile signal of touch point i is The force f corresponding to the contact point i With the following relationship:

[0068]

[0069] By making the simulation environment and the real environment i The values ​​of are equal, and the linear parameter k after calibration of the simulation environment based on the real environment can be obtained. i ′ and b′ i .

[0070] The method proposed in this disclosure for processing images of the robot's surroundings captured in a simulated environment and tactile feedback signals collected by a tactile sensor can effectively reduce the difference between the simulated environment and the real environment. A control strategy prediction network trained based on the adjusted simulated environment can be directly and effectively applied to the real environment without the need for further training and adjustment of the trained control strategy prediction network.

[0071] Figure 5A is a schematic flowchart illustrating a method 500 for training a neural network model according to an embodiment of the present disclosure.

[0072] Among them, in step S510, for the first task, the first state information of the manipulator in the simulation environment when performing the first task is obtained, wherein the first state information of the manipulator includes: an image reflecting the environment surrounding the manipulator, a force feedback signal of the manipulator, and a tactile feedback signal when the manipulator contacts an object.

[0073] Training the model in a simulation environment does not require powering on and simulating the various states of the robot in a real environment, which can prevent damage to the robot in the real environment, thereby obtaining a neural network model with better performance for controlling the robot at a low cost.

[0074] The image of the environment surrounding the manipulator can be an image captured from different perspectives by multiple cameras (for example, two front and rear cameras), and the tactile feedback signal when the manipulator contacts an object can be a signal collected by multiple tactile sensors (for example, four tactile sensors respectively installed on the four claws of a four-claw manipulator). The force feedback signal of the manipulator can be a force signal of multiple dimensions (for example, 6 degrees of freedom). Optionally, in order to more comprehensively consider the information of the force signal, the most recent readings of the force sensor can be considered to obtain the force characteristics.

[0075] According to an embodiment of the present disclosure, before obtaining the first state information of the manipulator, the first state information obtained in the simulation environment may be preprocessed. For example, the environment parameters of the image may be randomized (for example, randomized lighting, color and texture parameters, etc.), and random shadows or noise may be added to the image (for example, applying Gaussian blur, adding white noise, etc.). In addition, the tactile feedback signal collected from the simulation environment may be calibrated based on the tactile feedback signal collected from the real environment. For example, Figures 4A-4CAs described, the transformation relationship between the tactile feedback signal and the tactile feedback force in the real environment and the transformation relationship between the tactile feedback signal and the tactile feedback force in the simulation environment can be obtained respectively; and based on the relationship between the tactile feedback signal and the tactile feedback force in the real environment (that is, the conversion relationship between the value of the tactile feedback signal and the force value), the relationship between the tactile feedback signal and the tactile feedback force in the simulation environment is calibrated.

[0076] The above preprocessing can effectively reduce the difference between the simulation environment and the real environment. The control strategy prediction network trained based on the adjusted simulation environment can be directly and effectively applied to the real environment without further training and adjustment of the trained control strategy prediction network.

[0077] In step S520 , the image, the force feedback signal, and the tactile feedback signal are encoded respectively to obtain image features, force features, and tactile features.

[0078] According to an embodiment of the present disclosure, the image may be encoded using an image feature encoding network to obtain the image features. The image feature encoding network may be a pre-trained image feature extraction network (e.g., a convolutional neural network (CNN), a recurrent neural network (RNN), a feature pyramid network, etc.).

[0079] The image feature encoding network can be trained simultaneously with the control strategy prediction network. To reduce the number of parameters required to update the neural network model each time and to avoid mutual interference between the image feature encoding network and the control strategy prediction network during training, the image feature encoding network and the control strategy prediction network can also be trained separately. For example, the parameters of the control strategy prediction network can be fixed first, and the parameters of the image feature encoding network can be optimized. Then, the parameters of the image feature encoding network can be fixed again, and the parameters of the control strategy prediction network can be optimized. This process can be repeated multiple times until the processing results of both models meet expectations.

[0080] Other neural network models can be used to assist in the training of the image feature encoding network. For example, if the task is an insertion task, a position prediction network can be used to predict a first positional relationship between the object to be inserted and the object being inserted based on the image features; a second positional relationship between the object to be inserted and the object being inserted can be obtained through the simulation environment; a loss function value can be calculated based on the difference between the first positional relationship and the second positional relationship; and the parameters of the image feature encoding network and the position prediction network can be optimized based on the loss function value.

[0081] It should be understood that the position prediction network here is only used to assist the training of the image feature encoding network. The neural network model ultimately used to control the manipulator includes the image feature encoding network and the control strategy prediction network, but does not include the position prediction network.

[0082] According to an embodiment of the present disclosure, the force feedback signal may be subjected to dimensionality transformation and regularization processing to obtain the force feature.

[0083] According to an embodiment of the present disclosure, dimension transformation and regularization processing may be performed on the tactile feedback signal to obtain the tactile feature.

[0084] Optionally, the force feedback signal and / or the tactile feedback signal may be filtered and effective data screened (such as the above combination Figure 3B and Figure 3C After pre-processing operations such as those described above, the processed force feedback signal and / or tactile feedback signal is encoded to obtain the force feature or the tactile feature.

[0085] It should be understood that the process of performing dimensional transformation and regularization processing on the force feedback signal, or performing dimensional transformation and regularization processing on the tactile feedback signal can be implemented using a simple operation matrix, which can have the structure of a neural network model, but its parameters do not need to be optimized during the training process.

[0086] In step S530 , the image feature, the force feature, and the tactile feature are fused to obtain a fused feature.

[0087] According to an embodiment of the present disclosure, the image feature, the force feature, and the tactile feature may be spliced ​​to obtain a fusion feature, but the present disclosure is not limited thereto.

[0088] In step S540 , based on the fusion features, a control strategy prediction network is used to predict a first control strategy for the manipulator.

[0089] According to an embodiment of the present disclosure, the process of using a control strategy prediction network to predict the first control strategy for the manipulator conforms to a Markov decision process. Specifically, the process of the manipulator performing the first task includes multiple states, and the first control strategy includes a first control action corresponding to each of the multiple states of the manipulator. For each of the multiple states of the manipulator, the control strategy prediction network can be used to predict the first control action for the manipulator, and based on the state and the first control action corresponding to the state, the next state of the manipulator in the simulation environment is determined (that is, the manipulator is controlled in the simulation environment to perform the control action corresponding to the state on the basis of the state to obtain the next state of the manipulator), wherein the multiple states in the process of executing the first task are determined in sequence by determining the next state for the determined next state in sequence until the first task is completed. In each of the multiple states of the manipulator, the simulation environment can be used to determine the first state information of the manipulator in the state, and the first control action includes at least one of a moving action and a twisting action.

[0090] For example, the fusion feature can be input into the control strategy prediction network π to predict the control action A to be executed, and the simulation software can be used to obtain the next state S' based on the current state S and the predicted control action A. Then, for the next state S', the control action A' to be executed is predicted, and based on the next state S' and the control action A' to be executed, the next state S' in the simulation environment is determined until the first task is completed (that is, the first task is successfully completed, or the first task fails to execute). Starting from the initial state of the first task, the control action to be executed at each step is predicted, and its next state is determined, and then the next state of the next state is determined in sequence for the determined next state until the first task is completed. In this process, the initial state and each determined next state constitute the multiple states.

[0091] In step S550, a first reward value corresponding to the first control strategy is calculated.

[0092] According to an embodiment of the present disclosure, a first reward value for executing a first control action corresponding to each of the multiple states of the manipulator can be determined; and a weighted sum of the first reward values ​​for the first control action corresponding to each state of the manipulator is performed to determine a first reward value corresponding to the first control strategy. For example, the first reward value can be calculated using the above formula (2).

[0093] According to an embodiment of the present disclosure, for an insertion task, the first reward value can be determined based on the insertion depth of the object, wherein the first reward value increases as the insertion depth of the object increases. For example, the first reward value can be calculated by the above formula (1). Optionally, in the event that the insertion task fails, the penalty value of the last first control action performed by the manipulator (i.e., the last control action before the task fails) can also be determined; and the first reward value of the last first control action performed by the manipulator is determined based on the insertion depth of the object and the penalty value. For example, the reward value r' can be determined based on the insertion depth of the object, and the difference between the reward value r' and the penalty value is used as the first reward value. Optionally, the reward value r' can also be determined based on the insertion depth of the object, and the product of the reward value r' and the penalty value is used as the first reward value. In fact, there are many ways to calculate the first reward value and the first return value, and the present disclosure does not impose specific restrictions on this.

[0094] In step S560 , the parameters of the control strategy prediction network are optimized based on the first reward value.

[0095] Before executing method 500 to optimize the parameters of the control strategy prediction network, second state information of the manipulator in a simulation environment when performing the second task can also be obtained for the second task, wherein the second state information of the manipulator includes: an image reflecting the environment surrounding the manipulator (the second state information may not include the force feedback signal of the manipulator, or the tactile feedback signal when the manipulator contacts an object); encoding the image to obtain image features; based on the image features, using the control strategy prediction network to predict the second control strategy for the manipulator; calculating the second reward value corresponding to the second control strategy; and optimizing the parameters of the control strategy prediction network based on the second reward value.

[0096] According to an embodiment of the present disclosure, the first task may be an oblique insertion task, and the second task may be a straight insertion task. Specifically, the parameters of the control strategy prediction network can be optimized using an image reflecting the robot's surroundings, enabling the network to handle simple straight insertion tasks. Force and tactile feedback signals from the image reflecting the robot's surroundings can then be used to optimize the parameters of the control strategy prediction network, enabling it to further handle complex oblique insertion tasks.

[0097] It should be noted that the direct insertion task in the present disclosure is a task of inserting a first object into a second object at a vertical angle, including the following: Figure 2A The task of inserting the pin-shaped component vertically into the hole as shown, and Figure 2BThe task of inserting the ring part vertically outside the plug-shaped part is shown. The oblique insertion task in the present disclosure is the task of inserting the first object into the second object at a non-vertical angle, including the following: Figure 2C The task of inserting a plug-shaped component into a hole at a non-perpendicular angle as shown, and Figure 2D The task shown is to place the annular member over the plug-shaped member at a non-perpendicular angle.

[0098] According to an embodiment of the present disclosure, the process of using a control strategy prediction network to predict the second control strategy for the manipulator also conforms to the Markov decision process. Specifically, the process of the manipulator performing the second task may include multiple states, and the second control strategy includes a second control action corresponding to each of the multiple states of the manipulator. For each of the multiple states of the manipulator, the control strategy prediction network is used to predict the second control action for the manipulator, and based on the state and the second control action corresponding to the state, the next state of the manipulator in the simulation environment is determined (that is, the manipulator is controlled in the simulation environment to perform the control action corresponding to the state on the basis of the state to obtain the next state of the manipulator), wherein the next state of the determined next state is determined in sequence until the second task is completed, thereby determining the multiple states in the second task in sequence. In each of the multiple states of the manipulator, the simulation environment can be used to determine the second state information of the manipulator in the state, and the second control action includes a moving action.

[0099] For example, the image features can be input into the control strategy prediction network π to predict the control action A to be executed, and the simulation software can be used to obtain the next state S' based on the current state S and the predicted control action A. Then, for the next state S', the control action A' to be executed is predicted, and based on the next state S' and the control action A' to be executed, the next state S' in the simulation environment is determined until the second task is completed (that is, the second task is successfully completed, or the second task fails to execute). Starting from the initial state of the second task, the control action to be executed at each step is predicted, and its next state is determined, and then the next state of the next state is determined in sequence for the determined next state until the second task is completed. In this process, the initial state and each determined next state constitute the multiple states.

[0100] According to an embodiment of the present disclosure, a second reward value for executing the second control action corresponding to each of the multiple states of the manipulator can be determined; the second reward values ​​for the second control actions corresponding to each state of the manipulator are weighted and summed to determine the second reward value corresponding to the second control strategy. For example, the second reward value can be calculated using the above formulas (1) and (2). In fact, there are many ways to calculate the second reward value and the second reward value, and this disclosure does not specifically limit this.

[0101] The following combination Figure 5B The process of training the neural network model is further explained.

[0102] Figure 5A In the training process of the neural network model for the first task and the training process of the neural network model for the second task have Figure 5B Steps shown.

[0103] Among them, for the neural network model training process for the first task, the first state information of the manipulator in the simulation environment when performing the first task can be obtained, wherein the first state information of the manipulator includes: an image reflecting the environment around the manipulator, a force feedback signal of the manipulator, and a tactile feedback signal when the manipulator contacts an object (step S511). Then, the image, the force feedback signal, and the tactile feedback signal can be encoded respectively to obtain image features, force features, and tactile features (step S512). By fusing the image features, the force features, and the tactile features, a fusion feature can be obtained (this step S513 only exists in the neural network model training process for the first task, and is not required for the neural network model training process for the second task, so it is shown in a dotted box). Based on the fusion feature, the control strategy prediction network can be used to predict the first control action for the current state (step S514). Then, the next state of the manipulator in the simulation environment can be determined based on the current state and the first control action corresponding to the current state (step S515). When the next state of the manipulator is determined, it can be determined whether the first task has been completed (step S516). If the first task has not been completed, the above steps S511 to S516 are repeated. If the first task is completed, the reward value corresponding to each predicted first control action can be used to calculate the first reward value corresponding to the first control strategy (step S517), and the first control strategy includes all predicted first control actions from the start of the first task to the completion of the first task. Furthermore, the parameters of the control strategy prediction network can be optimized based on the first reward value (step S518) so that the performance of the control strategy prediction network meets the requirements (that is, the control accuracy of the manipulator meets the requirements).

[0104] During the neural network model training process for the second task, second state information of the manipulator in the simulation environment while performing the second task can be obtained, wherein the second state information of the manipulator includes an image reflecting the environment surrounding the manipulator (step S511). Subsequently, the image can be encoded to obtain image features (step S512). Based on the image features, a control strategy prediction network can be used to predict a second control action for the current state (step S514). Then, based on the current state and the second control action corresponding to the current state, the next state of the manipulator in the simulation environment can be determined (step S515). Once the next state of the manipulator is determined, it can be determined whether the second task has been completed (step S516). If the second task has not been completed, steps S511, S512, S514, S515, and S516 are repeated. If the second task has been completed, the reward value corresponding to each predicted second control action can be used to calculate a second reward value corresponding to the second control strategy (step S517), where the second control strategy includes all predicted second control actions from the start of the second task to the completion of the second task. Furthermore, the parameters of the control strategy prediction network may be optimized based on the second reward value (step S518 ) so that the performance of the control strategy prediction network meets the requirements (ie, the control accuracy of the manipulator meets the requirements).

[0105] Figure 6 is a schematic flow chart illustrating a method 600 for controlling a robot according to an embodiment of the present disclosure.

[0106] In step S610, the state information of the manipulator when performing the task is obtained, wherein the state information of the manipulator includes: an image reflecting the environment around the manipulator, a force feedback signal of the manipulator, and a tactile feedback signal when the manipulator contacts an object.

[0107] According to an embodiment of the present disclosure, the tasks may be various tasks of manipulating objects using a manipulator. For example, the tasks may include grasping objects (of varying shapes and sizes) (e.g., rings, nails, gears, etc.), rotating an object to a specific angle, inserting an object into a hole or column, etc.

[0108] In step S620 , the image, the force feedback signal, and the tactile feedback signal are encoded respectively to obtain image features, force features, and tactile features.

[0109] According to an embodiment of the present disclosure, an image feature encoding network may be used to encode the image to obtain the image features.

[0110] According to an embodiment of the present disclosure, the force feedback signal may be subjected to dimensionality transformation and regularization processing to obtain the force feature.

[0111] According to an embodiment of the present disclosure, dimension transformation and regularization processing may be performed on the tactile feedback signal to obtain the tactile feature.

[0112] In step S630 , the image feature, the force feature, and the tactile feature are fused to obtain a fused feature.

[0113] In step S640 , based on the fusion features, a control strategy prediction network is used to predict a control strategy for the manipulator.

[0114] The process of using a control strategy prediction network to predict the control strategy for the manipulator conforms to a Markov decision process. Specifically, the process of the manipulator performing the task includes multiple states, the control strategy includes a control action corresponding to each of the multiple states of the manipulator, and the control action includes at least one of a moving action and a twisting action, wherein the process of using a control strategy prediction network to predict the control strategy for the manipulator includes: for each of the multiple states of the manipulator, using the control strategy prediction network to predict the control action for the manipulator. wherein, in each state, after the control action predicted for the state is executed on the manipulator, the state of the manipulator changes to the next state, wherein the control action for the manipulator is predicted for the next state in turn, and the predicted control action is executed on the manipulator until the task is completed; wherein, the multiple states are generated sequentially by the manipulator under the control of the control action predicted by the control strategy prediction network during the execution of the task.

[0115] In step S650, the movement of the manipulator is controlled based on the control strategy.

[0116] According to an embodiment of the present disclosure, the computer or controller for controlling the manipulator can convert the control strategy into a control signal for the manipulator, thereby controlling the manipulator to perform multiple steps to complete the task. For example, in a real environment, the control action to be performed can be predicted based on the current state of the manipulator, and the computer or controller for controlling the manipulator can be used to make the manipulator perform the predicted control action, thereby changing the manipulator from the current state to the next state. And so on, until the task is completed, the manipulator will appear in the process of performing the task.

[0117] It should be understood that the processing of steps S610 to S640 in method 600 is similar to that of steps S510 to S540 in method 500 and will not be repeated here. According to an embodiment of the present disclosure, the control strategy prediction network can be trained using the above-mentioned method 500. The method 500 is implemented in a simulation environment (e.g., a virtual manipulator control environment), and the method 600 can be implemented in a real environment corresponding to the simulation environment (i.e., based on a real manipulator control structure).

[0118] Method 600 fuses the features of vision, force, and touch, and uses a control strategy prediction network to predict the control strategy based on the fused features, which is more accurate and can be applied to oblique insertion tasks, insertion tasks that require rotating the object by a certain angle, insertion tasks for small objects, and other refined insertion tasks.

[0119] The control strategy prediction network based on the present disclosure can realize a series of control operations from the approach of the object to be inserted and the inserted object to the completion of the insertion, without switching the model in different stages (for example, using a vision-based model before the object to be inserted and the inserted object come into contact, and using a force-based model after the object to be inserted and the inserted object come into contact), with better universality and system stability.

[0120] Figure 7 2 is a schematic diagram showing the composition of an apparatus 700 for training a neural network model according to an embodiment of the present disclosure.

[0121] According to an embodiment of the present disclosure, the apparatus 700 for training a neural network model may include: an information acquisition module 710 , a feature encoding module 720 , a feature fusion module 730 , a strategy prediction module 740 , a calculation module 750 , and a model training module 760 .

[0122] Among them, the information acquisition module 710 can be configured to: for a first task, obtain the first state information of the manipulator in a simulation environment when performing the first task, wherein the first state information of the manipulator includes: an image reflecting the environment surrounding the manipulator, a force feedback signal of the manipulator, and a tactile feedback signal when the manipulator contacts an object.

[0123] The feature encoding module 720 may be configured to respectively encode the image, the force feedback signal, and the tactile feedback signal to obtain image features, force features, and tactile features.

[0124] The feature fusion module 730 may be configured to fuse the image feature, the force feature, and the tactile feature to obtain a fused feature.

[0125] The strategy prediction module 740 may be configured to: predict a first control strategy for the manipulator based on the fusion features using a control strategy prediction network.

[0126] As previously described, in a simulation environment, the manipulator is controlled based on the first control strategy, causing the manipulator to sequentially execute multiple control actions until the first task is completed. Specifically, starting from an initial state at the start of the first task, the control strategy prediction network predicts the control actions to be executed at each step, and determines the next state within the simulation environment. Then, the next state of each next state is sequentially determined for each determined next state until the first task is completed.

[0127] The calculation module 750 may be configured to calculate a first reward value corresponding to the first control strategy.

[0128] The model training module 760 may be configured to optimize the parameters of the control strategy prediction network based on the first reward value.

[0129] According to an embodiment of the present disclosure, before optimizing the parameters of the control strategy prediction network based on the first reward value, the information acquisition module 710 may also be configured to: for the second task, obtain the second state information of the manipulator in the simulation environment when performing the second task, wherein the second state information of the manipulator includes: an image reflecting the environment surrounding the manipulator. The feature encoding module 720 may also be configured to: encode the image to obtain image features. The strategy prediction module 740 may also be configured to: predict the second control strategy for the manipulator using the control strategy prediction network based on the image features. The calculation module 750 may also be configured to: calculate the second reward value corresponding to the second control strategy. The model training module 760 may also be configured to: optimize the parameters of the control strategy prediction network based on the second reward value.

[0130] According to an embodiment of the present disclosure, the first task may be a task based on vision, force, and touch, and the second task may be a task based on vision. For example, the first task may be a more complex oblique insertion task, and the second task may be a simpler straight insertion task.

[0131] It should be understood that Figure 7 The apparatus 700 for training a neural network model can be implemented as shown in FIG. Figure 5AThe various methods 500 for training a neural network model described herein include the information acquisition module 710, the feature encoding module 720, the feature fusion module 730, the strategy prediction module 740, the calculation module 750, and the model training module 760, which can respectively implement the processing of step S510, step S520, step S530, step S540, step S550, and step S560, and are not further described herein.

[0132] Figure 8 FIG. 8 is a schematic diagram showing the composition of an apparatus 800 for controlling a robot according to an embodiment of the present disclosure.

[0133] According to an embodiment of the present disclosure, an apparatus 800 for controlling a manipulator may include: an information acquisition module 810 , a feature encoding module 820 , a feature fusion module 830 , a strategy prediction module 840 and a control module 850 .

[0134] Among them, the information acquisition module 810 can be configured to: obtain the status information of the manipulator when performing a task, wherein the status information of the manipulator includes: an image reflecting the environment around the manipulator, the force feedback signal of the manipulator, and the tactile feedback signal when the manipulator contacts an object.

[0135] The feature encoding module 820 may be configured to respectively encode the image, the force feedback signal, and the tactile feedback signal to obtain image features, force features, and tactile features.

[0136] The feature fusion module 830 may be configured to fuse the image feature, the force feature, and the tactile feature to obtain a fused feature.

[0137] The strategy prediction module 840 may be configured to: predict the control strategy for the manipulator based on the fusion features using a control strategy prediction network.

[0138] The control module 850 may be configured to control the movement of the manipulator based on the control strategy.

[0139] In a real-world environment, the robot can be controlled based on the control strategy, causing it to sequentially execute multiple control actions until the task is completed. Specifically, starting from the initial state at the start of the task, the control strategy prediction network predicts the control actions to be executed at each step. The robot is then controlled to execute the control actions to determine its next state. The next state of the next state is then determined sequentially based on the determined next state until the task is completed.

[0140] It should be understood that the device 800 for controlling the manipulator can be implemented as follows Figure 6Various methods 600 for controlling a manipulator are described. The control strategy prediction network can be trained using the above-described method 500. The information acquisition module 810, feature encoding module 820, feature fusion module 830, strategy prediction module 840, and control module 850 can respectively implement the processing of steps S610, S620, S630, S640, and S650, which will not be described in detail here.

[0141] The neural network model for controlling the manipulator provided in the embodiment of the present disclosure can be deployed in a controller for controlling the manipulator, which can be integrated with the manipulator or structurally independent of the manipulator (for example, the manipulator can be controlled based on an electronic device independent of the manipulator (such as a server, terminal, etc.)). According to the embodiment of the present disclosure, the neural network model for controlling the manipulator can be integrated in the terminal. The terminal can be a mobile phone, tablet computer, laptop computer, desktop computer, personal computer (PC), smart speaker or smart watch, etc., but is not limited to this. For another example, the neural network model for controlling the manipulator can also be integrated in the server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0142] The neural network model for controlling the manipulator provided by the embodiment of the present disclosure may also involve artificial intelligence cloud services in the field of cloud technology. Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network in a wide area network or a local area network to realize the calculation, storage, processing and sharing of data. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the application of cloud computing business model. It can form a resource pool that can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The background services of technical network systems require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the rapid development and application of the Internet industry, in the future, each item may have its own identification mark, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately. All kinds of industry data require strong system backing support, which can only be achieved through cloud computing.

[0143] Among them, AI cloud services, also commonly referred to as AIaaS (AI as a Service), are a mainstream AI platform service model. Specifically, AIaaS platforms separate several common AI services and provide them as standalone or packaged services in the cloud. This service model is similar to opening an AI-themed mall: all developers can access one or more AI services provided by the platform through the application programming interface (API). Some experienced developers can also use the platform's AI framework and AI infrastructure to deploy and operate their own cloud AI services.

[0144] During the training phase of the neural network model for controlling the manipulator, the server can obtain training samples from the simulation environment and train the neural network model for controlling the manipulator based on the training samples. After training is completed, the trained neural network model for controlling the manipulator can be deployed to one or more controllers or servers (or cloud services) to control the manipulator in a real environment to complete various tasks.

[0145] It is worth noting that all training samples used in this disclosure comply with the legality, ethics, and privacy requirements of laws and regulations. Specifically, all training samples are sourced from legitimate sources and have been collected with the explicit permission of the user. Furthermore, all training samples used in this disclosure adhere to privacy protection principles, have been rigorously screened and cleaned, and will not be disclosed to any unauthorized third party.

[0146] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.

[0147] For example, the method or apparatus according to the embodiment of the present disclosure may also be implemented by Figure 9 The architecture of the computing device 3000 shown in FIG. Figure 9As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the method provided in the present disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 9 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 9 One or more components of a computing device are shown.

[0148] According to another aspect of the present disclosure, a computer-readable storage medium is also provided. Computer-readable instructions are stored on the computer storage medium. When the computer-readable instructions are executed by a processor, the method according to the embodiment of the present disclosure described with reference to the above figures can be executed. The computer-readable storage medium in the embodiment of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. It should be noted that memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0149] Embodiments of the present disclosure also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a method according to an embodiment of the present disclosure.

[0150] In summary, the embodiments of the present disclosure provide methods, devices, computer program products, and storage media for training a neural network model, as well as methods, devices, computer program products, and storage media for controlling a robotic arm.

[0151] The method for training a neural network model disclosed herein includes: for a first task, obtaining first state information of a manipulator in a simulation environment when performing the first task, wherein the first state information of the manipulator includes: an image reflecting the environment surrounding the manipulator, a force feedback signal of the manipulator, and a tactile feedback signal when the manipulator contacts an object; encoding the image, the force feedback signal, and the tactile feedback signal respectively to obtain image features, force features, and tactile features; fusing the image features, the force features, and the tactile features to obtain fused features; based on the fused features, using a control strategy prediction network to predict a first control strategy for the manipulator; calculating a first reward value corresponding to the first control strategy; and optimizing parameters of the control strategy prediction network based on the first reward value.

[0152] By integrating visual, force, and tactile features and utilizing a control strategy prediction network to predict control strategies based on these integrated features, this method achieves higher accuracy and is applicable to various specialized insertion tasks, such as oblique insertion, insertion tasks requiring rotation of an object at a certain angle, and insertion tasks targeting small objects. The control strategy prediction network based on this disclosure can implement a series of control operations from the approach of the object to be inserted to the completion of insertion, without requiring stage-by-stage model switching, resulting in improved universality and system stability.

[0153] Furthermore, the method proposed in this disclosure for processing images of the robot's surroundings captured in a simulated environment and tactile feedback signals collected by a tactile sensor can effectively reduce the difference between the simulated environment and the real environment. The control strategy prediction network trained based on the adjusted simulated environment can be directly and effectively applied to the real environment without the need for further training and adjustment of the trained control strategy prediction network.

[0154] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of the code, and the module, program segment, or a part of the code contains at least one executable instruction for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.

[0155] This disclosure uses specific terms to describe the embodiments of the present disclosure. For example, "first / second embodiment," "one embodiment," and / or "some embodiments" refer to a certain feature, structure, or characteristic associated with at least one embodiment of the present disclosure. Therefore, it should be emphasized and noted that "one embodiment," "an embodiment," or "an alternative embodiment" mentioned two or more times in different locations in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of the present disclosure may be appropriately combined.

[0156] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0157] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. It should also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology and should not be interpreted in an idealized or highly formal sense, unless expressly defined as such herein.

[0158] The above is an illustration of the present invention and should not be considered as limiting thereof. Although several exemplary embodiments of the present invention have been described, it will be readily understood by those skilled in the art that many modifications may be made to the exemplary embodiments without departing from the novel teachings and advantages of the present invention. Therefore, all such modifications are intended to be included within the scope of the present invention as defined by the claims. It should be understood that the above is an illustration of the present invention and should not be considered as being limited to the specific embodiments disclosed, and modifications to the disclosed embodiments and other embodiments are intended to be included within the scope of the appended claims. The present invention is defined by the claims and their equivalents.

Claims

1. A method for training a neural network model, comprising: For a first task, obtaining first state information of the manipulator in a simulation environment when performing the first task, wherein the first state information of the manipulator includes: an image reflecting the environment surrounding the manipulator, a force feedback signal of the manipulator, and a tactile feedback signal when the manipulator contacts an object; Encoding the image, the force feedback signal, and the tactile feedback signal respectively to obtain image features, force features, and tactile features; fusing the image feature, the force feature, and the tactile feature to obtain a fused feature; Based on the fusion features, using a control strategy prediction network to predict a first control strategy for the manipulator; Calculating a first reward value corresponding to the first control strategy; and Optimizing parameters of the control strategy prediction network based on the first reward value.

2. The method according to claim 1, wherein The process of the manipulator performing the first task includes multiple states, and the first control strategy includes a first control action corresponding to each of the multiple states of the manipulator. Predicting the first control strategy for the manipulator using a control strategy prediction network includes: For each of the plurality of states of the manipulator, predict a first control action for the manipulator using the control strategy prediction network, and determine a next state of the manipulator in the simulation environment based on the state and the first control action corresponding to the state, wherein the plurality of states in the first task are sequentially determined by sequentially determining the next state until the first task is completed. Wherein, in each of the multiple states of the manipulator, the simulation environment is used to determine first state information of the manipulator in the state, and the first control action includes at least one of a moving action and a twisting action.

3. The method according to claim 2, wherein: Calculating a first reward value corresponding to the first control strategy includes: For each of the plurality of states of the manipulator, determining a first reward value for executing a first control action corresponding to the state; A weighted sum is performed on the first reward values ​​of the first control actions corresponding to each state of the manipulator to determine a first reward value corresponding to the first control strategy.

4. The method according to claim 3, wherein: The first task is an insert task, For each of the multiple states of the manipulator, determining a first reward value for executing a first control action corresponding to the state includes: determining the first reward value based on an insertion depth of the object, wherein the first reward value increases as the insertion depth of the object increases.

5. The method according to claim 4, wherein: In the event that the insertion task fails to be executed, determining a penalty value of the last first control action executed by the manipulator; and The first reward value of the last first control action performed by the manipulator is determined based on the insertion depth of the object and the penalty value.

6. The method according to claim 1, before optimizing the parameters of the control strategy prediction network based on the first reward value, the method further comprises: For the second task, obtaining second state information of the manipulator when performing the second task in the simulation environment, wherein the second state information of the manipulator includes: an image reflecting the environment surrounding the manipulator; Encoding the image to obtain image features; Based on the image features, using a control strategy prediction network to predict a second control strategy for the manipulator; Calculating a second reward value corresponding to the second control strategy; Optimize parameters of the control strategy prediction network based on the second reward value.

7. The method according to claim 6, wherein: The process of the manipulator performing the second task includes multiple states, and the second control strategy includes a second control action corresponding to each of the multiple states of the manipulator. The second control strategy for the manipulator predicted by the control strategy prediction network includes: For each of the plurality of states of the manipulator, predict a second control action for the manipulator using the control strategy prediction network, and determine a next state of the manipulator in the simulation environment based on the state and the second control action corresponding to the state, wherein the plurality of states in the second task are sequentially determined by sequentially determining the next state until the second task is completed. Wherein, in each of the multiple states of the manipulator, the simulation environment is used to determine second state information of the manipulator in the state, and the second control action includes a moving action.

8. The method of claim 7, wherein: Calculating the second reward value corresponding to the second control strategy includes: For each of the plurality of states of the manipulator, determining a second reward value for executing a second control action corresponding to the state; A weighted sum is performed on the second reward values ​​of the second control actions corresponding to each state of the manipulator to determine a second reward value corresponding to the second control strategy.

9. The method of claim 6, wherein: The first task is an oblique insertion task, and the second task is a straight insertion task.

10. The method of claim 1, wherein: Before obtaining the first state information of the manipulator, the method further includes: preprocessing the first state information obtained in the simulation environment, Preprocessing the first state information obtained in the simulation environment includes at least one of the following: Randomize the environment parameters for the image, Add random shading or noise to the image, The tactile feedback signal collected from the simulation environment is calibrated based on the tactile feedback signal collected from the real environment.

11. The method according to claim 10, wherein: Based on the tactile feedback signal collected from the real environment, calibrating the tactile feedback signal collected from the simulation environment includes: respectively obtaining a transformation relationship between a tactile feedback signal and a tactile feedback force in a real environment and a transformation relationship between a tactile feedback signal and a tactile feedback force in a simulated environment; Based on the relationship between the tactile feedback signal and the tactile feedback force in the real environment, the relationship between the tactile feedback signal and the tactile feedback force in the simulation environment is calibrated.

12. The method of claim 1, wherein: Encoding the force feedback signal to obtain a force characteristic includes: performing dimensionality transformation and regularization processing on the force feedback signal to obtain the force feature; Encoding the tactile feedback signal to obtain a tactile feature includes: Performing dimension transformation and regularization processing on the tactile feedback signal to obtain the tactile feature.

13. The method of claim 1, wherein: Encoding the image to obtain image features includes: encoding the image using an image feature encoding network to obtain the image features, In the case where the task is an insert task, the method further includes: Based on the image features, using a position prediction network to predict a first positional relationship between the object to be inserted and the object to be inserted; Acquiring a second positional relationship between the object to be inserted and the object to be inserted through the simulation environment; calculating a value of a loss function based on a difference between the first positional relationship and the second positional relationship; and Parameters of the image feature encoding network and the position prediction network are optimized based on the value of the loss function.

14. A method for controlling a robot, comprising: Acquiring state information of the manipulator when performing a task, wherein the state information of the manipulator includes: an image reflecting the environment surrounding the manipulator, a force feedback signal of the manipulator, and a tactile feedback signal when the manipulator contacts an object; Encoding the image, the force feedback signal, and the tactile feedback signal respectively to obtain image features, force features, and tactile features; fusing the image feature, the force feature, and the tactile feature to obtain a fused feature; Based on the fusion features, using a control strategy prediction network to predict a control strategy for the manipulator; and The movement of the manipulator is controlled based on the control strategy.

15. The method of claim 14, wherein: The process of the manipulator performing the task includes multiple states, the control strategy includes a control action corresponding to each of the multiple states of the manipulator, and the control action includes at least one of a movement action and a twisting action. The control strategy prediction network is used to predict the control strategy for the manipulator, including: For each of a plurality of states of the manipulator, predicting a control action for the manipulator using the control strategy prediction network; In each state, after the robot performs the control action predicted for the state, the state of the robot changes to the next state. wherein, predicting a control action for the manipulator for the next state in sequence, and executing the predicted control action on the manipulator until the task is completed; The multiple states are generated sequentially by the manipulator during the execution of the task under the control of the control action predicted by the control strategy prediction network.

16. The method of claim 14, wherein: Encoding the image to obtain image features includes: Encoding the image using an image feature encoding network to obtain the image features; Encoding the force feedback signal to obtain a force characteristic includes: Performing dimension transformation and regularization processing on the force feedback signal to obtain the force feature; Encoding the tactile feedback signal to obtain a tactile feature includes: Performing dimension transformation and regularization processing on the tactile feedback signal to obtain the tactile feature.

17. A device for training a neural network model, comprising: an information acquisition module configured to: for a first task, acquire first state information of a manipulator in a simulation environment when performing the first task, wherein the first state information of the manipulator includes: an image reflecting the environment surrounding the manipulator, a force feedback signal of the manipulator, and a tactile feedback signal when the manipulator contacts an object; a feature encoding module, configured to: encode the image, the force feedback signal, and the tactile feedback signal respectively to obtain image features, force features, and tactile features; a feature fusion module, configured to: fuse the image feature, the force feature, and the tactile feature to obtain a fused feature; A strategy prediction module is configured to: predict a first control strategy for the manipulator using a control strategy prediction network based on the fusion feature; a calculation module, configured to: calculate a first reward value corresponding to the first control strategy; and The model training module is configured to optimize the parameters of the control strategy prediction network based on the first reward value.

18. A device for controlling a manipulator, comprising: an information acquisition module configured to: acquire state information of the manipulator when performing a task, wherein the state information of the manipulator includes: an image reflecting the environment surrounding the manipulator, a force feedback signal of the manipulator, and a tactile feedback signal when the manipulator contacts an object; a feature encoding module, configured to: encode the image, the force feedback signal, and the tactile feedback signal respectively to obtain image features, force features, and tactile features; a feature fusion module, configured to: fuse the image feature, the force feature, and the tactile feature to obtain a fused feature; A strategy prediction module is configured to: predict a control strategy for the manipulator based on the fusion features using a control strategy prediction network; and The control module is configured to control the movement of the manipulator based on the control strategy.

19. A computer program product, comprising computer software code, wherein the computer software code is configured to implement the method according to any one of claims 1 to 16 when executed by a processor.

20. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the instructions are used to implement the method according to any one of claims 1 to 16 when executed by a processor.