A control model training method and a robot arm control method
By acquiring task and state sample information, and using the adjudication network to calculate the responsibility weight signal and update the core network parameters, the adaptability problem of traditional robotic arm control methods in complex environments and tasks is solved, achieving rapid learning and efficient decision-making.
Patent Information
- Application Number
- CN202310996590.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-08
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2043-08-08
AI Technical Summary
Traditional robotic arm control methods lack adaptability to changes in environment or task, resulting in singular decision-making and low training efficiency.
By acquiring task and state sample information, the responsibility weight signal of each processing layer is calculated using the adjudication network, and then input into the learning layer of the core network to update the parameters of relevant sub-modules until the core network converges, thereby achieving rapid learning and decision-making.
It accelerates the learning and decision-making processes, and improves the adaptability and training efficiency of the robotic arm in complex environments and tasks.
Smart Images

Figure CN117067202B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot control, and more particularly to a method for training a control model and a method for controlling a robotic arm. Background Technology
[0002] Robotic arm control methods have always been a hot research topic, and they have gradually moved from the initial industrial field into people's daily lives. However, most of them are only applied to situations where the working environment is simple or the work task is fixed. Traditional robotic arm control methods lack adaptability when the environment or task changes.
[0003] Critics' evaluation models can solve the problem of fixed task assignments mentioned above. However, after each interaction, the state abstraction function in the selected module transforms the world observations into a module-specific subset of world states. Modules are only associated with the world they operate on and are not coupled with other modules or arbitrators. This makes the system's learning dependent on a single module, leading to decision uniqueness. Furthermore, the number of modules in its architecture is predetermined, which gradually limits its performance as the complexity of the task increases, resulting in decision uniqueness and reduced training efficiency. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a training method for a control model and a control method for a robotic arm, which can accelerate the learning and decision-making processes.
[0005] To address the aforementioned problems, firstly, the present invention provides a method for training a control model, comprising the following steps:
[0006] Acquire task sample information and state sample information; the state sample information includes the state sample information of the robotic arm and the state sample information of the target object, and the task sample information includes the actions to be performed;
[0007] The task sample information and the state sample information are input into the adjudication network to obtain the current responsibility weight signal of each processing layer; the adjudication network includes several processing layers;
[0008] The responsibility weight signal of each processing layer is input into the corresponding learning layer in the core network to obtain the predicted action; the core network includes several learning layers, and the learning layers of the core network correspond one-to-one with the processing layers of the adjudication network. The subsequent learning layers of the core network are updated according to the preceding learning layers.
[0009] The network parameters of the core network are updated based on the current layer responsibility weight signal, the predicted action, and the state sample information.
[0010] Based on the predicted action and the executed action, determine whether the core network has converged. If the core network has not converged, repeatedly acquire task sample information and state sample information until the core network converges.
[0011] Optionally, the step of inputting the task sample information and the state sample information into the adjudication network to obtain the current responsibility weight signal for each processing layer specifically includes:
[0012] The task sample information is processed through a fully connected layer to obtain a first vector;
[0013] The state sample information is then processed through two fully connected layers to obtain a second vector.
[0014] The current responsibility weight signal for each processing layer is obtained based on the first vector and the second vector.
[0015] Optionally, obtaining the current responsibility weight signal for each processing layer based on the first vector and the second vector specifically includes:
[0016] Multiply the first vector and the second vector to obtain the first input of the adjudication network, and obtain the initial responsibility weight signal of each processing layer based on the first input;
[0017] The initial responsibility weight signal of each processing layer is normalized to obtain the current responsibility weight signal of each processing layer.
[0018] Optionally, obtaining the initial responsibility weight signal for each processing layer based on the first input specifically includes:
[0019] The first input is passed through an activation function layer and a fully connected layer to obtain the initial responsibility weight signal of the first processing layer;
[0020] For each processing layer after the first layer of the adjudication network, the first input and the initial responsibility weight signal of the previous processing layer are passed through two fully connected layers and an activation function layer to obtain the initial responsibility weight signal of the current processing layer, until the initial responsibility weight signal of each processing layer is obtained.
[0021] Optionally, the step of inputting the current responsibility weight signal of each processing layer into the corresponding learning layer in the core network to obtain the predicted action specifically includes:
[0022] The output of the first learning layer of the core network is obtained based on the state sample information;
[0023] The output of the second learning layer of the core network is obtained based on the output of the first learning layer and the current responsibility weight signal of the first learning layer.
[0024] For each learning layer after the second layer of the core network, the prediction action is based on the current learning layer's responsibility weight signal and the output of the current learning layer obtained from the output of the previous learning layers, until the output of the last learning layer; the output of the previous learning layers includes the outputs of all learning layers before the current learning layer.
[0025] Optionally, updating the network parameters of the core network based on the current layer responsibility weight signal, the prediction action, and the state sample information specifically includes:
[0026] The module responsibility weight signal of each sub-module in the current layer is obtained based on the current layer responsibility weight signal; the current layer responsibility weight signal includes the module responsibility weight signals of several sub-modules;
[0027] The new state and reward are obtained based on the predicted action and the state sample information;
[0028] When the module responsibility weight signal of any submodule is greater than a preset value, the network parameters of the submodule are updated according to the predicted action, the state sample information, the new state, and the reward; the network parameters include policy parameters and parameters of the state action value function.
[0029] Optionally, updating the network parameters of any submodule based on the predicted action, the state sample information, the new state, and the reward specifically includes:
[0030] Update the temperature parameters based on the predicted action and the state sample information;
[0031] Calculate the action state value function based on the predicted action and the state sample information, and update the parameters of the state action value function based on the action state value function, the new state, and the reward.
[0032] The policy parameters are updated based on the predicted action, the state sample information, and the action state value function.
[0033] To address the aforementioned problems, in a first aspect, the present invention provides a control method for a robotic arm, the control method comprising:
[0034] Acquire task information and status information; the status information includes the status information of the robotic arm and the status information of the target object, and the task information includes multiple actions to be performed;
[0035] The task information and state information are input into the control model to obtain the actual execution action; the control model is trained by a control model training method.
[0036] To address the aforementioned problems, in a first aspect, the present invention provides an electronic device, characterized in that the electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement a training method for a control model and / or a control method for a robotic arm.
[0037] To address the aforementioned problems, in a first aspect, the present invention provides a computer-readable storage medium, characterized in that it stores a processor-executable program, which, when executed by a processor, is used to perform a training method for a control model and / or a control method for a robotic arm.
[0038] Implementing this invention offers the following advantages: This invention acquires task sample information and state sample information; the state sample information includes the state sample information of the robotic arm and the target object; the task sample information includes the actions to be performed. The task sample information and the state sample information are input into an adjudication network to obtain the current-layer responsibility weight signal for each processing layer. The adjudication network includes several processing layers. Based on the current-layer responsibility weight signal, the predicted action, and the state sample information, the network parameters of the core network are updated. Even as the complexity of the task increases, the core network can update the parameters of task-related sub-modules based on the responsibility weight signal, without requiring all modules to participate in the update. The responsibility weight signal of each processing layer is input into the corresponding learning layer in the core network to obtain the predicted action. The core network includes several learning layers, and the learning layers of the core network correspond one-to-one with the processing layers of the adjudication network. The subsequent learning layers of the core network are updated according to the preceding learning layers. The convergence of the core network is determined based on the predicted action and the executed action. If the core network has not converged, the task sample information and state sample information are repeatedly obtained until the core network converges. This allows the core network to select relevant sub-modules for learning based on the responsibility weight signal calculated by the adjudication network, and all relevant sub-modules learn the tasks related to them, thus accelerating the learning and decision-making processes. Attached Figure Description
[0039] Figure 1 This is a flowchart of a training method for a control model provided by the present invention;
[0040] Figure 2 This is a schematic diagram of the structure of a control model according to a training method for a control model provided by the present invention;
[0041] Figure 3 This is a schematic diagram of the adjudication network structure of a training method for a control model provided by the present invention;
[0042] Figure 4This is a schematic diagram of the core network and memory module structure of a control model training method provided by the present invention;
[0043] Figure 5 This is a schematic diagram of the virtual layer structure of a training method for a control model provided by the present invention;
[0044] Figure 6 This is a schematic diagram of the structure of an electronic device provided by the present invention. Detailed Implementation
[0045] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are only for ease of explanation and do not limit the order of the steps. The execution order of each step in the embodiments can be adapted according to the understanding of those skilled in the art.
[0046] To address the above problems, in some embodiments, such as Figure 1 and Figure 2 As shown, Figure 1 This is a flowchart of a method for training a control model. Figure 2 This invention provides a method for training a control model, comprising the following steps: (The diagram shows the structure of a control model.)
[0047] S1100, Obtain task sample information and status sample information.
[0048] The state sample information includes the state sample information of the robotic arm and the state sample information of the target object, and the task sample information includes the actions to be performed.
[0049] The tasks of the robotic arm may include, but are not limited to, opening and closing doors and windows, grasping objects and pins. The state sample information of the robotic arm may include, but is not limited to, the pose state and coordinate parameters of the robotic arm. The state sample information of the target object may include, but is not limited to, the position information, size information and quantity information of the target object.
[0050] Specifically, the action to be performed can be one of the following: opening and closing a door, opening and closing a window, grabbing an object, or using a latch. Each training session acquires task sample information and state sample information related to that task.
[0051] S1200: Input the task sample information and the state sample information into the adjudication network to obtain the current responsibility weight signal of each processing layer.
[0052] The adjudication network comprises several processing layers.
[0053] The adjudication network includes at least one processing layer, and each processing layer includes several sub-modules.
[0054] The layer responsibility weight signal can be a vector that includes the responsibility weight signals of all sub-modules and tasks in each layer.
[0055] Specifically, during training, each submodule calculates the correlation between the submodule and the task, i.e., the responsibility weight signal. When any processing layer outputs the responsibility weight signal of each submodule and the task to the next processing layer, it outputs it in the form of a vector and uses it as an input to the next processing layer, i.e., the responsibility weight signal of that layer.
[0056] S1300: Input the current responsibility weight signal of each processing layer into the corresponding learning layer in the core network to obtain the predicted action.
[0057] The core network includes several learning layers, and each learning layer of the core network corresponds one-to-one with the processing layer of the adjudication network. The subsequent learning layers of the core network are updated based on the preceding learning layers.
[0058] Each learning layer of the core network comprises several sub-modules, and each learning layer corresponds one-to-one with each processing layer of the aforementioned adjudication network, meaning the depth of the adjudication network corresponds to the depth of the core network.
[0059] Specifically, the responsibility weight signal of each processing layer of the adjudication network is input to the corresponding learning layer of the core network. For example, the responsibility weight signal of the first processing layer of the adjudication network is input to the first learning layer of the core network; the responsibility weight signal of the second processing layer of the adjudication network is input to the second learning layer of the core network.
[0060] Similarly, the responsibility weight signal obtained from each processing layer of the core network can be input into the corresponding learning layer of the core network, all the way to the last layer of the core network, where the last layer outputs the predicted action for the task.
[0061] The inputs to all learning layers after the second learning layer of the core network, in addition to the current responsibility weight signal of the second processing layer of the adjudication network, also include relevant information from other learning layers. For example, the input to the third learning layer also includes relevant information from the first learning layer, and the input to the fourth learning layer, in addition to the current responsibility weight signal of the third processing layer of the adjudication network, also includes relevant information from the first and second learning layers. This relevant information includes, but is not limited to, relevant sub-modules in previous learning layers, responsibility weight signals of relevant sub-modules, and parameters of relevant sub-modules.
[0062] S1400: Update the network parameters of the core network according to the current layer responsibility weight signal, the prediction action, and the state sample information.
[0063] Specifically, when the responsibility weight signal of any submodule in the current layer responsibility weight signal is greater than the preset value, it is considered that the submodule is highly relevant to the task being trained. The parameters of the submodule are updated according to the execution action and state sample information of the task. In this way, only the submodules that are highly relevant to the task being trained are updated, instead of the parameters of all submodules, which improves training efficiency and allows the most relevant submodule to be selected to learn the strategy of the task.
[0064] S1500. Determine whether the core network has converged based on the predicted action and the executed action. If the core network has not converged, repeatedly acquire task sample information and state sample information until the core network converges.
[0065] Specifically, the core network receives immediate rewards by predicting actions for the task. When immediate return When the rate of change tends to stabilize, the core network is considered to have converged for the task, and training begins for the next task. (Immediate feedback is required.) If the rate of change is greater than the preset value, the core network has not converged. Continue to repeat steps S1100-S1500 until the core network converges for the task.
[0066] To address the above problems, in some embodiments, such as Figure 3 As shown, Figure 3 This is a schematic diagram of the adjudication network structure for a training method of a control model. The step of inputting the task sample information and the state sample information into the adjudication network to obtain the current responsibility weight signal for each processing layer specifically includes:
[0067] S1210. The task sample information is processed through a fully connected layer to obtain a first vector.
[0068] Specifically, the task sample information A vector representation of dimension d is also obtained after passing through a fully connected layer. It is used as the input to the first processing layer of the adjudication network.
[0069] S1220. The state sample information is processed through two fully connected layers to obtain a second vector.
[0070] Specifically, the state sample information of the training task A vector of dimension d is obtained after passing through two fully connected layers. It is used as the input to the first processing layer of the adjudication network.
[0071] S1230. Obtain the current responsibility weight signal for each processing layer based on the first vector and the second vector.
[0072] Specifically, the first vector and the second vector are multiplied together, and their product is used as the first input to each processing layer of the adjudication network. The output of the processing layer above each processing layer is used as the second input to each processing layer of the adjudication network, thus obtaining the current responsibility weight signal of each processing layer.
[0073] To address the aforementioned issues, in some embodiments, obtaining the current-layer responsibility weight signal for each processing layer based on the first vector and the second vector specifically includes:
[0074] S1231. Multiply the first vector and the second vector to obtain the first input of the adjudication network, and obtain the initial responsibility weight signal of each processing layer based on the first input.
[0075] Specifically, the first vector and the second vector are multiplied, and their product is used as the first input to each processing layer of the adjudication network. This yields the initial responsibility weight signal for each processing layer.
[0076] To address the aforementioned issues, in some embodiments, step S1231, which involves obtaining the initial responsibility weight signal for each processing layer based on the first input, specifically includes:
[0077] S1231a. The first input layer of activation function layer and the first fully connected layer are used to obtain the initial responsibility weight signal of the first processing layer.
[0078] The first processing layer of the adjudication network includes, but is not limited to, an activation function ReLU layer and a fully connected layer.
[0079] Specifically, in this embodiment, since the first processing layer does not have the initial responsibility weight signal of the previous processing layer, the output of the first processing layer is to convert the aforementioned first input... The output of the first processing layer is obtained after passing through a ReLU activation function layer and a fully connected layer. This is the initial responsibility weight signal of the first processing layer, as shown in Equation (1):
[0080] (1)
[0081] in, This represents the parameters of a fully connected layer.
[0082] S1231b: Repeat the process of passing the first input and the initial responsibility weight signal of the previous processing layer through two fully connected layers and an activation function layer to obtain the initial responsibility weight signal of the current processing layer, until the initial responsibility weight signal of each processing layer is obtained.
[0083] Specifically, the processing layers following the first processing layer include, but are not limited to, a ReLU activation function layer and two fully connected layers. In this embodiment, the initial responsibility weight signal of the first processing layer and the first input are... As the input to the second layer, it passes through a fully connected layer and a ReLU activation function layer, and then through another fully connected layer to obtain the output of the second processing layer, which is the initial responsibility weight signal of the second processing layer, as shown in Equation (1):
[0084] (2)
[0085] in, These are the parameters of the first fully connected layer in the second processing layer. These are the parameters for the second fully connected layer of the second processing layer.
[0086] The input dimension of the first fully connected layer is The output dimension is d, and its purpose is to transform the weight responsibility vector output from the previous layer. The dimension is transformed into the same dimension d as the state sample information and task sample information, then passed through a ReLU function layer, and finally through a fully connected layer. The second fully connected layer has an input dimension of d and an output dimension of d. .
[0087] Similarly, the initial responsibility weight signals of the remaining processing layers are similar to those of the second processing layer. The initial responsibility weight signals of the first processing layer can be replaced with the initial responsibility weight signals of the previous layer, as shown in equation (3).
[0088] (3)
[0089] in, For any processing layer in an arbitrary decision network, For the first The layer preceding the layer of the layer processing, For the first The parameters of the first fully connected layer in the processing layer are... For the first The parameters of the second fully connected layer of the layer processing layer.
[0090] S1232. Normalize the initial responsibility weight signal of each processing layer to obtain the current responsibility weight signal of each processing layer.
[0091] Specifically, from the perspective of dimensions The initial responsibility weight signal of each processing layer is normalized by formula (4) for the elements in the initial responsibility weight signal vector, that is, the responsibility weight signal corresponding to the sub-module of the core network in each processing layer, to obtain the responsibility weight signal of the current layer.
[0092] (4)
[0093] in, Represents any processing layer in the adjudication network. Indicates the first in the core network The i-th submodule of the learning layer directs to the i-th... The responsibility weight signal input to the j-th submodule of the +1 learning layer. It is the responsibility weight signal of the i-th submodule in the initial responsibility weight signal of the l-th processing layer of the adjudication network.
[0094] To address one of the aforementioned problems, in some embodiments, such as Figure 4 As shown, Figure 4 This is a schematic diagram of the core network and memory module structure of a training method for a control model. Step S1300, which involves inputting the current responsibility weight signal of each processing layer into the learning layer of the core network corresponding to the current responsibility weight signal to obtain the predicted action, specifically includes:
[0095] S1310. Obtain the output of the first learning layer of the core network based on the state sample information.
[0096] Specifically, the input vector of the first learning layer of the core network is The input vector consists of state sample information. After passing through two fully connected layers, we obtain a vector of dimension d, represented as... Memory module saves , where vector Including the input of N sub-modules in the first learning layer .
[0097] Based on the input of the first learning layer The output of the first learning layer is obtained through equation (5).
[0098] (5)
[0099] in, These are the fully connected layer parameters for the j-th submodule of the first learning layer. This represents the responsibility weight signal for the i-th submodule of the first learning layer. From the dimension The responsibility weight signal is obtained from the current layer responsibility weight signal, such as the responsibility weight signal of the i-th sub-module in the first learning layer, which is the above. .
[0100] get Then, the memory module stores... .
[0101] S1320. The output of the second learning layer of the core network is obtained based on the output of the first learning layer and the current responsibility weight signal of the first learning layer.
[0102] Specifically, the output of the first learning layer and the current responsibility weight signal of the first learning layer are combined through a virtual layer between the first and second learning layers to obtain the input vector of the second learning layer. ,in Includes N elements That is, the input of the N sub-modules of the second learning layer.
[0103] The synthesis process of the virtual layer between the first learning layer and the second learning layer is shown in equation (6).
[0104] (6)
[0105] in, This represents the output vector of the i-th module in the virtual layer, with dimension d, and is also an input vector of the corresponding module in the virtual layer. The superscript AR indicates a virtual layer reconstructed from the output of the first learning layer of the core network, the current weight signal of the first learning layer of the core network, and the memory module. When calculating the first learning layer of the core network, the output of the memory module is zero, therefore J=0. Only the current weight signal of the second processing layer of the core network, such as Figure 4 As shown.
[0106] We obtain an input to the second learning layer. Then, the output of the second learning layer is calculated according to equation (7).
[0107] (7)
[0108] in, For the fully connected layer parameters of the j-th submodule of the second learning layer, This represents the responsibility weight signal for the i-th submodule of the second learning layer. From the dimension The responsibility weight signal is obtained from the current layer responsibility weight signal, such as the responsibility weight signal of the i-th sub-module in the second learning layer, which is the above. .
[0109] After receiving the output of the second learning layer, the memory module saves it. Based on this, storage .
[0110] S1330, Each learning layer after the second layer of the core network outputs the current learning layer based on the current learning layer's responsibility weight signal and the outputs of all previous layers, until the last learning layer outputs the predicted action; the outputs of all previous layers include the outputs of all learning layers before the current learning layer.
[0111] Specifically, the input vector of the third learning layer and each subsequent learning layer is: , vector is Includes N elements That is, the input of N sub-modules in each learning layer.
[0112] The process of integrating the virtual layer between the current learning layer and the previous learning layer is shown in Equation (8).
[0113] (8)
[0114] in, Represents any layer of the core network. superscript This represents the virtual layer reorganized from the core network and memory modules, such as... Figure 5 As shown, Figure 5 This is a schematic diagram of the virtual layer structure in a training method for a control model. The memory module stores... The previous layer included the first The learning layer contains significant relevant information regarding its responsibility weight signals. This information includes task-related sub-modules, their responsibility weight signals, and so on, and is related to the core network's... The outputs of the learning layers are used together as the first layer. The input to the layer. That is, the memory module stores module information that meets the conditions for learning layers less than or equal to the (l-2)th layer. Let d represent the input vector of the i-th module in the virtual layer, with dimension d.
[0115] At this time, the core network is... The input vector for the learning layer is provided by N+J sub-modules of the recombined virtual layer, where J is the number of sub-modules stored in the memory module that are less than the first sub-module. -2 Number of sub-modules in the learning layer. This represents the fully connected layer parameters of the i-th module in the virtual layer. This indicates that the i-th submodule of the virtual layer is connected to the i-th submodule of the core network. The "responsibility weight signal" output by the j-th submodule of the learning layer. The higher the value, the greater the influence of the submodule on the decision-making process for that task. Also from the core network -1 layer learning layer's current responsibility weight signal and memory module output The responsibility weight signal is obtained by recombining the responsibility weight signals of all learning layers before the -1 learning layer, where the core network's first layer is... The current responsibility weight signal for the -1 learning layer is: The responsibility weight signal for all learning layers corresponding to the memory module storage module is: .
[0116] The above-mentioned number Input of the learning layer and the The input of each submodule in the learning layer Then, calculate the first according to formula (9). The output of the learning layer.
[0117] (9)
[0118] in, For the first The fully connected layer parameters of the j-th submodule of the learning layer are obtained. For the first The responsibility weight signal of the i-th submodule in the learning layer. From the dimension The weight of responsibility at the current level is obtained from the signal, such as the first level. The responsibility weight signal of the i-th submodule of the learning layer, as described above. .
[0119] After obtaining the number After learning the output of each learning layer, the memory module stores the output of all previous learning layers. This continues until the last learning layer of the core network outputs the predicted action.
[0120] To address one of the aforementioned problems, in some embodiments, step S1400 of the present invention, which involves updating the network parameters of the core network based on the current layer responsibility weight signal, the prediction action, and the state sample information, specifically includes:
[0121] S1410. Obtain the module responsibility weight signal of each sub-module in the current layer according to the current layer responsibility weight signal.
[0122] The current layer responsibility weight signal includes the module responsibility weight signals of several sub-modules.
[0123] Specifically, when the layer responsibility weight signal is a dimension... From the vector, we obtain the module responsibility weight signal for each sub-module in the current layer, such as the responsibility weight signal of the i-th sub-module in the first learning layer. That is, the above .
[0124] S1420. Obtain a new state and reward based on the predicted action and the state sample information.
[0125] Specifically, the core network outputs predicted actions. Obtain new state sample information and instant returns .
[0126] S1430. When the module responsibility weight signal of any submodule is greater than a preset value, the network parameters of the submodule are updated according to the predicted action, the state sample information, the new state, and the reward; the network parameters include policy parameters and parameters of the state action value function.
[0127] Specifically, when the responsibility weight of the i-th module in the l-th layer during the training task is greater than or equal to a preset value, it indicates that the sub-module is strongly correlated with the task, and the network parameters of the sub-module are updated; otherwise, they are not updated. After the update, the memory module stores the sub-module information; otherwise, it does not store it. The sub-module information includes, but is not limited to, the position of the core network of the sub-module and the responsibility weight signal of the sub-module.
[0128] If the preset value is not limited to 80%, then the update condition is: ,
[0129] To address one of the aforementioned problems, in some embodiments, step S1430, which involves updating the network parameters of any submodule based on the predicted action, the state sample information, the new state, and the reward, specifically includes:
[0130] S1431. Update the temperature parameters based on the predicted action and the state sample information.
[0131] Specifically, the core network uses SAC to train the policy. The SAC algorithm is a maximum entropy deep reinforcement learning algorithm based on stochastic policies, trained using the Actor-Critic framework, with temperature parameters... It is an automatic update, not subject to the above update conditions, and updates according to formula (10).
[0132] (10)
[0133] in, This represents the minimum expected entropy required. Indicates a predicted action. Indicates state sample information, Representation strategy.
[0134] S1432. Calculate the action state value function based on the predicted action and the state sample information, and update the parameters of the state action value function based on the action state value function, the new state, and the reward.
[0135] Specifically, as shown in equation (11), the action state value function is updated.
[0136] (11)
[0137] in, The parameters of the action state value function that have not been updated in the Critic value network. The target network parameters are the action state value function.
[0138] The process of obtaining it is as follows: calculate the action state value function. The parameters of the updated state action value function are determined based on the action state value function, the new state, and the reward. Specifically, the network parameters of the Critic value function for a single task are calculated according to formulas (12) and (13).
[0139] (12)
[0140] (13)
[0141] in, Indicates a predicted action. Indicates state sample information, The strategy is defined as follows: D represents the experience replay pool, and a portion of the data is extracted from the experience replay pool D for network updates.
[0142] The network parameters of the Critic value function for each individual task are calculated to obtain the network parameters of the Critic value function for the task set of multiple tasks, as shown in Equation (14).
[0143] (14)
[0144] Here, T represents a single task, and the network parameters for the Critic value function in this paper's algorithm are based on the task distribution. Above, a single task T is distributed... Mid-sampling, parameter optimization maximizes from The objective is the average expected return of all sampled tasks.
[0145] S1433. Update the policy parameters based on the predicted action, the state sample information, and the action state value function.
[0146] Specifically, the strategy for a single task is as follows: The objective of strategy optimization is shown in equation (15).
[0147] (15)
[0148] Then update the strategy parameters according to equation (16).
[0149] (16)
[0150] Here, T represents a single task, and the parameter updates of the Actor policy network in this paper are based on the task distribution. Above, task T is distributed Mid-sampling, optimizing policy parameters to maximize from The objective is the average expected return of all sampled tasks.
[0151] To address one of the aforementioned problems, in some embodiments, the present invention provides a control method for a robotic arm, the control method comprising:
[0152] S2100, Obtain task information and status information; the status information includes the status information of the robotic arm and the status information of the target object, and the task information includes multiple actions to be performed;
[0153] S2200, Input the task information and status information into the control model to obtain the actual execution action; the control model is trained by the aforementioned control model training method.
[0154] Specifically, the robotic arm inputs the task information and state information into the control model, which has been trained. The control model outputs an action and causes the robotic arm to execute that action.
[0155] To address the above problems, in some embodiments, such as Figure 6 As shown, Figure 6 This is a schematic diagram of the structure of an electronic device provided by the present invention. The present invention also provides an electronic device, which includes a processor 10 and a memory 20. The memory 20 stores a computer program, and the processor 10 executes the computer program to implement any of the methods described in the above method embodiments.
[0156] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. The memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include remote memory located remotely relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0157] Furthermore, this application also discloses a computer program product or computer program stored in a computer-readable storage medium. A processor of a computer device can read the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the computer device to perform the described method. Similarly, the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0158] The present invention also provides a computer-readable storage medium storing a processor-executable program, which, when executed by a processor, is used to perform any of the methods described in the above method embodiments.
[0159] It is understood that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0160] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
Claims
1. A method for training a control model, characterized in that, Includes the following steps: Acquire task sample information and state sample information; the state sample information includes the state sample information of the robotic arm and the state sample information of the target object, and the task sample information includes the actions to be performed; The task sample information and the state sample information are input into the adjudication network to obtain the current responsibility weight signal of each processing layer; the adjudication network includes several processing layers; The responsibility weight signal of each processing layer is input into the corresponding learning layer in the core network to obtain the predicted action; the core network includes several learning layers, and the learning layers of the core network correspond one-to-one with the processing layers of the adjudication network. The subsequent learning layers of the core network are updated according to the preceding learning layers. The network parameters of the core network are updated based on the current layer responsibility weight signal, the predicted action, and the state sample information. Based on the predicted action and the executed action, determine whether the core network has converged. If the core network has not converged, repeatedly acquire task sample information and state sample information until the core network converges. The step of inputting the task sample information and the state sample information into the adjudication network to obtain the current responsibility weight signal for each processing layer specifically includes: The task sample information is processed through a fully connected layer to obtain a first vector; The state sample information is then processed through two fully connected layers to obtain a second vector. The current responsibility weight signal of each processing layer is obtained based on the first vector and the second vector; The step of obtaining the current responsibility weight signal for each processing layer based on the first vector and the second vector specifically includes: Multiply the first vector and the second vector to obtain the first input of the adjudication network, and obtain the initial responsibility weight signal of each processing layer based on the first input; The initial responsibility weight signal of each processing layer is normalized to obtain the current responsibility weight signal of each processing layer. The step of obtaining the initial responsibility weight signal for each processing layer based on the first input specifically includes: The first input is passed through an activation function layer and a fully connected layer to obtain the initial responsibility weight signal of the first processing layer; For each processing layer after the first layer of the adjudication network, the first input and the initial responsibility weight signal of the previous processing layer are passed through two fully connected layers and an activation function layer to obtain the initial responsibility weight signal of the current processing layer, until the initial responsibility weight signal of each processing layer is obtained. The step of inputting the responsibility weight signal of each processing layer into the corresponding learning layer of the core network to obtain the predicted action specifically includes: The output of the first learning layer of the core network is obtained based on the state sample information; The output of the second learning layer of the core network is obtained based on the output of the first learning layer and the current responsibility weight signal of the first learning layer. For each learning layer after the second layer of the core network, the prediction action is based on the current learning layer's responsibility weight signal and the output of the current learning layer obtained from the output of the previous learning layers, until the output of the last learning layer; the output of the previous learning layers includes the outputs of all learning layers before the current learning layer.
2. The method according to claim 1, characterized in that, The step of updating the network parameters of the core network based on the current layer responsibility weight signal, the predicted action, and the state sample information specifically includes: The module responsibility weight signal of each sub-module in the current layer is obtained based on the current layer responsibility weight signal; the current layer responsibility weight signal includes the module responsibility weight signals of several sub-modules; The new state and reward are obtained based on the predicted action and the state sample information; When the module responsibility weight signal of any submodule is greater than a preset value, the network parameters of the submodule are updated according to the predicted action, the state sample information, the new state, and the reward; the network parameters include policy parameters and parameters of the state action value function.
3. The method according to claim 2, characterized in that, The step of updating the network parameters of any submodule based on the predicted action, the state sample information, the new state, and the reward specifically includes: Update the temperature parameters based on the predicted action and the state sample information; Calculate the action state value function based on the predicted action and the state sample information, and update the parameters of the state action value function based on the action state value function, the new state, and the reward. The policy parameters are updated based on the predicted action, the state sample information, and the action state value function.
4. A control method for a robotic arm, characterized in that, The control method includes: Acquire task information and status information; the status information includes the status information of the robotic arm and the status information of the target object, and the task information includes multiple actions to be performed; The task information and status information are input into the control model to obtain the actual execution action; the control model is trained by the training method described in any one of claims 1-3.
5. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1-4.
6. A computer-readable storage medium, characterized in that, It contains a processor-executable program, which, when executed by a processor, is used to perform the method as described in any one of claims 1-4.