Shaft hole assembly method and related apparatus
The shaft and hole assembly model trained by meta-reinforcement learning and prior knowledge solves the problem of universality of existing shaft and hole assembly methods when the size or number changes, and realizes the efficient completion of multiple types of shaft and hole assembly tasks.
Patent Information
- Application Number
- CN202310518497.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-09
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2043-05-09
AI Technical Summary
Existing shaft and hole assembly methods lack versatility and practicality when shaft and hole size or number vary, making it difficult to adapt to various types of shaft and hole assembly tasks.
The shaft and hole assembly model is trained using meta-reinforcement learning and prior knowledge. Through multiple iterations and optimization of the operation evaluation network and action selection network, assembly actions suitable for different types of shaft and hole assembly tasks are obtained.
This improves the versatility and practicality of shaft and hole assembly, enabling the assembly model to be applied to various types of shaft and hole assembly tasks, thereby increasing assembly efficiency and success rate.
Smart Images

Figure CN116766178B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automated assembly technology, and in particular to a shaft hole assembly method and related equipment. Background Technology
[0002] In the industrial sector, there is an increasing demand for robots to automatically complete high-precision assembly tasks. Shaft and hole assembly is the core operation of most assembly tasks, and common types include single-axis hole assembly, dual-axis hole assembly, and tri-axis hole assembly.
[0003] Existing shaft and hole assembly methods are often only applicable to the same type of assembly task. The shaft and hole size and number of shaft and holes are the same in the same type of assembly task. When the shaft and hole size or number of shaft and holes change, the shaft and hole assembly method is not applicable, resulting in poor universality and practicality of the shaft and hole assembly method.
[0004] Therefore, it is very important to provide a shaft and hole assembly method that enables robots to complete multiple types of shaft and hole assembly tasks. Summary of the Invention
[0005] This invention provides a shaft hole assembly method and related equipment to solve the shortcomings of the existing shaft hole assembly in terms of its lack of versatility and practicality, and to improve the versatility and practicality of shaft hole assembly.
[0006] This invention provides a shaft hole assembly method, comprising:
[0007] Perform multiple iterations; each iteration includes:
[0008] Obtain the current state of the shaft; the current state includes the current force information;
[0009] The current state is input into the shaft-hole assembly model to obtain the current assembly action; the current assembly action includes translation and rotation; the shaft-hole assembly model is trained on multiple different types of shaft-hole assembly tasks based on prior knowledge and meta-reinforcement learning.
[0010] Perform the current assembly action on the axis;
[0011] If the depth of the shaft insertion hole does not meet the preset assembly requirements, the next iteration operation is performed.
[0012] In some embodiments, the shaft-hole assembly model is obtained based on the following steps:
[0013] Construct multiple shaft and hole assembly tasks of different categories;
[0014] Perform multiple iterations of training operations; each iteration of training operations includes:
[0015] One shaft and hole assembly task is randomly selected from the multiple different categories of shaft and hole assembly tasks and used as the current training task.
[0016] Execute the current training task and update the parameters of the shaft-hole assembly model;
[0017] Upon completion of the current training task, obtain the cumulative reward corresponding to the current training task;
[0018] Based on the cumulative reward corresponding to the current training task and the cumulative reward corresponding to the previous training task, determine the convergence status of the cumulative reward corresponding to the training task.
[0019] If the cumulative reward corresponding to the training task fails to converge, the next iteration of training will be executed.
[0020] In some embodiments, performing the current training task and updating the parameters of the shaft-hole assembly model includes:
[0021] The current training task is executed through multiple iterations, and the parameters of the shaft-hole assembly model are updated in each iteration; wherein each iteration includes:
[0022] Obtain the similarity between the previous assembly training action and the previous demonstration action, and the current training state; the previous assembly training action is the training action output by the network selection network in the shaft-hole assembly model based on the previous training state in the previous iteration training round; the previous demonstration action is the demonstration action output by the assembly demonstration model based on the previous training state; the current training state includes the current force information of the training shaft.
[0023] The similarity and the current training state are input into the action evaluation network in the shaft hole assembly model to obtain the current first loss function value;
[0024] Update the parameters of the action evaluation network based on the current first loss function value;
[0025] The similarity and the current training state are input into the action selection network in the shaft hole assembly model to obtain the current second loss function value;
[0026] The parameters of the action selection network are updated based on the current second loss function value;
[0027] If the current training task is not completed, the next iteration training round will be executed.
[0028] In some embodiments, inputting the similarity and the current training state into the action evaluation network of the shaft-hole assembly model to obtain the current first loss function value includes:
[0029] The similarity and the current training state are input into the action evaluation network to obtain the current state value and the current reward value.
[0030] Based on the current state value, the current reward value, and the value of the next state, obtain the current advantage value;
[0031] Based on the current advantage value, obtain the current first loss function value.
[0032] In some embodiments, inputting the similarity and the current training state into the action selection network of the shaft-hole assembly model to obtain the current second loss function value includes:
[0033] The similarity and the current training state are input into the action selection network to obtain the current assembly training action;
[0034] Based on the probability of the current assembly training action and the probability of the previous assembly training action, as well as the current advantage value, the current second loss function value is obtained.
[0035] In some embodiments, before obtaining the similarity between the previous assembly training action and the previous demonstration action, the method further includes:
[0036] Acquire teaching status and teaching actions;
[0037] Model the teaching state and the teaching action to obtain the assembly demonstration model.
[0038] The present invention also provides a shaft hole assembly device, comprising:
[0039] The execution module is used to perform multiple iterations; each iteration includes:
[0040] Obtain the current state of the shaft; the current state includes the current force information;
[0041] The current state is input into the shaft-hole assembly model to obtain the current assembly action; the current assembly action includes translation and rotation; the shaft-hole assembly model is trained on multiple different types of shaft-hole assembly tasks based on prior knowledge and meta-reinforcement learning.
[0042] Perform the current assembly action on the axis;
[0043] If the depth of the shaft insertion hole does not meet the preset assembly requirements, the next iteration operation is performed.
[0044] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the shaft hole assembly method as described above.
[0045] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the shaft hole assembly method as described above.
[0046] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the shaft hole assembly method as described above.
[0047] The present invention provides a shaft and hole assembly method and related equipment, which learns multiple different types of shaft and hole assembly tasks through meta-reinforcement learning. By using prior knowledge, the learning efficiency of meta-reinforcement learning is improved, and the resulting shaft and hole assembly model can be applied to different types of shaft and hole assembly tasks, thereby improving the versatility and practicality of shaft and hole assembly. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0049] Figure 1 This is one of the schematic flowcharts of a shaft hole assembly method provided in an exemplary embodiment of the present invention;
[0050] Figure 2 This is a second schematic flowchart of a shaft hole assembly method provided in an exemplary embodiment of the present invention;
[0051] Figure 3 This is a third schematic flowchart of a shaft hole assembly method provided in an exemplary embodiment of the present invention;
[0052] Figure 4 This is a fourth schematic flowchart of a shaft hole assembly method provided in an exemplary embodiment of the present invention;
[0053] Figure 5 This is the fifth flowchart illustrating the shaft hole assembly method provided in an exemplary embodiment of the present invention;
[0054] Figure 6 This is a sixth schematic flowchart of a shaft hole assembly method provided in an exemplary embodiment of the present invention;
[0055] Figure 7This is a schematic diagram of the shaft hole assembly device provided in an exemplary embodiment of the present invention;
[0056] Figure 8 This is a schematic diagram of the physical structure of an electronic device provided in an exemplary embodiment of the present invention. Detailed Implementation
[0057] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. It should be noted that, unless otherwise specified, the embodiments and features of the embodiments of this invention can be combined with each other. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.
[0059] It should be further noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0060] In this invention, "at least one" means one or more, and "more than one" means two or more. The terms "first," "second," "third," "fourth," etc. (if present) in this invention are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0061] In embodiments of the present invention, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" or "for example" in embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0062] Please refer to Figure 1 , Figure 1This is one of the flowcharts illustrating a shaft hole assembly method provided in an exemplary embodiment of the present invention. The embodiment of the present invention provides a shaft hole assembly method, the executing entity of which can be an intelligent agent, such as a robot. The shaft hole assembly method includes the following steps:
[0063] Step 110: Perform multiple iterations.
[0064] Specifically, the agent can obtain one assembly action by performing one iteration operation. Since one assembly action is often insufficient to successfully assemble the shaft and hole, the agent needs to perform multiple iteration operations to obtain multiple assembly actions, and achieve successful shaft and hole assembly through multiple assembly actions. Each iteration operation may include steps 1101 to 1104.
[0065] Step 1101: Obtain the current state of the axis.
[0066] In the current iteration, the agent obtains the current state of the shaft through a force sensor mounted on the shaft. The current state includes the current force information.
[0067] The expression for state s is as follows:
[0068] s = [F x ,F y ,F z M x M y M z ]
[0069] In the formula, s represents the state, and F x F y and F z M represents the force measured by the force sensor along the x, y, and z directions, respectively. x M y and M z These represent the torques along the x, y, and z directions, respectively. The x-direction represents the forward / backward movement of the axis, the y-direction represents the left / right movement of the axis, and the z-direction represents the up / down movement of the axis.
[0070] In some embodiments, the number of shafts and holes is the same, with one shaft corresponding to one hole. The number of shafts can be single, double, or triple, etc., and the present invention does not impose a specific limitation on the number of shafts.
[0071] Step 1102: Input the current state into the shaft hole assembly model to obtain the current assembly action.
[0072] Specifically, in related technologies, the Actor-Critic algorithm is used to construct the shaft-hole assembly model. The Actor-Critic algorithm is a standard Deep Reinforcement Learning (DRL) algorithm.
[0073] Standard DRLs are designed for specific Markov Decision Processes (MDPs) and use certain learning algorithms to find an optimal policy. This optimal policy guides the agent to make the best decisions in a specific task (what action to take in what state). The learning algorithms in standard DRLs rely heavily on the interaction between the agent and the environment, resulting in high training costs. Once the environment changes, the previously learned optimal policy becomes inapplicable, requiring retraining for the new environment, leading to low learning efficiency.
[0074] Meta-learning is introduced into the standard DRL to form meta-reinforcement learning. Meta-reinforcement learning is a learning algorithm that can quickly adapt to different new tasks and obtain the corresponding optimal policy.
[0075] The shaft-hole assembly model in this invention is trained on multiple different types of shaft-hole assembly tasks based on prior knowledge and meta-reinforcement learning.
[0076] The agent applies meta-reinforcement learning to learn multiple different categories of shaft and hole assembly tasks. Prior knowledge is used to improve the learning efficiency of meta-reinforcement learning, resulting in a well-trained shaft and hole assembly model. Therefore, the trained shaft and hole assembly model can be applied to different categories of shaft and hole assembly tasks. After obtaining the trained shaft and hole assembly model, the agent inputs the acquired current state into the model, and the model outputs the current assembly action. The current assembly action includes translation and rotation.
[0077] The expression for assembly action 'a' is as follows:
[0078] a=[Δ x ,Δ y ,Δ z ,α x ,α y ,α z ]
[0079] In the formula, 'a' represents the assembly action, and 'Δ' represents the assembly action. x Δ y and Δ z Let α represent the translation along the x, y, and z directions, respectively. x α y and α z These represent the rotation amounts along the x, y, and z directions, respectively.
[0080] In some embodiments, the agent inputs the current state into the shaft-hole assembly model, the shaft-hole assembly model outputs the optimal action strategy, and the action with the highest probability in the optimal action strategy is selected for inverse normalization to obtain the current assembly action.
[0081] Step 1103: Perform the current assembly action on the axis.
[0082] Specifically, after receiving the current assembly action, the intelligent agent executes the current assembly action on the axis, that is, it translates and rotates the axis according to the translation and rotation amounts in the current assembly action.
[0083] Step 1104: Check whether the depth of the shaft insertion hole meets the preset assembly requirements.
[0084] Specifically, in some embodiments, the preset assembly requirements may include a preset minimum insertion depth and a preset maximum insertion depth, etc.
[0085] After the current assembly action is completed on the shaft, the agent detects whether the depth of the shaft insertion hole meets the preset assembly requirements, that is, whether the depth of the shaft insertion hole is between the preset minimum insertion depth and the preset maximum insertion depth.
[0086] If the depth of the shaft insertion hole does not meet the preset assembly requirements, that is, if the depth of the shaft insertion hole is detected to be less than the preset minimum insertion depth or greater than the preset maximum insertion depth, the agent executes the next iteration operation, that is, the agent executes steps 1101 to 1104 again.
[0087] If the depth of the shaft insertion hole meets the preset assembly requirements, that is, if the depth of the shaft insertion hole is detected to be between the preset minimum insertion depth and the preset maximum insertion depth, it indicates that the shaft hole assembly is successful and the agent ends the iteration.
[0088] The shaft and hole assembly method provided in this embodiment learns multiple different types of shaft and hole assembly tasks through meta-reinforcement learning. By using prior knowledge, the learning efficiency of meta-reinforcement learning is improved, and the resulting shaft and hole assembly model can be applied to different types of shaft and hole assembly tasks, thereby improving the versatility and practicality of shaft and hole assembly.
[0089] Please refer to Figure 2 , Figure 2 This is a second schematic flowchart of a shaft-hole assembly method provided in an exemplary embodiment of the present invention. This embodiment is a further improvement on the foregoing embodiment, mainly in the specific process of the intelligent agent training the shaft-hole assembly model. The flowchart of this embodiment is as follows: Figure 2 As shown, it includes the following steps:
[0090] Step 210: Construct multiple different types of shaft and hole assembly tasks.
[0091] Specifically, the automatic assembly process performed by the intelligent agent can be viewed as an MDP process, where each assembly task T... n It can be represented as a quintuple n A n ,P n ,R n ,γ>。 Where, S n Let s be a finite set of states in the nth assembly task. Let s be a state in the set. k ∈S n A n Let a be a finite set of actions in the nth assembly task. k ∈A n a k For state s k The following actions can be performed; p n For the state transition equation of the nth assembly task, P k Indicates that in state s k Perform action a k Then P(s′) k |s k ,a k The probability of jumping to state s′ is given by ) k ;R n Here, γ is the reward function; γ is the discount coefficient, 0 ≤ γ ≤ 1. How the shaft and hole assembly model learns the Markov decision process corresponding to each assembly task is key to achieving automated assembly of multi-category shaft and hole assembly tasks. The training process of the shaft and hole assembly model is explained in detail below.
[0092] The intelligent agent constructs multiple different environments, such as physical simulation environments, mathematical simulation environments, and realistic assembly environments. In each environment, it constructs multiple different categories of shaft and hole assembly tasks based on the size and number of shaft holes.
[0093] In some embodiments, multiple different types of shaft hole assembly tasks may include: tasks that assemble shaft holes of different sizes and numbers in different environments; tasks that assemble shaft holes of the same size but different numbers in different environments; tasks that assemble shaft holes of different sizes but the same number in different environments; tasks that assemble shaft holes of the same size and number in different environments; tasks that assemble shaft holes of the same size but different numbers in the same environment; and tasks that assemble shaft holes of different sizes but the same number in the same environment.
[0094] In some embodiments, the number of shaft holes can be a single shaft hole, a double shaft hole, or a triple shaft hole, etc.
[0095] In the mathematical simulation environment, it is assumed that there is point contact between the shaft and the hole, and that the elastic deformation principle is satisfied.
[0096] The expression for the three-dimensional force acting on the end of the shaft is as follows:
[0097]
[0098] In the formula, F represents the three-dimensional force on the end of the shaft, K represents the number of shaft holes, δ represents the elastic deformation coefficient, and p i 1 p represents the maximum deformation vector of the i-th axis at the top of the hole. i 2 The maximum deformation vector of the i-th axis at the bottom of the hole.
[0099] The expression for the three-dimensional torque acting on the end of the shaft is as follows:
[0100]
[0101] In the formula, M represents the three-dimensional torque experienced at the end of the shaft, K represents the number of shaft holes, δ represents the elastic deformation coefficient, and p i 1 Let l represent the maximum deformation vector of the i-th axis at the top of the hole. i 1 p represents the offset vector from the center point of the i-th axis to the point of maximum deformation at the top of the hole. i 2 The maximum deformation vector of the i-th axis at the bottom of the hole, l i 2 This represents the offset vector from the center point of the i-th axis to the point of maximum deformation at the bottom of the hole.
[0102] The state of the shaft can be determined by the three-dimensional force and three-dimensional torque acting on its end.
[0103] Step 220: Perform multiple iterations of training.
[0104] Specifically, in order for the shaft-hole assembly model to learn from multiple different types of shaft-hole assembly tasks, the agent performs multiple iterative training operations. Each iterative training operation may include the following steps 2201 to 2205.
[0105] Step 2201: Randomly select one shaft and hole assembly task from multiple different categories of shaft and hole assembly tasks as the current training task.
[0106] Specifically, in each training iteration, the agent randomly selects one shaft-hole assembly task from multiple different categories of shaft-hole assembly tasks, with each task having an equal probability of being selected. For example, if there are 6 different categories of shaft-hole assembly tasks, then each task has a 1 / 6 probability of being selected. The agent then uses the selected shaft-hole assembly task as its current training task.
[0107] In some embodiments, since each training task is randomly selected from multiple different categories of shaft and hole assembly tasks, the current training task may be the same as or different from the training task in the previous iteration.
[0108] Step 2202: Execute the current training task and update the parameters of the shaft-hole assembly model.
[0109] Specifically, in each iteration training round, the shaft-hole assembly model outputs an assembly training action, executes the current training task through multiple iteration training rounds, and updates the parameters of the shaft-hole assembly model in each iteration training round.
[0110] Step 2203: Check whether the current training task has been completed.
[0111] Specifically, whether the current training task has been completed can be detected by checking whether the shaft hole has been successfully assembled, or whether the force on the shaft exceeds the maximum radial force.
[0112] If the detection shaft hole is successfully assembled, or if the force on the shaft exceeds the maximum radial force, it indicates that the current training task is complete, and the agent executes step 2204.
[0113] If the detection shaft hole is not successfully assembled and the force on the shaft does not exceed the maximum radial force, it indicates that the current training task is not completed, and the agent executes step 2203.
[0114] Step 2204: Obtain the cumulative reward corresponding to the current training task.
[0115] Specifically, each training iteration corresponds to a reward, and the rewards corresponding to multiple training iterations are summed to obtain the cumulative reward for the current training task.
[0116] Step 2205: Check whether the cumulative reward corresponding to the training task has converged.
[0117] Specifically, the training task includes the current training task and the previous training tasks. The agent obtains the cumulative reward corresponding to the previous training tasks in the previous iterations of training.
[0118] The agent determines the convergence of the cumulative reward for the current training task based on the cumulative reward for previous training tasks. Specifically, this can be achieved by starting with the cumulative reward for the current training task and checking if the cumulative rewards for a preset number of training tasks are close. If the cumulative rewards for the preset number of training tasks are close, it indicates that the cumulative reward for the training task has converged; if the cumulative rewards for the preset number of training tasks differ significantly, it indicates that the cumulative reward for the training task has not converged.
[0119] If the cumulative reward corresponding to the training task converges, the agent ends the iterative training and obtains the trained shaft-hole assembly model; if the cumulative reward corresponding to the training task does not converge, the agent executes the next iterative training, that is, it executes steps 2201 to 2205 again.
[0120] The shaft and hole assembly method provided in this embodiment constructs multiple shaft and hole assembly tasks of different categories, randomly selects one shaft and hole assembly task for training, so that the trained shaft and hole assembly model can adapt to different categories of shaft and hole assembly tasks, thereby improving the versatility and practicality of the shaft and hole assembly model.
[0121] Please refer to Figure 3 , Figure 3 This is the third flowchart illustrating an exemplary embodiment of the shaft-hole assembly method provided by the present invention. This embodiment is a detailed description of the foregoing embodiments, mainly illustrating the specific process by which the agent executes the current training task and updates the parameters of the shaft-hole assembly model. The flow of this embodiment is as follows: Figure 3 As shown, it includes the following steps:
[0122] Step 310: Execute the current training task through multiple iterations of training rounds, and update the parameters of the shaft-hole assembly model in each iteration of training rounds.
[0123] Specifically, the shaft-hole assembly model includes an action selection network and an action evaluation network. The action selection network is mainly used to train the action selection strategy to determine the optimal assembly action. The action evaluation network is used to score the assembly actions and guide the selection of the optimal assembly action.
[0124] In some embodiments, the action selection network can be an Actor network, and the action evaluation network can be a Critic network.
[0125] Each training iteration corresponds to one assembly action. Through multiple training iterations, multiple assembly actions are obtained, and the current training task is executed using these multiple assembly actions. The parameters of the action selection network and the action evaluation network are updated in each training iteration.
[0126] Each iteration of the training round may include the following steps 3101 to 3106.
[0127] Step 3101: Obtain the similarity between the previous assembly training action and the previous demonstration action, and the current training state.
[0128] Specifically, the agent obtains its current training state through force sensors. The current training state includes the current force information on the training axes.
[0129] The agent acquires the previous assembly training action and the previous demonstration action. The previous assembly training action is the assembly training action output by the action selection network based on the previous training state. The previous demonstration action is the demonstration action output by the assembly demonstration model based on the previous training state. The assembly demonstration model is a Gaussian model built based on prior knowledge.
[0130] The agent calculates the similarity between the previous assembly training action and the previous demonstration action, and obtains the similarity score.
[0131] In some embodiments, the expression for similarity is as follows:
[0132]
[0133] In the formula, m t-1 Indicates assembly training actions With demonstration actions The degree of similarity between them This indicates that the action selection network is trained based on the training state s corresponding to the (t-1)th iteration. t-1 Output assembly training actions, This indicates that the assembly demonstration model is based on the training state s corresponding to the (t-1)th iteration training round. t-1 The output demonstration action.
[0134] Step 3102: Input the similarity and current training state into the action evaluation network in the shaft hole assembly model to obtain the current first loss function value.
[0135] Specifically, the agent inputs the similarity and current training state into the action evaluation network in the shaft hole assembly model to obtain the data output by the action evaluation network.
[0136] In some embodiments, the data output by the action evaluation network may include data for evaluating the value of the current state, similarity, and the reward value corresponding to the current training state.
[0137] The agent obtains the current first loss function value based on the data output by the action evaluation network.
[0138] Step 3103: Update the parameters of the action evaluation network based on the current first loss function value.
[0139] Specifically, the agent sets a preset first loss function value based on the allowable accuracy of the action evaluation network, and compares the current first loss function value with the preset first loss function value. If the current first loss function value is greater than the preset first loss function value, the parameters of the action evaluation network are updated; if the current first loss function value is less than the preset first loss function value, the parameters of the action evaluation network are already optimal and do not need to be updated.
[0140] In some embodiments, the optimal parameters of the action evaluation network include the weight parameters θ. V .
[0141] Step 3104: Input the similarity and current training state into the action selection network in the shaft hole assembly model to obtain the current second loss function value.
[0142] Specifically, the agent inputs the similarity and current training state into the action selection network in the shaft hole assembly model, and obtains the data output by the action selection network.
[0143] In some embodiments, the data output by the action selection network may include: the current action selection strategy and the current assembled training action, etc.
[0144] The agent obtains the current second loss function value based on the data output by the action selection network and the data output by the action evaluation network.
[0145] Step 3105: Update the parameters of the action selection network based on the current second loss function value.
[0146] Specifically, the agent sets a preset second loss function value based on the allowable accuracy of the action selection network, and compares the current second loss function value with the preset second loss function value. If the current second loss function value is greater than the preset second loss function value, the parameters of the action evaluation network are updated; if the current second loss function value is less than the preset second loss function value, the parameters of the action evaluation network are already optimal and do not need to be updated.
[0147] In some embodiments, the optimal parameters for the action selection network include the weight parameters θ. A .
[0148] Step 3106: Check whether the current training task has been completed.
[0149] Specifically, whether the current training task has been completed can be detected by checking whether the shaft hole has been successfully assembled, or whether the force on the shaft exceeds the maximum radial force.
[0150] If the detection shaft hole is successfully assembled, or if the force on the shaft exceeds the maximum radial force, the current training task is completed, and the agent executes the next training task.
[0151] If the detection shaft hole is not successfully assembled and the force on the shaft does not exceed the maximum radial force, it indicates that the current training task is not completed, and the agent executes the next iteration training round, namely steps 3101 to 3106.
[0152] The shaft-hole assembly method provided in this embodiment obtains the current first loss function value by inputting similarity and current training state into the action evaluation network, and updates the parameters of the action evaluation network based on the current first loss function value. It also obtains the current second loss function value by inputting similarity and current training state into the action selection network, and updates the parameters of the action selection network based on the current second loss function value, thereby achieving accurate acquisition of the action evaluation network and the action selection network.
[0153] Please refer to Figure 4 , Figure 4 This is the fourth flowchart illustrating the shaft-hole assembly method provided in an exemplary embodiment of the present invention. This embodiment is a detailed description of the foregoing embodiments, mainly illustrating the specific process of inputting similarity and current training state into the action evaluation network in the shaft-hole assembly model to obtain the current first loss function value. The flow of this embodiment is as follows: Figure 4 As shown, it includes the following steps:
[0154] Step 410: Input the similarity and current training state into the action evaluation network to obtain the current state value and current reward value.
[0155] Specifically, the agent inputs the similarity score and the current training state into the action evaluation network, which outputs the current state value and the current reward value. The current state value corresponds to the similarity score and the current training state, and the current reward value also corresponds to the similarity score and the current training state.
[0156] In the t-th iteration of training, the state value output by the action evaluation network is in, Let T represent the state value function. i Let represent the i-th training task, p(T) represent the probability of selecting a training task, and s t Let m represent the training state corresponding to the t-th iteration training round. t-1 Indicates assembly training actions With demonstration actions The degree of similarity between them This indicates that the action selection network is based on the training state s corresponding to the (t-1)th iteration training round. t-1 Output assembly training actions, This indicates that the assembly demonstration model is based on the training state s corresponding to the (t-1)th iteration training round. t-1 The output demonstration action, θ V These represent the weight parameters of the action evaluation network. The state value is derived from the fusion of similarity scores. It can more effectively represent the value of a state.
[0157] The expression for the reward value output by the action evaluation network in the t-th iteration of training is as follows:
[0158] r t =R(s) t ,m t-1 )
[0159] In the formula, r t Let represent the reward value corresponding to the t-th iteration of training, R() represent the reward function, and s t Let m represent the training state corresponding to the t-th iteration training round. t-1 Indicates assembly training actions With demonstration actions The degree of similarity between them This indicates that the action selection network is based on the training state s corresponding to the (t-1)th iteration training round. t-1 Output assembly training actions, This indicates that the assembly demonstration model is based on the training state s corresponding to the (t-1)th iteration training round. t-1 The output demonstration action.
[0160] Step 420: Obtain the current advantage value based on the current state value, the current reward value, and the value of the next state.
[0161] Specifically, the agent inputs the similarity between the current assembly training action and the current demonstration action, along with the next training state, into the action evaluation network, which outputs the value of the next state. Here, the current assembly training action is the assembly training action output by the action selection network based on the current training state, and the current demonstration action is the demonstration action output by the assembly demonstration model based on the current training state.
[0162] In the t-th training iteration, the expression for the advantage value is as follows:
[0163]
[0164] In the formula, A t Let r represent the advantage value corresponding to the t-th training iteration. t Let γ represent the reward value corresponding to the t-th iteration of training, and let γ represent the discount factor, where 0 ≤ γ ≤ 1. Let represent the state value corresponding to the (t+1)th iteration of training, where Let T represent the state value function. i Let represent the i-th training task, p(T) represent the probability of selecting a training task, and s t+1 Let m represent the training state corresponding to the (t+1)th iteration training round. t Indicates assembly training actions With demonstration actions The degree of similarity between them This indicates that the action selection network is based on the training state s corresponding to the training round t.t Output assembly training actions, This indicates that the assembly demonstration model is based on the training state s corresponding to the t-th iteration training round. t The output demonstration action, θ V These represent the weight parameters of the action evaluation network. Let s represent the state value corresponding to the t-th iteration of training. t Let m represent the training state corresponding to the t-th iteration training round. t-1 Indicates assembly training actions With demonstration actions The degree of similarity between them.
[0165] After obtaining the value of the next state, the agent substitutes the current state value, the current reward value, and the value of the next state into the expression for the advantage value to obtain the current advantage value.
[0166] Step 430: Obtain the current first loss function value based on the current advantage value.
[0167] Specifically, the expression for the first loss function is as follows:
[0168]
[0169] In the formula, L(θ) V Let A represent the first loss function, L() represent the loss function, and A t T represents the advantage value corresponding to the t-th training iteration. i Let represent the i-th training task, and p(T) represent the probability of selecting a training task. This indicates that the execution is carried out with probability p(T) on the selected T. i The expected value of the task with the number of iteration training rounds t as the variable.
[0170] After obtaining the current advantage value, the agent substitutes the current advantage value into the expression of the first loss function to obtain the current value of the first loss function.
[0171] The shaft-hole assembly method provided in this embodiment obtains the current state value and the current reward value by inputting the similarity and the current training state into the action evaluation network. Based on the current state value, the current reward value, and the value of the next state, the current advantage value is obtained. Based on the current advantage value, the current first loss function value is obtained, which realizes the accurate acquisition of the current first loss function value. This is beneficial for accurately updating the parameters of the action evaluation network and further facilitates the acquisition of an accurate shaft-hole assembly model.
[0172] Please refer to Figure 5 , Figure 5This is the fifth flowchart illustrating the shaft-hole assembly method provided in an exemplary embodiment of the present invention. This embodiment is a detailed description of the foregoing embodiments, mainly illustrating the specific process of inputting similarity and current training state into the action selection network of the shaft-hole assembly model to obtain the current second loss function value. The flow of this embodiment is as follows: Figure 5 As shown, it includes the following steps:
[0173] Step 510: Input the similarity and current training state into the current action selection network to obtain the current assembly training action.
[0174] Specifically, the agent inputs the similarity score and the current training state into the action selection network, which then outputs the current action selection policy. The action selection policy is a set of assembled actions and their corresponding probabilities, and it follows a Gaussian distribution, with the horizontal axis representing the assembled actions and the vertical axis representing the probabilities.
[0175] In the t-th iteration of training, the action selection strategy is: Where, θ A This indicates the weight parameters of the network selected for the action, a. t s represents the assembly training action corresponding to the t-th iteration training round. t Let m represent the training state corresponding to the t-th iteration training round. t-1 Indicates assembly training actions With demonstration actions The degree of similarity between them.
[0176] The agent selects the assembly action with the highest probability in the current action selection strategy as the current assembly training action.
[0177] Step 520: Based on the probability of the current assembly training action and the probability of the previous assembly training action, as well as the current advantage value, obtain the current second loss function value.
[0178] Specifically, the expression for the second loss function is as follows:
[0179]
[0180] In the formula, L(θ) V ) represents the second loss function, L() represents the loss function, and θ A T represents the weight parameters of the network used to select actions. i Let represent the i-th training task, and p(T) represent the probability of selecting the current training task. This indicates that the execution is carried out with probability p(T) on the selected T. i The expected value, r, in the task is the number of training iterations t. t (θ) represents the ratio of the highest probability in the new and old action selection strategies, A tLet represent the advantage value corresponding to the t-th iteration of training, where ∈ is a hyperparameter, typically with a value of 0.2.
[0181] in,
[0182]
[0183] In the formula, r t (θ) represents the ratio of the probability of assembling the training action in the new and old action selection strategies. θ represents the probability of incorporating a training action into a new action selection strategy. A This indicates the weight parameters of the network used to select actions. This represents the probability of assembling a training action in the old action selection strategy. Represents relative to θ A In this regard, the weight parameters have not been updated.
[0184] The probability of assembling a training action at the current time is the same as the probability of assembling a training action in the new action selection strategy; the probability of assembling a training action at the previous time is the same as the probability of assembling a training action in the old action selection strategy.
[0185] The agent inputs the probability of the current assembly training action, the probability of the previous assembly training action, and the current advantage value into the second loss function to obtain the current value of the second loss function.
[0186] The shaft-hole assembly method provided in this embodiment obtains the current assembly training action by inputting the similarity and current training state into the current action selection network. Based on the probability of the current assembly training action and the probability of the previous assembly training action, as well as the current advantage value, the current second loss function value is obtained. This achieves accurate acquisition of the current second loss function value, which is beneficial for accurately updating the parameters of the action selection network and further beneficial for obtaining an accurate shaft-hole assembly model.
[0187] Please refer to Figure 6 , Figure 6 This is the sixth flowchart illustrating a shaft-hole assembly method provided in an exemplary embodiment of the present invention. This embodiment is a detailed description of the foregoing embodiments, mainly illustrating the specific process of obtaining the assembly demonstration model. The flow of this embodiment is as follows: Figure 6 As shown, it includes the following steps:
[0188] Step 610: Obtain the teaching status and teaching actions.
[0189] Specifically, prior knowledge can be teaching data during the shaft and hole assembly process, which includes teaching states and teaching actions.
[0190] The intelligent agent acquires teaching states and teaching actions through demonstration learning algorithms, that is, it imitates the shaft and hole assembly process to acquire teaching states and teaching actions.
[0191] Step 620: Model the teaching state and teaching actions to obtain the assembly demonstration model.
[0192] Specifically, the expression for the initial assembly demonstration model is as follows:
[0193]
[0194] In the formula, α represents the teaching action, ζ represents the teaching state, and λ represents the teaching state. k μ represents the weight coefficient of the k-th Gaussian model. k Let ∑ represent the mean of the k-th Gaussian model. k Let θ represent the covariance of the k-th Gaussian model, where θ = {θ1, θ2, ..., θ} K}, θ k These are the parameters of the Gaussian mixture model, θ k ={λ k ,μ k ,∑ k}
[0195] The intelligent agent inputs the teaching state and teaching actions into the initial assembly demonstration model, determines the parameters of the model, and thus obtains the constructed assembly demonstration model.
[0196] The shaft and hole assembly method provided in this embodiment constructs an assembly demonstration model through teaching states and teaching actions, which is beneficial for obtaining similarity by applying the assembly demonstration model, thereby accelerating the learning efficiency of the shaft and hole assembly module for multiple types of skills.
[0197] The shaft hole assembly device provided by the present invention is described below. The shaft hole assembly device described below and the shaft hole assembly method described above can be referred to in correspondence.
[0198] Figure 7 This is a schematic diagram of the structure of a shaft hole assembly device provided in an exemplary embodiment of the present invention, as shown below. Figure 7 As shown, the shaft hole assembly device includes: an execution module 710.
[0199] Execution module 710 is used to perform multiple iteration operations; each iteration operation includes:
[0200] Obtain the current state of the shaft; the current state includes the current force information;
[0201] The current state is input into the shaft-hole assembly model to obtain the current assembly action; the current assembly action includes translation and rotation; the shaft-hole assembly model is trained on multiple different types of shaft-hole assembly tasks based on prior knowledge and meta-reinforcement learning.
[0202] Perform the current assembly action on the axis;
[0203] If the depth of the shaft insertion hole does not meet the preset assembly requirements, the next iteration operation is performed.
[0204] In some embodiments, the shaft hole assembly device further includes a construction module and a training module.
[0205] The building module is used to build multiple different types of shaft and hole assembly tasks;
[0206] The training module is used to perform multiple iterative training operations; the training module includes: extracting sub-modules, executing sub-modules, obtaining sub-modules, and determining sub-modules, wherein:
[0207] The extraction submodule is used to randomly select one shaft and hole assembly task from the multiple different categories of shaft and hole assembly tasks as the current training task.
[0208] The execution submodule is used to execute the current training task and update the parameters of the shaft-hole assembly model;
[0209] The acquisition submodule is used to acquire the cumulative reward corresponding to the current training task after the current training task is completed.
[0210] The determining submodule is used to determine the convergence status of the cumulative reward corresponding to the training task based on the cumulative reward corresponding to the current training task and the cumulative reward corresponding to the previous training task.
[0211] The execution submodule is also used to execute the next iteration of training if the cumulative reward corresponding to the training task has not converged.
[0212] In some embodiments, the execution submodule is specifically used to: execute the current training task through multiple iterations of training rounds, and update the parameters of the shaft-hole assembly model in each iteration of training rounds; the execution submodule includes: a first acquisition unit, a second acquisition unit, a second acquisition unit, a first update unit, a third acquisition unit, a second update unit, and an execution unit.
[0213] The first acquisition unit is used to acquire the similarity between the previous assembly training action and the previous demonstration action, and the current training state; the previous assembly training action is the training action output by the network selection network in the shaft-hole assembly model according to the previous training state in the previous iteration training round; the previous demonstration action is the demonstration action output by the assembly demonstration model according to the previous training state; the current training state includes the current force information of the training shaft.
[0214] The second acquisition unit is used to input the similarity and the current training state into the action evaluation network in the shaft hole assembly model to obtain the current first loss function value;
[0215] The first update unit is used to update the parameters of the action evaluation network based on the current first loss function value;
[0216] The third acquisition unit inputs the similarity and the current training state into the action selection network in the shaft hole assembly model to obtain the current second loss function value.
[0217] The second update unit is used to update the parameters of the action selection network based on the current second loss function value;
[0218] The execution unit is used to execute the next iteration training round if the current training task has not been completed.
[0219] In some embodiments, the second acquisition unit includes: a first acquisition subunit, a second acquisition subunit, and a third acquisition subunit. Wherein:
[0220] The first acquisition subunit is used to input the similarity and the current training state into the action evaluation network to obtain the current state value and the current reward value;
[0221] The second acquisition subunit is used to acquire the current advantage value based on the current state value, the current reward value, and the next state value;
[0222] The third acquisition subunit is used to acquire the current first loss function value based on the current advantage value.
[0223] In some embodiments, the third acquisition unit includes: a fourth acquisition subunit and a fifth acquisition subunit. Wherein:
[0224] The fourth acquisition subunit is used to input the similarity and the current training state into the action selection network to obtain the current assembly training action.
[0225] The fifth acquisition subunit is used to acquire the current second loss function value based on the probability of the current assembly training action, the probability of the previous assembly training action, and the current advantage value.
[0226] In some embodiments, the shaft hole assembly device further includes: a first acquisition module and a second acquisition module. Wherein:
[0227] The first acquisition module is used to acquire the teaching status and teaching actions;
[0228] The second acquisition module is used to model the teaching state and the teaching action to obtain the assembly demonstration model.
[0229] It should be noted that the shaft hole assembly device provided by the present invention can realize all the method steps implemented in the above method embodiments and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiments and the beneficial effects will not be described in detail.
[0230] Figure 8 This is a schematic diagram of the physical structure of an electronic device provided in an exemplary embodiment of the present invention, as shown below. Figure 8 As shown, the electronic device may include a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute a shaft-hole assembly method, which includes: performing multiple iterative operations; each iterative operation includes: obtaining the current state of the shaft; the current state includes current force information; inputting the current state into the shaft-hole assembly model to obtain the current assembly action; the current assembly action includes translation and rotation; the shaft-hole assembly model is trained on multiple different types of shaft-hole assembly tasks based on prior knowledge and meta-reinforcement learning; executing the current assembly action on the shaft; and if the depth of the shaft insertion hole does not meet the preset assembly requirements, executing the next iterative operation.
[0231] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0232] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the shaft-hole assembly method provided by the above methods. The method includes: performing multiple iterative operations; each iterative operation includes: obtaining the current state of the shaft; the current state includes current force information; inputting the current state into the shaft-hole assembly model to obtain the current assembly action; the current assembly action includes translation and rotation; the shaft-hole assembly model is trained on multiple different types of shaft-hole assembly tasks based on prior knowledge and meta-reinforcement learning; performing the current assembly action on the shaft; and if the depth of the shaft insertion hole does not meet the preset assembly requirements, performing the next iterative operation.
[0233] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the shaft-hole assembly method provided by the above methods. The method includes: performing multiple iterative operations; each iterative operation includes: obtaining the current state of the shaft; the current state includes current force information; inputting the current state into a shaft-hole assembly model to obtain a current assembly action; the current assembly action includes translation and rotation; the shaft-hole assembly model is trained on multiple different types of shaft-hole assembly tasks based on prior knowledge and meta-reinforcement learning; performing the current assembly action on the shaft; and if the depth of the shaft insertion hole does not meet the preset assembly requirements, performing the next iterative operation.
[0234] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0235] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0236] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A shaft hole assembly method characterized by, The method comprises the following steps: performing multiple iteration operations; each iteration operation comprises: obtaining a current state of the shaft; the current state comprises current force information; inputting the current state into a shaft-hole assembly model to obtain a current assembly action; the current assembly action comprises a translation amount and a rotation amount; the shaft-hole assembly model is trained based on prior knowledge and meta-reinforcement learning in multiple different categories of shaft-hole assembly tasks; the different categories of shaft-hole assembly tasks comprise shaft-hole assembly tasks with different combinations of shaft-hole sizes and numbers of shaft holes in different environments; performing the current assembly action on the shaft; if the depth of the shaft in the hole does not meet the preset assembly requirement, performing the next iteration operation; the shaft-hole assembly model is trained based on the following steps: constructing multiple different categories of shaft-hole assembly tasks; performing multiple iteration training operations; each iteration training operation comprises: randomly selecting a shaft-hole assembly task from the multiple different categories of shaft-hole assembly tasks as a current training task; performing the current training task and updating the parameters of the shaft-hole assembly model; after the completion of the current training task, obtaining the cumulative reward corresponding to the current training task; determining the convergence of the cumulative reward corresponding to the training task according to the cumulative reward corresponding to the current training task and the cumulative rewards corresponding to the previous training tasks; if the cumulative reward corresponding to the training task does not converge, performing the next iteration training; the performing of the current training task and the updating of the parameters of the shaft-hole assembly model comprises: performing the current training task through multiple iteration training rounds and updating the parameters of the shaft-hole assembly model in each iteration training round; wherein each iteration training round comprises: obtaining a similarity between a previous assembly training action and a previous demonstration action, and a current training state; the previous assembly training action is a training action output by a motion selection network in the shaft-hole assembly model according to a previous training state in a previous iteration training round; the previous demonstration action is a demonstration action output by an assembly demonstration model according to the previous training state; the current training state comprises current force information of a training shaft; inputting the similarity and the current training state into a motion evaluation network in the shaft-hole assembly model to obtain a current first loss function value; updating the parameters of the motion evaluation network based on the current first loss function value; inputting the similarity and the current training state into a motion selection network in the shaft-hole assembly model to obtain a current second loss function value; updating the parameters of the motion selection network based on the current second loss function value; if the current training task is not completed, performing the next iteration training round.
2. The shaft hole assembly method according to claim 1, characterized by, the inputting of the similarity and the current training state into the motion evaluation network in the shaft-hole assembly model to obtain the current first loss function value comprises: inputting the similarity and the current training state into the motion evaluation network to obtain a current state value and a current reward value; According to the current state value and the current reward value, and a next state value, an advantage value is obtained; According to the current advantage value, the current first loss function value is obtained.
3. The shaft hole assembly method according to claim 2, characterized by, The action selection network in the shaft hole assembly model is inputted with the similarity and the current training state, and a current second loss function value is obtained. The action selection network is inputted with the similarity and the current training state, and a current assembly training action is obtained; The current second loss function value is obtained based on a probability of the current assembly training action and a probability of a previous assembly training action, and the current advantage value.
4. The shaft hole assembly method according to claim 1, characterized by, Before obtaining the similarity between the previous assembly training action and the previous demonstration action, the method further includes: A teaching state and a teaching action are obtained; The teaching state and the teaching action are modeled to obtain an assembly demonstration model.
5. A shaft hole assembly apparatus, characterized by, The method includes: An execution module is configured to perform a plurality of iteration operations; Each iteration operation includes: A current state of the shaft is obtained; the current state includes current force information; The current state is inputted into a shaft hole assembly model to obtain a current assembly action; the current assembly action includes a translation amount and a rotation amount; the shaft hole assembly model is trained based on prior knowledge and meta-reinforcement learning in a plurality of different categories of shaft hole assembly tasks; the different categories of shaft hole assembly tasks include shaft hole assembly tasks with varying combinations of shaft hole sizes and numbers of shaft holes in different environments; The current assembly action is performed on the shaft; If the depth of the shaft in the hole does not meet a preset assembly requirement, a next iteration operation is performed; The shaft hole assembly device further includes a construction module and a training module; The construction module is configured to construct a plurality of different categories of shaft hole assembly tasks; The training module is configured to perform a plurality of iteration training operations; The training module includes an extraction submodule, an execution submodule, an obtaining submodule, and a determination submodule, wherein: the extraction submodule is configured to extract a shaft hole assembly task from the plurality of different categories of shaft hole assembly tasks as a current training task; the execution submodule is configured to execute the current training task and update parameters of the shaft hole assembly model; the obtaining submodule is configured to obtain a cumulative reward corresponding to the current training task after the current training task is completed; and the determination submodule is configured to determine a convergence condition of the cumulative reward corresponding to the training task according to the cumulative reward corresponding to the current training task and a cumulative reward corresponding to a previous training task; and the execution submodule is further configured to perform a next iteration training if the cumulative reward corresponding to the training task does not converge. The execution submodule is specifically configured to execute the current training task through a plurality of iteration training rounds, and update the parameters of the shaft hole assembly model in each iteration training round; and the execution submodule includes a first obtaining unit, a second obtaining unit, a second obtaining unit, a first updating unit, a third obtaining unit, a second updating unit, and an execution unit. The first obtaining unit is configured to obtain a similarity between a previous assembly training action and a previous demonstration action, and a current training state; the previous assembly training action is a training action output by a motion selection network in the shaft-hole assembly model according to a previous training state in a previous iteration training round; the previous demonstration action is a demonstration action output by an assembly demonstration model according to the previous training state; and the current training state includes current force information of a training shaft; The second obtaining unit is configured to input the similarity and the current training state into a motion evaluation network in the shaft-hole assembly model, and obtain a current first loss function value; The first updating unit is configured to update parameters of the motion evaluation network based on the current first loss function value; The third obtaining unit is configured to input the similarity and the current training state into the motion selection network in the shaft-hole assembly model, and obtain a current second loss function value; The second updating unit is configured to update parameters of the motion selection network based on the current second loss function value; The execution unit is configured to perform a next iteration training round in a case where the current training task is not completed.
6. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The computer program is executed by the processor to implement the shaft-hole assembly method in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the shaft-hole assembly method in any one of claims 1 to 4.
8. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the shaft-hole assembly method in any one of claims 1 to 4.
Citation Information
Patent Citations
Intelligent robot grabbing method based on action demonstration teaching
CN111890357A
Robot shaft hole assembling method based on deep reinforcement learning and admittance control
CN115674204A