Relay and base assembly control method and device, electronic equipment and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INST OF AUTOMATION CHINESE ACAD OF SCI
- Filing Date
- 2023-11-27
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]本发明提供一种继电器与底座的装配控制方法、装置、电子设备及介质,用以解决现有技术中继电器与底座装配困难的问题
[0032] The present invention provides a method, device, electronic device, and medium for assembling a relay and its base. By combining a force-position hybrid controller and a reinforcement learning network, the automatic assembly of the relay and its base is achieved. The relevant control parameters in the force-position hybrid controller are optimized by the reinforcement learning training strategy, which can improve the adaptability of the force-position hybrid controller. At the same time, the force-position hybrid controller can reduce the ineffective exploration of the reinforcement learning network during the training process and improve the training efficiency.
Smart Images

Figure CN117826577B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of reinforcement learning technology, and in particular to an assembly control method, device, electronic device, and medium for a relay and a base. Background Technology
[0002] A relay is an electronic control device commonly used in automatic control systems. It is typically mounted on a fixed base and then connected to the circuit through the base. Because relays have many pins, and the pins and holes are assembled with an interference fit, fine adjustments are required during the assembly process to avoid jamming. Currently, this assembly is mainly done manually. Using robots to complete this interference fit assembly task is challenging. First, in traditional control methods, the controller parameters need to be finely adjusted manually according to the assembly object and remain constant throughout the assembly process, making adaptive adjustments impossible based on changes in assembly depth. Second, during assembly, axial control requires simultaneous position and force control; however, due to the complexity of the actual contact model between the relay and the base, it is difficult to establish an accurate force model, thus making it impossible to determine the appropriate desired force. Summary of the Invention
[0003] This invention provides a method, device, electronic device, and medium for assembling a relay and its base, in order to solve the problem of difficult assembly of relays and bases in the prior art.
[0004] This invention provides a method for assembling a relay and a base, applied to a force-position hybrid controller. The force-position hybrid controller includes a proportional-derivative controller and a proportional controller. The force-position hybrid controller is used to control the position of an end effector, and the end effector is used to assemble the relay onto the base. The method includes:
[0005] The radial force on the end effector is controlled by the proportional-derivative controller to obtain a first position increment of the end effector. The first position increment is used to adjust the position of the relay in the x-axis direction and the position in the y-axis direction.
[0006] The proportional controller is used to control the axial position of the relay to obtain the second position increment of the end effector. The second position increment is used to adjust the insertion depth of the relay in the z-axis direction.
[0007] Wherein, the x-axis, y-axis, and z-axis are coordinate axes in the workpiece coordinate system, which is a coordinate system constructed with the geometric center of the relay as the origin; the proportional coefficient of the proportional-derivative controller is adjusted based on the output value of the reinforcement learning network, which is trained based on the forces acting on the end effector in different coordinate axis directions at historical moments and the loading depth of the relay in the z-axis direction at historical moments.
[0008] In some embodiments, controlling the radial force on the end effector using the proportional-derivative controller to obtain a first position increment of the end effector includes:
[0009] When the radial force deviation input to the proportional-derivative controller exceeds the lower limit threshold of the radial force deviation, the first position increment is obtained based on the proportional coefficient of the proportional-derivative controller, the derivative coefficient of the proportional-derivative controller, the radial force deviation, and the change in the radial force deviation. The first position increment is used to indicate the position increment of the relay in the x-axis direction and the position increment in the y-axis direction.
[0010] Based on the first position increment, the position of the relay in the x-axis direction and the position in the y-axis direction are adjusted.
[0011] In some embodiments, controlling the axial position of the relay using the proportional controller to obtain a second position increment of the end effector includes:
[0012] Based on the axial position deviation of the relay in the z-axis direction and the proportional coefficient of the proportional controller, a first control quantity for the axial position of the relay is determined;
[0013] Based on the axial position increment threshold in the z-axis direction and the first control quantity, a second control quantity for the axial position of the relay is determined;
[0014] Based on the second control quantity and the preset radial force coefficient, the second position increment is obtained, and the second position increment is used to indicate the position increment of the relay in the z-axis direction;
[0015] The preset radial force coefficient is determined based on the force on the relay pin in the x-axis direction, the force on the relay pin in the y-axis direction, and the upper limit threshold of the radial force.
[0016] In some embodiments, determining a second control quantity for the axial position of the relay based on the axial position increment threshold in the z-axis direction and the first control quantity includes:
[0017] If the first control quantity exceeds the axial position increment threshold, the second control quantity is determined as the axial position increment threshold.
[0018] If the first control quantity does not exceed the axial position increment threshold, the second control quantity is determined as the first control quantity.
[0019] In some embodiments, the reinforcement learning network is trained in the following manner:
[0020] Based on the state actions at multiple historical moments, a state sequence is constructed, and each state action is constructed based on the force on the end effector in different coordinate axis directions and the insertion depth of the relay in the z-axis direction;
[0021] Based on the state sequence, the reinforcement learning network is trained, and the action performance is evaluated using a reward function. The reward function is used to reduce the force on the end effector in the x-axis direction, the force in the y-axis direction, and the number of control steps in the assembly process.
[0022] In some embodiments, the reward function is expressed as:
[0023]
[0024] Where c1 is the first weighting coefficient, c2 is the second weighting coefficient, c3 is the third weighting coefficient, k1 is the distribution coefficient, and p zg p is the target insertion depth of the relay in the z-axis direction. z(t+1) f is the insertion depth of the relay in the t-th control time step in the z-axis direction. xt f is the force exerted on the end effector in the x-axis direction at the t-th control time step. yt Let n(a) be the force exerted on the end effector in the y-axis direction at the t-th control time step. t ) represents the control step number of the currently executed state action, n max s(a) represents the maximum number of control steps in a round. t This is used to indicate whether the assembly was successfully completed.
[0025] The present invention also provides an assembly control device for a relay and a base, applied to a force-position hybrid controller. The force-position hybrid controller includes a proportional-derivative controller and a proportional controller. The force-position hybrid controller is used to control the position of an end effector, and the end effector is used to assemble the relay onto the base. The device includes:
[0026] The first control module is used to control the radial force on the end effector using the proportional-derivative controller to obtain a first position increment of the end effector. The first position increment is used to adjust the position of the relay in the x-axis direction and the position in the y-axis direction.
[0027] The second control module is used to control the axial position of the relay using the proportional controller to obtain the second position increment of the end effector. The second position increment is used to adjust the insertion depth of the relay in the z-axis direction.
[0028] Wherein, the x-axis, y-axis, and z-axis are coordinate axes in the workpiece coordinate system, which is a coordinate system constructed with the geometric center of the relay as the origin; the proportional coefficient of the proportional-derivative controller is adjusted based on the output value of the reinforcement learning network, which is trained based on the forces acting on the end effector in different coordinate axis directions at historical moments and the loading depth of the relay in the z-axis direction at historical moments.
[0029] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the assembly control method of the relay and the base as described above.
[0030] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the assembly control method of the relay and base as described above.
[0031] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the assembly control method for the relay and base as described above.
[0032] The present invention provides a method, device, electronic device, and medium for assembling a relay and its base. By combining a force-position hybrid controller and a reinforcement learning network, the automatic assembly of the relay and its base is achieved. The relevant control parameters in the force-position hybrid controller are optimized by the reinforcement learning training strategy, which can improve the adaptability of the force-position hybrid controller. At the same time, the force-position hybrid controller can reduce the ineffective exploration of the reinforcement learning network during the training process and improve the training efficiency. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0034] Figure 1 This is a flowchart illustrating the assembly control method for the relay and base provided by the present invention.
[0035] Figure 2 This is a schematic diagram of the control system framework for the assembly control method of the relay and the base provided by the present invention;
[0036] Figure 3 This is a schematic diagram of the assembly control device for the relay and the base provided by the present invention;
[0037] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0038] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0039] The following is combined with Figures 1-4 The present invention describes the assembly control method, apparatus, electronic device, and medium for the relay and base.
[0040] The execution subject of the relay and base assembly control method provided by this invention can be an electronic device, a component in an electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, handheld computers, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This invention does not impose specific limitations.
[0041] The technical solution of the present invention will be described in detail below using the computer execution of the assembly control method for the relay and base provided by the present invention as an example.
[0042] Figure 1 This is a flowchart illustrating the assembly control method for the relay and base provided by the present invention. (Refer to...) Figure 1 The present invention provides a relay and base assembly control method, applied to a force-position hybrid controller. The force-position hybrid controller includes a proportional-derivative controller and a proportional controller. The force-position hybrid controller is used to control the position of the end effector, and the end effector is used to assemble the relay onto the base. The method includes:
[0043] Step 110: Use a proportional-derivative controller to control the radial force on the end effector to obtain the first position increment of the end effector. The first position increment is used to adjust the position of the relay in the x-axis direction and the position in the y-axis direction.
[0044] Step 120: Use a proportional controller to control the axial position of the relay to obtain the second position increment of the end effector. The second position increment is used to adjust the insertion depth of the relay in the z-axis direction.
[0045] The x-axis, y-axis, and z-axis are the coordinate axes in the workpiece coordinate system, which is constructed with the geometric center of the relay as the origin. The proportional coefficient of the proportional-derivative controller is adjusted based on the output value of the reinforcement learning network, which is trained based on the forces acting on the end effector in different coordinate axis directions at historical moments and the loading depth of the relay in the z-axis direction at historical moments.
[0046] In actual implementation, a workpiece coordinate system is established with the geometric center point of the relay as the origin. The positive direction of the z-axis is set vertically downward, the positive direction of the x-axis is set to the right parallel to the horizontal surface of the hole, and the direction of the y-axis is determined by the right-hand rule.
[0047] It should be noted that, in this invention, an end effector refers to any tool connected to the edge of a robot and possessing a certain function. This may include robot grippers, robot tool quick-change devices, robot collision sensors, robot rotary connectors, or robot pressure tools, etc. A robot's end effector is generally considered to be a robot peripheral device, a robot accessory, a robot tool, or an end-effector tool.
[0048] The end effector in this invention is used to control the relay, enabling the relay to be successfully assembled with the base.
[0049] The control system framework of the present invention is as follows Figure 2 As shown, the force-position hybrid controller uses the selection matrix S f and S p The position of the end effector and the force acting on the end effector are separated into two independent and decoupled subspaces for processing.
[0050] The force control channel (F) in the force-position hybrid controller uses a proportional-derivative (PD) controller to control the radial force on the end effector. Based on the feedback radial force deviation, the first position increment of the end effector is obtained. The relay's position can then be adjusted along the x-axis and y-axis in the workpiece coordinate system based on this first position increment.
[0051] The position control channel (P) in the force-position hybrid controller uses a proportional (P) controller to control the axial position of the relay, obtaining a second position increment of the end effector. Based on the second position increment, the relay is inserted to the desired depth in the z-axis direction.
[0052] To improve the adaptability of the force-position hybrid controller, this invention also employs a reinforcement learning network to adaptively adjust the proportional gain of the PD controller in the force-position hybrid controller based on force information obtained from the force sensor and information such as the relay insertion depth. The output value of the reinforcement learning network is the proportional gain of the PD controller. The proportional gain of the PD controller includes the proportional gain Δk in the x-axis direction. fpx The proportional gain Δk in the y-axis direction fpy .
[0053] The relay and base assembly control method provided by this invention achieves automatic assembly of the relay and base by combining a force-position hybrid controller and a reinforcement learning network. The relevant control parameters in the force-position hybrid controller are optimized by the reinforcement learning training strategy, which can improve the adaptability of the force-position hybrid controller. At the same time, the force-position hybrid controller can reduce the invalid exploration of the reinforcement learning network during the training process and improve the training efficiency.
[0054] In some embodiments, step 110 may include:
[0055] When the radial force deviation input to the proportional-derivative controller exceeds the lower limit threshold of the radial force deviation, a first position increment is obtained based on the proportional coefficient of the proportional-derivative controller, the derivative coefficient of the proportional-derivative controller, the radial force deviation, and the change in the radial force deviation. The first position increment is used to indicate the position increment of the relay in the x-axis direction and the position increment in the y-axis direction.
[0056] Based on the first position increment, adjust the position of the relay in the x-axis direction and the position in the y-axis direction.
[0057] In actual operation, when the radial force deviation input to the PD controller does not exceed the lower limit threshold of the radial force deviation, the position of the end effector is not adjusted.
[0058] When the radial force deviation input to the PD controller exceeds the lower limit threshold of the radial force deviation, the position of the end effector is adjusted according to the obtained first position increment.
[0059] The first position increment can include the position increment of the relay in the x-axis direction and the position increment in the y-axis direction. The first position increment is obtained by formula (1):
[0060]
[0061] Where, ΔP x Let ΔP be the position increment of the relay in the x-axis direction. y The radial force deviation is the position increment of the relay in the y-axis direction, including the force deviation e of the end effector in the x-axis direction. fx and the force deviation e in the y-axis direction fy Δe fx Let x be the change in force deviation along the x-axis, and Δe be the value of Δe. fy e represents the change in force deviation along the y-axis. fT0 This represents the lower threshold value for radial force deviation. k fpx k is the proportional coefficient of the PD controller in the x-axis direction. fpy k is the proportional coefficient of the PD controller in the y-axis direction. fdx Let k be the differential coefficient in the x-axis direction. fdydenoted as the differential coefficient in the y-axis direction.
[0062] In some embodiments, step 120 may include:
[0063] Based on the axial position deviation of the relay in the z-axis direction and the proportional coefficient of the proportional controller, the first control quantity for the axial position of the relay is determined.
[0064] Based on the axial position increment threshold in the z-axis direction and the first control quantity, a second control quantity for the axial position of the relay is determined.
[0065] Based on the second control quantity and the preset radial force coefficient, the second position increment is obtained. The second position increment is used to indicate the position increment of the relay in the z-axis direction.
[0066] The preset radial force coefficient is determined based on the force on the relay pins in the x-axis direction, the force on the relay pins in the y-axis direction, and the upper limit threshold of the radial force.
[0067] In actual implementation, a proportional (P) controller is used to control the axial position of the relay to obtain the first control quantity, so that the insertion depth of the relay reaches the desired depth. The first control quantity is expressed by formula (2).
[0068] ΔP z1 =k piz e pz (2)
[0069] Where, ΔP z1 k is the first control variable of the end effector on the axial position of the relay. piz e is the proportional coefficient of the P controller in the z-axis direction. pz This represents the axial positional deviation of the relay in the z-axis direction.
[0070] To ensure that the parts are not damaged during assembly, the obtained axial position control value needs to be limited, as shown in formula (3), to obtain the second control value ΔP of the end effector on the axial position of the relay. z2 .
[0071] In some embodiments, determining a second control quantity for the axial position of the relay based on an axial position increment threshold in the z-axis direction and a first control quantity includes:
[0072] If the first control quantity exceeds the axial position increment threshold, the second control quantity is determined as the axial position increment threshold.
[0073] If the first control quantity does not exceed the axial position increment threshold, the second control quantity is determined as the first control quantity.
[0074] In actual implementation, the second control variable can be expressed as:
[0075]
[0076] Where, ΔP zT This is the threshold value for the axial position increment in the z-axis direction. Therefore, it can be seen that in the first control variable ΔP... z1 Exceeding the axial position increment threshold ΔP zT At that time, the second control quantity ΔP z2 The axial position increment threshold ΔP zT ; in the first control quantity ΔP z1 Not exceeding the axial position increment threshold ΔP zT At that time, the second control quantity ΔP z2 The first control variable ΔP z1 .
[0077] Determine the second control quantity ΔP z2 Then, combined with the preset radial force coefficient k in formula (4) f Adjust the output of the P controller in the position control channel, i.e., the second position increment. The second position increment is used to indicate the position increment ΔP of the relay in the z-axis direction. z To adjust the insertion depth of the relay in the z-axis direction, as shown in formula (5).
[0078]
[0079] ΔP z =k f ΔP z2 (5)
[0080] Among them, F x F is the force F exerted on the relay pin in the x-axis direction. y F is the force F exerted on the relay pin in the y-axis direction. r F is the radial force acting on the shaft. rT This is the upper limit threshold for radial force.
[0081] Understandably, when the radial force on the relay shaft exceeds the upper limit threshold of the radial force, the assembly is considered to have failed.
[0082] In some embodiments, reinforcement learning (RL) networks are trained as follows:
[0083] Based on the state actions at multiple historical moments, a state sequence is constructed. Each state action is constructed based on the force on the end effector in different coordinate axis directions and the insertion depth of the relay in the z-axis direction.
[0084] Based on the state sequence, a reinforcement learning network is trained, and a reward function is used to evaluate the action performance. The reward function is used to reduce the force on the end effector in the x-axis direction, the force in the y-axis direction, and the number of control steps in the assembly process.
[0085] In practice, due to the latency in feedback control, the state should be stored as a continuous sequence of states within a certain window size during reinforcement learning policy training. For example, the window size can be set to 4, i.e., the state action is [f x(t-3) f y(t-3) f z(t-3) p z(t-3) , ..., f xt ,,f yt f zt p zt ] T
[0086] State-action pairs are stored in an experience replay pool. When the amount of stored data exceeds 200 sets, 56 state-action pairs are randomly selected each time. An offline reinforcement learning algorithm, such as the Deep Deterministic Policy Gradient (DDPG) algorithm, is used to train the reinforcement learning network to update the policy until the reinforcement learning network converges.
[0087] In network training, a reward function can be set, which is designed to encourage the agent to complete the assembly task with smaller radial forces and fewer control steps during the training process.
[0088] In some embodiments, the reward function is expressed as:
[0089]
[0090] Where c1 is the first weighting coefficient, c2 is the second weighting coefficient, c3 is the third weighting coefficient, k1 is the distribution coefficient, and p zg p represents the target insertion depth of the relay in the z-axis direction. z(t+1) f represents the insertion depth of the relay at the t-th control time step in the z-axis direction. xt Let f be the force exerted on the end effector in the x-axis direction at the t-th control time step. yt Let n(a) be the force exerted on the end effector in the y-axis direction at the t-th control time step. t ) represents the control step number of the currently executed state action, n max s(a) represents the maximum number of control steps in a round. t This is used to indicate whether the assembly process has been successfully completed.
[0091] Among them, s(a tA value of 0 indicates successful assembly. t A value of -1 indicates assembly failure. k1 can be determined based on the correlation coefficient of the normal distribution, and its specific value can be set based on the actual situation; no specific limitation is made here.
[0092] This invention proposes an assembly control method for relays and bases, enabling automated assembly. Force-position hybrid controllers are commonly used in tasks involving extensive contact. While they offer some robustness to external disturbances, they lack adaptability and struggle to achieve optimal compliant control during actual assembly. Reinforcement learning methods exhibit strong adaptability to changes in environmental states, but require extensive interaction with the environment, resulting in low sample efficiency. This invention combines the advantages of both methods, proposing a force-position hybrid controller that optimizes controller parameters through reinforcement learning training. This control method leverages reinforcement learning to improve controller adaptability while utilizing the force-position controller to reduce ineffective exploration during reinforcement learning network training, thereby improving training efficiency.
[0093] The assembly control device for the relay and the base provided by the present invention will be described below. The assembly control device for the relay and the base described below can be referred to in correspondence with the assembly control method for the relay and the base described above.
[0094] Figure 3 This is a structural schematic diagram of the assembly control device for the relay and base provided by the present invention. (Refer to...) Figure 3 The relay and base assembly control device provided by the present invention is applied to a force-position hybrid controller. The force-position hybrid controller includes a proportional-derivative controller and a proportional controller. The force-position hybrid controller is used to control the position of the end effector. The end effector is used to assemble the relay onto the base. The device includes a first control module 310 and a second control module 320.
[0095] The first control module 310 is used to control the radial force on the end effector using the proportional-derivative controller to obtain a first position increment of the end effector. The first position increment is used to adjust the position of the relay in the x-axis direction and the position in the y-axis direction.
[0096] The second control module 320 is used to control the axial position of the relay using the proportional controller to obtain a second position increment of the end effector. The second position increment is used to adjust the insertion depth of the relay in the z-axis direction.
[0097] Wherein, the x-axis, y-axis, and z-axis are coordinate axes in the workpiece coordinate system, which is a coordinate system constructed with the geometric center of the relay as the origin; the proportional coefficient of the proportional-derivative controller is adjusted based on the output value of the reinforcement learning network, which is trained based on the forces acting on the end effector in different coordinate axis directions at historical moments and the loading depth of the relay in the z-axis direction at historical moments.
[0098] The relay and base assembly control device provided by this invention achieves automatic assembly of the relay and base by combining a force-position hybrid controller and a reinforcement learning network. The relevant control parameters in the force-position hybrid controller are optimized by the reinforcement learning training strategy, which can improve the adaptability of the force-position hybrid controller. At the same time, the force-position hybrid controller can reduce the invalid exploration of the reinforcement learning network during the training process and improve the training efficiency.
[0099] In one embodiment, the first control module 310 is specifically used for:
[0100] When the radial force deviation input to the proportional-derivative controller exceeds the lower limit threshold of the radial force deviation, the first position increment is obtained based on the proportional coefficient of the proportional-derivative controller, the derivative coefficient of the proportional-derivative controller, the radial force deviation, and the change in the radial force deviation. The first position increment is used to indicate the position increment of the relay in the x-axis direction and the position increment in the y-axis direction.
[0101] Based on the first position increment, the position of the relay in the x-axis direction and the position in the y-axis direction are adjusted.
[0102] In one embodiment, the second control module 320 is specifically used for:
[0103] Based on the axial position deviation of the relay in the z-axis direction and the proportional coefficient of the proportional controller, a first control quantity for the axial position of the relay is determined;
[0104] Based on the axial position increment threshold in the z-axis direction and the first control quantity, a second control quantity for the axial position of the relay is determined;
[0105] Based on the second control quantity and the preset radial force coefficient, the second position increment is obtained, and the second position increment is used to indicate the position increment of the relay in the z-axis direction;
[0106] The preset radial force coefficient is determined based on the force on the relay pin in the x-axis direction, the force on the relay pin in the y-axis direction, and the upper limit threshold of the radial force.
[0107] In one embodiment, the second control module 320 is specifically used for:
[0108] If the first control quantity exceeds the axial position increment threshold, the second control quantity is determined as the axial position increment threshold.
[0109] If the first control quantity does not exceed the axial position increment threshold, the second control quantity is determined as the first control quantity.
[0110] In one embodiment, the reinforcement learning network is trained as follows:
[0111] Based on the state actions at multiple historical moments, a state sequence is constructed, and each state action is constructed based on the force on the end effector in different coordinate axis directions and the insertion depth of the relay in the z-axis direction;
[0112] Based on the state sequence, the reinforcement learning network is trained, and the action performance is evaluated using a reward function. The reward function is used to reduce the force on the end effector in the x-axis direction, the force in the y-axis direction, and the number of control steps in the assembly process.
[0113] In one embodiment, the reward function is expressed as:
[0114]
[0115] Where c1 is the first weighting coefficient, c2 is the second weighting coefficient, c3 is the third weighting coefficient, k1 is the distribution coefficient, and p zg p is the target insertion depth of the relay in the z-axis direction. z(t+1) f is the insertion depth of the relay in the t-th control time step in the z-axis direction. xt f is the force exerted on the end effector in the x-axis direction at the t-th control time step. yt Let n(a) be the force exerted on the end effector in the y-axis direction at the t-th control time step. t ) represents the control step number of the currently executed state action, n max s(a) represents the maximum number of control steps in a round. t This is used to indicate whether the assembly was successfully completed.
[0116] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call logic instructions in the memory 430 to execute an assembly control method for the relay and its base, applied to a force-position hybrid controller. The force-position hybrid controller includes a proportional-derivative controller and a proportional controller. The force-position hybrid controller is used to control the position of the end effector, which is used to assemble the relay onto the base. The method includes:
[0117] The radial force on the end effector is controlled by the proportional-derivative controller to obtain a first position increment of the end effector. The first position increment is used to adjust the position of the relay in the x-axis direction and the position in the y-axis direction.
[0118] The proportional controller is used to control the axial position of the relay to obtain the second position increment of the end effector. The second position increment is used to adjust the insertion depth of the relay in the z-axis direction.
[0119] Wherein, the x-axis, y-axis, and z-axis are coordinate axes in the workpiece coordinate system, which is a coordinate system constructed with the geometric center of the relay as the origin; the proportional coefficient of the proportional-derivative controller is adjusted based on the output value of the reinforcement learning network, which is trained based on the forces acting on the end effector in different coordinate axis directions at historical moments and the loading depth of the relay in the z-axis direction at historical moments.
[0120] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0121] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, the computer program being executed by a processor, the computer being able to execute the relay and base assembly control method provided by the above methods, applied to a force-position hybrid controller, the force-position hybrid controller including a proportional-derivative controller and a proportional controller, the force-position hybrid controller being used to control the position of an end effector, the end effector being used to assemble the relay onto the base, the method including:
[0122] The radial force on the end effector is controlled by the proportional-derivative controller to obtain a first position increment of the end effector. The first position increment is used to adjust the position of the relay in the x-axis direction and the position in the y-axis direction.
[0123] The proportional controller is used to control the axial position of the relay to obtain the second position increment of the end effector. The second position increment is used to adjust the insertion depth of the relay in the z-axis direction.
[0124] Wherein, the x-axis, y-axis, and z-axis are coordinate axes in the workpiece coordinate system, which is a coordinate system constructed with the geometric center of the relay as the origin; the proportional coefficient of the proportional-derivative controller is adjusted based on the output value of the reinforcement learning network, which is trained based on the forces acting on the end effector in different coordinate axis directions at historical moments and the loading depth of the relay in the z-axis direction at historical moments.
[0125] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the assembly control method for a relay and a base provided by the methods described above, applied to a force-position hybrid controller, the force-position hybrid controller including a proportional-derivative controller and a proportional controller, the force-position hybrid controller being used to control the position of an end effector, the end effector being used to assemble the relay onto the base, the method comprising:
[0126] The radial force on the end effector is controlled by the proportional-derivative controller to obtain a first position increment of the end effector. The first position increment is used to adjust the position of the relay in the x-axis direction and the position in the y-axis direction.
[0127] The proportional controller is used to control the axial position of the relay to obtain the second position increment of the end effector. The second position increment is used to adjust the insertion depth of the relay in the z-axis direction.
[0128] Wherein, the x-axis, y-axis, and z-axis are coordinate axes in the workpiece coordinate system, which is a coordinate system constructed with the geometric center of the relay as the origin; the proportional coefficient of the proportional-derivative controller is adjusted based on the output value of the reinforcement learning network, which is trained based on the forces acting on the end effector in different coordinate axis directions at historical moments and the loading depth of the relay in the z-axis direction at historical moments.
[0129] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0130] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for assembling and controlling a relay and a base, characterized in that, A method for using a force-position hybrid controller, the force-position hybrid controller including a proportional-derivative controller and a proportional controller, the force-position hybrid controller being used to control the position of an end effector, the end effector being used to assemble a relay onto a base, the method comprising: The radial force on the end effector is controlled by the proportional-derivative controller to obtain a first position increment of the end effector. This first position increment is used to adjust the relay's position. x Position in the axial direction and in y Position along the axis; The proportional controller is used to control the axial position of the relay to obtain a second position increment of the end effector. This second position increment is used to adjust the relay's position. z Insertion depth in the axial direction; Among them, the x Axial direction, the y Axial direction and the z The axes are coordinate axes in the workpiece coordinate system, which is constructed with the geometric center of the relay as the origin; the proportional coefficient of the proportional-derivative controller is adjusted based on the output value of the reinforcement learning network, which is based on the forces acting on the end effector in different coordinate axis directions at historical moments and the relay's position at those historical moments. z Loading depth training in the axial direction; The reinforcement learning network is trained in the following manner: Based on the state actions at multiple historical moments, a state sequence is constructed, where each state action is based on the force exerted on the end effector in different coordinate axis directions and the relay's... z Components are constructed for the loading depth in the axial direction. Based on the state sequence, the reinforcement learning network is trained, and a reward function is used to evaluate action performance. This reward function is used to reduce the performance of the end effector during the action sequence. x Force in the axial direction, in the y Forces in the axial direction and the number of control steps during assembly; The reward function is expressed as follows: ; in, c 1 is the first weighting coefficient. c 2 is the second weighting coefficient. c 3 is the third weighting coefficient. k 1 is the distribution coefficient. p zg For the relay in z Target loading depth in the axial direction p z(t+1) For the relay in z The loading depth at the t-th control time step in the axial direction. f xt For the first t The end effector in each control time step x Force in the axial direction, f yt For the first t The end effector in each control time step y Force in the axial direction, n ( a t ) represents the control step number of the currently executed state action. n max The maximum number of control steps in a round. s ( a t This is used to indicate whether the assembly was successfully completed.
2. The assembly control method for the relay and the base according to claim 1, characterized in that, The step of using the proportional-derivative controller to control the radial force on the end effector to obtain the first position increment of the end effector includes: When the radial force deviation input to the proportional-derivative controller exceeds the lower limit threshold of the radial force deviation, the first position increment is obtained based on the proportional coefficient of the proportional-derivative controller, the derivative coefficient of the proportional-derivative controller, the radial force deviation, and the change in the radial force deviation. The first position increment is used to indicate the position of the relay at the specified location. x The position increment in the axial direction and in the y Position increment in the axial direction; Based on the first position increment, adjust the relay at the position. x Position in the axial direction and in the y Position along the axis.
3. The assembly control method for the relay and the base according to claim 1, characterized in that, The step of using the proportional controller to control the axial position of the relay to obtain the second position increment of the end effector includes: Based on the relay in z The axial position deviation in the axial direction and the proportional coefficient of the proportional controller are used to determine the first control quantity for the axial position of the relay. Based on the above z The axial position increment threshold in the axial direction and the first control quantity are used to determine the second control quantity for the axial position of the relay; Based on the second control quantity and the preset radial force coefficient, the second position increment is obtained. The second position increment is used to indicate the position of the relay. z Position increment in the axial direction; Wherein, the preset radial force coefficient is based on the force exerted on the relay pins by the x The force in the axial direction, the force exerted on the pins of the relay y The upper limit thresholds for axial force and radial force are determined.
4. The assembly control method for the relay and the base according to claim 3, characterized in that, The basis of z The second control quantity for determining the axial position of the relay is determined by using the axial position increment threshold in the axial direction and the first control quantity, including: If the first control quantity exceeds the axial position increment threshold, the second control quantity is determined as the axial position increment threshold. If the first control quantity does not exceed the axial position increment threshold, the second control quantity is determined as the first control quantity.
5. An assembly control device for a relay and a base, characterized in that, A force-position hybrid controller is applied to a force-position hybrid controller, the force-position hybrid controller including a proportional-derivative controller and a proportional controller, the force-position hybrid controller being used to control the position of an end effector, the end effector being used to assemble a relay onto a base, the device comprising: The first control module is used to control the radial force on the end effector using the proportional-derivative controller, and to obtain a first position increment of the end effector. This first position increment is used to adjust the relay's position. x Position in the axial direction and in y Position along the axis; The second control module is used to control the axial position of the relay using the proportional controller to obtain a second position increment of the end effector. This second position increment is used to adjust the relay's position. z Insertion depth in the axial direction; Among them, the x Axial direction, the y Axial direction and the z The axes are coordinate axes in the workpiece coordinate system, which is constructed with the geometric center of the relay as the origin; the proportional coefficient of the proportional-derivative controller is adjusted based on the output value of the reinforcement learning network, which is based on the forces acting on the end effector in different coordinate axis directions at historical moments and the relay's position at those historical moments. z Loading depth training in the axial direction; The reinforcement learning network is trained in the following manner: Based on the state actions at multiple historical moments, a state sequence is constructed, where each state action is based on the force exerted on the end effector in different coordinate axis directions and the relay's... z Components are constructed for the loading depth in the axial direction. Based on the state sequence, the reinforcement learning network is trained, and a reward function is used to evaluate action performance. This reward function is used to reduce the performance of the end effector during the action sequence. x Force in the axial direction, in the y Forces in the axial direction and the number of control steps during assembly; The reward function is expressed as follows: ; in, c 1 is the first weighting coefficient. c 2 is the second weighting coefficient. c 3 is the third weighting coefficient. k 1 is the distribution coefficient. p zg For the relay in z Target loading depth in the axial direction p z(t+1) For the relay in z The loading depth at the t-th control time step in the axial direction. f xt For the first t The end effector in each control time step x Force in the axial direction, f yt For the first t The end effector in each control time step y Force in the axial direction, n ( a t ) represents the control step number of the currently executed state action. n max The maximum number of control steps in a round. s ( a t This is used to indicate whether the assembly was successfully completed.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the assembly control method for the relay and the base as described in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the assembly control method for the relay and the base as described in any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the assembly control method for the relay and the base as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Grinding constant force control method based on deep reinforcement learning PPO algorithm
CN114660940A
KR20220013884A