Single-mechanical-arm furniture assembly motion planning method based on noise pulse hybrid reinforcement learning

By introducing a noise pulse hybrid reinforcement learning method in the robotic arm motion planning, the design strategy-value network and noise pulse neural network solve the problems of low flexibility and high energy consumption in furniture assembly, and realize efficient and low-energy-consuming action planning, which improves the application capabilities of robotic arm in intelligent manufacturing.

CN120023815AActive Publication Date: 2025-05-23ANHUI UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510341347.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-05-23
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The robotic arm has problems of low flexibility and high energy consumption in furniture assembly motion planning, making it difficult to achieve efficient and low-energy action planning in complex environments.

Method used

Using a method based on noise pulse hybrid reinforcement learning, a single robotic arm furniture assembly motion planning method is designed, and efficient motion planning of the robotic arm is achieved through strategy-value networks and noise pulse neural networks.

Benefits of technology

It has achieved a high degree of flexibility and low energy consumption in furniture assembly tasks, improved the continuous operation ability of the robotic arm in resource-constrained scenarios, and promoted the practical development of the field of intelligent manufacturing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120023815A_ABST
    Figure CN120023815A_ABST
Patent Text Reader

Abstract

The invention provides a single-mechanical-arm furniture assembly motion planning method based on noise pulse hybrid reinforcement learning, and the method specifically comprises the following steps: collecting observation information, building a strategy-value network to guide a mechanical arm to carry out furniture assembly, and designing a noise pulse neural network as a strategy network; designing a coding layer based on group coding, and converting observation information into a pulse sequence; designing a pulse neural layer with noise, and outputting final pulse activity according to the pulse sequence; converting the pulse activity into action expression through a non-pulse decoding layer with noise; designing a dynamic adjustment noise level optimization method, and modifying a loss function to realize noise-strategy balance; and network parameters are optimized, an optimal model is obtained, and the mechanical arm is guided to complete an assembly task. A guarantee is provided for continuous operation of the mechanical arm in a resource limited scene, and practical development of intelligent equipment in the field of intelligent manufacturing is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot arm motion planning, and in particular to a single robot arm furniture assembly motion planning method based on noise pulse hybrid reinforcement learning. Background Art

[0002] Robotic arms have shown many significant advantages in furniture assembly tasks, which not only improve production efficiency, but also optimize product quality and production safety. First of all, robotic arms can significantly improve the efficiency and quality of furniture assembly. Its high precision and repeatability enable it to quickly and accurately complete the tasks of grasping, positioning and assembling parts, avoiding fatigue and errors that may occur in manual operations. For example, in the IKEA furniture assembly project, the robotic arm can efficiently complete the assembly task in a short time through 3D camera recognition and precise force control technology. Secondly, the flexibility and programmability of the robotic arm enable it to adapt to furniture parts of different sizes and shapes. By replacing the end effector or adjusting the program, the robotic arm can easily switch between different assembly tasks to meet diverse production needs. This flexibility not only improves the adaptability of the production line, but also reduces the adjustment costs caused by product changes.

[0003] In addition, the application of robotic arms has significantly reduced labor costs and labor intensity. In traditional furniture manufacturing, assembly work is usually labor-intensive and repetitive, which can easily lead to worker fatigue and work-related injuries. The introduction of robotic arms can replace manual labor to complete these heavy tasks, reduce dependence on manpower, and reduce material waste and product defects caused by human errors. For example, the Universal Robots collaborative robotic arm not only reduces assembly time but also reduces the risk of employee injury through precise control and high-speed movement capabilities. In terms of quality control, the high-precision operation of the robotic arm can ensure the consistency and stability of furniture assembly. Its repeated positioning accuracy can reach millimeters or even micrometers, which can effectively reduce product quality problems caused by assembly errors. For example, some advanced robotic arms use six-dimensional force sensors to achieve precise force control feedback to ensure precise operation of each assembly link. In short, the use of robotic arms provides great flexibility and adaptability for industrial production.

[0004] As a neural network architecture that simulates the working mode of biological nervous systems, the spiking neural network has many unique advantages. It transmits information through discrete spike signals, which can more realistically imitate the behavior patterns of neurons in the brain. This architecture not only performs well in processing spatiotemporal information, but also achieves efficient computing with lower energy consumption due to its characteristic of triggering pulse emission only when the threshold is reached. When noise is injected into the spiking neural network, the performance and robustness of the network can be further improved. This noise injection mechanism is similar to the randomness in the biological nervous system, which can enhance the network's sensitivity to input signals and enable it to better adapt and learn in complex or uncertain environments. At the same time, the introduction of noise can also improve the network's training process, and by increasing the diversity of training, it can help the network avoid falling into local optimality, thereby improving the generalization ability of the model. Summary of the invention

[0005] In response to the problems of low flexibility and high energy consumption in robot arm motion planning, the present invention provides a single robot arm furniture assembly motion planning method based on noise pulse hybrid reinforcement learning, aiming to take into account both the high flexibility and low energy consumption of robot arm motion planning, and promote the practical development of intelligent equipment in the field of intelligent manufacturing.

[0006] In order to achieve the above technical objectives, the present invention provides the following technical solutions:

[0007] A single robot arm furniture assembly motion planning method based on noise pulse hybrid reinforcement learning specifically includes the following steps:

[0008] S1. Collect observation information of the robot arm in the furniture assembly task scenario, build a strategy-value network to guide the robot arm to assemble furniture, and design a noise pulse neural network as a strategy network; the noise pulse neural network includes an encoding layer, a pulse neural layer and a decoding layer;

[0009] S2. Design the encoding layer based on population coding to convert the observed information into a deterministic pulse sequence. The pulse sequence is a set of discrete time pulses generated by neurons during the encoding process, which represents the time and intensity characteristics of the observed information.

[0010] S3, designing a pulse neural layer with noise, and outputting the final pulse activity according to the pulse sequence obtained in step S2, where the pulse activity is the overall dynamic behavior of neurons in the neural layer, reflecting the activation state of neurons;

[0011] S4, then through the noisy non-pulse decoding layer, the input pulse activity is converted into action expression;

[0012] S5. Design a method to dynamically adjust the noise level. By modifying the loss function, the noise can be dynamically adjusted to improve the training efficiency while ensuring full exploration.

[0013] S6. Based on the deep reinforcement learning framework SAC, the parameters of the noise pulse neural network are optimized to obtain the optimal model; the trained optimal model is used to guide the robotic arm to complete the furniture assembly task.

[0014] Furthermore, the observation information in step S1 is specifically:

[0015] The robot arm state r and furniture component state f at each time step t; the observation information o is expressed by the formula:

[0016] o={r(h 1 ,h 2 ; g 1 ;e 1 ,e 2 ,e 3 ,e 4 ),f(p n ,q n )};

[0017] Among them, h 1 ,h 2 Respectively represent the position and velocity of the robot arm joint; g 1 Indicates the position of the robot gripper; e 1 ~e 4 They represent the position, quaternion, linear velocity and angular velocity of the end effector of the robot arm respectively; p n ,q n are the 3D coordinates and orientation of the nth furniture component respectively.

[0018] Furthermore, step S2 specifically includes:

[0019] S21, observation data o for observation information o at time i i Any data o of dimension j in ij , calculate the group activation value p of the observation data of the jth dimension through the Gaussian function j , the formula is:

[0020]

[0021] Among them, u j It represents the response center of the observed data in the jth dimension, σ j It represents the response width of the observation data in the jth dimension;

[0022] S22. At each time step t, the neuron's membrane potential cumulative population activation value p j , and then update the membrane potential state; the update rules are as follows:

[0023] v j(t) = v j (t-1)+p j ;

[0024]

[0025] v j (t) = v j (t)-V th ·s j (t);

[0026] Among them, v j (t) represents the neuron membrane potential state related to the j-th dimension observation data at the current time t, which is initially 0; To determine whether the voltage accumulation exceeds the threshold V th If the threshold is exceeded, it takes 1, otherwise it takes 0, s j (t) represents the generated pulse train.

[0027] Furthermore, step S3 specifically includes:

[0028] S31. Design a pulse neural layer with a circuit leakage-charge-discharge working principle with noise, in which the current of any pulse neuron Update according to the following rules:

[0029]

[0030] Among them, δ represents the current decay rate, X t represents external input; l represents the layer number of the neural network where the neuron is located, and m represents the mth neuron in the layer;

[0031] S32. If noise is introduced into the neuron charging process, the charging voltage of any spiking neuron with noise will be Update according to the following rules:

[0032]

[0033] Where v represents the membrane voltage after discharge, η represents the voltage decay rate, and σ v represents the noise parameter, ε v represents noise extracted from any random process;

[0034] S33, according to the updated charging voltage, any pulse neuron pulse signal Update and issue according to the following rules:

[0035]

[0036] Among them, Θ(·) processes each element in the input charging point cloud. If it exceeds the threshold V th , it takes 1; otherwise, it takes 0. σ s represents the noise parameter, and ε s represents the noise extracted from any random process;

[0037] S34. After emitting a pulse signal, reset the membrane potential. The update rule of the membrane potential v(t) after discharging is as follows:

[0038]

[0039] Among them, V rest represents the resting potential.

[0040] Furthermore, step S4 specifically includes:

[0041] S41. Design a neuron with a non-pulse integration-discharge working principle that injects noise as the decoding layer; at each time step t, the membrane potential v r (t) of the r-th neuron in the decoding layer is updated according to the pulse signal emitted by the pulse neural layer. The update rule is as follows:

[0042] v r (t) = v r (t - 1) + x t ;

[0043] Among them, x t represents the pulse signal input at the current time step; if the model is currently in the training mode, noise is also added to the input. The update method of the membrane potential after adding noise is:

[0044] v r (t) = v r (t - 1) + x t + σ s ⊙ ε s ;

[0045] Among them, σ s represents the noise parameter, and ε s represents the noise extracted from any random process;

[0046] S42. Obtain an intuitive action expression through the membrane potentials [v r1 , v r2 , …, v rT of the neurons in the decoding layer. The specific decoding method is to calculate the average membrane potential of the entire time series. The formula expression is:

[0047]

[0048] Among them, out represents the obtained action expression.

[0049] More specifically, the action expression obtained after decoding in step S42 is specifically:

[0050] At each time step t, the robotic arm action a = {a 1 , a 2 , …, a 9};

[0051] Among them, a 1 ~a 7 respectively represent the speeds of seven robotic arm joints; a 8 represents selection, that is, selecting one from many furniture components for grasping; a 9 represents connection and is used to represent interaction with the environment and objects.

[0052] Furthermore, step S5 specifically includes:

[0053] S51. Use the constructed noisy pulse neural network as the policy network and calculate its original loss L old ; Then, take the sum of the original loss and the noise variances of the non-pulse neurons in the decoding layer as the new loss function L new of the policy network, and the formula is expressed as:

[0054]

[0055] Among them, N A represents the action dimension, σ i is the noise standard deviation of the i-th non-pulse neuron, and k is a coefficient that controls the update ratio between the original gradient update direction and the noise reduction direction, and its value is dynamically adjusted according to the current reward and task;

[0056] S52. When the reward value is low, a small k value encourages the robotic arm to explore new strategies using noise; when the reward value is high, a large k value controls the robotic arm to reduce the use of noise and focus on using existing strategies to obtain greater rewards; the formula for k is expressed as:

[0057]

[0058] Among them, k 0 is a factor that controls the size of the loss term; R mean represents the average reward value after multiple trainings; R max represents the highest reward value; R min represents the lowest reward value.

[0059] Furthermore, step S6 specifically includes:

[0060] S61. Use the double Q network to minimize the loss function to train the value network; the formula is expressed as:

[0061] L critic =(Q target -Q real1 ) 2 +(Q target -Q real2 ) 2 ;

[0062] Among them, Q target Represents the prediction of future rewards, which is calculated by the target network and the policy network; Q real1 and Q real2 are two different value estimates of the same state-action pair by the value network; Q target The calculation formula is as follows:

[0063] Q target =γ·(1-d t )·(minC target (o t+1 ,a t+1 )-α·logπ(o t+1 ))+r t ;

[0064] Where γ is the discount factor for future rewards, d t Is a Boolean value; if d t =1, it means the current step is the final state, otherwise d t =0; C target (o t+1 ,a t+1 ) is the target network for the next state o t+1 and action a t+1 The estimated Q value of ; α is the temperature parameter for adjusting the entropy of the strategy; logπ(o t+1 ) is the policy entropy of the next state; r t is the reward at time t;

[0065] S62. The policy network in the deep reinforcement learning framework is composed of a noise pulse neural network, while the value network is composed of two MLP instances, each of which consists of three fully connected layers;

[0066] S63. Randomly initialize the parameters of the entire reinforcement learning framework, calculate the gradient of the loss with respect to the policy network parameters, and use the Adam optimizer to noise-spike the weights of the neural network.

[0067] S64. Deploy the optimally configured noise pulse neural network to the furniture assembly task scenario, execute steps S1-S4, and guide the robot arm in the task scenario to complete the furniture assembly task.

[0068] Based on the above technical solution, the present invention has at least the following beneficial effects:

[0069] 1. The present invention designs a strategy-value network for guiding the motion planning of robotic arm furniture assembly. A noisy pulse neural network is introduced into the strategy network, while a deep neural network is still used in the value network. The introduced pulse neural network triggers the pulse emission only when the threshold is reached, thereby achieving efficient calculation with lower energy consumption; while the value network still retains the deep neural network, which can make the performance of the entire strategy-value network not inferior to the original one;

[0070] 2. By designing a spiking neural layer with noisy neurons, time-related noise is introduced during the charging and transmission of neurons, which improves the exploration efficiency in diverse environments and helps to make diverse decisions in complex environments. By designing a non-spiking decoding module with noise, the output pulse signal is converted into a more intuitive action space expression. The dynamic adjustment of the noise level method is adopted to achieve a balance between strategy and noise.

[0071] 3. Finally, the network parameter configuration is optimized based on the deep reinforcement learning algorithm framework, and an intelligent decision-making model that can be transferred to the robotic arm furniture assembly task is finally constructed; this method ultimately achieves the construction of an energy-efficient motion generation system for the robotic arm, enabling it to maintain continuous operation capabilities in resource-constrained scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 A schematic diagram of a strategy-value network constructed in the method proposed in the present invention for guiding the motion planning of robotic arm furniture assembly;

[0073] Figure 2 A schematic diagram of a spiking neuron model with noise in the method proposed in the present invention;

[0074] Figure 3 It is a simulation schematic diagram of the robot arm in the present invention performing family assembly. DETAILED DESCRIPTION

[0075] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0076] Although the steps in the present invention are arranged with reference numerals, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It can be understood that the term "and / or" used herein relates to and encompasses any and all possible combinations of one or more of the associated listed items.

[0077] With the development of robotic arm technology, more and more robotic arms are applied to practical tasks with complex high-dimensional observation and action spaces. These robotic arms usually rely on limited on-board energy sources, so it is necessary to develop energy-saving continuous control solutions to extend their operating time. Although deep reinforcement learning methods based on policy gradients have successfully learned optimal control policies in complex tasks, these methods are often accompanied by the disadvantage of high energy consumption, which limits their application in various scenarios. In contrast, spiking neural networks provide a low-energy alternative that can effectively solve this problem. On this basis, the present application further adds relevant noise to the spiking neural network. An appropriate noise level can help the network avoid unstable phenomena during the learning process, generate more diverse decisions, and be closer to the working mechanism of biological neural networks. Therefore, the present application designs a spiking neural network model framework with noise, including an encoding layer, a spiking neural layer with noisy neurons, and a non-spiking decoding module with noise. In addition, in order to balance the relationship between the policy and the noise, the present application also proposes a noise reduction method for non-spiking neurons. Optimize the parameters of the above modules based on the deep reinforcement learning algorithm framework, and finally construct an intelligent decision-making model that can be transferred to the robotic arm furniture assembly task.

[0078] This study aims to build an action generation system with the optimal energy efficiency for the robotic arm, enabling it to maintain continuous operation ability in resource-constrained scenarios and promoting the practical development of intelligent equipment in the field of intelligent manufacturing. As Figure 1 shown, a single robotic arm furniture assembly motion planning method based on noise-pulse hybrid reinforcement learning of the present invention is shown, which specifically includes the following steps:

[0079] S1. Collect the observation information of the robotic arm in the furniture assembly task scenario, build a policy-value network to guide the robotic arm for furniture assembly, and design a spiking neural network with noise as the policy network; the spiking neural network with noise includes an encoding layer, a spiking neural layer, and a decoding layer;

[0080] As a preferred embodiment, the observation information in step S1 is specifically:

[0081] The robotic arm state r and the furniture component state f at each time step t; the observation information o is expressed by the formula:

[0082] o = {r(h1 ,h 2 ; g 1 ;e 1 ,e 2 ,e 3 ,e 4 ),f(p n ,q n )};

[0083] Among them, h 1 ,h 2 Respectively represent the position and velocity of the robot arm joint; g 1 Indicates the position of the robot gripper; e 1 ~e 4 They represent the position, quaternion, linear velocity and angular velocity of the end effector of the robot arm respectively; p n ,q n are the 3D coordinates and orientation of the nth furniture component respectively.

[0084] S2. Design the encoding layer based on population coding to convert the observed information into a deterministic pulse sequence. The pulse sequence is a set of discrete time pulses generated by neurons during the encoding process, which represents the time and intensity characteristics of the observed information.

[0085] As a preferred implementation, step S2 specifically includes:

[0086] S21, observation data o for observation information o at time i i Any data o of dimension j in ij In this embodiment, j∈{1,2,…,64} indicates that there are 64 dimensions of data in the robot state r and the furniture component state; the group activation value p of the observation data of the jth dimension is calculated by the Gaussian function j (The group activation value indicates the response intensity of the observation information in the group encoding), and the formula is expressed as:

[0087]

[0088] Among them, u j It represents the response center of the observed data in the jth dimension, σ j It represents the response width of the observed data of the jth dimension. In this embodiment, 10 neurons are set for each dimension for group encoding.

[0089] S22. At each time step t, the neuron's membrane potential cumulative population activation value p j , and then update the membrane potential state; the update rules are as follows:

[0090] v j (t) = v j(t-1)+p j ;

[0091]

[0092] v j (t) = v j (t)—V th ·s j (t);

[0093] Among them, v j (t) represents the neuron membrane potential state related to the j-th dimension observation data at the current time t, which is initially 0; To determine whether the voltage accumulation exceeds the threshold V th If the threshold is exceeded, it takes 1, otherwise it takes 0, s j (t) represents the pulse sequence generated; the above three formulas respectively describe the accumulation of group activation values, the judgment of pulse release, and the adjustment of membrane potential based on whether a pulse is generated.

[0094] S3, design a noisy spike neural layer, and output the final spike activity according to the spike sequence obtained in step S2. The spike activity is the overall dynamic behavior of neurons in the neural layer, reflecting the activation state of neurons. An appropriate noise level can help the network avoid instability during the learning process, and can also generate more diverse decisions, and is closer to the working mechanism of biological neural networks.

[0095] As a preferred embodiment, Figure 2 As shown, step S3 specifically includes:

[0096] S31. Design a pulse neural layer with a circuit leakage-charge-discharge working principle with noise, in which the current of any pulse neuron Update according to the following rules:

[0097]

[0098] Among them, δ represents the current decay rate, X t represents external input; l represents the layer number of the neural network where the neuron is located, and m represents the mth neuron in the layer;

[0099] S32. If noise is introduced into the neuron charging process, the charging voltage of any spiking neuron with noise will be Update according to the following rules:

[0100]

[0101] Where v represents the membrane voltage after discharge, η represents the voltage decay rate, and σ vrepresents the noise parameter, ε v represents noise extracted from any random process;

[0102] S33, according to the updated charging voltage, any pulse neuron pulse signal Update and issue according to the following rules:

[0103]

[0104] Among them, Θ(·) processes each element in the input charging point cloud, and if it exceeds the threshold V th , then it takes 1, otherwise it takes 0; σ s represents the noise parameter, ε s represents noise extracted from any random process;

[0105] S34, after the pulse signal is released, the membrane potential is reset, and the updating rule of the membrane potential v(t) after discharge is as follows:

[0106]

[0107] Among them, V rest Represents the resting potential.

[0108] In addition, in this embodiment, three sequentially connected pulse neural layers with noise are selected as the final neural layer of the noise pulse neural network, such as Figure 1 As shown: Layer 1 and Layer 2 are hidden layers with 256 pulse neurons, and Layer 3 is the output layer with 90 pulse neurons (considering that the robot arm has a total of nine actions, 10 neurons are used for each action in this embodiment); at the same time, the input data of each layer is converted into a format suitable for pulse neuron processing through a linear layer in advance.

[0109] The present application introduces time-related noise in the process of neuron charging and discharging-neuron firing, thereby improving the exploration efficiency in diverse environments; and after the above-mentioned leakage-charging-discharging process, the pulse neural layer finally outputs pulse activity that reflects the overall dynamic behavior.

[0110] S4, then through the noisy non-pulse decoding layer, the input pulse activity is converted into action expression; the design of the non-pulse decoding layer is based on continuous numerical calculation, which can decode the pulse activity more accurately and reduce the information loss caused by pulse sparsity or time step limitation;

[0111] As a preferred implementation, step S4 specifically includes:

[0112] S41. Design a neuron with a non-pulse integration-discharge working principle that injects noise as the decoding layer; at each time step t, the membrane potential v of the rth neuron in the decoding layer r (t) Update according to the pulse signal sent by the pulse neural layer. The update rules are as follows:

[0113] v r (t) = v r (t-1)+x t ;

[0114] Among them, x t It represents the pulse signal input of the current time step; if the model is currently in training mode, noise is also added to the input. After adding noise, the update method of the membrane potential is:

[0115] v r (t) = v r (t-1)+x t +σ s ⊙ε s ;

[0116] Among them, σ s represents the noise parameter, ε s represents noise extracted from any random process;

[0117] S42, through the membrane potential of neurons in the decoding layer [v r1 ,v r2 ,…,v rT ] to obtain intuitive action expression. The specific decoding method is to calculate the average membrane potential of the entire time series, which is expressed as:

[0118]

[0119] Among them, out represents the obtained action expression; this decoding method assumes that the most useful information is contained in the average dynamics of the entire time series, so the output is the average of the membrane potential of all time steps; the action expression obtained by the decoding layer is more intuitive, and in the furniture assembly task of the robot arm, the action expression obtained is specifically:

[0120] The robot action a at each time step t is a={a 1 ,a 2 ,…,a 9};

[0121] Among them, a 1 ~a 7 Respectively represent the speeds of the seven robotic arm joints; a 8 represents selection, i.e., picking up one piece of furniture from many; a 9Represents a connection, used to represent interaction with the environment and objects.

[0122] At this point, the pulse signal is converted into the action to be performed by the robot arm. The next two steps mainly discuss the balance between noise and strategy, as well as the optimization of network parameters.

[0123] S5. Design a method to dynamically adjust the noise level. By modifying the loss function, the noise can be dynamically adjusted to improve the training efficiency while ensuring full exploration.

[0124] As a preferred implementation, step S5 specifically includes:

[0125] S51. Use the constructed noise pulse neural network as the policy network and calculate its original loss L old ; Then the sum of the original loss and the noise variance of the non-spiking neurons in the decoding layer is used as the new loss function L of the policy network new , the formula is:

[0126]

[0127] Among them, N A represents the action dimension, σ i is the noise standard deviation of the i-th non-spiking neuron, k is the coefficient that controls the update ratio between the original gradient update direction and the noise reduction direction, and its value is dynamically adjusted according to the current reward and task;

[0128] S52. When the reward value is low, a smaller k value indicates that the noise variance term has little impact on the loss function, and the robot arm is encouraged to use noise to explore new strategies; when the reward value is high, a larger k value indicates that the noise variance term has a strong impact on the loss function, and the robot arm is controlled to reduce the use of noise and focus on using existing strategies to obtain greater rewards; thus achieving a balance between noise use and strategy rewards; the formula for k is expressed as:

[0129]

[0130] Among them, k 0 is the factor that controls the size of the loss term; R mean Represents the average reward value after multiple trainings; R max Indicates the maximum reward value; R min Indicates the minimum reward value.

[0131] S6. Optimize the parameters of the noise pulse neural network based on the deep reinforcement learning framework SAC to obtain the optimal model; use the trained optimal model to guide the robotic arm to complete the furniture assembly task;

[0132] As a preferred implementation, step S6 specifically includes:

[0133] S61. Use the double Q network to minimize the loss function to train the value network; the formula is:

[0134] L critic =(Q target -Q real1 ) 2 +(Q target -Q real2 ) 2 ;

[0135] Among them, Q target Represents the prediction of future rewards, which is calculated by the target network and the policy network; Q real1 and Q real2 are two different value estimates of the same state-action pair by the value network; Q target The calculation formula is as follows:

[0136] Q target =γ·(1-d t )·(minC target (o t+1 ,a t+1 )-α·logπ(o t+1 ))+r t ;

[0137] Where γ is the discount factor for future rewards, d t Is a Boolean value; if d t =1, it means the current step is the final state, otherwise d t =0; C target (o t+1 ,a t+1 ) is the target network for the next state o t+1 and action a t+1 The estimated Q value of ; α is the temperature parameter for adjusting the entropy of the strategy; logπ(o t+1 ) is the policy entropy of the next state; r t is the reward at time t;

[0138] S62. The policy network in the deep reinforcement learning framework is composed of a noise pulse neural network, while the value network is composed of two MLP instances, each of which consists of three fully connected layers;

[0139] S63. Randomly initialize the parameters of the entire reinforcement learning framework, calculate the gradient of the loss with respect to the policy network parameters, and use the Adam optimizer to noise-spike the weights of the neural network.

[0140] S64, deploying the optimally configured noise pulse neural network to the furniture assembly task scenario, and executing steps S1-S4, such as Figure 3As shown, the robot arm in the task scenario is guided to complete the task of assembling furniture.

[0141] So far, through the method proposed in the present invention, a single robotic arm furniture assembly motion planning method based on noise pulse hybrid reinforcement learning, a low-energy strategy value network for robotic arm furniture assembly tasks has been constructed; this method significantly reduces energy consumption when performing mechanical motion planning by designing a noise pulse neural network to replace the traditional deep network as the strategy network; at the same time, due to the addition of noise, this method can also make diversified decisions in complex environments; in this way, this study can build an action generation system with the best energy efficiency for the robotic arm, so that it can maintain continuous operation capabilities in resource-constrained scenarios, and promote the practical development of intelligent equipment in the field of intelligent manufacturing.

[0142] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

[0143] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A single robot arm furniture assembly motion planning method based on noise pulse hybrid reinforcement learning, characterized in that: The specific steps include: S1. Collect observation information of the robot arm in the furniture assembly task scenario, build a strategy-value network to guide the robot arm to assemble furniture, and design a noise pulse neural network as a strategy network; The noise pulse neural network includes an encoding layer, a pulse neural layer and a decoding layer; S2. Design the encoding layer based on population coding to convert the observed information into a deterministic pulse sequence. The pulse sequence is a set of discrete time pulses generated by neurons during the encoding process, which represents the time and intensity characteristics of the observed information. S3, designing a pulse neural layer with noise, and outputting the final pulse activity according to the pulse sequence obtained in step S2, where the pulse activity is the overall dynamic behavior of neurons in the neural layer, reflecting the activation state of neurons; S4, then through the noisy non-pulse decoding layer, the input pulse activity is converted into action expression; S5. Design a method to dynamically adjust the noise level. By modifying the loss function, the noise can be dynamically adjusted to improve the training efficiency while ensuring full exploration. S6. Based on the deep reinforcement learning framework SAC, the parameters of the noise pulse neural network are optimized to obtain the optimal model; the trained optimal model is used to guide the robotic arm to complete the furniture assembly task.

2. According to the single robot arm furniture assembly motion planning method based on noise pulse hybrid reinforcement learning according to claim 1, it is characterized in that: The observation information in step S1 is specifically: The robot arm state r and furniture component state f at each time step t; the observation information o is expressed by the formula: o={r(h1,h2;g1;e1,e2,e3,e4),f(p n ,q n )}; Among them, h1 and h2 represent the position and velocity of the robot joint respectively; g1 represents the position of the robot gripper; e1~e4 represent the position, quaternion, linear velocity and angular velocity of the robot end effector respectively; p n ,q n are the 3D coordinates and orientation of the nth furniture component respectively.

3. According to the single robot arm furniture assembly motion planning method based on noise pulse hybrid reinforcement learning in claim 1, it is characterized in that: Step S2 specifically includes: S21, observation data o for observation information o at time i i Any data o of dimension j in ij , the activation value p of the observation data of the jth dimension in the population neurons is calculated by the Gaussian function j , the formula is: Among them, u j It represents the response center of the observed data in the jth dimension, σ j It represents the response width of the observation data in the jth dimension; S22. At each time step t, the neuron's membrane potential cumulative population activation value p j , and then update the membrane potential state; the update rules are as follows: v j (t)=v j (t-1)+p j ; v j (t)=v j (t)-V th ·s j (t); Among them, v j (t) represents the neuron membrane potential state related to the j-th dimension observation data at the current time t, which is initially 0; To determine whether the voltage accumulation exceeds the threshold V th If the threshold is exceeded, it takes 1, otherwise it takes 0, s j (t) represents the generated pulse train.

4. According to the single robot arm furniture assembly motion planning method based on noise pulse hybrid reinforcement learning according to claim 1, it is characterized in that: Step S3 specifically includes: S31. Design a pulse neural layer with a circuit leakage-charge-discharge working principle with noise, in which the current of any pulse neuron Update according to the following rules: Among them, δ represents the current decay rate, X t represents external input; l represents the layer number of the neural network where the neuron is located, and m represents the mth neuron in the layer; S32. If noise is introduced into the neuron charging process, the charging voltage of any noisy spiking neuron will be Update according to the following rules: Where v represents the membrane voltage after discharge, η represents the voltage decay rate, and σ v represents the noise parameter, ε v represents noise extracted from any random process; S33, according to the updated charging voltage, any pulse neuron pulse signal Update and issue according to the following rules: Among them, Θ(·) processes each element in the input charging point cloud, and if it exceeds the threshold V th , then it takes 1, otherwise it takes 0; σ s represents the noise parameter, ε s represents noise extracted from any random process; S34, after the pulse signal is released, the membrane potential is reset, and the updating rule of the membrane potential v(t) after discharge is as follows: Among them, V rest Represents the resting potential.

5. The single robot arm furniture assembly motion planning method based on noise pulse hybrid reinforcement learning according to claim 1 is characterized in that: Step S4 specifically includes: S41. Design a neuron with a non-pulse integration-discharge working principle that injects noise as the decoding layer; at each time step t, the membrane potential v of the rth neuron in the decoding layer r (t) Update according to the pulse signal sent by the pulse neural layer. The update rules are as follows: v r (t)=v r (t-1)+x t ; Among them, x t It represents the pulse signal input of the current time step; if the model is currently in training mode, noise is also added to the input. After adding noise, the update method of the membrane potential is: v r (t)=v r (t-1)+x t +σ s ⊙ε s ; Among them, σ s represents the noise parameter, ε s represents noise extracted from any random process; S42, through the membrane potential of neurons in the decoding layer [v r1 ,v r2 ,…,v rT ] to obtain intuitive action expression. The specific decoding method is to calculate the average membrane potential of the entire time series, which is expressed as: Among them, out represents the obtained action expression.

6. The single-manipulator furniture assembly motion planning method based on noise pulse hybrid reinforcement learning according to claim 5 is characterized in that: The action expression obtained after decoding in step S42 is specifically: The robot action a={a1,a2,…,a9} at each time step t; Among them, a1~a7 represent the speeds of the seven robot arm joints respectively; a8 represents selection, that is, selecting one furniture part from many parts to grasp; a9 represents connection, which is used to represent the interaction with the environment and objects.

7. The single-manipulator furniture assembly motion planning method based on noise pulse hybrid reinforcement learning according to claim 1 is characterized in that: Step S5 specifically includes: S51. Use the constructed noise pulse neural network as the policy network and calculate its original loss L old ; Then the sum of the original loss and the noise variance of the non-spiking neurons in the decoding layer is used as the new loss function L of the policy network new , the formula is: Among them, N A represents the action dimension, σ i is the noise standard deviation of the i-th non-spiking neuron, k is the coefficient that controls the update ratio between the original gradient update direction and the noise reduction direction, and its value is dynamically adjusted according to the current reward and task; S52. When the reward value is low, a smaller k value encourages the robot to use noise to explore new strategies; when the reward value is high, a larger k value controls the robot to reduce the use of noise and focus on using existing strategies to obtain greater rewards; the formula for k is: Among them, k0 is the factor that controls the size of the loss term; R mean Represents the average reward value after multiple trainings; R max Indicates the maximum reward value; R min Indicates the minimum reward value.

8. The single robot arm furniture assembly motion planning method based on noise pulse hybrid reinforcement learning according to claim 1 is characterized in that: Step S6 specifically includes: S61. Use the double Q network to minimize the loss function to train the value network; the formula is: L critic =(Q target -Q real1 ) 2 +(Q target -Q real2 ) 2 ; Among them, Q target Represents the prediction of future rewards, which is calculated by the target network and the policy network; Q real1 and Q real2 are two different value estimates of the same state-action pair by the value network; Q target The calculation formula is as follows: Q target =γ·(1-d t )·(minC target (the t+1 ,to t+1 )-α·logπ(or t+1 ))+r t 4 Where γ is the discount factor for future rewards, d t Is a Boolean value; if d t =1, it means the current step is the final state, otherwise d t =0; C target (o t+1 ,a t+1 ) is the target network for the next state o t+1 and action a t+1 The estimated Q value of ; α is the temperature parameter for adjusting the entropy of the strategy; logπ(o t+1 ) is the policy entropy of the next state; r t is the reward at time t; S62. The policy network in the deep reinforcement learning framework is composed of a noise pulse neural network, while the value network is composed of two MLP instances, each of which consists of three fully connected layers; S63. Randomly initialize the parameters of the entire reinforcement learning framework, calculate the gradient of the loss with respect to the policy network parameters, and use the Adam optimizer to noise-spike the weights of the neural network. S64. Deploy the optimally configured noise pulse neural network to the furniture assembly task scenario, execute steps S1-S4, and guide the robot arm in the task scenario to complete the furniture assembly task.

Citation Information

Patent Citations

  • Robot self-adaptive grabbing control method and system based on pulse neural network

    CN113743287A

  • Pulse neural network simulation method based on GPU

    CN114186665A

  • Single mechanical arm motion planning method based on pulse mixing reinforcement learning assembly task

    CN118438457A

  • Mechanical arm grabbing strategy optimization method based on time sequence task continuous reinforcement learning

    CN118578396A

  • Impulse noise parameter estimation method based on hybrid neural network

    CN119341658A