Mechanical command control method and system based on sample learning
By combining a dual-channel heterogeneous neural network and a safety assessment network, safe movements of the robotic arm are generated, solving the problem of performance degradation of the robotic arm in complex tasks, realizing flexible and efficient control, and ensuring reliability and safety in different environments.
Patent Information
- Application Number
- CN202511211919.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Robotic arms are prone to performance degradation in complex tasks, making it difficult to achieve flexible control and resulting in inefficient and unreliable performance in different people and environments.
The robot arm is controlled by a dual-channel heterogeneous neural network that observes its status in real time. It combines a safety assessment network and an IQL-attention mechanism to generate the final safe action through historical action extraction and exploration action generation channels. Dynamic hybrid control is then used to control the robot arm.
This avoids performance degradation of the robotic arm in complex situations, achieves flexible and efficient control, and ensures reliability and safety in different environments.
Smart Images

Figure CN120715911B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robotics and upper limb rehabilitation technology, and in particular to a mechanical command control method and system based on sample learning. Background Technology
[0002] In recent years, reinforcement learning (RL) has made significant progress in the field of robot control, especially in the precise manipulation and task execution of robotic arms. Traditional control theory-based robotic arm control methods typically rely on accurate physical modeling and strict environmental assumptions, while data-driven reinforcement learning methods can learn from large amounts of sample data without relying on accurate modeling. The actor-critic method of reinforcement learning has shown high performance in this scenario. However, off-policy methods face challenges in terms of sample efficiency and the exploration-exploitation balance, especially prone to performance degradation in complex tasks.
[0003] On the other hand, in the actual control of robotic arms, sample-based instruction control methods are gradually replacing traditional rule-based and model-driven methods. This method can utilize a large amount of collected sample data and use reinforcement learning algorithms to enable the robotic arm to autonomously learn efficient instruction control strategies, thereby achieving flexible operation control. Especially in applications requiring high precision and adaptability, sample-based learning methods enable robotic arms to perform more efficiently and reliably in different tasks and environments.
[0004] In summary, existing technologies often suffer from performance degradation in complex tasks and struggle to achieve flexible control of robotic arms, resulting in inefficient and unreliable performance of robotic arms in different people and environments. Summary of the Invention
[0005] Therefore, the purpose of this invention is to provide a mechanical command control method and system based on sample learning, so as to overcome the shortcomings of the prior art.
[0006] In a first aspect, the present invention provides a mechanical command control method based on sample learning, the method comprising:
[0007] Establish an environmental model for the robotic arm;
[0008] The robotic arm's own state is observed in real time based on a dual-channel heterogeneous neural network, and the robotic arm action in the current state is selected to interact with the environment model to obtain a reward value. The dual-channel heterogeneous neural network includes a historical action extraction channel, an exploration action generation channel, and a dynamic mixing layer.
[0009] A security assessment network based on a neural network is constructed, and a security threshold is generated based on the security assessment network, and the security threshold is incorporated into the reward value;
[0010] The robot arm's historical movements are output using IQL combined with an attention mechanism;
[0011] The optimal action in the historical action extraction channel is selected by combining the implicit strategy in the historical action extraction channel with the attention mechanism, and the final action fusion output is obtained by dynamic hybrid control.
[0012] Candidate actions are generated based on the reward value, and the fusion output of the final action is combined to generate the safe action of the robotic arm.
[0013] Compared with the prior art, the beneficial effects of the present invention are: by using the reward value obtained by the interaction between the current state of the robotic arm action and the environment model to transform the current state to the next state of the robotic arm, the performance degradation of the robotic arm control under complex conditions can be avoided. Furthermore, by using the optimal action from the historical actions and obtaining the fusion output of the final action through dynamic hybrid control, the robotic arm control can be performed flexibly and efficiently, avoiding unreliable situations in different environments.
[0014] Furthermore, the environment model includes the state of the robotic arm, the movements of the robotic arm, and the target position.
[0015] Furthermore, the step of observing the robotic arm's own state in real time based on a dual-channel heterogeneous neural network includes:
[0016] The joint torques and end effector force data of the robotic arm are extracted based on the multimodal perception layer and the dual-channel heterogeneous neural network.
[0017] By combining the attention mechanism in the historical action extraction channel and IQL to extract the robotic arm action from the replay buffer, the robotic arm's own state is obtained based on the joint torque of the robotic arm, the force data of the end effector, and the robotic arm action.
[0018] Furthermore, the expression for the security assessment network is:
[0019] ;
[0020] In the formula, Indicates the risk value of the action. This represents the sigmoid function. Indicates the splicing status. Indicates an action, This represents the concatenation of vectors representing states and actions. This represents the weight matrix of the security assessment network. This represents the bias term for the security assessment network.
[0021] Furthermore, the step of obtaining the fused output of the final action through dynamic hybrid control includes:
[0022] The weights are calculated by dynamically adjusting the mixing coefficients based on Bellman error and environmental change rate to obtain the final fused output of the action.
[0023] Secondly, the present invention also provides a mechanical command control system based on sample learning, the system comprising:
[0024] A module is created to build the environment model of the robotic arm;
[0025] An observation module is used to observe the state of the robotic arm in real time based on a dual-channel heterogeneous neural network, and select the current state of the robotic arm action in its own state to interact with the environment model to obtain a reward value. The dual-channel heterogeneous neural network includes a historical action extraction channel, an exploration action generation channel, and a dynamic mixing layer.
[0026] A construction module is used to construct a security evaluation network based on a neural network, generate a security threshold based on the security evaluation network, and incorporate the security threshold into the reward value;
[0027] The output module is used to output the historical movements of the robotic arm using IQL combined with an attention mechanism;
[0028] The filtering module is used to filter out the optimal action from the historical actions by combining the implicit strategy in the historical action extraction channel with the attention mechanism, and to obtain the final action fusion output by dynamic hybrid control.
[0029] The generation module is used to generate candidate actions based on the reward value, and combine the fusion output of the final action to generate the safe action of the robotic arm.
[0030] Furthermore, the observation module includes:
[0031] The first extraction unit is used to extract the joint torque and end effector force data of the robotic arm based on the multimodal perception layer and through the dual-channel heterogeneous neural network;
[0032] The second extraction unit is used to extract the robotic arm's movements from the replay buffer by combining the attention mechanism in the historical action extraction channel and IQL, and to obtain the robotic arm's own state based on the joint torque of the robotic arm, the force data of the end effector, and the robotic arm's movements.
[0033] Furthermore, the filtering module includes:
[0034] The calculation output unit is used to perform weight calculations by dynamically adjusting the mixing coefficients based on Bellman error and environmental change rate to obtain the final fused output of the action.
[0035] Thirdly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described sample-based mechanical instruction control method.
[0036] Fourthly, the present invention also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described sample-based mechanical instruction control method. Attached Figure Description
[0037] Figure 1 This is a flowchart of the mechanical command control method based on sample learning in the first embodiment of the present invention;
[0038] Figure 2 This is a structural block diagram of a mechanical command control system based on sample learning in the second embodiment of the present invention;
[0039] Figure 3 This is a schematic diagram of the hardware structure of the electronic device in the third embodiment of the present invention.
[0040] Explanation of key component symbols:
[0041] 10. Establishment Module; 20. Observation Module; 30. Construction Module; 40. Output Module; 50. Filtering Module; 60. Generation Module;
[0042] 70. Bus; 71. Processor; 72. Memory; 73. Communication interface.
[0043] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation
[0044] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.
[0045] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.
[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0047] Example 1
[0048] Please see Figure 1 The image shows a sample-based mechanical command control method according to the first embodiment of the present invention, the method comprising steps S1 to S6:
[0049] S1, Establish the environmental model of the robotic arm;
[0050] Understandably, an environmental model of the robotic arm is established, which includes the state of the robotic arm, the movements of the robotic arm, and the target position.
[0051] It is worth noting that the current state of the robotic arm and the user commands constitute the state space of the strategy, including the position and speed of the robotic arm, as well as environmental constraints. Environmental constraints include the rotational speed of the motors, the direction of the end effector represented by 6ternions, the positions and velocities of activated and deactivated joints that can be measured in hardware, the user command input, the steering rate input, and the relative offset of the current end effector position from its position at the next moment in polar coordinates. It is also worth noting that in this embodiment, the positions and velocities of activated and deactivated joints that can be measured in hardware include the lengths and angles of the robotic arm's five motors and six links, as well as the acceleration of the five motors; the user command input represents the required swing support ratio; and the steering rate input is the absolute angle by which the robotic arm deviates from its target position.
[0052] It is worth noting that in this embodiment, the motion space includes motor commands, end effector control commands, motion sequence commands, and upper-level motion strategies updated at a frequency of 40Hz, which means that the control strategy is updated 40 times per second. The lower-level motions are updated at a fixed frequency of 200Hz and are updated in conjunction with the upper-level controller. The upper-level controller updates the strategy at a limit of 40, but the lower-level actuators execute the actual joint control actions at a higher frequency to ensure smooth joint movement and stability.
[0053] Among them, the motor commands include angle control and speed control, the end effector control commands include position control and end attitude control, and the action sequence commands are for the robotic arm motors to execute each action in the corresponding order.
[0054] S2, the robot arm's own state is observed in real time based on a dual-channel heterogeneous neural network, and the robot arm's current state action is selected from its own state to interact with the environment model to obtain a reward value. The dual-channel heterogeneous neural network includes a historical action extraction channel, an exploration action generation channel, and a dynamic mixing layer.
[0055] Specifically, step S2 includes steps S21 to S22:
[0056] S21, Based on the multimodal perception layer and through the dual-channel heterogeneous neural network, extract the joint torque and end effector force data of the robotic arm;
[0057] S22, combine the attention mechanism in the historical action extraction channel and IQL to extract the robotic arm action from the replay buffer, and obtain the robotic arm's own state based on the joint torque of the robotic arm, the force data of the end effector and the robotic arm action;
[0058] It should be noted that the dual-channel heterogeneous neural network consists of three parts: a historical action extraction channel, an exploratory action generation channel, and a dynamic mixing layer. The historical action extraction channel extracts high-quality historical actions based on historical experience, the exploratory action generation channel generates preferred and safe actions by trying to generate new actions, and the dynamic mixing layer mixes the current state and preferred and safe actions to output the final action.
[0059] Understandably, during implementation, the process involves: inputting joint torques of the robotic arm and force data from the end effector through a multimodal perception layer for feature extraction; generating a dual-channel strategy by extracting actions from the replay buffer using an attention mechanism and IQL method in the historical action channel; exploring the action channel by generating target-oriented noise or a regularization term based on entropy; implementing hybrid control decision-making by adjusting weights through three mechanisms: environmental change perception, strategy entropy calculation, and safety constraint coverage; conducting safety verification and execution by evaluating whether the action threshold deviates from the safety boundary using a risk prediction model, discarding the action if it does, and sampling the nearest safe action from the replay buffer; and implementing closed-loop learning by incorporating the calculation of the action risk value for safety verification into the reward function and reconstructing the function, sampling using a priority-based experience replay method during online learning.
[0060] During real-time control, the historical channel runs on the CPU core (Intel TBB parallel computing), the exploration channel is deployed on the GPU (CUDA accelerates noise generation), and the hybrid controller uses FPGA to achieve hard real-time (<100μs response).
[0061] It's worth noting that in the historical action extraction channel, using the IQL method combined with an attention mechanism to output historical actions can effectively capture the temporal dependencies and key information in the sequence of actions, thereby enhancing the understanding and representation of historical actions. The core of the attention mechanism is calculating the importance weight of each time step in the historical action sequence. Given a historical action sequence... and current state , , , , Let T represent the action at time 1, the action at time 2, and the action at time T in the sequence, respectively, where T is the length of the historical action sequence.
[0062] The attention weights are calculated as follows: First, the current state is... And every historical action Mapped to the feature space, the state encoding is represented as: Action coding is represented as: ,in This indicates the use of a multilayer perceptron network. The attention score is obtained by calculating the similarity between the query and the key: Query... ,in It is a learnable weight matrix; the key ,in It is a learnable weight matrix; attention score Where d is the vector dimension, Represents the query vector, derived from the current state features. Through the weight matrix The mapping is used to measure the correlation between the current state and historical actions; Represents the key vector, given by the first... A historical action Features Through the weight matrix The mapping is obtained, and the query vector is obtained. Similarity is calculated to generate an attention score. T represents the length of the historical action sequence, i.e., the total number of actions contained in the sequence. Scaling factor. To mitigate gradient vanishing, the attention score is converted into a probability distribution using a softmax function. ,in Indicates historical actions Importance weights. In the operator, the input to the Q-value function integrates the current state and attention-weighted historical action representations, utilizing attention weights. Weighted fusion of historical action features: The current state features are concatenated with historical action representations and then input into the Q-network. , where [;] denotes vector concatenation. Two Q-networks are used in the IQL method to mitigate overestimation. In the action selection strategy, this is implicitly achieved by maximizing the Q-value. The strategy samples actions based on the Q-value using a Boltzmann distribution, typically selecting the action with the largest Q-value during the inference phase. An extended form of this attention mechanism is also considered, taking into account the temporal dependencies of historical actions and introducing positional encoding, which can be a sine function or learnable parameters. Multiple attention heads can also be introduced for parallel computation to capture relationships between different subspaces.
[0063] It should be noted that the reward value settings include the position of the robotic arm, the speed of the robotic arm, the acceleration of the robotic arm, and the motor offset angle, expressed as:
[0064] ;
[0065] In the formula, Represents the reward function, This indicates the distance between the end effector of the robotic arm and the target position; This indicates the speed of the end effector of the robotic arm; This indicates the acceleration of the end effector of the robotic arm; A penalty term indicating the motor's offset angle; , , , These represent the weights of position, velocity, acceleration, and offset angle, respectively, used to balance the impact of each factor on the reward.
[0066] At the same time, in order to adapt to actual robotic arm scenarios, while considering Simultaneously, the orientation conditions of the robotic arm's end effector are considered. Furthermore, joint parameters, damping, friction, and the target position within the workspace are randomized to simulate various environmental disturbances. This is specifically represented as follows:
[0067] ;
[0068] in, With simulation environment Consistent, only adding the time dimension to the representation; This indicates the upright reward of the robotic arm, which is used to detect whether the posture of the end effector is stable and to ensure that the end effector is within a reasonable workspace. This indicates the alignment between the end effector's orientation and the target orientation, measured by calculating the angle between the target orientation and the actual orientation, specifically expressed as... , Indicates the included angle; Indicating a successful task reward: (1) If the robotic arm reaches the target area but is not fully aligned with the orientation, a smaller reward will be given. (2) If the robotic arm not only reaches the target but also successfully aligns with the target orientation, a larger reward will be given. .
[0069] S3, construct a security evaluation network based on a neural network, generate a security threshold based on the security evaluation network, and incorporate the security threshold into the reward value;
[0070] The expression for the security assessment network is:
[0071] ;
[0072] In the formula, Indicates the risk value of the action. This represents the sigmoid function. Indicates the splicing status. Indicates an action, This represents the concatenation of vectors representing states and actions. This represents the weight matrix of the security assessment network;
[0073] Adjustable parameters learned during neural network training are used to perform a linear transformation on the input concatenated state-action vector, extracting risk-related features between states and actions. This represents the bias term of the safety assessment network, which is also a learnable parameter used to adjust the output offset after linear transformation, helping the safety assessment network fit the mapping relationship of action risk values. The output is compressed into the [0,1] interval using the sigmoid function to obtain the action risk value m.
[0074] The motion risk value obtained by fitting the mapping relationship of motion risk value through a safety assessment network can evaluate variables and thus determine whether the motion of the robotic arm is within the safe range.
[0075] Understandably, the reward value can transition the state to the next moment of the robotic arm's state, thus avoiding the performance degradation of the robotic arm control in complex situations.
[0076] It's important to explain that the safety threshold represents the boundary between safety and danger for the robotic arm, while the reward value serves as an incentive. Incorporating the safety threshold into the reward value, through a combination of punishment and incentive mechanisms, transforms behaviors that violate the safety threshold into negative incentives, thus enabling the robotic arm to choose safer actions. It's worth noting that safety-related variables are extracted from the robotic arm's environmental state, and evaluated through a safety assessment network to derive the safety threshold. In the dual-channel heterogeneous neural network, reward calculation occurs during the experience generation phase and influences the training of the historical action extraction channel and the exploratory action generation channel. By incorporating the safety threshold into the reward value to obtain a safe reward value, potential safety risks are avoided, leading to safer, more reliable, and more robust decision-making.
[0077] S4, using IQL combined with an attention mechanism to output the historical movements of the robotic arm;
[0078] It should be noted that IQL is an algorithm specifically designed for offline reinforcement learning. It learns from a fixed and pre-collected dataset, which contains a subset of state-action pairs. By combining IQL with an attention mechanism, it selects stable actions to output historical actions, thus ensuring the relative quality of historical actions.
[0079] S5, the optimal action in the historical action extraction channel is selected by combining the implicit strategy in the historical action extraction channel with the attention mechanism, and the final action fusion output is obtained by dynamic hybrid control.
[0080] It should be noted that implicit strategy refers to an indirect comparison method, which selects the optimal action for the current moment from historical actions by comparing the action content in the historical actions.
[0081] Specifically, step S5 includes step S51:
[0082] S51 uses a dynamic adjustment of the mixing coefficients based on Bellman error and environmental change rate to calculate the weights in order to obtain the final fusion output of the action;
[0083] Understandably, the historical action channel processing receives state information s and employs IQL-attention joint computation: implicit Q-learning (IQL) combined with an attention mechanism is used to select the best historical action, and historical actions are weighted according to attention weights. In the exploration action channel processing, after receiving state information s, generative gradient-sensitive noise can be added, or the exploration intensity can be adjusted using a SAC entropy term. Other optional embodiments also provide selection methods for other entropy terms, such as state entropy and state-action occupancy entropy. Dynamic hybrid control is the core decision-making module, using dynamically adjusted hybrid coefficients based on Bellman error and environmental change rate for weight calculation; finally, the fused action output is obtained.
[0084] The historical action channel extracts high-value historical actions through an attention mechanism, while the exploration channel generates gradient-sensitive noise. The two operate independently to meet real-time requirements (40Hz update frequency). IQL calculation for the historical channel can be deployed on the CPU, while noise generation for the exploration channel is accelerated by the GPU. The hybrid controller achieves microsecond-level response through the FPGA.
[0085] S6, Generate candidate actions based on the reward value, and combine the fusion output of the final action to generate the safe action of the robotic arm;
[0086] It should be noted that when a robotic arm performs training tasks, it needs to continuously interact with the human body and the training environment through command control. The safety of its movements directly affects the user's personal safety and the effectiveness of rehabilitation training. Therefore, before the system officially interacts with the environment, strict safety constraints must be used to pre-assess and control the risks of the movements.
[0087] Understandably, the raw motion commands output by the robotic arm's intelligent agent policy network include information such as joint angles and speed control values. The input is a concatenated vector of the current state and candidate actions. A safety assessment network calculates the risk value, and the obtained safe action risk value is compared with a threshold comparator to determine whether the current action meets the normal threshold and is a safe action. The nearest neighbor safe action is searched in the replay buffer, prioritizing safe actions that have a historical risk value < 0.3 and a deviation from the current target position < 5cm. Finally, the physical constraints of the action are output. Execution is carried out at a frequency of 200Hz by the underlying DSP chip, and the historical best action is fused with the current strategy during safe action generation to achieve a smooth transition.
[0088] In practical implementation, a security assessment network is constructed to predict the risk value of actions, expressed by the formula as follows: ,in This represents the sigmoid function, with an output range of [0,1]. The calculation of the action risk value *m* is essentially a quantitative estimate of the degree of danger of the state-action pair. Based on the neural network-based risk assessment function, at the input layer, the states are concatenated... and actions In the environment configuration section, in this embodiment, the state is set. , Represents the state vector. , This represents the action vector. In the hidden layers of the multilayer perceptron, a nonlinear transformation is performed using weight matrices and biases, followed by... The function compresses values to the [0,1] interval. Experiments showed that a threshold m=0.7 is used as the boundary between safety and danger. Simultaneously, this threshold is incorporated into the reward function:
[0089] ,in, This represents the risk penalty / reward term, which is the negative reward in the reward function used to constrain the safety of actions. This indicates the actual risk value of the action. Because it primarily targets the command control of rehabilitation robots, the system is more inclined to trigger a safety retreat mechanism, making it more suitable for high-risk scenarios. When the safety retreat is triggered, the corresponding action... Represented as:
[0090] , This indicates a safe action, which is an action that meets the risk threshold requirements and is selected through a safety verification mechanism. This is the action that the robotic arm ultimately performs. This represents candidate security actions retrieved from the replay buffer, used for comparison and filtering against actions generated by the current policy. Indicates candidate actions The risk value (corresponding to m in the output of the security assessment network) is used to measure the security of candidate actions, and D represents the replay buffer.
[0091] In the collaborative mechanism of the dual-channel architecture, the historical channel uses the replay buffer state-action density formula as the prior distribution for attention calculation:
[0092] ;
[0093] Let represent the probability distribution of choosing action 'a' in state 's', used to describe the robot arm's action selection strategy. This indicates proportionality, used to simplify the expression of probability distributions, i.e. The value of and The values of are directly proportional. Represents the action value function. State-action occupancy density is used to measure policy. Next state and actions The access probability reflects the importance of that state-action pair in historical experience. The exploration channel achieves dynamic control of the exploration intensity through the linkage adjustment of the entropy regularization term and the noise parameter.
[0094] In summary, the sample-based mechanical command control method in the above embodiments of the present invention transforms the current state to the next state of the robotic arm by using the reward value obtained from the interaction between the current state of the robotic arm action and the environment model. This avoids the performance degradation problem of robotic arm control in complex situations. Furthermore, by using the optimal action from historical actions and obtaining the fused output of the final action through dynamic hybrid control, the robotic arm can be controlled flexibly and efficiently, avoiding unreliable situations in different environments.
[0095] Example 2
[0096] Please see Figure 2 The figure shows a sample-learning-based mechanical command control system according to a second embodiment of the present invention. The system includes:
[0097] Module 10 is used to create the environment model of the robotic arm;
[0098] The observation module 20 is used to observe the state of the robotic arm in real time based on a dual-channel heterogeneous neural network, and select the current state of the robotic arm action in its own state to interact with the environment model to obtain a reward value. The dual-channel heterogeneous neural network includes a historical action extraction channel, an exploration action generation channel, and a dynamic mixing layer.
[0099] Module 30 is used to construct a security evaluation network based on a neural network, generate a security threshold based on the security evaluation network, and incorporate the security threshold into the reward value. The expression of the security evaluation network is:
[0100] ;
[0101] In the formula, Indicates the risk value of the action. This represents the sigmoid function. Indicates the splicing status. Indicates an action, This represents the concatenation of vectors representing states and actions. This represents the weight matrix of the security assessment network. This represents the bias term for the security assessment network;
[0102] Output module 40 is used to output the historical movements of the robotic arm using IQL combined with an attention mechanism;
[0103] The filtering module 50 is used to filter out the optimal action from the historical actions by combining the implicit strategy in the historical action extraction channel with the attention mechanism, and to obtain the final action fusion output by dynamic hybrid control.
[0104] The generation module 60 is used to generate candidate actions based on the reward value, and combine the fusion output of the final action to generate the safe action of the robotic arm.
[0105] In some alternative embodiments, the observation module 20 includes:
[0106] The first extraction unit is used to extract the joint torque and end effector force data of the robotic arm based on the multimodal perception layer and through the dual-channel heterogeneous neural network;
[0107] The second extraction unit is used to extract the robotic arm's movements from the replay buffer by combining the attention mechanism in the historical action extraction channel and IQL, and to obtain the robotic arm's own state based on the joint torque of the robotic arm, the force data of the end effector, and the robotic arm's movements.
[0108] In some alternative embodiments, the filtering module 50 includes:
[0109] The calculation output unit is used to perform weight calculations by dynamically adjusting the mixing coefficients based on Bellman error and environmental change rate to obtain the final fused output of the action.
[0110] The functions or operation steps implemented by the above modules and units are largely the same as those in the above method embodiments, and will not be repeated here.
[0111] The mechanical command control system based on sample learning provided in this embodiment of the invention has the same implementation principle and technical effects as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the system embodiment can be referred to the corresponding content in the aforementioned method embodiment.
[0112] Example 3
[0113] The present invention also provides an electronic device, please refer to [link / reference]. Figure 3 The figure shown is a schematic diagram of the hardware structure of the electronic device in the third embodiment of the present invention.
[0114] The electronic device may include a processor 71 and a memory 72 storing computer program instructions.
[0115] Specifically, the processor 71 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement this application.
[0116] The memory 72 may include a mass storage device for data or instructions. For example, and not limitingly, the memory 72 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 72 may include removable or non-removable (or fixed) media. Where appropriate, the memory 72 may be internal or external to a data processing device. In a particular embodiment, the memory 72 is non-volatile memory. In a particular embodiment, the memory 72 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), an electrically alterable read-only memory (EAROM), or flash memory, or a combination of two or more of these. Where appropriate, the RAM can be Static Random-Access Memory (SRAM) or Dynamic Random-Access Memory (DRAM). DRAM can be Fast Page Mode Dynamic Random-Access Memory (FPMDRAM), Extended Data Out Dynamic Random-Access Memory (EDODRAM), Synchronous Dynamic Random-Access Memory (SDRAM), etc.
[0117] The memory 72 can be used to store or cache various data files that need to be processed and / or communicated, as well as possible computer program instructions executed by the processor 71.
[0118] The processor 71 reads and executes the computer program instructions stored in the memory 72 to implement the sample learning-based mechanical instruction control method of Embodiment 1 described above.
[0119] In some embodiments, the electronic device may further include a communication interface 73 and a bus 70. For example, Figure 3 As shown, the processor 71, memory 72, and communication interface 73 are connected through bus 70 and complete communication with each other.
[0120] The communication interface 73 is used to enable communication between the various modules, devices, units, and / or equipment in this application. The communication interface 73 can also enable data communication with other components such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations.
[0121] Bus 70 includes hardware, software, or both, that couples the components of a device together. Bus 70 includes, but is not limited to, at least one of the following: data bus, address bus, control bus, expansion bus, and local bus. For example, and not as a limitation, bus 70 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 70 may include one or more buses. Although this application describes and illustrates a specific bus, this application considers any suitable bus or interconnection.
[0122] The electronic device can acquire a mechanical command control system based on sample learning and execute the mechanical command control method based on sample learning in this embodiment.
[0123] Furthermore, in conjunction with the sample-based mechanical instruction control method in Embodiment 1 above, this application can provide a storage medium for implementation. This storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement the sample-based mechanical instruction control method of Embodiment 1 above.
[0124] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0125] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A mechanical command control method based on sample learning, characterized in that, The method includes: Establish an environmental model for the robotic arm; The robotic arm's own state is observed in real time based on a dual-channel heterogeneous neural network, and the robotic arm action in the current state is selected to interact with the environment model to obtain a reward value. The dual-channel heterogeneous neural network includes a historical action extraction channel, an exploration action generation channel, and a dynamic mixing layer. A security assessment network based on a neural network is constructed, and a security threshold is generated based on the security assessment network, and the security threshold is incorporated into the reward value; The robot arm's historical movements are output using IQL combined with an attention mechanism; The optimal action in the historical action extraction channel is selected by combining the implicit strategy in the historical action extraction channel with the attention mechanism, and the final action fusion output is obtained by dynamic hybrid control. Candidate actions are generated based on the reward value, and the fusion output of the final action is combined to generate the safe action of the robotic arm.
2. The machine command control method based on sample learning according to claim 1, characterized in that, The environment model includes the state of the robotic arm, the movements of the robotic arm, and the target position.
3. The machine command control method based on sample learning according to claim 1, characterized in that, The step of observing the robotic arm's own state in real time based on a dual-channel heterogeneous neural network includes: The joint torques and end effector force data of the robotic arm are extracted based on the multimodal perception layer and the dual-channel heterogeneous neural network. By combining the attention mechanism in the historical action extraction channel and IQL to extract the robotic arm action from the replay buffer, the robotic arm's own state is obtained based on the joint torque of the robotic arm, the force data of the end effector, and the robotic arm action.
4. The machine command control method based on sample learning according to claim 1, characterized in that, The expression for the security assessment network is: ; In the formula, Indicates the risk value of the action. This represents the sigmoid function. Indicates the splicing status. Indicates an action, This represents the concatenation of vectors representing states and actions. This represents the weight matrix of the security assessment network. This represents the bias term for the security assessment network.
5. The machine command control method based on sample learning according to claim 1, characterized in that, The steps described above for obtaining the final motion fusion output through dynamic hybrid control include: The weights are calculated by dynamically adjusting the mixing coefficients based on Bellman error and environmental change rate to obtain the final fused output of the action.
6. A mechanical command control system based on sample learning, characterized in that, The system includes: A module is created to build the environment model of the robotic arm; An observation module is used to observe the state of the robotic arm in real time based on a dual-channel heterogeneous neural network, and select the current state of the robotic arm action in its own state to interact with the environment model to obtain a reward value. The dual-channel heterogeneous neural network includes a historical action extraction channel, an exploration action generation channel, and a dynamic mixing layer. A construction module is used to construct a security evaluation network based on a neural network, generate a security threshold based on the security evaluation network, and incorporate the security threshold into the reward value; The output module is used to output the historical movements of the robotic arm using IQL combined with an attention mechanism; The filtering module is used to filter out the optimal action from the historical actions by combining the implicit strategy in the historical action extraction channel with the attention mechanism, and to obtain the final action fusion output by dynamic hybrid control. The generation module is used to generate candidate actions based on the reward value, and combine the fusion output of the final action to generate the safe action of the robotic arm.
7. The mechanical command control system based on sample learning according to claim 6, characterized in that, The observation module includes: The first extraction unit is used to extract the joint torque and end effector force data of the robotic arm based on the multimodal perception layer and through the dual-channel heterogeneous neural network; The second extraction unit is used to extract the robotic arm's movements from the replay buffer by combining the attention mechanism in the historical action extraction channel and IQL, and to obtain the robotic arm's own state based on the joint torque of the robotic arm, the force data of the end effector, and the robotic arm's movements.
8. The mechanical command control system based on sample learning according to claim 6, characterized in that, The filtering module includes: The calculation output unit is used to perform weight calculations by dynamically adjusting the mixing coefficients based on Bellman error and environmental change rate to obtain the final fused output of the action.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the machine instruction control method based on sample learning as described in any one of claims 1 to 5.
10. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the machine instruction control method based on sample learning as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Pedestrian accompanying control method and device of robot, mobile robot and medium
CN113467462A
Off-policy control policy evaluation
US20200304545A1