Stacked workpiece grabbing method and system based on SAC mechanical arm vision

By proposing a stacked workpiece grasping method based on SAC robotic arm vision, combining the candidate set of grasping poses and the stacking topology, and using the SAC reinforcement learning algorithm to optimize grasping decisions, the problem of low grasping stability and success rate in stacked workpiece scenarios is solved, and higher robustness and generalization performance are achieved.

CN121893253APending Publication Date: 2026-04-21SUZHOU JIYUAN INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU JIYUAN INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2025-12-30
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies for robotic arm grasping methods in stacked workpiece scenarios lack stability and overall success rate. They struggle to produce reasonable grasping decisions in situations with multiple target candidates, significant occlusion and overlay, and strong spatial interference. Furthermore, traditional reinforcement learning suffers from training difficulties and inappropriate reward design, leading to strategy stagnation or suboptimal results.

Method used

A stacked workpiece grasping method based on SAC robotic arm vision is adopted. By acquiring a candidate set of grasping poses for the workpieces, the grasping suitability score is calculated based on the stacking topology. The SAC reinforcement learning algorithm is used to control the robotic arm to perform grasping actions. The training process is optimized by combining reward mechanism and experience replay mechanism, and grasping consequence evaluation and topological relationship are incorporated.

Benefits of technology

It improves the success rate and stability of grasping in industrial settings, reduces the probability of collisions and grasping failures, enhances the system's robustness and generalization ability in complex environments, and reduces the problem of low training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121893253A_ABST
    Figure CN121893253A_ABST
Patent Text Reader

Abstract

The invention relates to a stacked workpiece grabbing method and system based on SAC mechanical arm vision. The method comprises the steps that S1, a grabbing pose candidate set of each workpiece is obtained according to depth images of stacked workpieces in a material box; s2, on the basis of the current workpiece stacking topology, the grabbing suitability degree score of each workpiece in the pose candidate set is obtained; the grabbing suitability score is used for evaluating the grabability and the grabbing consequence; s3, on the basis of an SAC reinforcement learning algorithm, a mechanical arm is controlled to execute a grabbing action on the workpiece with the highest grabbing suitability degree score; s4, scoring is conducted according to a reward mechanism after the current workpiece is grabbed; the grabbing record of the mechanical arm is stored in an SAC training playback pool, and an SAC reinforcement learning algorithm is updated; and the step S1 to the step S4 are executed repeatedly till the grabbing operation of the workpieces in the material box is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robotic arm control technology, specifically to a method and system for stacked workpiece gripping based on SAC robotic arm vision. Background Technology

[0002] In industrial sorting, depalletizing and loading, and bin handling applications, workpieces are often stacked randomly in bins or workbenches, exhibiting typical characteristics such as random stacking, mutual occlusion, tilted placement, tight compression, and partial suspended support. In these scenarios, the workpiece surface textures are not very different, the boundaries are incomplete, depth data is missing or noisy, and multiple graspable candidates may exist simultaneously in the same frame, making the visual perception results significantly uncertain. Especially when workpieces occlude each other or have overlapping relationships, some targets are only exposed in local areas, leading to increased fluctuations in the confidence of the detection box and grasping pose prediction. The number and quality of candidate grasping poses are difficult to guarantee consistently, thus directly affecting the reliability of subsequent grasping strategies.

[0003] Existing grasping systems mostly adopt a serial process of "perception - candidate generation - target selection - trajectory planning - grasping execution". Target selection is usually based on rules or heuristic strategies (such as selecting the closest, most confident, or closest target to the center), followed by grasping posture and trajectory planning. However, in stacked scenarios, heuristic target selection often fails to explicitly express key constraints such as "being crushed, insufficient accessibility, occlusion of the gripper feed path, and collapse of the upper layer after grasping". It is easy to select workpieces that seem "easy to grasp" but are actually crushed by the upper layer, cannot be entered by the gripper, or will cause instability of the stacked structure once grasped. This leads to failures such as missed grasp, collision, gripper jamming, workpiece slippage, and secondary collision, reducing production line cycle time and increasing equipment wear risk. Meanwhile, the stacked environment makes the contact relationships complex and the local space narrow. The grasping action usually relies on continuous visual planning and control to gradually approach the desired grasping pose. When there are factors such as image noise, calibration error, end effector compliance and control delay, the end trajectory is prone to deviation and accumulation of error, which further increases the probability of collision and grasping failure.

[0004] To enhance adaptability in complex environments, reinforcement learning methods have been introduced into grasping control in recent years. However, training difficulties remain in stacked object scenarios: on the one hand, successful grasping is a typical sparse reward event, and if only "success / failure" is used as feedback, the agent needs a large number of interactions to learn an effective policy; on the other hand, poorly designed process rewards can easily cause the policy to linger in local neighborhoods to "cheat," resulting in slow convergence or the formation of suboptimal policies. In addition, if traditional reinforcement learning does not incorporate "stack relationships and grasping consequences" into the state and target selection mechanism, it often can only learn "how to approach a certain target," but it is difficult to learn "which one should be grasped first," thus failing to achieve a stable overall success rate and robustness under multi-target stacking conditions.

[0005] Therefore, there is an urgent need for a unified method that can simultaneously perform stacked workpiece identification, stacking relationship analysis, optimal grasping target determination, and continuous control, enabling the system to output more reasonable grasping decisions even when multiple target candidates exist, occlusion and overlay are significant, and spatial interference is strong. Furthermore, by combining reinforcement learning algorithms suitable for continuous action spaces, while ensuring exploration capabilities, more effective reward guidance and constraint modeling can improve learning efficiency and execution safety, thereby enhancing the success rate, stability, and generalization ability of stacked workpiece grasping tasks in industrial settings. Summary of the Invention

[0006] In view of this, embodiments of the present invention provide a method and system for grasping stacked workpieces based on SAC robotic arm vision, in order to solve the problems of instability and low overall success rate of existing robotic arm grasping methods for stacked workpieces.

[0007] This invention provides a method for grasping stacked workpieces based on SAC robotic arm vision, including: S1, obtain a candidate set of gripping poses for each workpiece based on the depth image of the stacked workpieces in the bin; S2, based on the current workpiece stacking topology, obtain the grasping suitability score for each workpiece in the pose candidate set; the grasping suitability score is used to evaluate graspability and grasping consequences; S3, based on the SAC reinforcement learning algorithm, controls the robotic arm to perform a grasping action on the workpiece with the highest grasping suitability score; S4. After completing the grasping of the current workpiece, a score is given according to the reward mechanism; the record of this robotic arm grasping is stored in the SAC training playback pool to update the SAC reinforcement learning algorithm. Repeat steps S1 to S4 until the workpiece grabbing operation in the bin is completed.

[0008] Optionally, in S1, the step of acquiring the depth image of the stacked workpieces inside the bin includes: A 6-DOF robotic arm and parallel grippers are used to acquire depth images of stacked workpieces inside the bin by using a hand-eye mounted RGB-D camera. The steps to obtain the candidate set of gripping poses for each workpiece include: Based on the horizontal detection frame B of the i-th workpiece i Based on the category, obtain the set of tilted grab boxes r=r k , where each element r k Includes the center point coordinates, tilt angle, tilt width, and confidence level of the corresponding workpiece; For each element r k Calculate whether the coordinates of its center point fall within the corresponding horizontal detection box B. i If so, then the corresponding element r k The candidate set R to which the corresponding workpiece is assigned i middle.

[0009] Optionally, in S2, the construction of the current workpiece stack topology includes: Depth point cloud data is obtained from the depth image of the stacked workpieces in the bin. Based on the depth point cloud data and occlusion relationships, a stacking topology graph G=(V,E) is constructed, where V is the set of workpiece instances and E contains at least one of the following relationship edges: support / overlay relationship, occlusion relationship and potential interference relationship.

[0010] Optionally, in S2, the construction of the current workpiece stack topology also includes: If the point cloud of workpiece j is above workpiece i, and the projected overlap area exceeds a preset threshold, then record that workpiece j covers workpiece i. If there is an adjacent point cloud occlusion in the gripper feed direction and the minimum distance is less than the preset distance, then the accessibility is determined to be reduced. If, after a workpiece is grasped, the length of the overlying link above it is greater than 0, and the projection of the upper center of gravity may extend beyond the support area, then the stability risk increases.

[0011] Optionally, in S2, the calculation of the gripping suitability score for each workpiece in the pose candidate set includes: S i =ω1*Q i +ω2*L i -ω3*C i -ω4*T i -ω5*O i ; Wherein, the subscript i indicates that the current workpiece is the i-th workpiece; Q i The quality of the pose candidate grasping is indicated by evaluation metrics including candidate confidence, gripper width matching, and angle reasonableness. L iThis indicates the reachability of the crawled data, and the evaluation metrics include inverse kinematics feasibility, feed unobstructedness, and workspace constraint satisfaction. C i The collision risk is indicated by evaluation metrics including the minimum clearance with adjacent workpieces / boxes and the estimated collision trajectory. T i The evaluation indicators for stability risk include the probability estimate of collapse after capture and the level of overburden. O i The degree of occlusion is indicated by evaluation metrics including the proportion of visible area and the proportion of depth loss. The weights ω1 to ω5 are taken as non-negative values ​​and normalized.

[0012] Optionally, in S3, the steps of the SAC reinforcement learning algorithm controlling the robotic arm to perform a grasping action on the workpiece with the highest grasping suitability score include: The initial state s of the SAC reinforcement learning algorithm t Let it be: Horizontal detection box B i eigenvector F i , capture suitability score S i The topological adjacency of workpiece i, the initial end-effector pose / velocity of the gripper, and the gripper opening / closing degree; where the feature vector F i It is based on the candidate set R i The extracted feature vectors include the center point coordinates, tilt box angle, tilt box width, and confidence level. Action a of the SAC reinforcement learning algorithm t Let: The end velocity of the gripper be v = {v x v y v z w x w y w z}, Gripper opening and closing u grip ∈[0,1], target score vector g t ; Set the target selection of the SAC reinforcement learning algorithm as: i * =argmax(g t ).

[0013] Optionally, in S4, the reward mechanism includes: r t =r succ +r fail +r col +r stab +r prog ; If the capture and placement are successful, then r succ The value is increased by 1.0; if the catch misses / falls / times out, then r... failThe value of r is -1.0; if a collision occurs, then r... col The value is -0.2; if it causes significant slippage / fall in the upper layer, then r stab The value is -0.5; the process-guided difference reward r prog =α*(d t-1 -d t ), d t α represents the image similarity or pose error distance, and α is the scaling factor.

[0014] Optionally, in S4, the steps of storing the current robotic arm grasping record into the SAC training playback pool and updating the SAC reinforcement learning algorithm include: Sample (s) t a t r t s t+1 Store it in the replay pool and update the Actor and Critic of the SAC reinforcement learning algorithm; Among them, r t Samples with a value ≥ 0 are stored in the Good pool, and r is... t Samples with a value <0 are stored in the Bad pool. Each update involves sampling at a ratio of Good:Bad=0.8:0.2, and an upper limit for the capacity of each pool is set.

[0015] Optionally, it also includes: The SAC reinforcement learning algorithm's decision system includes a safety action library. When an abnormal state is detected, the reinforcement learning output is paused, an action from the safety action library is executed, and then the reinforcement learning control is switched back.

[0016] This invention also provides a stacked workpiece gripping system based on SAC robotic arm vision, comprising: At least one processor; and, A memory that is communicatively connected to at least one processor; wherein, The memory stores instructions that can be executed by at least one processor, which enables the at least one processor to perform the aforementioned SAC robotic arm vision-based stacked workpiece gripping method.

[0017] The beneficial effects of this invention are: 1. This invention provides a stacked workpiece grasping method based on SAC robotic arm vision. During the target selection phase, a stacking topology is explicitly introduced, and a grasping suitability score for each workpiece is calculated based on the stacking topology. The workpiece stacking topology is incorporated into the decision-making process, transforming target selection from a "nearest / highest confidence" strategy focused solely on local visual indicators to a "task-optimal" strategy that comprehensively considers accessibility, interference risk, and grasping consequences. This improves the safety and cycle stability of continuous operations in industrial settings. This invention also employs the SAC algorithm for continuous motion learning, naturally adapting to the continuous end-effector pose / velocity control requirements in visual scenarios. Simultaneously, entropy regularization maintains the exploratory nature of the strategy during the training phase, avoiding early trapping in conservative or suboptimal action patterns. It maintains good robustness and generalization performance even in the presence of noise, calibration errors, and environmental changes. Compared to discrete motion or solutions heavily reliant on manual planning, this invention can learn smoother, more stable trajectories within a continuous control space, thereby reducing control jitter and error accumulation.

[0018] 2. The embodiments of the present invention incorporate the overlapping, obstruction and potential interference relationships between workpieces into the decision-making basis. By quantifying and punishing or suppressing the risks of upper layer slippage, falling and structural instability after gripping, the probability of secondary collisions, workpiece scattering and gripper jamming during the gripping process can be significantly reduced.

[0019] 3. This invention integrates multi-target workpiece recognition and grasping pose prediction into a joint model, and establishes a correspondence between "workpiece and grasping candidate set" through matching reasoning. This ensures that each workpiece retains multiple feasible grasping pose options even under occlusion and partial visibility conditions. Compared to schemes that only output a single grasping point or rely solely on the center of the detection box, this structure improves the coverage and fault tolerance of candidate grasping poses: when a candidate becomes unavailable due to occlusion, depth loss, or local interference, the system can quickly switch between candidate sets for the same or different targets, increasing the actual graspable ratio and adaptability to complex stacking patterns.

[0020] 4. This invention introduces a difference-guided term in the reward design, enabling the agent to receive consistent and dense feedback even when the grasping process is not yet successful. This significantly alleviates the low training efficiency problem caused by the "sparse success reward" in stacked grasping tasks. By rewarding or penalizing the difference in state evaluation values ​​between adjacent time steps, the strategy tends to continuously reduce the distance to the desired grasping state, decrease the hesitant behavior of repeatedly making small movements in the local neighborhood, and improve the efficiency of reaching the graspable posture area, thereby accelerating the overall convergence speed and improving the final success rate.

[0021] 5. This invention employs an experience replay mechanism to improve sample utilization and reduce online interaction costs. It also introduces a pooling sampling strategy to balance "reinforcement learning from successful experiences" and "constraint learning from failed samples" during the training process. By controlling the ratio of positive to negative samples, it can stably increase the proportion of effective strategies while maintaining continuous penalties for high-risk behaviors such as collisions and collapses. This ensures learning efficiency while improving training stability, reducing policy oscillations, and enhancing the reliability of the final strategy in complex stacking environments. Attached Figure Description

[0022] The features and advantages of the invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the invention in any way. In the drawings: Figure 1 A flowchart of a stacked workpiece grasping method based on SAC robotic arm vision is shown in an embodiment of the present invention; Figure 2 A general flowchart of a stacked workpiece grasping method based on SAC robotic arm vision is shown in an embodiment of the present invention; Figure 3 A structural diagram of a stacked workpiece gripping system based on SAC robotic arm vision is shown in an embodiment of the present invention. Detailed Implementation

[0023] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0024] This invention provides a method for grasping stacked workpieces based on SAC (Soft Actor-Critic) robotic arm vision, such as... Figure 1 As shown, it includes: S1, obtain a candidate set of gripping poses for each workpiece based on the depth image of the stacked workpieces in the bin.

[0025] In this embodiment, a 6-DOF robotic arm and parallel grippers are used, and an RGB-D camera is mounted in a hand-eye configuration to acquire depth images of the stacked workpieces inside the bin. In a specific embodiment, the robotic arm's joint angles, end effector pose, end effector velocity, and gripper opening / closing state are also acquired simultaneously. Before acquiring the images, a calibration relationship is established between the camera coordinate system and the robot's base coordinate system or end effector coordinate system.

[0026] The steps to obtain the candidate set of gripping poses for each workpiece include: Based on the horizontal detection frame B of the i-th workpiece i Get the set of tilted grab boxes r=r k , where each element r kIncludes the center point coordinates, tilt angle, width, and confidence level of the corresponding workpiece; For each element r k Calculate whether the coordinates of its center point fall within the corresponding horizontal detection box B. i If so, then the corresponding element r k The candidate set R to which the corresponding workpiece is assigned i middle.

[0027] Specifically, an example of a candidate association rule is: (1) Only retain the one with the highest confidence in the detection box as B. i (2) For each r k Calculate whether its center point falls within B. i Inside; if it falls in, then r k The candidate set R belonging to workpiece i i (3) If there are too many candidates for the same workpiece, the Top candidates can be selected by confidence level. M For example, M=1.

[0028] In a specific implementation, the depth image is input into a multi-target workpiece recognition network, which outputs the category information and target detection box for each workpiece. A grasping pose prediction network then outputs tilted grasping boxes / grasping pose parameters, forming a candidate set of grasping poses for each workpiece. The candidates are then filtered and matched using inference to obtain the association result of "workpiece-grasping candidate". The grasping pose candidates at least include: the grasping center point (u, v), the grasping angle θ, the gripper opening width W, the grasping confidence q, and / or their equivalent 6D pose representation.

[0029] S2, based on the current workpiece stack topology, obtain the grasping suitability score for each workpiece in the pose candidate set.

[0030] In this embodiment, the grasp suitability score evaluates the graspability and grasping consequences of the workpiece in the current stacking state.

[0031] S3, based on the SAC reinforcement learning algorithm, controls the robotic arm to perform a grasping action on the workpiece with the highest grasping suitability score.

[0032] S4. After completing the grasping of the current workpiece, a score is given according to the reward mechanism; the record of this robotic arm grasping is stored in the SAC training playback pool to update the SAC reinforcement learning algorithm.

[0033] In this embodiment, the replay learning of the SAC algorithm is divided into two stages. The first stage is the model training stage before it is officially put into production. Historical data or training samples are used for training. After the system reaches the expected stability, it is put into production and enters the second stage. In actual use, the grabbing operations and scores are used as samples for the experience replay mechanism to continuously optimize the SAC algorithm.

[0034] Repeat steps S1 to S4 until the workpiece grabbing operation in the bin is completed.

[0035] In this embodiment, a stacking topology is explicitly introduced during the target selection stage, and the grasp suitability score of each workpiece is calculated based on the stacking topology. The stacking topology relationship of the workpieces is incorporated into the decision-making basis, so that the target selection is transformed from focusing only on the "nearest / highest confidence" of local visual indicators to a comprehensive consideration strategy, thereby improving the safety and cycle stability of continuous operation in industrial sites.

[0036] This invention also employs the SAC algorithm for continuous motion learning, naturally adapting to the need for continuous end-effector pose / velocity control in visual scenarios. Simultaneously, it utilizes entropy regularization to maintain the exploratory nature of the strategy during the training phase, avoiding early pitfalls into conservative or suboptimal action patterns. This ensures good robustness and generalization performance even in the presence of noise, calibration errors, and environmental changes. Compared to discrete motion or solutions heavily reliant on manual planning, the SAC-based robotic arm vision-based stacked workpiece grasping method provided by this invention can learn smoother, more stable trajectories within a continuous control space, thereby reducing control jitter and error accumulation.

[0037] As an optional implementation, in S2, the construction of the current workpiece stack topology includes: Depth point cloud data is obtained from the depth image of the stacked workpieces in the bin. Based on the depth point cloud data and occlusion relationships, a stacking topology graph G=(V,E) is constructed, where V is the set of workpiece instances and E contains at least one of the following relationship edges: support / overlay relationship, occlusion relationship and potential interference relationship.

[0038] As an optional implementation, in S2, the construction of the current workpiece stack topology further includes: If the point cloud of workpiece j is above workpiece i, and the projected overlap area exceeds a preset threshold A. th For example, 0.15×min(area) i area j If workpiece j covers workpiece i, then record that workpiece j is pressed against workpiece i. If the gripper feed direction is obscured by a nearby point cloud and the minimum distance is less than the preset distance d occ For example, if the distance is 10mm, then the accessibility is considered reduced. If, after a workpiece is grasped, the length of the overlying link above it is greater than 0, and the projection of the upper center of gravity may extend beyond the support area, then the stability risk increases.

[0039] As an optional implementation, in S2, the calculation of the gripping suitability score for each workpiece in the pose candidate set includes: S i =ω1*Q i +ω2*Li -ω3*C i -ω4*T i -ω5*O i ; Wherein, the subscript i indicates that the current workpiece is the i-th workpiece; Q i The quality of the pose candidate grasping is indicated by evaluation metrics including candidate confidence, gripper width matching, and angle reasonableness. L i This indicates the reachability of the crawled data, and the evaluation metrics include inverse kinematics feasibility, feed unobstructedness, and workspace constraint satisfaction. C i The collision risk is indicated by evaluation metrics including the minimum clearance with adjacent workpieces / boxes and the estimated collision trajectory. T i The evaluation indicators for stability risk include the probability estimate of collapse after capture and the level of overburden. O i The degree of occlusion is indicated by evaluation metrics including the proportion of visible area and the proportion of depth loss. The weights ω1 to ω5 are taken as non-negative values ​​and normalized. For example, ω1=0.35, ω2=0.15, ω3=0.2, ω4=0.2, and ω5=0.1. In a specific embodiment, the weights can be adjusted according to the task's offline parameter tuning or online learning method.

[0040] As an optional implementation, in S3, the step of the SAC reinforcement learning algorithm controlling the robotic arm to perform a gripping action on the workpiece with the highest gripping suitability score includes: The initial state s of the SAC reinforcement learning algorithm t Let it be: Horizontal detection box B i eigenvector F i , capture suitability score S i The topological adjacency of workpiece i, the initial end-effector pose / velocity of the gripper, and the gripper opening / closing degree; wherein, the feature vector F i It is based on the candidate set R i The extracted feature vectors include the center point coordinates, tilt box angle, tilt box width, and confidence level. Action a of the SAC reinforcement learning algorithm t Let: The end velocity of the gripper be v = {v x v y v z w x w y w z}, Gripper opening and closing u grip ∈[0,1], target score vector g t ; Set the target selection of the SAC reinforcement learning algorithm as: i * =argmax(g t In this embodiment, the strategy or rule for target selection can be set to select the candidate with the highest graspability importance and the second highest grasping consequence importance as the desired grasping pose.

[0041] As an optional implementation, in S4, the reward mechanism includes: r t =r succ +r fail +r col +r stab +r prog ; If the capture and placement are successful, then r succ The value is increased by 1.0; if the catch misses / falls / times out, then r... fail The value of r is -1.0; if a collision occurs, then r... col The value is -0.2; if it causes significant slippage / fall in the upper layer, then r stab The value is -0.5; the process-guided difference reward r prog =α*(d t-1 -d t ), d t α represents the image similarity or pose error distance, and α is the scaling factor.

[0042] As an optional implementation, in S4, the step of storing the current robotic arm grasping record into the SAC training playback pool and updating the SAC reinforcement learning algorithm includes: Sample (s) t a t r t s t+1 Store it in the replay pool and update the Actor and Critic of the SAC reinforcement learning algorithm; Among them, r t Samples with a value ≥ 0 are stored in the Good pool, and r is... t Samples with a value <0 are stored in the Bad pool. Each update involves sampling at a ratio of Good:Bad=0.8:0.2, and an upper limit for the capacity of each pool is set.

[0043] In specific applications, such as Figure 2 As shown, the deployment and training completion strategy follows the following system flow: cyclic perception—candidate generation—topology evaluation—target selection—trajectory control—grasping closure—transfer and placement.

[0044] As an optional implementation, it also includes: setting a safety action library in the decision system of the SAC reinforcement learning algorithm; wherein, when an abnormal state is detected, such as continuous collisions, no progress for a long time in a round, or a joint approaching its limit, the reinforcement learning (RL) output is paused, and an action in the safety action library is executed, such as raising, retreating, resetting and re-observing, before switching back to RL control, so as to improve system safety and industrial availability.

[0045] In a specific embodiment, the hyperparameter settings for training the decision network are shown in Table 1.

[0046] Table 1. Hyperparameter settings for decision network training ; Finally, a training model with 20,000 epochs was used as the test network. The test results are shown in Table 2. The actual success rate of this invention reached a considerable 74%.

[0047] Table 2 Comparison of Decision Network Test Results under Simulation and Real-World Conditions ; This invention also provides a stacked workpiece gripping system based on SAC robotic arm vision, such as... Figure 3 As shown, the system may include a processor 31 and a memory 32, wherein the processor 31 and the memory 32 can be connected via a bus or other means. Figure 3 Taking the example of a connection between China and Israel via a bus.

[0048] Processor 31 can be a central processing unit (CPU). Processor 31 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips.

[0049] The memory 32, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the stacked workpiece grasping method based on SAC robotic arm vision in the embodiments of the present invention. The processor 31 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory 32, thereby realizing the stacked workpiece grasping method based on SAC robotic arm vision in the above method embodiments.

[0050] The memory 32 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the processor 31, etc. Furthermore, the memory 32 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 32 may optionally include memory remotely located relative to the processor 31, and these remote memories may be connected to the processor 31 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0051] The one or more modules are stored in the memory 32, and when executed by the processor 31, they perform actions such as... Figure 1 , Figure 2 The stacked workpiece gripping method based on SAC robotic arm vision in the illustrated embodiment.

[0052] For specific details regarding the above system, please refer to the relevant documentation. Figure 1 , Figure 2 The relevant descriptions and effects in the illustrated embodiments are for understanding purposes only and will not be repeated here.

[0053] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc.; the storage medium can also include combinations of the above types of memory.

[0054] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A method for grasping stacked workpieces based on SAC robotic arm vision, characterized in that, include: S1, obtain a candidate set of gripping poses for each workpiece based on the depth image of the stacked workpieces in the bin; S2, Based on the current workpiece stacking topology, obtain the grasping suitability score for each workpiece in the pose candidate set; the grasping suitability score is used to evaluate graspability and grasping consequences; S3, based on the SAC reinforcement learning algorithm, controls the robotic arm to perform a grasping action on the workpiece with the highest grasping suitability score; S4. After completing the grasping of the current workpiece, a score is given according to the reward mechanism; the record of this robotic arm grasping is stored in the SAC training playback pool to update the SAC reinforcement learning algorithm. Repeat steps S1 to S4 until the workpiece grabbing operation in the bin is completed.

2. The stacked workpiece gripping method based on SAC robotic arm vision according to claim 1, characterized in that, In S1, the steps for obtaining the depth image of the stacked workpieces inside the bin include: A 6-DOF robotic arm and parallel grippers are used, and an RGB-D camera is mounted in a hand-eye configuration to acquire depth images of the stacked workpieces inside the bin. The steps to obtain the candidate set of gripping poses for each workpiece include: Based on the horizontal detection frame B of the i-th workpiece i Get the set of tilted grab boxes r=r k , where each element r k Includes the center point coordinates, tilt angle, tilt width, and confidence level of the corresponding workpiece; For each element r k Calculate whether the coordinates of its center point fall within the corresponding horizontal detection box B. i If so, then the corresponding element r k The candidate set R to which the corresponding workpiece is assigned i middle.

3. The stacked workpiece gripping method based on SAC robotic arm vision according to claim 1, characterized in that, In S2, the construction of the current workpiece stack topology includes: Depth point cloud data is obtained from the depth image of the stacked workpieces in the bin. Based on the depth point cloud data and occlusion relationships, a stacking topology graph G=(V,E) is constructed, where V is the set of workpiece instances and E contains at least one of the following relationship edges: support / overlay relationship, occlusion relationship and potential interference relationship.

4. The stacked workpiece gripping method based on SAC robotic arm vision according to claim 3, characterized in that, In S2, the construction of the current workpiece stack topology also includes: If the point cloud of workpiece j is above workpiece i, and the projected overlap area exceeds a preset threshold, then record that workpiece j covers workpiece i. If there is an adjacent point cloud occlusion in the gripper feed direction and the minimum distance is less than the preset distance, then the accessibility is determined to be reduced. If, after a workpiece is grasped, the length of the overlying link above it is greater than 0, and the projection of the upper center of gravity may extend beyond the support area, then the stability risk increases.

5. The stacked workpiece gripping method based on SAC robotic arm vision according to claim 3, characterized in that, In S2, the calculation of the gripping suitability score for each workpiece in the pose candidate set includes: S i =ω1*Q i +ω2*L i -ω3*C i -ω4*T i -ω5*O i ; Wherein, the subscript i indicates that the current workpiece is the i-th workpiece; Q i The quality of the pose candidate grasping is indicated by evaluation metrics including candidate confidence, gripper width matching, and angle reasonableness. L i This indicates the reachability of the crawled data, and the evaluation metrics include inverse kinematics feasibility, feed unobstructedness, and workspace constraint satisfaction. C i The collision risk is indicated by evaluation metrics including the minimum clearance with adjacent workpieces / boxes and the estimated collision trajectory. T i The evaluation indicators for stability risk include the probability estimate of collapse after capture and the level of overburden. O i The degree of occlusion is indicated by evaluation metrics including the proportion of visible area and the proportion of depth loss. The weights ω1 to ω5 are taken as non-negative values ​​and normalized.

6. The stacked workpiece gripping method based on SAC robotic arm vision according to claim 5, characterized in that, In S3, the steps of the SAC reinforcement learning algorithm controlling the robotic arm to perform the grasping action on the workpiece with the highest grasping suitability score include: The initial state s of the SAC reinforcement learning algorithm t Let it be: Horizontal detection box B i eigenvector F i , capture suitability score S i The topological adjacency of workpiece i, the initial end-effector pose / velocity of the gripper, and the gripper opening / closing state; wherein, the feature vector F i It is based on the candidate set R i The extracted feature vectors include the center point coordinates, tilt box angle, tilt box width, and confidence level. Action a of the SAC reinforcement learning algorithm t Let: The end velocity of the gripper be v = {v x v y v z w x w y w z }, Gripper opening / closing degree u grip ∈[0,1], target score vector g t ; Set the target selection of the SAC reinforcement learning algorithm as: i * =argmax(g t ).

7. The stacked workpiece gripping method based on SAC robotic arm vision according to claim 6, characterized in that, In S4, the reward mechanism includes: r t =r succ +r fail +r col +r stab +r prog ; If the capture and placement are successful, then r succ The value is increased by 1.0; if the catch misses / falls / times out, then r... fail The value of r is -1.0; if a collision occurs, then r... col The value is -0.2; if it causes significant slippage / fall in the upper layer, then r stab The value is -0.5; the process-guided difference reward r prog =α*(d t-1 -d t ), d t α represents the image similarity or pose error distance, and α is the scaling factor.

8. The stacked workpiece gripping method based on SAC robotic arm vision according to claim 7, characterized in that, In S4, the steps for updating the SAC reinforcement learning algorithm include storing the record of the robotic arm's grasping action in the SAC training playback pool: Sample (s) t a t r t s t+1 Store it in the replay pool and update the Actor and Critic of the SAC reinforcement learning algorithm; Among them, r t Samples with a value ≥ 0 are stored in the Good pool, and r is... t Samples with a value <0 are stored in the Bad pool. Each update involves sampling at a ratio of Good:Bad=0.8:0.2, and an upper limit for the capacity of each pool is set.

9. The stacked workpiece gripping method based on SAC robotic arm vision according to claim 1, characterized in that, Also includes: The SAC reinforcement learning algorithm's decision system includes a safety action library; when an abnormal state is detected, the reinforcement learning output is paused, the action in the safety action library is executed, and then the reinforcement learning control is switched back.

10. A stacked workpiece gripping system based on SAC robotic arm vision, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the stacked workpiece gripping method based on SAC robotic arm vision as described in any one of claims 1 to 9.