A positioning method and system based on visual perception

By identifying light and shadow interference and constructing a digital twin model, the physical accessibility and stability of grasping options are evaluated, solving the problems of light and shadow interference and occlusion in visual positioning in stacked scenarios, and achieving high-precision and efficient workpiece positioning.

CN121074123BActive Publication Date: 2026-04-07JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing vision-based positioning technologies struggle to handle light and shadow interference and occlusion issues in stacked scenarios, leading to inaccurate positioning and potentially causing workpieces to scatter.

Method used

By receiving the definition information of the target workpiece, identifying and locating its spatial region, separating light and shadow interference, constructing a scene digital twin model, evaluating the physical accessibility and stability of the grasping options, generating grasping instructions with priority ranking, and optimizing the positioning process by iteratively updating visual data.

Benefits of technology

It improves positioning accuracy and efficiency in stacked scenarios, reduces the collision risk of robot operations, and adapts to the needs of automated operations in multiple fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074123B_ABST
    Figure CN121074123B_ABST
Patent Text Reader

Abstract

The application is suitable for the field of computer vision and robot control, and provides a positioning method and system based on visual perception. The method comprises the following steps: receiving target workpiece definition information, positioning the spatial region thereof in initial visual scene data; separating light and shadow interference, and restoring intrinsic visual features of the workpiece; constructing a scene digital twin model based on the intrinsic features and physical simulation; generating a grasping option with priority and dynamic confidence; identifying a key occlusion removal option; and iteratively performing a removal operation until the target is grasped. The system comprises corresponding functional modules. The application improves positioning accuracy and operation safety, adapts to dynamic stacked scenes, efficiently realizes target workpiece positioning and grasping, and is suitable for multi-field automated operation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of computer vision and robot control, and particularly relates to a positioning method and system based on visual perception. BACKGROUND

[0002] With the rapid development of industrial automation and intelligent logistics and warehousing, robots are increasingly widely used in workpiece grabbing, sorting, assembly and other scenarios, and the positioning technology based on visual perception is the core support for realizing precise operation of robots. Currently, visual positioning technology has expanded from a single static scene to a complex dynamic stacked scene, and the demand in the fields of mechanical manufacturing, automobile part processing, logistics and express sorting is continuously growing. The industry has higher requirements for the accuracy, anti-interference ability, scene adaptability and operation safety of visual positioning, and needs to solve key problems such as occlusion of stacked workpieces, light interference, and grabbing decision optimization, in order to improve the automation level and efficiency of robot operation.

[0003] The existing positioning scheme based on visual perception mainly receives two-dimensional image information of the target workpiece, matches the target area in the visual scene data and converts it into three-dimensional coordinates to realize positioning, and cannot remove the light interference component. In the stacked scene, only the visible part of the workpiece can be geometrically reconstructed, and it is difficult to deduce the pose of the occluded part, and simple grabbing commands may cause the scattered workpieces to fall. SUMMARY

[0004] The purpose of the present application is to provide a positioning method and system based on visual perception, which aims to solve the technical problems existing in the prior art determined in the background art.

[0005] The present application is implemented as follows: a positioning method based on visual perception, the method comprising:

[0006] receiving definition information of a target workpiece, and identifying and positioning the spatial region of the target workpiece in the stacked structure in the initial visual scene data based on the definition information;

[0007] analyzing the initial visual scene data, calculating and separating the light interference component in the scene, and restoring the intrinsic visual features of the surfaces of all stacked and occluded workpieces;

[0008] based on the intrinsic visual features, identifying the visible parts of the workpieces in the scene, and taking the geometric information of these visible parts as constraint conditions to perform dynamic simulation of the stacking stability in a virtual physical environment, deducing the most likely three-dimensional geometric shape and spatial pose of the occluded part through iterative calculation, and constructing a scene digital twin model containing complete workpieces;

[0009] Based on the scene digital twin model, the physical accessibility, the grasping stability and the influence on the overall stacked structure of each potential grasping point are analyzed, a plurality of grasping options with priority ranking are generated, and a dynamic confidence reflecting the accuracy of the pose corresponding to the grasping option and the safety of the operation process of the grasping option is calculated for each grasping option;

[0010] According to the grasping option and the dynamic confidence, the gain of the visual area of the target workpiece after each grasping option is executed and the disturbance to the overall stacked stability are evaluated, and a key occlusion removal option is identified from the grasping options;

[0011] The robot is preferentially guided to execute the key occlusion removal option, and after each grasping operation, updated visual scene data is acquired and iteration is performed until the target workpiece is successfully positioned and grasped.

[0012] As a further scheme of the present application, the identification and positioning of the spatial region of the target workpiece in the stacked structure specifically comprises:

[0013] Receiving a two-dimensional template image of the target workpiece, a three-dimensional design model and a labeled region in the initial visual scene data as definition information;

[0014] Based on the definition information, a target region with similar visual features to the target workpiece is matched in the initial visual scene data;

[0015] For the target region matched successfully, the target region is converted from two-dimensional image coordinates to a three-dimensional space coordinate system to obtain the spatial region of the target workpiece.

[0016] As a further scheme of the present application, the restoration of the intrinsic visual features of the surfaces of all stacked and occluded workpieces specifically comprises:

[0017] An optical reflection model library containing a plurality of typical lighting conditions and workpiece surface materials is constructed;

[0018] Based on the optical reflection model library, the initial visual scene data is analyzed to solve the inverse problem of the lighting equation, and the direct lighting, indirect lighting and mutual reflection components in the initial visual scene are calculated;

[0019] The direct lighting, indirect lighting and mutual reflection components are separated from the initial visual scene data, the inherent color and texture information of the surfaces of all stacked and occluded workpieces is retained, and the intrinsic visual features are obtained.

[0020] As a further scheme of the present application, the construction of the scene digital twin model containing the complete workpiece specifically comprises:

[0021] extracting visible edge contours and surface keypoints of all stacked, occluded workpieces from the intrinsic visual features, and reconstructing three-dimensional point cloud data of visible parts of all stacked, occluded workpieces;

[0022] locally registering the three-dimensional point cloud data with a structure model of the corresponding stacked, occluded workpiece, determining the pose of the visible part of the stacked, occluded workpiece in a global coordinate system as a constraint condition;

[0023] constructing a simulation environment containing gravity, friction and collision detection, and randomly generating a plurality of possible pose configurations of the occluded part based on the constraint condition;

[0024] performing physical stability evaluation on each possible pose configuration, selecting a configuration that meets the static equilibrium condition and is consistent with the constraint of the visible part as an optimal solution, and constructing a scene digital twin model.

[0025] As a further scheme of the present application, the dynamic confidence reflecting the accuracy of the pose corresponding to each grasping option and the safety of the operation process of the grasping option is calculated for each grasping option, specifically including:

[0026] generating potential grasping points for each stacked, occluded workpiece on the surface of each stacked, occluded workpiece in the scene digital twin model based on curvature analysis and attitude stability criteria;

[0027] evaluating the physical accessibility of each potential grasping point on the stacked, occluded workpiece through a collision detection algorithm;

[0028] for the physically accessible grasping points, calculating the relative position relationship between the potential grasping point and the centroid of the corresponding stacked, occluded workpiece, the normal vector of the contact surface and the friction cone, and quantitatively evaluating the grasping stability of each potential grasping point;

[0029] for each potential grasping point, simulating the process of removing the stacked, occluded workpiece corresponding to the potential grasping point, and predicting the degree of influence of the removal action on the stability of the stack structure;

[0030] comprehensively evaluating the results of the physical accessibility, grasping stability and degree of influence on the stability of the stack structure, generating a priority ranking of the grasping options, and calculating a dynamic confidence.

[0031] As a further scheme of the present application, the key occlusion removal option is identified from the grasping options, specifically including:

[0032] for each grasping option in the priority ranking, simulating the execution of the removal operation corresponding to the grasping option in the scene digital twin model, and calculating the change in visible area of the target workpiece after removal under multiple perspectives;

[0033] Evaluate the stability changes of the stacked structure after performing a removal operation in a scene digital twin model, including calculating the changes in the center of gravity offset and contact force distribution of the remaining stack and the occluded workpiece;

[0034] By combining changes in visible area and stability, the obstacle clearing efficiency score for each grabbing option is calculated;

[0035] Based on the obstacle removal effectiveness score and dynamic confidence level, the key obstruction removal options are selected from the crawling options.

[0036] As a further aspect of the present invention, the step of acquiring updated visual scene data and iterating after each grasping operation until the target workpiece is successfully located and grasped specifically includes:

[0037] Based on the key occlusion removal options, generate a collision-free motion trajectory that satisfies robot dynamics constraints.

[0038] Control the robot to execute the motion trajectory and complete the stacking, grasping and moving of the obstructed workpieces corresponding to the key obstruction removal options;

[0039] After each grabbing and moving operation is completed, visual data of the work area is re-acquired to obtain updated visual scene data;

[0040] Based on updated visual scene data, update the scene digital twin model, grasping options, and key occlusion removal options;

[0041] When the target workpiece is determined to be fully exposed and directly graspable, a final grasping instruction is generated for the target workpiece, controlling the robot to complete the grasping.

[0042] Another object of the present invention is to provide a positioning system based on visual perception, the system comprising:

[0043] The target workpiece identification and positioning module is used to receive the definition information of the target workpiece and, based on the definition information, identify and locate the spatial region of the target workpiece in the stacked structure in the initial visual scene data.

[0044] The light and shadow interference removal module is used to analyze the initial visual scene data, calculate and analyze and separate the light and shadow interference components in the scene, and restore the intrinsic visual features of all stacked and occluded workpiece surfaces.

[0045] The scene digital twin model construction module is used to identify the visible parts of the workpiece in the scene based on intrinsic visual features, and use the geometric information of these visible parts as constraints to perform dynamic simulation of stacking stability in a virtual physical environment. Through iterative calculation, the most likely three-dimensional geometric shape and spatial pose of the occluded part are deduced to construct a scene digital twin model containing the complete workpiece.

[0046] The grab option generation module is used to analyze the physical reachability, grab stability and impact on the overall stacking structure of each potential grab point based on the scene digital twin model, generate several grab options with priority order, and calculate a dynamic confidence level for each grab option that comprehensively reflects the accuracy of the pose corresponding to the grab option and the safety of the operation process of executing the grab option.

[0047] The critical obstruction identification module is used to evaluate the gain on the visible area of ​​the target workpiece and the disturbance to the overall stacking stability after each gripping option is executed, based on the gripping options and dynamic confidence, and to identify the critical obstruction removal option from the gripping options.

[0048] The robot operation execution module is used to prioritize guiding the robot to execute the key occlusion removal option, and after each grasping operation, it reacquires updated visual scene data and iterates until the target workpiece is successfully located and grasped.

[0049] As a further embodiment of the present invention, the key occlusion recognition module includes:

[0050] The potential gripping point generation module is used to generate potential gripping points for each stacked and occluded workpiece surface in the digital twin model of the scene based on curvature analysis and attitude stability criteria.

[0051] The collision detection and evaluation module is used to evaluate the physical reachability of each potential gripping point on a stacked or occluded workpiece using a collision detection algorithm.

[0052] The stability quantification module is used to calculate the relative positional relationship between potential gripping points and the centroids of corresponding stacked and occluded workpieces, as well as the normal vector and friction cone of the contact surface for physically reachable gripping points, and to quantitatively evaluate the gripping stability of each potential gripping point.

[0053] The removal process simulation module is used to simulate the process of removing the stacked and occluded workpieces corresponding to each potential gripping point, and predict the impact of the removal action on the stability of the stack structure.

[0054] The priority ranking module is used to generate a priority ranking of the grabbing options by comprehensively considering the evaluation results of physical reachability, grabbing stability and the degree of impact on heap structure stability, and to calculate the dynamic confidence level.

[0055] As a further embodiment of the present invention, the robot operation execution module includes:

[0056] The motion trajectory generation module is used to generate a collision-free motion trajectory that meets robot dynamics constraints based on the key occlusion removal options.

[0057] The motion trajectory execution module is used to control the robot to execute motion trajectories and complete the stacking, grasping and moving of workpieces corresponding to the key obstruction removal options.

[0058] The visual data re-acquisition module is used to re-acquire visual data of the work area after each grabbing and moving operation to obtain updated visual scene data.

[0059] The model update module is used to update the scene digital twin model, grasping options, and key occlusion removal options based on updated visual scene data.

[0060] The gripping control module is used to generate the final gripping command for the target workpiece when the target workpiece is determined to be fully exposed and can be gripped directly, and to control the robot to complete the gripping.

[0061] The beneficial effects of this invention are:

[0062] The visual perception-based positioning method and system provided by this invention solves the core problems existing in the positioning of stacked scenes through multi-stage collaborative optimization, and has significant technical advantages and application value. The scheme first accurately locates the initial spatial region of the target workpiece using multi-source definition information, and then restores the intrinsic visual features of the workpiece through light and shadow interference separation technology, providing an accurate visual foundation for subsequent positioning. Next, based on the intrinsic features and physical simulation, a digital twin model of the scene containing the complete workpiece is constructed, filling the technical gap in the pose inference of occluded parts and achieving the integrity of positioning information. When generating grasping options, physical accessibility, grasping stability, and impact on the stacked structure are comprehensively evaluated, and key obstructions are screened by combining dynamic confidence and obstacle removal efficiency scores to ensure the reliability and relevance of grasping decisions. Finally, through a dynamic iteration mechanism, visual data and models are updated after each operation to adapt to scene changes and avoid static decision-making bias. Overall, this solution significantly improves the accuracy and efficiency of target workpiece positioning in stacked scenarios, reduces the collision risk and stack collapse probability of robot operation, and does not require reconstruction of the core algorithm for different scenarios. It can be widely adapted to the automation operation needs of multiple fields such as industrial manufacturing and logistics warehousing, providing comprehensive and reliable technical support for precise robot operation. Attached Figure Description

[0063] Figure 1 A flowchart illustrating a vision-based localization method provided in an embodiment of the present invention;

[0064] Figure 2 A flowchart for identifying and locating a target workpiece in a stacked structure, provided by an embodiment of the present invention;

[0065] Figure 3A flowchart for restoring the intrinsic visual features of all stacked and obscured workpiece surfaces, provided in an embodiment of the present invention;

[0066] Figure 4 A flowchart for constructing a scene digital twin model containing a complete workpiece, provided in an embodiment of the present invention;

[0067] Figure 5 A flowchart for calculating dynamic confidence level provided in an embodiment of the present invention;

[0068] Figure 6 A flowchart for identifying key obstruction removal options provided in an embodiment of the present invention;

[0069] Figure 7 This is a flowchart provided by an embodiment of the present invention until the target workpiece is successfully located and captured;

[0070] Figure 8 A structural block diagram of a positioning system based on visual perception provided in an embodiment of the present invention;

[0071] Figure 9 This is a structural block diagram of a key occlusion recognition module provided in an embodiment of the present invention;

[0072] Figure 10 This is a structural block diagram of the robot operation execution module provided in an embodiment of the present invention. Detailed Implementation

[0073] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0074] Figure 1 A flowchart of a vision-based localization method provided in an embodiment of the present invention is shown below. Figure 1 As shown, the method includes:

[0075] S100: Receive definition information for the target workpiece, and in the initial visual scene data, identify and locate the spatial region of the target workpiece in the stacked structure based on the definition information;

[0076] The target workpiece definition information received by the system includes a two-dimensional template image, a three-dimensional design model, and an initial visual scene annotation area. The two-dimensional template image provides key visual feature benchmarks for the target workpiece, including surface texture, color distribution, and typical contours. These features are the core basis for subsequent visual matching. The three-dimensional design model supplements the complete geometric structural parameters of the workpiece, including size, volume, and spatial positional relationships of key parts, providing geometric references for subsequent conversion from two-dimensional visual data to three-dimensional spatial positioning. The annotation area in the initial visual scene plays a range constraint role, narrowing the spatial range of target search, improving recognition efficiency, and reducing misjudgments caused by interference from irrelevant regional features.

[0077] After obtaining the definition information, the system enters the target area matching stage. Based on the two-dimensional visual features in the definition information, the system performs layer-by-layer feature extraction and comparison on the initial visual scene data within the marked area, locks the potential target areas that meet the conditions, and initially filters out areas with visual features consistent with the target workpiece from the complex stacked background, eliminating interference from non-target workpieces.

[0078] The conversion from two-dimensional image coordinates to three-dimensional spatial coordinates achieves a key leap from visual perception to physical positioning. The system combines the intrinsic and extrinsic parameters of the camera, as well as the dimensional parameters of the target workpiece's three-dimensional design model, and uses algorithms such as perspective projection transformation to convert the pixel coordinates of the successfully matched target area in the two-dimensional image into physical coordinates in the three-dimensional spatial coordinate system, while simultaneously determining the initial posture of the target workpiece.

[0079] In the scenario of stacked mechanical parts, the two-dimensional image region of the target workpiece can be transformed by coordinates to obtain its specific position and tilt angle in the workbench coordinate system; in the scenario of stacked cardboard boxes in logistics warehousing, the target cardboard box can be transformed by coordinates to obtain its specific position and tilt angle in the warehouse rack coordinate system, thereby clarifying the spatial region of the target workpiece in the stacking structure, establishing the connection between visual data and the physical world, and providing a clear spatial reference for subsequent digital twin model construction, robot grasping path planning and other links.

[0080] S200, the initial visual scene data is analyzed, and the light and shadow interference components in the scene are separated by calculation and analysis to restore the intrinsic visual features of all stacked and occluded workpiece surfaces.

[0081] We need to build an optical reflection model library that includes several typical lighting conditions and workpiece surface materials. In reality, the lighting environment of stacked scenes varies, and the workpiece materials also have different reflection characteristics. Different combinations of lighting and materials will make the workpiece present different visual effects. If there is no reference for the corresponding reflection rules, the subsequent light and shadow separation will be inaccurate due to the lack of standard basis. Therefore, the model library needs to cover common lighting types and material properties to form a standardized analysis framework to ensure that matching reflection rules can be found in different application scenarios.

[0082] Based on an optical reflection model library, the initial visual scene data is analyzed and the inverse problem of the illumination equation is solved. The brightness value of each pixel in the initial visual scene data is a superposition of direct illumination, indirect illumination, and mutual reflection components. These superimposed components can mask the inherent visual features of the workpiece, making it impossible to accurately identify key information such as edges and textures. The purpose of solving the inverse problem of the illumination equation is to decompose the mixed light signal into individual illumination components. By calling the reflection model in the model library that matches the current scene, the contribution of each illumination component in each pixel is calculated, clarifying which visual changes are caused by external illumination and which belong to the characteristics of the workpiece itself. In the scenario of stacked mechanical parts, strong light spots and shadows may exist on the surface of metal workpieces; this step can accurately separate these illumination interferences. In the scenario of stacked electronic components, the surface of plastic workpieces may have uneven brightness due to ambient light reflection; this process can also remove irrelevant illumination components.

[0083] The three illumination components mentioned above are separated from the initial visual scene data. The inherent color and texture information of all stacked and occluded workpiece surfaces are retained to obtain intrinsic visual features. Intrinsic color and texture are the essential visual attributes of the workpiece and the core basis for subsequent identification of the visible part of the workpiece and extraction of geometric information. Only by obtaining accurate intrinsic visual features can we ensure that the edge contour extraction and surface key point positioning of the visible part of the workpiece are accurate in subsequent steps, and provide reliable visual input for dynamic simulation of stacking stability and construction of digital twin model.

[0084] S300, based on intrinsic visual features, identifies the visible parts of the workpiece in the scene and uses the geometric information of these visible parts as constraints to perform dynamic simulation of stacking stability in a virtual physical environment. Through iterative calculation, it deduces the most likely three-dimensional geometric shape and spatial pose of the occluded part and constructs a scene digital twin model containing the complete workpiece.

[0085] The visible edge contours and surface key points of all stacked and occluded workpieces are extracted from the intrinsic visual features. The intrinsic visual features have been stripped of light and shadow interference and retain the inherent visual attributes of the workpieces. The extracted edge contours and key points can accurately reflect the geometric shape of the visible parts, avoiding contour blurring or key point offset caused by light and shadow. At the same time, the three-dimensional point cloud data of the visible parts are reconstructed, transforming the two-dimensional visual features into geometric data in three-dimensional space, upgrading the shape of the visible parts from planar information to a three-dimensional structure, and providing a three-dimensional basis for subsequent registration with the design model.

[0086] The 3D point cloud data is locally registered with the corresponding structural model of the workpiece. The structural model contains the design geometric parameters of the workpiece. The registration process can align the 3D point cloud of the visible part with the corresponding area of ​​the design model, thereby determining the pose of the visible part in the global coordinate system. The core of this step is to establish the spatial constraints of the visible part to ensure that the subsequent simulation of the occluded part will not deviate from the actual spatial position of the workpiece, and to avoid the simulation results from being out of touch with the real scene.

[0087] A simulation environment incorporating gravity, friction, and collision detection is constructed. The stability of workpieces in a real stacking scenario is directly affected by these physical factors; ignoring these factors would lead to the inferred pose of the occluded portion not conforming to the actual equilibrium state. Based on the pose constraints of the visible portion, the system randomly generates several possible pose configurations for the occluded portion. Since the occluded portion cannot be directly observed visually, multiple sets of possible configurations are needed to cover the potential spatial state, providing a basis for subsequent selection of the optimal solution. Finally, the physical stability of each possible pose configuration is evaluated. Static equilibrium calculations are used to determine whether the configuration satisfies the equilibrium conditions of gravity and friction, and to verify whether the configuration is consistent with the pose constraints of the visible portion. Configurations that conflict with the visible portion geometrically or are physically unstable are eliminated. The optimal solution is selected and integrated into the model, ultimately constructing a scene digital twin model containing the complete shape and spatial pose of all workpieces.

[0088] S400, based on a scene digital twin model, analyzes the physical reachability, grasping stability and impact on the overall stacking structure of each potential grasping point, generates several grasping options with priority order, and calculates a dynamic confidence level for each grasping option that comprehensively reflects the accuracy of the pose corresponding to the grasping option and the safety of the operation process of executing the grasping option.

[0089] Potential gripping points are generated on the surfaces of stacked, occluded workpieces in the scene's digital twin model. This generation process relies on curvature analysis and attitude stability criteria. Curvature analysis filters areas with smaller surface curvature, as these areas have a larger contact area, reducing the probability of workpiece slippage during gripping. The attitude stability criterion focuses on the compatibility of the gripping point with the direction of gravity, ensuring that the workpiece remains stable under its own weight after gripping, avoiding gripping failure due to attitude imbalance. Next, a collision detection algorithm evaluates the physical accessibility of each potential gripping point. This algorithm simulates the complete motion path of the robot's end effector from its initial position to the gripping point, checking for potential collisions with other workpieces, worktables, or the robot itself. Only gripping points without collision risk are retained for subsequent stages, aiming to eliminate invalid options that the robot cannot actually reach, thus avoiding mechanical collision accidents.

[0090] To quantify the gripping stability of physically accessible gripping points, it is necessary to calculate the relative position of the gripping point to the corresponding workpiece's center of mass. The closer the relative position is to the center of mass, the smaller the torque generated during gripping, and the less likely the workpiece is to flip. At the same time, the normal vector and friction cone of the contact surface are analyzed. The normal vector determines the effective direction of the gripping force, while the friction cone clarifies the range that the gripping force must satisfy to prevent the workpiece from sliding. By quantifying these parameters, the ability of each gripping point to maintain workpiece stability can be accurately determined.

[0091] The process of removing the workpiece corresponding to each potential gripping point is simulated. Based on the physical parameters of the digital twin model, the changes in the center of gravity distribution and contact force adjustment of the remaining stack after removal are calculated to determine whether the removal action will cause the remaining workpiece to become unbalanced and collapse, thus avoiding damage to the overall stack structure by a single gripping operation and causing obstacles to subsequent operations. Finally, considering physical accessibility, gripping stability, and impact on the stack structure, a priority ranking of gripping options is generated. Options with high accessibility, strong stability, and minimal impact on the stack are prioritized. Simultaneously, the dynamic confidence level of each option is calculated. This dynamic confidence level, by integrating reliability data from various evaluation dimensions, quantitatively reflects the overall performance of the option in terms of location accuracy and operational safety; the higher the confidence level, the greater the probability of successful execution.

[0092] S500, based on the gripping options and dynamic confidence, evaluates the gain on the visible area of ​​the target workpiece and the disturbance to the overall stacking stability after each gripping option is executed, and identifies the key obstruction removal option from the gripping options.

[0093] For each grasping option in the priority ranking, the corresponding removal operation is simulated in the scene's digital twin model, and the change in the visible area of ​​the target workpiece after removal is calculated from multiple perspectives. Multi-view evaluation aims to avoid the limitations of a single perspective, ensuring that regardless of which working direction the robot subsequently acquires visual data from, it can accurately determine the actual exposure effect of the target workpiece after the occlusion is removed. Virtual simulation, on the other hand, eliminates the need for trial-and-error operations on the real stacked structure, avoiding the risk of stacking imbalance or workpiece damage that may result from actual removal from the outset.

[0094] The changes in stack stability after a removal operation are evaluated in a digital twin model, specifically by calculating the center of gravity shift and contact force distribution changes of the remaining stacked obscured workpieces. The center of gravity shift directly reflects the overall stack's equilibrium state, while the change in contact force distribution indicates whether the forces between local workpieces remain within safe support limits. This step aims to prevent focusing solely on the exposure of the target workpiece while neglecting the safety of the stack structure, avoiding subsequent stack collapse caused by the removal operation, and laying a safe foundation for subsequent robot operations.

[0095] By combining changes in visible area and stability, a clearance benefit score is calculated for each grabbing option. The core function of this score is to balance the benefits (exposure of the target workpiece) and risks (stack stability disturbances) of the removal operation. It avoids selecting options that only slightly increase the visible area but seriously damage stack stability, and also excludes invalid options that do not affect stability but do not substantially help expose the target, ensuring that each candidate option has practical operational value.

[0096] The selection of key obstruction removal options is determined by comprehensively screening based on obstacle removal efficiency scores and dynamic confidence scores. Dynamic confidence scores already cover the pose accuracy and operational safety of the grasping options. Combining them with obstacle removal efficiency ensures that the selected options not only have outstanding obstacle removal effects but can also be executed accurately and safely by the robot, avoiding obstacle removal operation failures due to unreliable grasping options themselves.

[0097] S600 prioritizes guiding the robot to execute the key obstruction removal option, and after each grasping operation, it reacquires updated visual scene data and iterates until the target workpiece is successfully located and grasped.

[0098] Motion trajectories are generated based on the key obstruction removal options. The generation process must strictly adhere to robot dynamics constraints, including limit thresholds for joint movement speed and acceleration, to prevent mechanical failure due to motion overload. Simultaneously, collision detection algorithms validate the trajectory to ensure the robot's end effector does not come into contact with stacked workpieces, worktables, or its own structure during the entire movement, guaranteeing operational safety from a path planning perspective. Next, the robot executes the generated motion trajectory, driving the end effector to complete the grasping and movement of the key obstruction. This step translates the decisions from the digital twin model into actual actions in the physical world, achieving precise removal of the obstruction and creating conditions for further exposure of the target workpiece.

[0099] After each grabbing and moving operation is completed, the system re-collects visual data of the work area. Since key obstructions are removed, the original stacked scene will change, and some previously obscured workpiece areas may be revealed. The new visual data can truly reflect the current state of the scene and provide accurate input for subsequent model updates.

[0100] Subsequently, based on the updated visual scene data, the scene digital twin model is corrected, including adjusting the spatial pose of the remaining workpieces and supplementing the geometric information of the newly exposed areas; at the same time, the accessibility and stability of potential gripping points are re-analyzed, the priority ranking and dynamic confidence of gripping options are updated, and a new round of key occlusion removal options are re-identified to ensure that the next operation decision is always based on the latest scene state.

[0101] The robot continuously determines whether the target workpiece is fully exposed and can be directly grasped. When all key grasping areas of the target workpiece are unobstructed and its spatial pose meets the conditions for direct robot grasping, the robot generates the final grasping command for the target workpiece and controls the robot to accurately grasp the target workpiece.

[0102] In a scenario of stacked mechanical parts, after the robot executes the removal trajectory of key shaft-type parts, it re-collects visual data and finds that part of the teeth of the target gear is still obscured by another shim. It then updates the digital twin model, identifies the shim as the new key obstruction, generates a new motion trajectory, and executes the removal until the target gear is fully exposed, at which point the final grasping command is generated. In a scenario of stacked logistics cardboard boxes, after the robot removes the upper layer of obstructing cardboard boxes, it re-collects data and finds that the top surface of the target cardboard box is still obscured by the edge of the adjacent cardboard box. The system updates the model and key obstructions, continues to execute the removal operation, and completes the grasping until the target cardboard box is fully exposed.

[0103] like Figure 2 As shown, the identification and positioning of the target workpiece in the spatial region of the stacked structure specifically includes:

[0104] S110, receive the two-dimensional template image of the target workpiece, the three-dimensional design model, and the marked area in the initial visual scene data as definition information;

[0105] S120, based on the defined information, match a target region in the initial visual scene data that has similar visual features to the target workpiece;

[0106] S130: For the successfully matched target area, the target area is transformed from two-dimensional image coordinates to three-dimensional spatial coordinate system to obtain the spatial area of ​​the target workpiece.

[0107] like Figure 3 As shown, the restoration of the intrinsic visual features of all stacked and obscured workpiece surfaces specifically includes:

[0108] S210, Construct an optical reflection model library containing several typical lighting conditions and workpiece surface materials;

[0109] S220, Based on the optical reflection model library, analyze the initial visual scene data, solve the inverse problem of the illumination equation, and calculate the direct illumination, indirect illumination and mutual reflection components in the initial visual scene;

[0110] ;

[0111] in: : in image coordinates The total pixel intensity value observed at that location is the initial visual scene data.

[0112] Target point The diffuse reflectance coefficient at a given location characterizes the inherent color and absorption properties of the stacked and obscured workpiece surface.

[0113] The number of main point light sources in the scene.

[0114] : No. The intensity of a point light source.

[0115] Target point The surface normal vector at a given location is derived from the geometric data in the scene's digital twin model.

[0116] From the target point Pointing to the The direction vector of a point light source.

[0117] Clamping operation of the dot product. Ensures that no illumination contribution occurs when the light source is on the back side of the surface (dot product is negative).

[0118] Target point The specular reflection coefficient at a certain point represents the gloss of the workpiece surface material.

[0119] : No. The reflection direction vector of incident light from a point light source about the normal.

[0120] From the target point The observation direction vector pointing towards the center of the camera.

[0121] Specular reflection specular index controls the size of the specular spot and is related to surface roughness.

[0122] Ambient light reflectance.

[0123] Ambient light intensity.

[0124] Target point The mutual reflection component at the point is formed by the light rays bouncing multiple times in the gaps between the stacked workpieces.

[0125] The inverse problem of the above rendering equation is solved by optimizing the algorithm. It is decomposed into direct illumination (diffuse reflection + specular reflection), indirect illumination (ambient light), and mutual reflection components, thereby separating the light and shadow interference.

[0126] S230, the direct illumination, indirect illumination and mutual reflection components are separated from the initial visual scene data, and the inherent color and texture information of all stacked and occluded workpiece surfaces are retained to obtain intrinsic visual features.

[0127] like Figure 4 As shown, the construction of a scene digital twin model containing complete workpieces specifically includes:

[0128] S310, extract the visible edge contours and surface key points of all stacked and occluded workpieces from the intrinsic visual features, and reconstruct the three-dimensional point cloud data of the visible parts of all stacked and occluded workpieces.

[0129] S320, The three-dimensional point cloud data is locally registered with the structural model of the corresponding stacked and occluded workpieces to determine the pose of the visible part of the stacked and occluded workpieces in the global coordinate system as a constraint condition.

[0130] S330, Construct a simulation environment that includes gravity, friction and collision detection, and randomly generate several possible pose configurations of the occluded parts based on the constraints.

[0131] S340 performs a physical stability assessment on each possible pose configuration, selects the configuration that satisfies the static equilibrium condition and is consistent with the constraints of the visible part as the optimal solution, and constructs a scene digital twin model.

[0132] like Figure 5 As shown, the calculation of a dynamic confidence level for each grasping option, which comprehensively reflects the accuracy of the pose corresponding to the grasping option and the safety of the operation process of executing the grasping option, specifically includes:

[0133] S410, on the surfaces of each stacked and occluded workpiece in the digital twin model of the scene, potential gripping points are generated for each stacked and occluded workpiece based on curvature analysis and attitude stability criteria.

[0134] ;

[0135] in:

[0136] Candidate points The overall score. The higher the score, the more suitable the point is as a potential crawling point.

[0137] , Candidate points The principal and secondary curvatures at a point. Used to quantify the degree of curvature of the local surface at that point; low curvature areas (flat areas) are more conducive to stable gripping.

[0138] : The three-dimensional coordinates of the centroid of the stacked and obscured workpiece.

[0139] Candidate points to the workpiece's center of mass The Euclidean distance. The closer the distance, the smaller the torque generated during grasping, and the more stable the operation.

[0140] Candidate points The surface normal vector at that location.

[0141] : The unit vector in the direction of gravity.

[0142] : The absolute value of the dot product of the normal vector and the direction of gravity. The larger this value is, the closer the gripping surface is to horizontal, and the better it can utilize friction to prevent the workpiece from slipping.

[0143] , , Weighting coefficients are used to balance the three factors—curvature, center-of-mass distance, and attitude stability—in the overall score. The relative importance of [the subject / method].

[0144] A large number of candidate points are sampled on the surface of the workpiece's 3D model. Calculate each point Select Points exceeding the threshold are subjected to non-maximum suppression in space, and the final output is a set of dispersed and stable potential grab points.

[0145] S420 uses a collision detection algorithm to evaluate the physical reachability of each potential gripping point on a stacked, occluded workpiece.

[0146] ;

[0147] in: Potential crawling points The accessibility assessment results are given, where 1 indicates accessibility and 0 indicates inaccessibility.

[0148] : Describes the movement of the robot's end effector from the starting point to the grasping point. The parameterized function of the trajectory, This is the normalized time parameter.

[0149] : Represents the geometry of all obstacles in the environment, including other stacks, occluding artifacts, and the robot itself.

[0150] : Moments on the trajectory The Euclidean distance between the end effector and the nearest obstacle.

[0151] : Safety distance threshold. This threshold is determined by the robot's absolute positioning error, the size of the end effector, and the control system delay, and is used to ensure collision-free movement.

[0152] For each potential crawl point A candidate trajectory is generated by the motion planner. And check whether the entire trajectory satisfies If satisfied, then This point is physically reachable.

[0153] S430 calculates the relative positional relationship between potential gripping points and the centroids of corresponding stacked and occluded workpieces, the normal vector of the contact surface, and the friction cone for physically reachable gripping points, and quantitatively evaluates the gripping stability of each potential gripping point.

[0154] ;

[0155] in: Potential crawling points The capture stability quality index. The larger the value, the stronger the ability to resist external disturbances and the higher the stability.

[0156] : The perturbation force applied to the centroid of the workpiece in any unit direction.

[0157] : Norm of the perturbation vector, constrained to 1 here, indicating that the worst-case scenario under a unit perturbation is being evaluated.

[0158] : Grasp Jacobian matrix. It is a matrix that positions the gripper at potential gripping points. The transformation matrix that maps the contact force applied at the point of contact to the resultant force and resultant moment at the center of mass of the workpiece.

[0159] This indicates that in order to maintain workpiece balance under a unit disturbance force, the gripper needs to be positioned at the contact point. The norm of the minimum contact force provided at that point.

[0160] This formula calculates stability under the worst-case scenario (i.e., when the required contact force is maximum). By finding... The direction of minimum disturbance force, The value is the amplitude of this minimum force. The constraint of the friction cone is implicit in... The construction of the matrix ensures that the calculated contact force is physically realizable (i.e., no slippage occurs).

[0161] S440, for each potential gripping point, simulate the process of removing the stacked and occluded workpieces corresponding to that potential gripping point, and predict the degree of impact of the removal action on the stability of the stack structure;

[0162] S450, taking into account the evaluation results of physical accessibility, grabbing stability and the degree of impact on heap structure stability, generates a priority ranking of the grabbing options and calculates dynamic confidence.

[0163] ;

[0164] in: : The final calculated dynamic confidence score. The higher the value, the higher the overall confidence level in the crawling option.

[0165] The index of the scoring metric is traversed through the set {R,S,I}, which represent the three evaluation dimensions of accessibility, stability, and structural impact, respectively.

[0166] : Represents the original score.

[0167] Normalized and adjusted score. For reachability and stability.

[0168] : The uncertainty quantification value. The larger the value, the lower the reliability of the score.

[0169] The preset weighting coefficients of each indicator in the confidence score synthesis reflect their relative importance.

[0170] : Structural impact score, a quantitative value calculated from the predicted impact of removal actions on the stability of the heap structure.

[0171] The core idea of ​​this formula is a weighted harmonic average. It multiplies the normalized value of each rating indicator by the inverse of its uncertainty. This means:

[0172] The higher the score, the greater the positive contribution to the confidence level.

[0173] The smaller (i.e.) The larger the score, the more weight the score carries in the final confidence level; that is, the more reliable the score, the greater its impact on the outcome.

[0174] like Figure 6 As shown, identifying the key obstruction removal option from the grabbing options specifically includes:

[0175] S510, for each grabbing option in the priority ranking, simulate the execution of the removal operation corresponding to the grabbing option in the scene digital twin model, and calculate the change in the visible area of ​​the target workpiece under multiple views after removal;

[0176] S520 evaluates the stability changes of the stacked structure after performing a removal operation in a scene digital twin model, including calculating the changes in the center of gravity offset and contact force distribution of the remaining stack and the occluded workpiece.

[0177] S530, taking into account changes in visible area and stability, calculates the obstacle clearing efficiency score for each grabbing option;

[0178] ;

[0179] in:

[0180] Obstacle removal benefit score. The higher the score, the greater the benefit of executing this crawling option.

[0181] The average net increase in the visible area of ​​the target workpiece under multiple preset viewing angles after the workpiece is removed in the simulation.

[0182] Total surface area of ​​the target workpiece. Used to normalize the visible area gain, making it a dimensionless proportion.

[0183] : The change in the overall stability of the stack structure after a simulated removal operation. This is a non-positive number; the smaller the value (the more negative), the more severe the stability degradation.

[0184] Stability attenuation coefficient. A constant greater than 0, used to control the intensity of the negative impact of stability disturbances on the benefit score. The larger the value, the more sensitive the system is to a decrease in stability.

[0185] This formula balances the benefits of seeing more targets with the risks of disrupting the structure. This ensured that when When it is too large, even The final obstacle clearing effectiveness score is very high. It will also be significantly suppressed, thereby avoiding the selection of dangerous operations that could lead to a collapse.

[0186] S540, based on obstacle removal efficiency score and dynamic confidence, filters out key obstruction removal options from the crawling options.

[0187] like Figure 7As shown, the step of acquiring updated visual scene data and iterating after each grasping operation until the target workpiece is successfully located and grasped specifically includes:

[0188] S610 generates a collision-free motion trajectory that satisfies robot dynamics constraints, based on the key occlusion removal option;

[0189] S620 controls the robot to execute motion trajectories and complete the stacking, grasping and moving of workpieces corresponding to the key obstruction removal options;

[0190] After completing each grasping and moving operation, the S630 re-collects visual data of the working area to obtain updated visual scene data.

[0191] S640 updates the scene digital twin model, grasping options, and key occlusion removal options based on updated visual scene data;

[0192] S650: When the target workpiece is determined to be fully exposed and directly graspable, a final grasping instruction is generated for the target workpiece, controlling the robot to complete the grasping.

[0193] Figure 8 A structural block diagram of a vision-based positioning system provided in an embodiment of the present invention is shown below. Figure 8 As shown, the system includes:

[0194] The target workpiece identification and positioning module 100 is used to receive definition information of the target workpiece and, based on the definition information, identify and locate the spatial region of the target workpiece in the stacked structure in the initial visual scene data.

[0195] The light and shadow interference removal module 200 is used to analyze the initial visual scene data, calculate and analyze and separate the light and shadow interference components in the scene, and restore the intrinsic visual features of all stacked and occluded workpiece surfaces.

[0196] The scene digital twin model construction module 300 is used to identify the visible parts of the workpiece in the scene based on intrinsic visual features, and use the geometric information of these visible parts as constraints to perform dynamic simulation of stacking stability in a virtual physical environment. Through iterative calculation, the most likely three-dimensional geometric shape and spatial pose of the occluded part are deduced to construct a scene digital twin model containing the complete workpiece.

[0197] The grasping option generation module 400 is used to analyze the physical reachability, grasping stability and impact on the overall stacking structure of each potential grasping point based on the scene digital twin model, generate several grasping options with priority order, and calculate a dynamic confidence level for each grasping option that comprehensively reflects the accuracy of the pose corresponding to the grasping option and the safety of the operation process of executing the grasping option.

[0198] The critical obstruction identification module 500 is used to evaluate the gain on the visible area of ​​the target workpiece and the disturbance to the overall stacking stability after each grasping option is executed, based on the grasping options and dynamic confidence, and to identify the critical obstruction removal option from the grasping options.

[0199] The robot operation execution module 600 is used to prioritize guiding the robot to execute the key occlusion removal option, and after each grasping operation, it re-acquires updated visual scene data and iterates until the target workpiece is successfully located and grasped.

[0200] like Figure 9 As shown, the key occlusion recognition module 500 includes:

[0201] The potential gripping point generation module 510 is used to generate potential gripping points for each stacked and occluded workpiece surface in the scene digital twin model based on curvature analysis and attitude stability criteria.

[0202] The collision detection and evaluation module 520 is used to evaluate the physical reachability of each potential gripping point on a stacked, occluded workpiece using a collision detection algorithm.

[0203] The stability quantification module 530 is used to calculate the relative positional relationship between potential gripping points and the centroids of corresponding stacked and occluded workpieces, the normal vector of the contact surface, and the friction cone for physically reachable gripping points, and to quantitatively evaluate the gripping stability of each potential gripping point.

[0204] The removal process simulation module 540 is used to simulate the process of removing the stacked and occluded workpieces corresponding to each potential gripping point, and predict the degree of impact of the removal action on the stability of the stack structure.

[0205] The priority ranking module 550 is used to generate a priority ranking of the grabbing options by comprehensively evaluating the physical reachability, grabbing stability and the degree of impact on the stability of the heap structure, and to calculate the dynamic confidence level.

[0206] like Figure 10 As shown, the robot operation execution module 600 includes:

[0207] The motion trajectory generation module 610 is used to generate a motion trajectory that satisfies robot dynamics constraints and is collision-free, based on the key occlusion removal option.

[0208] The motion trajectory execution module 620 is used to control the robot to execute motion trajectories and complete the stacking, grasping and moving operations of the stacked and obscured workpieces corresponding to the key obstruction removal options.

[0209] The visual data reacquisition module 630 is used to reacquire visual data of the work area after each grasping and moving operation to obtain updated visual scene data.

[0210] The model update module 640 is used to update the scene digital twin model, grasping options, and key occlusion removal options based on the updated visual scene data.

[0211] The gripping control module 650 is used to generate a final gripping instruction for the target workpiece when the target workpiece is determined to be fully exposed and can be gripped directly, and to control the robot to complete the gripping.

[0212] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0213] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

[0214] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A localization method based on visual perception, characterized in that, The method includes: Receive definition information for the target workpiece, and identify and locate the spatial region of the target workpiece in the stacked structure based on the definition information in the initial visual scene data; The initial visual scene data is analyzed, and the light and shadow interference components in the scene are separated by calculation and analysis to restore the intrinsic visual features of all stacked and occluded workpiece surfaces. Based on intrinsic visual features, the visible parts of the workpiece in the scene are identified, and the geometric information of these visible parts is used as a constraint to perform dynamic simulation of stacking stability in a virtual physical environment. The most likely three-dimensional geometric shape and spatial pose of the occluded part are deduced through iterative calculation, and a scene digital twin model containing the complete workpiece is constructed. Based on the scene digital twin model, the physical reachability, grasping stability and impact on the overall stacking structure of each potential grasping point are analyzed, and several grasping options with priority order are generated. For each grasping option, a dynamic confidence level is calculated that comprehensively reflects the accuracy of the pose corresponding to the grasping option and the safety of the operation process of executing the grasping option. Based on the gripping options and dynamic confidence, the gain on the visible area of ​​the target workpiece and the disturbance to the overall stacking stability after each gripping option is executed are evaluated, and the key occlusion removal options are identified from the gripping options. The robot is guided to perform the key occlusion removal option first, and after each grasping operation, the updated visual scene data is reacquired and iterated until the target workpiece is successfully located and grasped.

2. The method according to claim 1, characterized in that, The identification and location of the target workpiece in the spatial region of the stacked structure specifically includes: Receive the two-dimensional template image, three-dimensional design model, and marked areas in the initial visual scene data of the target workpiece as definition information; Based on the defined information, a target region with similar visual features to the target workpiece is matched in the initial visual scene data; For the target area that is successfully matched, the target area is transformed from two-dimensional image coordinates to three-dimensional spatial coordinate system to obtain the spatial area of ​​the target workpiece.

3. The method according to claim 2, characterized in that, The process of restoring the intrinsic visual features of all stacked and obscured workpiece surfaces specifically includes: Construct a library of optical reflection models that include several typical lighting conditions and workpiece surface materials; Based on the optical reflection model library, the initial visual scene data is analyzed, the inverse problem of the illumination equation is solved, and the direct illumination, indirect illumination and mutual reflection components in the initial visual scene are calculated. The direct illumination, indirect illumination, and mutual reflection components are separated from the initial visual scene data, and the intrinsic color and texture information of all stacked and occluded workpiece surfaces are preserved to obtain the intrinsic visual features.

4. The method according to claim 3, characterized in that, The construction of a scene digital twin model containing the complete workpiece specifically includes: Extract the visible edge contours and surface key points of all stacked and occluded workpieces from the intrinsic visual features, and reconstruct the three-dimensional point cloud data of the visible parts of all stacked and occluded workpieces. The three-dimensional point cloud data is locally registered with the structural model of the corresponding stacked and occluded workpieces to determine the pose of the visible part of the stacked and occluded workpieces in the global coordinate system, which serves as a constraint condition. Construct a simulation environment that includes gravity, friction, and collision detection, and randomly generate several possible pose configurations for the occluded parts based on the constraints. For each possible pose configuration, a physical stability assessment is performed, and the configuration that satisfies the static equilibrium condition and is consistent with the constraints of the visible part is selected as the optimal solution, thus constructing a digital twin model of the scene.

5. The method according to claim 4, characterized in that, The calculation of a dynamic confidence level for each grasping option, which comprehensively reflects the accuracy of the pose corresponding to the grasping option and the safety of the operation process of executing the grasping option, specifically includes: On the surfaces of each stacked and occluded workpiece in the digital twin model of the scene, potential gripping points are generated for each stacked and occluded workpiece based on curvature analysis and attitude stability criteria. The physical reachability of each potential gripping point on a stacked, occluded workpiece is evaluated using a collision detection algorithm. For physically reachable gripping points, calculate the relative positional relationship between potential gripping points and the centroids of corresponding stacked and occluded workpieces, the normal vector of the contact surface, and the friction cone, and quantitatively evaluate the gripping stability of each potential gripping point. For each potential gripping point, simulate the process of removing the stacked and occluded workpieces corresponding to that potential gripping point, and predict the degree of impact of the removal action on the stability of the stack structure; Based on the evaluation results of physical accessibility, grabbing stability, and the degree of impact on heap structure stability, a priority ranking of the grabbing options is generated, and dynamic confidence is calculated.

6. The method according to claim 5, characterized in that, The step of identifying the key obstruction removal option from the grabbing options specifically includes: For each grabbing option in the priority ranking, simulate the execution of the corresponding removal operation in the scene digital twin model, and calculate the change in the visible area of ​​the target workpiece under multiple views after removal; Evaluate the stability changes of the stacked structure after performing a removal operation in a scene digital twin model, including calculating the changes in the center of gravity offset and contact force distribution of the remaining stack and the occluded workpiece; By combining changes in visible area and stability, the obstacle clearing efficiency score for each grabbing option is calculated; Based on the obstacle removal effectiveness score and dynamic confidence level, the key obstruction removal options are selected from the crawling options.

7. The method according to claim 6, characterized in that, The step of acquiring updated visual scene data and iterating after each grasping operation until the target workpiece is successfully located and grasped specifically includes: Based on the key occlusion removal options, generate a collision-free motion trajectory that satisfies robot dynamics constraints. Control the robot to execute the motion trajectory and complete the stacking, grasping and moving of the obstructed workpieces corresponding to the key obstruction removal options; After each grabbing and moving operation is completed, visual data of the work area is re-acquired to obtain updated visual scene data; Based on updated visual scene data, update the scene digital twin model, grasping options, and key occlusion removal options; When the target workpiece is determined to be fully exposed and directly graspable, a final grasping instruction is generated for the target workpiece, controlling the robot to complete the grasping.

8. A positioning system based on visual perception, characterized in that, The system includes: The target workpiece identification and positioning module is used to receive the definition information of the target workpiece and, based on the definition information, identify and locate the spatial region of the target workpiece in the stacked structure in the initial visual scene data. The light and shadow interference removal module is used to analyze the initial visual scene data, calculate and analyze and separate the light and shadow interference components in the scene, and restore the intrinsic visual features of all stacked and occluded workpiece surfaces. The scene digital twin model construction module is used to identify the visible parts of the workpiece in the scene based on intrinsic visual features, and use the geometric information of these visible parts as constraints to perform dynamic simulation of stacking stability in a virtual physical environment. Through iterative calculation, the most likely three-dimensional geometric shape and spatial pose of the occluded part are deduced to construct a scene digital twin model containing the complete workpiece. The grab option generation module is used to analyze the physical reachability, grab stability and impact on the overall stacking structure of each potential grab point based on the scene digital twin model, generate several grab options with priority order, and calculate a dynamic confidence level for each grab option that comprehensively reflects the accuracy of the pose corresponding to the grab option and the safety of the operation process of executing the grab option. The critical obstruction identification module is used to evaluate the gain on the visible area of ​​the target workpiece and the disturbance to the overall stacking stability after each gripping option is executed, based on the gripping options and dynamic confidence, and to identify the critical obstruction removal option from the gripping options. The robot operation execution module is used to prioritize guiding the robot to execute the key occlusion removal option, and after each grasping operation, it reacquires updated visual scene data and iterates until the target workpiece is successfully located and grasped.

9. The system according to claim 8, characterized in that, The key obstruction identification module includes: The potential gripping point generation module is used to generate potential gripping points for each stacked and occluded workpiece surface in the digital twin model of the scene based on curvature analysis and attitude stability criteria. The collision detection and evaluation module is used to evaluate the physical reachability of each potential gripping point on a stacked or occluded workpiece using a collision detection algorithm. The stability quantification module is used to calculate the relative positional relationship between potential gripping points and the centroids of corresponding stacked and occluded workpieces, as well as the normal vector and friction cone of the contact surface for physically reachable gripping points, and to quantitatively evaluate the gripping stability of each potential gripping point. The removal process simulation module is used to simulate the process of removing the stacked and occluded workpieces corresponding to each potential gripping point, and predict the impact of the removal action on the stability of the stack structure. The priority ranking module is used to generate a priority ranking of the grabbing options by comprehensively considering the evaluation results of physical reachability, grabbing stability and the degree of impact on heap structure stability, and to calculate the dynamic confidence level.

10. The system according to claim 9, characterized in that, The robot operation execution module includes: The motion trajectory generation module is used to generate a collision-free motion trajectory that meets robot dynamics constraints based on the key occlusion removal options. The motion trajectory execution module is used to control the robot to execute motion trajectories and complete the stacking, grasping and moving of workpieces corresponding to the key obstruction removal options. The visual data re-acquisition module is used to re-acquire visual data of the work area after each grabbing and moving operation to obtain updated visual scene data. The model update module is used to update the scene digital twin model, grasping options, and key occlusion removal options based on updated visual scene data. The gripping control module is used to generate the final gripping command for the target workpiece when the target workpiece is determined to be fully exposed and can be gripped directly, and to control the robot to complete the gripping.

Citation Information

Patent Citations

  • Illumination model system and implementation method

    CN103699733A

  • Mechanical arm grabbing control method and system based on digital twinning

    CN119589673A