Trajectory planning system and method for shoe last gripping robot based on multi-modal perception
By employing multimodal perception and self-learning mechanisms, the recognition accuracy and stability of the shoe last gripping device have been improved, solving the problem of gripping failure in complex environments in traditional systems and enabling adaptive and flexible production.
Patent Information
- Application Number
- CN202511023512.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-24
AI Technical Summary
Existing shoe last gripping devices suffer from insufficient recognition accuracy, unstable gripping, and lack of self-learning ability when faced with complex lighting, occlusion, diverse materials, and dynamic conditions, resulting in a high gripping failure rate and making it difficult to achieve flexible production.
It employs multimodal perception to fuse visual images, structural acoustic waves, and infrared contours, and combines interference light illumination to determine the material. A stable window mechanism ensures the target remains stationary, a tactile sensing membrane is used for fine-tuning the gripper, and a self-learning module is introduced to optimize the grasping strategy.
It improves the recognition accuracy and clamping stability of shoe last grasping, reduces the grasping failure rate, realizes autonomous learning and strategy optimization, and adapts to complex industrial environments.
Smart Images

Figure CN120516725B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of intelligent manufacturing and robot control, and specifically relates to a shoe last grabbing robot trajectory planning system and method based on multi-modal perception. BACKGROUND
[0002] With the development of intelligent manufacturing and personalized customization, higher requirements are put forward for the accurate grabbing and flexible operation of special-shaped objects in automatic production lines. As a core intermediate part in the shoemaking industry, shoe lasts have complex structures, various materials, random postures, and surface reflection or surface mutation, which leads to long-term reliance on manual intervention in the automatic grabbing link, and becomes one of the key bottlenecks restricting the improvement of flexible production efficiency.
[0003] Existing shoe last grabbing devices are mostly based on traditional industrial robot arms combined with RGB cameras or depth vision modules for target recognition and path planning. Their working principles generally include image acquisition, contour extraction, grabbing point estimation, and fixed path execution. However, such systems generally face the following technical problems in actual application:
[0004] Single perception information, limited recognition accuracy: Traditional systems mostly rely on single-modal visual information (such as two-dimensional images or shallow depth maps) for target recognition, which is difficult to cope with complex scenes such as light interference, occlusion, stacking chaos, or material reflection, leading to inaccurate target positioning and large posture estimation errors.
[0005] Lack of dynamic judgment ability for physical state: Existing solutions generally assume that the target is in a static state, but in actual production lines, shoe lasts may be in a dynamic state due to rolling, unstable stacking, or vibration. Due to the lack of stability discrimination mechanism, the system often blindly executes grabbing when the target is not stable, which can easily cause grabbing failure or even equipment collision.
[0006] Coarse jaw control, insufficient fitting ability: Many jaws are only based on fixed path and fixed opening and closing control, and do not have the ability to adapt to the target surface in real time. Especially when facing curved shoe lasts or targets with uneven materials, problems such as clamping deviation, slipping, or insufficient clamping force often occur, affecting the reliability of grabbing.
[0007] Lack of learning mechanism, grabbing strategy cannot be optimized: Traditional systems are mostly offline planning mode, without self-feedback and experience accumulation function after operation, and cannot update parameters and improve strategies for repeated shoe types or failed scenes, making it difficult to achieve long-term autonomous performance improvement.
[0008] Therefore, in view of the above problems, it is urgent to provide a shoe last grabbing system with multi-modal fusion perception, target stability recognition ability, end fitting fine-tuning support, and grabbing strategy self-learning, which can improve the intelligent grabbing level of industrial special-shaped objects as a whole and promote the flexible process of shoe-making automatic production lines. SUMMARY
[0009] The present application aims to provide a shoe last grasping robot trajectory planning system and method based on multi-modal perception, which breaks through the statistical correlation limitations of traditional technology and realizes the upgrade of the diagnosis and treatment paradigm from "correlation" to "causality".
[0010] The technical solutions adopted by the present application are as follows:
[0011] A shoe last grasping robot trajectory planning system based on multi-modal perception, comprising:
[0012] a. a shoe last storage area for carrying disorderedly placed shoe lasts;
[0013] b. a mechanical arm grasping assembly, comprising an industrial robot arm with at least six degrees of freedom, a gripper, and a tactile sensing film and an interference perception device arranged on the surface of the gripper;
[0014] c. a multi-modal perception device, comprising a vision camera arranged above the operation area, a structural acoustic transceiver arranged on both sides of the operation area, and an infrared contour scanner;
[0015] d. a posture recognition module for extracting the toe point, heel point and instep point of the target shoe last based on the visual image and contour data, and generating a shoe last main shaft posture vector accordingly;
[0016] e. a surface interference perception module for irradiating and collecting interference fringes by double-wavelength interference light during the approach of the gripper to the shoe last, judging the surface roughness, reflection characteristics and material category of the shoe last, and outputting the clamping strategy parameters;
[0017] f. a grasping point optimization module for constructing a shoe last mechanical model according to the distribution of the toe point and the heel point, and optimizing the clamping point in the instep area to minimize the grasping torque, and adjusting the gripper fitting strategy combined with the surface characteristics;
[0018] g. a stable window generation module for determining whether the shoe last is in a stable and stationary state by structural acoustic wave echo difference analysis and high-frequency micro-vibration detection, and marking the graspable area as a stable window;
[0019] h. a path planning module for generating an optimal grasping path by backtracking in the stable window area with the grasping point as the target endpoint, combining obstacle voxel map and perception modal quality score;
[0020] i. a fine-tuning control module for completing surface fitting by tactile sensing film feedback and gripper posture fine-tuning action in the grasping end stage;
[0021] j. A self-learning module is used to record the failure cases of grabbing, including the stability area, surface state, clamping strategy and trajectory information, and update the local path strategy optimization table.
[0022] The visual camera is used to collect the overhead image of the operation area, and the shoe last edge profile, highlight area and possible overlapping occlusion boundary are extracted through image processing.
[0023] The interference sensing device includes a red and green laser cross emitter and a linear array photoelectric receiver. The receiver is used to collect the surface fringe pattern and judge the material and friction characteristics of the shoe last according to its periodicity, density and deformation.
[0024] The stability window generation module combines the edge profile jitter degree of the target object in the continuous image frame and the stability of the structural sound wave echo to judge whether the area is in a stable state.
[0025] The grabbing point optimization module preferentially avoids the surface area with high reflectivity or low roughness, and limits the clamping force within the safe threshold range for progressive adjustment.
[0026] The path planning module only plans the trajectory within the shoe last range marked by the stability window, and ignores the targets in the dynamic interference area or inclined stacked edge area.
[0027] The tactile sensing film on the surface of the clamping jaw is a deformable capacitive structure, which is used to sense the contact difference on both sides of the clamping jaw in real time, and drive the clamping jaw to perform a micro-motion action within 5° in the pitch or roll direction.
[0028] The fine-tuning control module is activated within 30mm range from the end of the target shoe last, and completes the micro-posture correction according to the tactile film data to ensure that the clamping jaw is in contact with the surface of the shoe last.
[0029] The self-learning module uses a lightweight local database to record the surface type, clamping effect and failure reason of each grabbing action, and calls the recommended parameter group for priority path and clamping strategy matching after identifying the same shoe last model.
[0030] A shoe last grabbing method of the system, comprising the following steps:
[0031] S1) Collect multi-modal sensing data of the shoe last area, including visual image, structural sound wave and infrared profile information;
[0032] S2) Extract the toe point, heel point and instep point, generate a posture vector, and construct a shoe last mechanical model;
[0033] S3) Activate the interference sensing module to obtain the interference fringe of the target shoe last surface and judge its material and friction characteristics;
[0034] S4) identifying whether the current target is in a stable state, if in a stable window area, entering a trajectory planning stage, otherwise switching to the next candidate target;
[0035] S5) based on the optimal grasping point in the stable window, backtracking to generate an obstacle avoidance path and output a clamping strategy;
[0036] S6) the robot executes the trajectory to approach the target, activates the tactile film at the end stage to fine-tune the posture, and completes the surface fitting;
[0037] S7) executing the grasping and judging success or failure, if failed, recording parameter information for subsequent optimization learning.
[0038] As described above, due to the adoption of the above technical scheme, the present application has the following beneficial effects:
[0039] The present application proposes a shoe last grasping robot trajectory planning system based on multi-modal perception, which constructs a multi-dimensional cooperative control system from perception, judgment, decision-making to execution around the actual problems of complex grasping objects, disordered stacking, variable material and high clamping stability requirements, and has a compact overall structure, reasonable function coupling and significant comprehensive improvement effect compared with the prior art.
[0040] Firstly, in the target recognition stage, the system breaks through the limitation of traditional single RGB vision or depth map, and innovatively integrates multiple modalities such as visual image, structural sound wave, infrared contour and dual-wavelength interference perception. Through principal component analysis to extract the main shaft direction of the shoe last, and combining with the point cloud data to locate the key points of the toe, instep and heel, high-precision posture recognition is realized. Especially the introduction of dual-wavelength interference light source and linear array photoelectric receiver, the target surface is analyzed by non-contact fringe, starting from the interference fringe period, density and deformation degree, the surface material type and friction grade are judged, which provides decision basis for intelligent matching of gripper strategy, and solves the problem that traditional system cannot distinguish the physical properties of the target.
[0041] In view of the implicit technical bottleneck of difficult target stability judgment and high failure rate of grasping, the present application further proposes a "stable window" mechanism, which uses the root mean square change of the structural sound wave echo and the Hausdorff distance calculation of the continuous image edge to judge whether the target is in a static and controllable state, and only in the stable window range to plan and execute the path, which reduces the mis-grasping, empty-grasping and other phenomena caused by target shaking, sliding or non-stable stacking from the source, and enhances the robustness of the system in unstructured environment.
[0042] In the clamping execution stage, the application integrates a capacitive tactile sensing film on the surface of the clamping jaw, combines the micro-pose adjustment capability in the end pitch and roll direction, and constructs a fine fitting control mechanism. When the clamping jaw enters the target within a range of 30 mm, the system automatically corrects the uneven clamping on both sides according to the tactile film feedback, realizes the pose fine adjustment within ±5°, effectively improves the clamping symmetry and fitting quality, and avoids the problems of clamping deviation and slipping caused by irregular target structure or uneven surface.
[0043] In addition, the application introduces a self-learning module to record the material type, clamping jaw strategy, path parameter and execution result of each grabbing operation, and automatically updates the optimization database for failed cases. In repeated tasks or similar last identification, the system can call the historical successful parameter group to execute preferentially, so as to realize the strategy evolution and long-term improvement of the grabbing quality.
[0044] In summary, through the system integration and optimization of four dimensions of perception multi-sourcing, judgment refinement, control fine-tuning and strategy self-learning, the performance of the last grabbing system in identification accuracy, clamping stability and execution success rate is comprehensively improved, especially suitable for real industrial application scenarios with stacked disorder, complex structure and mixed materials, has clear practical value and broad application prospect, and is obviously superior to the prior art. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 The figure is a schematic diagram of the system architecture of the application;
[0046] Figure 2 The figure is a schematic diagram of the method flow of the application. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical scheme and advantages of the application clearer and more understandable, the application will be further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the application and do not limit the application.
[0048] Referring to Figure 1 The application relates to a last grabbing robot trajectory planning system based on multi-modal perception, comprising:
[0049] a. a last storage area for carrying disordered placed lasts;
[0050] b. a mechanical arm grabbing component, comprising an industrial robot arm with at least six degrees of freedom, a clamping jaw, and a tactile sensing film and an interference sensing device arranged on the surface of the clamping jaw; the tactile sensing film on the surface of the clamping jaw is a deformable capacitive structure, which is used to sense the contact difference on both sides of the clamping jaw in real time, and drive the clamping jaw to perform a micro-motion action within 5° in the pitch or roll direction;
[0051] c. Multi-modal perception device, including a visual camera set above the operation area, a structural acoustic transceiver set on both sides of the operation area, and an infrared profile scanner; the visual camera is used to collect overhead images of the operation area, and extract the last edge profile, highlight area, and possible overlapping occlusion boundary through image processing;
[0052] d. Pose recognition module, used to extract the toe point, heel point, and instep point of the target last based on the visual image and profile data, and generate a last main shaft pose vector accordingly;
[0053] e. Surface interference perception module, used to judge the surface roughness, reflection characteristics, and material category of the last during the approach of the gripper, output the clamping strategy parameters by irradiating and collecting interference fringes through dual-wavelength interference light; the interference perception device includes a red and green laser cross emitter and a linear array photoelectric receiver, which is used to collect surface fringe patterns and judge the last material and friction characteristics according to the periodicity, density, and deformation;
[0054] f. Grasping point optimization module, used to construct a last mechanical model according to the distribution of the toe point and the heel point, and select the clamping point in the instep area that minimizes the grasping torque, and adjust the gripper fitting strategy combined with the surface characteristics; the grasping point optimization module preferentially avoids the surface area with high reflectivity or low roughness, and limits the progressive adjustment of the clamping force within the safety threshold range;
[0055] g. Stable window generation module, used to determine whether the last is in a stable and stationary state through structural acoustic echo difference analysis and high-frequency micro-vibration detection, and mark the graspable area as a stable window; the stable window generation module combines the edge profile jitter degree of the target object in the continuous image frame and the stability of the structural acoustic echo for real-time comparison to determine whether the area is in a stable state;
[0056] h. Path planning module, used to generate an optimal grasping path by backtracking in the stable window area with the grasping point as the target endpoint, combined with the obstacle voxel map and the perception modal quality score; the path planning module only plans the trajectory within the last range marked by the stable window, and ignores the target in the dynamic interference area or the inclined stacking edge area;
[0057] i. Fine-tuning control module, used to complete surface fitting through tactile sensing film feedback and gripper pose fine-tuning actions in the grasping end stage; the fine-tuning control module is activated within 30 mm of the target last end, and completes micro-pose correction according to the tactile film data to ensure that the gripper is in contact with the last surface;
[0058] j. a self-learning module for recording failed cases of grabbing, including stability area, surface state, clamping strategy and trajectory information, and updating a local path strategy optimization table; the self-learning module records the surface type, clamping effect and failure reason of each grabbing action by using a lightweight local database, and calls a recommended parameter group to perform priority path and gripper strategy matching after identifying the same shoe last model.
[0059] Further, referring to Figure 1 and 2 , a shoe last grabbing method of the system comprises the following steps:
[0060] S1) Before performing the shoe planting grabbing operation, the system first performs a multi-modal perception data acquisition process of the shoe planting area to ensure the accuracy and robustness of subsequent pose estimation and path planning.
[0061] Specifically, the system synchronously acquires visual images, structural sound wave depth data and infrared contour information, and realizes multi-source fusion through a unified space-time alignment mechanism. First, a top-down image is acquired by an industrial vision camera arranged above the operation area, the image resolution is 1280x1024 pixels, the frame rate is 30 frames per second, and x and y are two-dimensional spatial coordinates of a pixel point in the image. To suppress image noise and enhance edge stability, the image is first preprocessed by Gaussian filtering, and the processing formula is:
[0062] ;
[0063] wherein, represents the filtered image, is a two-dimensional Gaussian kernel, is a kernel standard deviation, and the value range is 0.8 to 1.5, and the symbol represents a two-dimensional convolution operation. The filtered image is extracted by a Canny edge detection algorithm to obtain the shoe planting boundary contour, and the area where there may be overlapping occlusion is identified.
[0064] At the same time, structural sound wave transceivers are arranged on both sides of the operation area to acquire the depth information of the shoe planting surface. The structural sound wave transceivers emit high-frequency pulse signals and receive echo signals, and calculate the depth value according to the echo time difference, and the calculation formula is:
[0065] ;
[0066] wherein, is the depth value of the shoe planting surface (unit: mm), is the propagation speed of the sound wave in the air, about 340 mm / ms, is the time difference between the round trip propagation of the structure-borne acoustic wave (unit: ms). To ensure the accuracy of the ranging, the emission angle of the structure-borne acoustic wave probe is limited to within ±30°, and dynamic registration is performed in conjunction with the occlusion area compensation strategy.
[0067] In order to further improve the integrity of spatial contour reconstruction, the system is also equipped with a line-scanning infrared contour scanner, which scans point by point along the axis of the operating area to obtain the vertical height of the shoe last surface. The infrared scanning output is based on the height value of the temperature difference mapping, and the conversion formula is:
[0068] ;
[0069] in, is the infrared profile height (unit: mm), is the infrared temperature difference comparison value of the scanning point (unit: °C), It is the temperature difference-height conversion coefficient, and its value range is generally 2.5 to 4.0 mm / °C. It is calibrated and corrected according to the surface material type and scanning distance.
[0070] The above three types of perception data are collected concurrently by the time synchronization controller of the main control system, and the time error is controlled within ±10ms. At the spatial level, the coordinate data output by each perception channel is calibrated through the spatial transformation matrix Perform unified projection and the conversion formula is:
[0071] ;
[0072] in, Represents the perception channel The original point position (homogeneous coordinate form) obtained under is the spatial transformation matrix between the channel and the main operation space, This unifies the fusion points in the operating coordinate system. This spatiotemporal alignment mechanism allows for complete recognition and regional construction of disordered shoe lasts in the operating area while maintaining perception accuracy, providing high-precision input data for subsequent posture recognition and grasping path planning.
[0073] S2) Extract the toe point, heel point, and waist point of the shoe, generate the posture vector, and construct the mechanical model of the shoe last;
[0074] After the multimodal perception data is collected and aligned, the system extracts the toe point, heel point and waist point based on the fused image and three-dimensional point cloud data, and generates a posture vector and shoe plant mechanical model to support subsequent grasping point optimization and path planning.
[0075] First, the system performs contour extraction on the fused image to obtain the shoe plant boundary point set , Two-dimensional coordinates of the i-th boundary point. The covariance matrix of the boundary point set is calculated by the least-enclosing rectangle fitting and principal component analysis (PCA) , and the maximum eigenvector direction is extracted as the initial estimation of the shoe last main axis direction . The mathematical expression of the main axis direction is:
[0076] ;
[0077] wherein, denotes the unit vector of the shoe last main axis direction, is the candidate direction vector, is the two-dimensional covariance matrix of the boundary point set.
[0078] After the main axis direction is determined, the system projects the boundary point set in the direction to find the point pair with the maximum projection distance, which are defined as the toe point and the heel point . The selection basis is:
[0079] ;
[0080] wherein, denotes the vector dot product operation, is the boundary point coordinate. This step ensures the stability and direction consistency of the positioning of the front and rear directions of the shoe last.
[0081] Then, on the main axis line segment connecting the toe point and the heel point , the system selects the middle part of the line segment about 30%-40% from the toe as the shoe waist candidate area , scans the width function in the vertical direction in this area, and selects the place with the minimum width as the waist point :
[0082] ;
[0083] wherein, denotes the left and right boundary spacing of the point in the direction perpendicular to the main axis direction. This strategy can effectively avoid the interference caused by structural asymmetry and ensure the reliability of the waist clamping.
[0084] Subsequently, the system constructs the pose vector based on the toe point and the heel point, which is defined as:
[0085] ;
[0086] wherein, is the unit pose vector, pointing to the front of the shoe last, Represents the vector modulus.
[0087] After completing the geometric posture extraction, the system further constructs a simplified mechanical model. As the clamping point, it is assumed that the center of gravity of the shoe last is borne by the toe and heel respectively 40% and 60% of the equivalent gravity and , calculate the resultant moment generated by the two ends on the waist :
[0088] ;
[0089] in, represents the vector cross product operation, The modulus of the resultant torque is used to evaluate the mechanical stability of the current shoe waist point. If the torque value is lower than a set threshold (e.g., 20 N-mm), the point is considered the preferred clamping position.
[0090] Through the above steps, the system completes the accurate identification of the key points of the shoe last and the modeling of its mechanical properties, providing a high-precision, low-interference structural foundation for subsequent grasping point optimization and strategic decision-making.
[0091] S3) After completing the shoe last posture extraction and preliminary screening of candidate grasping areas, the system automatically activates the interference perception module and performs a non-contact scan of the target shoe last surface to extract the interference fringe pattern and determine its material and friction characteristics, providing a basis for subsequent clamping strategies.
[0092] The interference sensing module uses dual-wavelength laser interference to illuminate the surface, including red light (wavelength nm) and green light (wavelength nm) cross-projection system, and high-precision linear array photoelectric receivers with a resolution of mm. When the gripper approaches the target shoe last When the target surface is within the range of mm, the system starts the interferometric measurement device, illuminates the target surface at an oblique angle and simultaneously collects fringe images. ,in represents the laser wavelength channel, Represents the spatial coordinate in the direction of the receiver.
[0093] After the image acquisition is completed, the system performs frequency domain and spatial analysis on the interference fringe pattern. First, the sampling length is identified by the counting algorithm. Number of complete fringes in the mm interval , calculate the interference fringe period :
[0094] ;in, For the stripe period, unit: mm, reflecting the micro-uniformity of the surface. Further, the stripe density is calculated The number of stripes per unit length is represented by:
[0095] This index can be used to evaluate the overall roughness of the surface, where The typical value range of stripes / mm.
[0096] In order to identify the degree of local distortion caused by the discontinuity of the microstructure, the system calculates the gray level curve of the image The second derivative is calculated and averaged to obtain the amount of stripe deformation :
[0097] ;
[0098] Where n is the number of sampling points, is the i-th point, reflecting the degree of local surface curvature. Low deformation value indicates that the surface is smooth and consistent, suitable for clamping; high deformation value indicates the presence of micro convex, cracks or edge reflection, which should be avoided as clamping sites.
[0099] The system makes a comprehensive judgment based on the established material identification database, combining the interference parameters T, D and three indicators. For example: if mm, stripes / mm, and , it is determined as high-friction materials such as leather or matte rubber, suitable for low-pressure clamps; if mm and , it is mostly smooth plastic or painted surface, and the system will increase the clamping force or adjust the fitting angle; if the periodicity is broken or there is strong reflection interference in the interference pattern, the system will automatically mark the area as "high-risk clamping area" to avoid subsequent selection of the grabbing path.
[0100] After completing the surface recognition, the system outputs the material category, friction level and clamping suggestion parameters, including the initial pressure of the clamps (such as 0.3-0.7N / mm 2 ) and the fine-tuning posture direction (pitch or roll angle priority), and transmits the parameter set to the clamp fine-tuning control module, providing an important guarantee for the success rate of end fitting and clamping.
[0101] Through this step, the system can quickly perceive the changes in the surface structure and physical properties of the shoe without contact, forming an important pre-judgment condition for adaptive grabbing of complex targets, thereby improving the overall robustness and precision.
[0102] S4) After the completion of the shoe material recognition, the system enters the target stability judgment stage, through the fusion of structural acoustic wave echo signal and continuous image edge jitter analysis, to determine whether the current target shoe last is in a stable and static state that can be grabbed. If the stability criterion is met, the target is marked as a "stable window" area, allowing it to enter the path planning stage; otherwise, the system automatically switches to the next candidate target for continuous judgment.
[0103] First, the system activates the bilateral structural acoustic wave transceiver, continuously transmits high-frequency pulses to the target shoe last surface, and records the echo intensity change sequence , where represents the time stamp, the sampling interval is 5ms, and the total recording time is 300ms. The root mean square of echo change is calculated :
[0104] ;
[0105] where, is the number of sampling points, is the th echo signal intensity, is its average value. If (unit normalized), it is determined that the structural acoustic wave signal is stable.
[0106] At the same time, the system synchronously analyzes the time sequence jitter degree of the target edge contour line in the visual image sequence, calculates the edge change value between the current frame and the previous frame using Hausdorff distance :
[0107] ;
[0108] where, is the coordinate point in the edge point set, represents the Euclidean distance. The system calculates the average edge jitter value of the continuous 10 frames of images, and if mm, it is considered that the target visual contour is static.
[0109] The system jointly judges the structural acoustic wave stability index and the image edge stability index, defines the stability function as:
[0110] ;
[0111] where, is the weight coefficient. If , the target state is determined to be "stable", and the system marks the area as a stable window for the path planning module to call.
[0112] If the stability function If the threshold is not reached, the system determines that the current target is disturbed by external forces, is in a state of rocking, sliding or tilting instability, does not enter the path generation process, and automatically switches to the next target to be grasped and repeats the above stability judgment process until an operable object is found or a "no graspable target" is reported.
[0113] Through this step, the system can effectively identify non-stationary targets in the physical environment, avoid grasping failure or misoperation, and significantly improve the operation safety and stability of the robot system in a dynamic environment.
[0114] S5) After the target shoe last is identified as a stable state and successfully marked as a "stable window" area, the system immediately enters the path planning phase. This phase takes the optimal grasping point output by the grasping point optimization module as the path endpoint, integrates the current operating area obstacle distribution and perception quality score, uses a backtracking path planning algorithm to generate an obstacle avoidance trajectory, and simultaneously outputs a set of clamping strategy parameters adapted to the surface characteristics.
[0115] First, the system obtains a grasping point in the stable window, which is located in the instep area and meets the minimum grasping torque and surface friction optimization conditions. Then, the system constructs a three-dimensional operating area voxel map with a resolution of 5mmx5mmx5mm, and marks the area above the background noise threshold in the point cloud as obstacle voxels for obstacle avoidance determination. To achieve path generation, the system takes the current robot end position as the starting point, uses heuristic backtracking search (such as A' or RRT*) to find the shortest collision-free path from , and calculates the path cost function :
[0116] ;
[0117] where: is the distance between two points (unit: mm); is the perception modality quality score at that path point position, with a high score indicating clear and unobstructed images; is the perception penalty weight, with an empirical value of 20-50; is the number of path segments.
[0118] During the path generation process, the system only allows path points to fall within the spatial area corresponding to the stable window and avoids low-confidence voxel areas marked as dynamic interference or stacking edges. If the path fails to meet this condition, the system automatically abandons the current target and switches to the next candidate path. After the generation is completed, the system calls the clamping strategy parameter set output by the surface interference perception module and binds it to the clamping action instruction at the end of the path. This parameter set includes:
[0119] Initial pressure of gripper , the range is generally 0.3-0.7N / mm 2 ;
[0120] Attitude alignment angle adjustment (pitch angle , roll angle ), controlled within ±5°;
[0121] Fine-tune the trigger distance ,The default setting is to start the fine-tuning module when the gripper is 30 mm away from the target;
[0122] Safety redundant action settings, such as retreat paths and retry instructions after grasping failure.
[0123] Finally, the path planning module outputs the complete grasping path trajectory The matching gripping control parameter group is pushed to the robot controller and fine-tuning module for subsequent grasping execution.
[0124] Through this step, the system realizes the coordinated optimization of paths and strategies coupled with dynamic perception quality and mechanical conditions in complex environments, ensuring that the grasping process is both obstacle-avoiding and precise, safe and adaptive, and improving the overall operation success rate.
[0125] S6) The robot executes the trajectory to approach the target and activates the tactile film at the end stage to fine-tune the posture and complete the surface bonding;
[0126] Specifically, the path planning module outputs the optimal trajectory After the corresponding gripping strategy is determined, the robot immediately executes the trajectory approach action and controls the gripper to move along the predetermined path to the target gripping point. Approach. During the approach process, the system continuously detects the trajectory progress. When the distance between the gripper and the target point is less than the set trigger distance, (The default value is 30mm), the tactile sensing film on the surface of the gripper is automatically activated and enters the end fine-tuning stage.
[0127] The tactile sensor film is a deformable capacitive structure embedded in the surface of the left and right arms inside the gripper. Its output signal is the capacitance change value. and , respectively representing the contact strength changes on the left and right sides of the gripper. The system calculates the difference between the two sides in real time. :
[0128] ;
[0129] wherein, if ( is set as the balance threshold (usually 0.02 pF), it is considered that the clamping jaw contact is symmetrical and the surface is well fitted without adjustment; if exceeds the threshold, the system starts the micro-posture correction mechanism.
[0130] The specific correction method is as follows: when : indicates that the left side contacts first, and the clamping jaw performs a small roll adjustment (the right side tilts forward), and the correction angle satisfies:
[0131] ;
[0132] When : indicates that the right side contacts first, and the clamping jaw performs a roll adjustment in the opposite direction, and the correction angle is calculated in the same way; wherein, is the amplification factor between the capacitance difference and the posture correction angle, and the value range is generally 60-100 deg / pF, to ensure that the fine adjustment is completed within ±5°.
[0133] At the same time, if the sensing film detects that both sides are uniformly contacted but the overall pressure is insufficient both <0.1 pF, the system can automatically adjust the pressing speed and pressing stroke of the clamping jaw according to the friction level label in the material identification stage, to avoid slipping on low-friction surfaces or rebounding on high-hardness surfaces. The posture fine adjustment process is controlled to be completed within 500 ms, and the system pauses for 100 ms after completing the correction to ensure mechanical stability, and then issues a clamping execution instruction to start the clamping jaw closing action and formally complete the surface fitting.
[0134] Through this step, the robot can make micro-response adjustments to the target surface state before actual grabbing, realize automatic compensation of contact deviation caused by irregular structures or uneven materials of the shoe tree, and thus improve the fitting consistency and grabbing stability, and reduce the risk of grabbing failure caused by clamping jaw angle error or local disengagement.
[0135] S7) After completing the end fine adjustment and achieving surface fitting, the robot immediately performs the clamping jaw closing action to complete the grabbing operation of the target shoe. During the grabbing execution, the system monitors the clamping force feedback, the touch film change trend and the mechanical arm load change in real time to determine whether the grabbing process is successful.
[0136] First, the system detects the actual contact capacitance value change comparison after the clamping jaw is closed, and if the left and right touch film feedback values , all greater than 0.2 pF and the difference between them , is considered to be clamped symmetrically; at the same time, the load change value ΔF measured by the end force sensor after grabbing needs to meet the following conditions:
[0137] ;
[0138] Wherein: is the actual load at the end of the mechanical arm after clamping (unit: N); is the empty load measurement value before clamping; is the minimum effective grabbing load, generally in the range of 0.8-1.2 N (calibrated according to the type of shoe).
[0139] If the above three conditions are met, the system determines that this time of grabbing is successful, and marks the current grabbing action as an effective operation.
[0140] If any condition is not met (such as clamping asymmetry, insufficient change in capacitance, no significant increase in load, etc.), it is determined that this time of grabbing fails. At this time, the self-learning module is automatically called to record the key parameters of the current grabbing task, including but not limited to the following:
[0141] Shoe number and shape feature code;
[0142] Current grabbing point position and surface material classification;
[0143] The stripe period T and the deformation degree Δφ output by the interference sensing module;
[0144] The stability score function S and the path cost function C(P);
[0145] The initial pressure of the gripper, the fine-tuning angle, and the contact value of the tactile film;
[0146] Grabbing failure type label (such as slipping, falling off, clamping deviation, empty clamping, etc.).
[0147] All information is stored in a local lightweight grabbing optimization database, and a mapping relationship between the current operation and the failure label is established. When the system encounters the same type of shoe last next time, it recognizes its shape and surface code, then calls the successful parameter group of the target of this type preferentially, and avoids the grabbing points or strategies that have been recorded as failures.
[0148] In addition, if grabbing of a certain type of shoe last fails for three consecutive times, the self-learning module will issue a risk reminder to the user and suggest re-verifying the interference sensing threshold, the path weight, or the clamping parameters of the gripper to avoid long-term repeated failures.
[0149] Through this step, the system builds a closed-loop grasping optimization mechanism, realizes learning from failure experience, and automatic adjustment from material adaptation to parameter recommendation, improves overall intelligence and stability, and enables the system to have continuous adaptation and evolution capabilities when facing complex heterogeneous targets.
[0150] Furthermore, building on the existing system, capacitive tactile sensing films have been added to the gripper surfaces to detect changes in contact capacitance between the left and right grippers and the shoe last. This invention further improves tactile signal processing by introducing a contact sequence encoding mechanism in the time domain to assist in identifying abnormal postures where the target shoe last appears "nearly upright" but actually supports asymmetrically.
[0151] Specifically, when the gripper approaches the target point, the system samples the capacitance change curves on the left and right sides at 1ms intervals, which are recorded as follows: Left tactile curve: Right tactile curve: ;
[0152] When it is detected that a curve on one side meets the following contact judgment conditions for the first time:
[0153] ;in, is the initial capacitance baseline of this side; The minimum contact capacitance threshold is usually 0.15pF. is the time point when the threshold is first reached.
[0154] Record the first contact time of the left and right sides , calculate the time difference: ;
[0155] If the following skew contact conditions are met:
[0156] ; The system determines that the current grasping point has an abnormal support posture, that is, the target shoe may be tilted, overturned, or embedded in an asymmetric support surface, although the visual main axis posture vector is normal, where is the contact time difference threshold, which is generally set to ; is the capacitance difference threshold, generally .
[0157] At this point, the system can automatically perform one of the following operations:
[0158] Send feedback to the path planning module and mark the current point as a "non-preferred grasping point";
[0159] If there are spare stable windows and grasping points, jump to the next grasping candidate point;
[0160] If there are no alternative points, the system prompts the operator to perform a visual review or adjust the target placement strategy.
[0161] In addition, the contact sequence code is also stored in the self-learning module synchronously as one of the "high-risk posture feature" identification conditions for subsequent similar shoe types.
[0162] Even if the posture vector is correct, the traditional visual recognition may still cause actual support deviation due to uneven stacking or differences in the hardness of the supports. The contact sequence and time difference collected by the tactile film can quickly perceive the physical contact abnormalities of the target, and make up for the visual misjudgment risk.
[0163] The first feedback habit of the human hand when contacting the object can also be simulated, so that the robot has the local abnormal identification ability of "first touch first knowledge", and the intelligent level of the grasping decision is improved.
[0164] In addition, eccentric targets are prone to clamping deviation or eccentric torsion, and the scheme can issue a correction signal in advance before the closing action, thereby avoiding operation accidents such as empty clamping or clamping turning of the clamping jaw in the weak structure area.
[0165] The above only describes the preferred embodiments of the present application and is not intended to limit the present application, and any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A shoe last grasping method based on multimodal perception, characterized in that: A trajectory planning system for a shoe last grasping robot based on multimodal perception is provided, the system comprising: a. Shoe last storage area, used to carry disorderly placed shoe lasts; b. A robotic arm grasping assembly comprising an industrial robot arm having at least six degrees of freedom, a gripper, and a tactile sensing film and an interference sensing device disposed on the surface of the gripper; c. A multimodal sensing device, comprising a visual camera positioned above the operating area, a structured acoustic wave transceiver, and an infrared profile scanner positioned on either side of the operating area; d. A posture recognition module, which extracts the toe, heel, and waist points of the target shoe last based on visual images and contour data, and generates the main axis posture vector of the shoe last accordingly; e. Surface interference sensing module, used to detect and analyze interference fringes generated by dual-wavelength interference light as the gripper approaches the shoe last, determine the shoe last's surface roughness, reflectivity, and material type, and output gripping strategy parameters. f. The gripping point optimization module is used to construct a mechanical model of the shoe last based on the distribution of the toe and heel points. It then selects the gripping point in the shoe waist area that minimizes the gripping torque and adjusts the gripper's fit strategy based on surface characteristics. g. A stable window generation module, which determines whether the shoe last is in a stable and stationary state through differential analysis of structural acoustic wave echoes and high-frequency micro-vibration detection, and marks the graspable area as a stable window; h. Path planning module, which uses the grasping point as the target destination within the stable window area and combines the obstacle avoidance voxel map with the perception modality quality score to backtrack and generate the optimal grasping path; i. Fine-tuning control module, used to complete surface bonding through tactile sensor film feedback and gripper posture fine-tuning at the end stage of grasping; j. Self-learning module, used to record grasping failure cases, including stability area, surface state, gripping strategy and trajectory information, and update the local path strategy optimization table; The shoe last grabbing method comprises the following steps: S1) Collect multimodal perception data of the shoe last area, including visual images, structural acoustic waves, and infrared profile information; S2) Extract the toe point, heel point, and waist point of the shoe, generate the posture vector, and construct the mechanical model of the shoe last; S3) activating the interference perception module to obtain interference fringes on the surface of the target shoe last and determine its material and friction characteristics; S4) Identify whether the current target is in a stable state. If it is in the stable window area, enter the trajectory planning stage; otherwise, switch to the next candidate target; S5) Based on the optimal grasping point in the stable window, backtrack to generate the obstacle avoidance path and output the gripping strategy; S6) The robot executes the trajectory to approach the target and activates the tactile film at the end stage to fine-tune the posture and complete the surface bonding; S7) Execute the crawl and determine success or failure. If it fails, record the parameter information for subsequent optimization and learning.
2. The shoe last grasping method based on multimodal perception according to claim 1, characterized in that: The visual camera is used to collect a top-view image of the operating area and extract the edge contour, highlight area and possible overlapping and occluded boundaries of the shoe last through image processing.
3. The shoe last grasping method based on multimodal perception according to claim 1, characterized in that: The interference sensing device includes red and green laser cross emitters and a linear array photoelectric receiver. The receiver is used to collect surface stripe patterns and judge the shoe last material and friction characteristics based on their periodicity, density and deformation.
4. The shoe last grasping method based on multimodal perception according to claim 1, characterized in that: The stable window generation module compares the jitter degree of the edge contour of the target object in the continuous image frames with the stability of the structural acoustic wave echo in real time to determine whether the area is in a stable state.
5. The shoe last grasping method based on multimodal perception according to claim 1, characterized in that: The gripping point optimization module preferentially avoids surface areas detected to have high reflectivity or low roughness, and limits the gripping force to be progressively adjusted within a safety threshold range.
6. The shoe last grasping method based on multimodal perception according to claim 1, characterized in that: The path planning module plans the trajectory only within the range of the shoe last marked by the stability window and ignores targets in the dynamic interference area or the tilted stack edge area.
7. The shoe last grasping method based on multimodal perception according to claim 1, characterized in that: The tactile sensing film on the surface of the gripper is a deformable capacitive structure, which is used to sense the contact difference between the two sides of the gripper in real time and drive the gripper to perform micro-repair actions within 5° in the pitch or roll direction.
8. The shoe last grasping method based on multimodal perception according to claim 1, characterized in that: The fine-tuning control module is activated within 30 mm of the target shoe last end and performs micro-posture correction based on the tactile film data to ensure that the gripper fits the shoe last surface.
9. The shoe last grasping method based on multimodal perception according to claim 1, characterized in that: The self-learning module uses a lightweight local database to record the surface type, clamping effect and failure reason of each grasping action, and calls the recommended parameter group to match the priority path and gripper strategy after identifying the same shoe last model.
Citation Information
Patent Citations
Mechanical arm active grabbing device and method based on multi-model fusion
CN109129474A
Industrial robot track generation method, system and device and storage medium
CN111055286A