Tray pose recognition method and device based on robot grabbing and storage medium

CN122798901APending Publication Date: 2026-09-22BEIJING GANGTIEXIA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611293240.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-25
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0004]本申请实施例提供了一种基于机器人抓取的托盘位姿识别方法、装置及存储介质,以至少解决相关技术中存在的由于托盘的位姿识别结果不准确,导致的机器人对托盘的抓取效果不理想的技术问题

Benefits of technology

[0009] According to another aspect of the embodiments of this application, a computer program product is provided, which, when executed on a data processing device, is adapted to perform the steps of a pallet pose recognition method based on robot grasping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122798901A_ABST
    Figure CN122798901A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, and storage medium for pallet pose recognition based on robot grasping. The method includes: acquiring a target color image and a target depth image of the target pallet; determining a target depth value based on the target depth image; determining the initial center position and target orientation of the target pallet based on the target color image; determining the target center position of the target pallet based on the target depth value and the initial center position, and defining the target orientation and target center position as the pose recognition result of the target pallet, wherein the target center position is a three-dimensional coordinate in the camera coordinate system; generating a robot grasping execution command based on the pose recognition result, and controlling the robot to grasp the target pallet according to the grasping execution command. This application solves the technical problem in related technologies where inaccurate pallet pose recognition results lead to unsatisfactory pallet grasping performance by the robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer vision and industrial automation, and more specifically, to a method, device, and storage medium for pallet pose recognition based on robot grasping. Background Technology

[0002] In scenarios such as warehousing and logistics, robotic loading and unloading, tooling positioning, and automated handling, the planar position, three-dimensional position, and orientation of large pallets directly affect the success rate of pallet grasping and cycle stability. Common approaches in related technologies include relying solely on deep learning, using bounding boxes for the entire image, or using a single color threshold to find markers (such as directional markers). These methods have the following drawbacks: first, the detection range is too large, making it susceptible to background interference; second, pose estimation is prone to drift when the target edge or is partially occluded; and third, they are highly data-dependent, requiring extensive training. Therefore, related technologies suffer from the technical problem of inaccurate pallet pose recognition, leading to unsatisfactory pallet grasping performance by the robot.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This application provides a method, apparatus, and storage medium for pallet pose recognition based on robot grasping, so as to at least solve the technical problem in the related art that the robot's grasping effect on the pallet is not ideal due to inaccurate pallet pose recognition results.

[0005] According to one aspect of the embodiments of this application, a pallet pose recognition method based on robot grasping is provided, comprising: acquiring a target color image and a target depth image of a target pallet, wherein the target pallet includes a target orientation marker with an asymmetrical structure for indicating the orientation of the target pallet; determining a target depth value based on the target depth image, wherein the target depth value refers to the depth value of the target pallet in a camera coordinate system; determining an initial center position and a target orientation of the target pallet based on the target color image, wherein the initial center position is a two-dimensional coordinate in an image coordinate system; determining a target center position of the target pallet based on the target depth value and the initial center position, and determining the target orientation and the target center position as the pose recognition result of the target pallet, wherein the target center position is a three-dimensional coordinate in a camera coordinate system; generating a robot grasping execution command based on the pose recognition result, and controlling the robot to grasp the target pallet according to the grasping execution command.

[0006] According to another aspect of the embodiments of this application, a pallet pose recognition device based on robot grasping is provided, comprising: a data acquisition module, configured to acquire a target color image and a target depth image of a target pallet, wherein the target pallet includes a target direction marker with an asymmetrical structure for indicating the orientation of the target pallet; a first determination module, configured to determine a target depth value based on the target depth image, wherein the target depth value refers to the depth value of the target pallet in the camera coordinate system; a second determination module, configured to determine an initial center position and a target direction of the target pallet based on the target color image, wherein the initial center position is a two-dimensional coordinate in the image coordinate system; a third determination module, configured to determine a target center position of the target pallet based on the target depth value and the initial center position, and determine the target direction and the target center position as the pose recognition result of the target pallet, wherein the target center position is a three-dimensional coordinate in the camera coordinate system; and a grasping module, configured to generate a robot grasping execution command based on the pose recognition result, and control the robot to grasp the target pallet according to the grasping execution command.

[0007] According to another aspect of the embodiments of this application, a non-volatile storage medium is provided, which stores multiple instructions, any one of which is adapted to be loaded by a processor for a pallet pose recognition method based on robot grasping.

[0008] According to another aspect of the embodiments of this application, an electronic device is provided, including: one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement any one of the following pallet pose recognition methods based on robot grasping.

[0009] According to another aspect of the embodiments of this application, a computer program product is provided, which, when executed on a data processing device, is adapted to perform the steps of a pallet pose recognition method based on robot grasping.

[0010] In this embodiment, a target color image and a target depth image of the target tray are acquired, wherein the target tray includes a target orientation marker with an asymmetrical structure for indicating the orientation of the target tray; based on the target depth image, a target depth value is determined, wherein the target depth value refers to the depth value of the target tray in the camera coordinate system; based on the target color image, the initial center position and target orientation of the target tray are determined, wherein the initial center position is a two-dimensional coordinate in the image coordinate system; based on the target depth value and the initial center position, the target center position of the target tray is determined, and the target orientation and target center position are determined as the pose recognition result of the target tray, wherein the target center position is a three-dimensional coordinate in the camera coordinate system; based on the pose recognition result, a robot grasping execution command is generated, and the robot is controlled to grasp the target tray according to the grasping execution command. This method achieves the goal of determining the target depth value by acquiring the target depth image of the target tray, determining the target center position of the target tray by combining the initial center position determined based on the target color image of the target tray, and jointly determining the pose recognition result of the target tray based on the target center position and the target direction determined based on the target color image. Based on the pose recognition result, a robot grasping execution command is generated to control the robot to grasp the target tray. This improves the accuracy of the pose recognition result of the target tray, thereby improving the success rate and stability of the robot's grasping of the target tray. It also solves the technical problem in related technologies where the robot's grasping effect on the tray is not ideal due to inaccurate tray pose recognition results. Attached Figure Description

[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0012] Figure 1 This is a flowchart of a pallet pose recognition method based on robot grasping, according to an embodiment of this application;

[0013] Figure 2 This is a flowchart of an optional pallet pose recognition method based on robot grasping, according to an embodiment of this application;

[0014] Figure 3 This is a schematic diagram of a tray pose recognition device based on robot grasping, according to an embodiment of this application;

[0015] Figure 4 This is a structural diagram of an electronic device provided according to an embodiment of this application. Detailed Implementation

[0016] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0017] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0018] It should be noted that the information (including but not limited to target color images, target depth images, and historical color images) and data (including but not limited to preset hue threshold ranges, preset saturation thresholds, preset brightness thresholds, and preset contrast thresholds) involved in this application are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation entry points are provided for users to choose to authorize or refuse. For example, interfaces are set up between this system and relevant users or institutions, providing users with corresponding operation entry points for them to choose to agree to or refuse automated decision results; if the user chooses to refuse, the process proceeds to the expert decision-making stage.

[0019] According to an embodiment of this application, a method embodiment for pallet pose recognition based on robot grasping is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0020] Figure 1 This is a flowchart of a pallet pose recognition method based on robot grasping, according to an embodiment of this application. Figure 1 As shown, the method includes the following steps:

[0021] Step S102: Obtain a target color image and a target depth image of the target tray, wherein the target tray includes a target orientation marker with an asymmetrical structure for indicating the orientation of the target tray;

[0022] It is understandable that by combining the target color image and the target depth image for collaborative perception, a unique, robust, and unambiguous visual recognition of the target tray pose can be achieved, thereby improving the accuracy of robot grasping and operational stability.

[0023] Optionally, the aforementioned target color image and target depth image are obtained after a series of preprocessing steps based on the acquired initial color image and initial depth image of the target tray. Specifically, this includes: acquiring the initial color image and initial depth image of the target tray, and performing temporal synchronization and spatial alignment on them to obtain the target color image and target depth image. Subsequently, spatial filtering, temporal filtering, and hole filling are performed on the aligned target depth image to improve the continuity and stability of the depth data. Based on the aligned target color image, an instance segmentation model (such as YOLOv5-Seg) is used for region extraction to obtain the detection region (i.e., detection box) and segmentation mask of the target tray, as well as the corresponding detection confidence. The detection confidence is used to characterize the reliability of the determination results of the detection region (i.e., detection box) and segmentation mask. YOLOv5-Seg is a real-time instance segmentation model based on deep learning, which has the function of simultaneously outputting the detection box, detection confidence, and pixel-level segmentation mask of the target object.

[0024] Optionally, the initial depth image can be filtered first, and then aligned with the initial color image; alternatively, depending on the sensor output format, filtering can be performed after alignment. The processing order can be adjusted according to the hardware interface and calibration procedure to ultimately obtain a target depth image aligned with the target color image.

[0025] Step S104: Based on the target depth image, determine the target depth value, where the target depth value refers to the depth value of the target tray in the camera coordinate system;

[0026] It is understandable that by extracting the target depth value from the target depth image, high-precision, noise-resistant, and robust depth perception of the target pallet's three-dimensional spatial position in the camera coordinate system can be achieved, providing the robot with a stable and reliable pose baseline and improving the positioning accuracy and motion safety of grasping and handling operations.

[0027] Optionally, all valid depth values ​​(i.e., depth readings of non-hole, non-invalid pixels in the depth image) can be extracted within the segmentation mask area of ​​the target tray. Subsequently, median filtering or other robust statistics (such as truncated mean, M-estimate, etc.) are used to aggregate the above valid depth values ​​to obtain representative depth values ​​(i.e., target depth values) characterizing the surface of the target tray.

[0028] Step S106: Based on the target color image, determine the initial center position and target orientation of the target tray, wherein the initial center position is a two-dimensional coordinate in the image coordinate system;

[0029] It is understandable that, based on the target color image, the segmentation mask of the target tray is obtained by region extraction and its centroid is calculated, so as to achieve highly robust two-dimensional localization of the initial center position of the target tray under complex background and local occlusion conditions, providing a stable spatial reference benchmark for subsequent local orientation recognition and reducing the localization error caused by detection box drift or edge loss.

[0030] In one optional embodiment, determining the initial center position and target orientation of the target tray based on the target color image includes: determining the detection bounding box and segmentation mask of the target tray based on the target color image; determining the centroid of the segmentation mask as the initial center position; determining the ROI region of the target tray based on the detection bounding box and segmentation mask, wherein the ROI region is used to define the search range of the target orientation identifier; determining the target region to which the target orientation identifier belongs based on the ROI region; and determining the target orientation of the target tray based on the target region.

[0031] The approach involves using an instance segmentation model to extract regions from the target color image, obtaining the bounding box and segmentation mask for the target tray. The centroid of the segmentation mask (in the image coordinate system) is identified and used as the initial center position of the target tray. Based on the bounding box and segmentation mask, a Region of Interest (ROI) is determined to define the search range for the target orientation marker. This ROI further defines the target region to which the orientation marker belongs. Geometric analysis of the target region determines the orientation of the orientation marker, thus determining the target orientation of the tray. By using the segmentation mask centroid instead of the bounding box center, more robust localization of the initial center position of the target tray is achieved. Furthermore, adaptive generation of local ROI regions based on the bounding box and segmentation mask allows for searching the target orientation marker only in high-confidence regions, effectively suppressing background interference and improving the accuracy and stability of target tray pose recognition in complex scenes.

[0032] Optionally, to improve the stability of coarse localization (i.e., initial center position), the center of the detection box is not directly used as the initial center position. Instead, the centroid of the segmentation mask is used as the initial center position (the centroid of the segmentation mask is calculated based on the actual foreground pixel distribution, which can truly reflect the geometric center of gravity of the target tray and avoid boundary drift of the detection box caused by edge occlusion or background interference), thereby reducing the impact of edge occlusion, local missing data and background interference on position estimation.

[0033] In one optional embodiment, determining the Region of Interest (ROI) of the target tray based on the detection box and the segmentation mask includes: determining the pixel short side length and first pixel area of ​​the detection box, and the second pixel area of ​​the segmentation mask; determining an initial expansion amount based on the pixel short side length, the first pixel area, and the second pixel area; determining the theoretical center position of the ROI based on the initial center position and a preset relative position coefficient, wherein the preset relative position coefficient is used to characterize the relative positional relationship between the target orientation marker and the target tray, and the theoretical center position is a two-dimensional coordinate in the image coordinate system; determining a historical offset based on the theoretical center position and a weighted average position, wherein the weighted average position is obtained based on the center position of the target orientation marker in the image coordinate system, determined from historical color images prior to the target color image; determining a scaling factor based on the target detection confidence, the first pixel area, and the second pixel area, wherein the target detection confidence is used to characterize the reliability of the detection box and segmentation mask determination results; determining the target expansion amount of the ROI based on the scaling factor, the initial expansion amount, and the historical offset; and determining the ROI based on the theoretical center position and the target expansion amount.

[0034] The ROI region of the target tray is determined as follows: First, the pixel short side length of the detection box (i.e., the short side length of the detection box, in pixels) and the first pixel area (i.e., the product of the pixel short side length and the pixel long side length of the detection box) are determined, along with the second pixel area of ​​the segmentation mask (i.e., the total number of pixels with a value of 1 in the segmentation mask), thus determining the initial expansion amount. Second, based on the initial center position of the target tray and the preset relative position coefficient, the theoretical center position of the ROI region in the image coordinate system is determined. Next, based on the historical color images preceding the target color image, the center position of the target direction marker in the image coordinate system is determined. Then, based on this center position, a weighted average is used to determine the weighted average position of the center of the target direction marker. Combined with the theoretical center position, the historical offset is determined. The actual position of the target tray corresponding to the historical frame of the historical color image is the same as the actual position of the target tray corresponding to the target frame of the target color image. Then, the mask integrity of the segmentation mask is determined based on the area of ​​the first and second pixels. A comprehensive confidence level is then determined by combining this with the target detection confidence level. The scaling factor is determined based on the comprehensive confidence level and the pre-set confidence-scaling factor mapping relationship. Next, the target expansion amount of the ROI region is determined based on the scaling factor, the initial expansion amount, and the historical offset. Combined with the theoretical center position of the ROI region, the ROI region of the target tray is obtained. Through an adaptive mechanism that fuses the geometric dimensions of the detection box and the segmentation mask, the historical position of the marker center, and the confidence level, a robust ROI region is dynamically constructed. This improves the detection rate and positioning stability of directional markers in complex occlusion, low-confidence, and multi-scale scenes, achieving high-precision recognition with "no missed detections, no drift, and no jitter."

[0035] Optionally, the theoretical center location The following methods can be used to determine this:

[0036]

[0037] in, ( ) is the initial center position. Let be the coordinates of the initial center position along the X-axis of the image coordinate system. Let be the initial center position coordinates along the Y-axis of the image coordinate system. This is the preset relative position coefficient for the longer side. The length of the longer side of the detection box, in pixels. This is the preset relative position coefficient for the shorter side. The shorter side length is the pixel.

[0038] Optionally, historical offset The following methods can be used to determine this:

[0039]

[0040]

[0041] in, For weighted average position, Let be the weight coefficient of the i-th historical color image. The number of historical color images, Let i be the center position of the identifier corresponding to the i-th historical color image. This indicates that a modulus calculation is being performed.

[0042] Optionally, overall confidence level The following methods can be used to determine this:

[0043]

[0044] in, The target detection confidence score is the result of the instance segmentation model simultaneously generating the bounding box and segmentation mask for the target tray. It characterizes the reliability of the output results of the bounding box and segmentation mask. The weighting coefficients for the target detection confidence. The weighting factor for mask integrity. The mask integrity is determined based on the ratio between the area of ​​the second pixel and the area of ​​the first pixel of the detection box.

[0045] Optionally, target expansion amount The following methods can be used to determine this:

[0046]

[0047] in, For the expected pixel length, Scaling factor This is the initial expansion amount.

[0048] Optionally, the final ROI region has a side length of 2. Based on the theoretical center location of the ROI region The ROI region of the target tray is determined. Setting the ROI region as a square has the following advantages: the pointing structure of the target direction indicator may be located on any side, and it presents any rotation angle in the target color image. The square can cover it isotropically, avoiding missed detections due to shape asymmetry; the square ROI supports direct cropping, standard morphological operations, and rotation-free feature extraction, and the processing time is reduced by more than 40% compared to the rectangular ROI, meeting the real-time requirements of industry.

[0049] In one optional embodiment, determining the initial expansion amount based on the short side length of the pixel, the area of ​​the first pixel, and the area of ​​the second pixel includes: determining the mask integrity of the segmentation mask based on the ratio between the area of ​​the second pixel and the area of ​​the first pixel; determining the expected pixel length of the target direction marker based on the short side length of the pixel and a preset ratio, wherein the preset ratio is used to characterize the relative size relationship between the target direction marker and the target tray; and determining the initial expansion amount based on the mask integrity and the expected pixel length.

[0050] It is understandable that the mask integrity of the segmentation mask is determined based on the ratio between the area of ​​the second pixel of the segmentation mask and the area of ​​the first pixel of the detection box. The expected pixel length of the target orientation marker is determined based on the pixel short side length of the detection box and a preset ratio used to characterize the relative size relationship between the target orientation marker and the target tray. An initial expansion amount is determined based on the mask integrity and the expected pixel length. By coordinating the initial expansion amount with the mask integrity and the expected pixel length, insufficient or redundant ROI areas caused by target occlusion or detection deviations are effectively compensated for, improving the robustness of the target orientation marker's capture rate in low-confidence, partially occluded scenarios.

[0051] Optionally, the expected pixel length of the target direction identifier The following methods can be used to determine this:

[0052]

[0053] in, For the preset ratio, The shorter side length is the pixel.

[0054] Optionally, initial expansion amount The following methods can be used to determine this:

[0055]

[0056] Optionally, a local Region of Interest (ROI) can be generated based on the detection bounding box of the target tray, and the target orientation identifier can be searched only within the ROI, thereby narrowing the search range and reducing computational cost. The target expansion amount of the ROI can be adaptively determined based on the target tray size, orientation identifier size, occlusion status, and historical recognition results. The target expansion amount of the ROI can be determined by the following factors: 1. The size of the target tray detection bounding box (e.g., the short side length of pixels and the area of ​​the first pixel); 2. The typical installation position of the target orientation identifier relative to the target tray body (represented by a preset relative position coefficient), and the approximate ratio of the target orientation identifier relative to the target tray (represented by a preset ratio); 3. Whether the edge of the target tray is occluded in the current frame (represented by mask integrity); 4. The stable position of the target orientation identifier in historical frames (represented by a weighted average position); 5. Target detection confidence and mask integrity. The ROI must satisfy the requirement that the target orientation identifier completely falls within the ROI, while the area is as small as possible to reduce computational cost.

[0057] Optionally, the color image (RGB (Red-Green-Blue Color Space) format) corresponding to the ROI region can be converted to the HSV (Hue-Saturation-Value Color Space) color space. Based on a specific color (such as the two hue intervals of red) or by contrast (the tray is usually dark green, and the direction is marked as a bright color with high saturation), threshold segmentation and morphological denoising can be performed on the ROI region to obtain multiple candidate contours (i.e., multiple initial candidate regions).

[0058] In one optional embodiment, determining the target region to which the target direction identifier belongs based on the ROI region includes: converting the image of the ROI region to the HSV color space to obtain three-channel image data corresponding to the ROI region, wherein the three-channel image data includes three channels: hue channel, saturation channel, and brightness channel; segmenting the ROI region based on the three-channel image data, a preset hue threshold range, a preset saturation threshold, and a preset brightness threshold to obtain a first region; determining the local contrast corresponding to each of the multiple pixels included in the ROI region; segmenting the ROI region based on the local contrast corresponding to each of the multiple pixels and a preset contrast threshold to obtain a second region; obtaining multiple initial candidate regions based on the intersection of the first region and the second region; and determining the target region from the multiple initial candidate regions based on the contour pixel area corresponding to each of the multiple initial candidate regions, the aspect ratio of the first minimum bounding rectangle corresponding to each of the multiple initial candidate regions, the solidity corresponding to each of the multiple initial candidate regions, the roundness corresponding to each of the multiple initial candidate regions, a preset area threshold range, a preset aspect ratio threshold range, a preset solidity threshold, and a preset roundness threshold.

[0059] The target region to which the target directional identifier belongs is determined as follows: First, the color image (RGB format) portion corresponding to the ROI region in the target color image is converted to the HSV color space to obtain three-channel image data: Hue (H) channel, representing color type; Saturation (S) channel, representing color purity; and Brightness (V) channel, representing brightness intensity. Second, based on the three-channel image data, preset hue threshold ranges, preset saturation thresholds, and preset brightness thresholds, the ROI region is segmented to obtain a first region. Next, the local contrast corresponding to each pixel within the ROI region is determined, and combined with a preset contrast threshold, the ROI region is segmented to obtain a second region. Finally, the intersection of the first and second regions is taken to obtain multiple initial candidate regions. Finally, based on the contour pixel area corresponding to each of the multiple initial candidate regions, the aspect ratio of the first minimum bounding rectangle corresponding to each of the multiple initial candidate regions, the solidity of each of the multiple initial candidate regions, the roundness of each of the multiple initial candidate regions, and preset threshold ranges for area, aspect ratio, solidity, and roundness, a target region that simultaneously satisfies the above four threshold conditions is determined from the multiple initial candidate regions. Multiple initial candidate regions are generated through HSV color and local contrast dual-modal fusion segmentation, and the target region is jointly selected by combining area, aspect ratio, solidity, and roundness four-dimensional geometric constraints. This achieves high-precision and robust recognition of target orientation markings under low resolution, strong background interference, illumination changes, and partial occlusion conditions.

[0060] Optionally, simultaneously satisfying the above four threshold conditions means that the area of ​​the contour pixels is within a preset area threshold range, the aspect ratio of the first minimum circumscribed rotating rectangle is within a preset aspect ratio threshold range, the solidity is greater than a preset solidity threshold, and the roundness is greater than a preset roundness threshold.

[0061] Optionally, the solidity is used to characterize the integrity of the contour structure of the corresponding initial candidate region, reflecting whether it is a solid closed shape rather than a hollow ring, broken edge or noise fragment; the roundness is used to characterize the degree to which the shape of the corresponding initial candidate region is close to a circle, in order to distinguish directional markers with pointed structures (low roundness) from circular interference objects (high roundness).

[0062] Optionally, the "first region" and "second region" can be represented using binary masks. Each may contain multiple unconnected connected components, and their intersection may also contain multiple independent connected regions, i.e., multiple initial candidate regions. A logical AND operation is performed on the binary masks of the first and second regions to obtain the intersection mask. Then, connected component analysis is performed on this intersection mask to extract all mutually separated connected white regions as multiple initial candidate regions. Since there may be multiple interfering objects in the industrial environment that are similar in color or local contrast to the target direction marker, the first and second regions may each contain multiple non-contiguous connected components. After logical intersection, only regions that simultaneously satisfy both color and texture conditions are retained. These regions typically appear as several independent, unconnected white patches. Through connected component analysis, each independent patch can be identified as an initial candidate region, providing a candidate pool for subsequent geometric feature selection.

[0063] Optionally, if the target region determined from multiple initial candidate regions is not unique based on the contour pixel area corresponding to each of the multiple initial candidate regions, the aspect ratio of the first minimum bounding rectangle corresponding to each of the multiple initial candidate regions, the solidity corresponding to each of the multiple initial candidate regions, the roundness corresponding to each of the multiple initial candidate regions, a preset area threshold range, a preset aspect ratio threshold range, a solidity threshold, and a roundness threshold, then the initial candidate region with the largest solidity is taken as the final target region. This is because solidity can most directly reflect the structural integrity of the direction marker and still has unique discrimination ability under conditions of occlusion, noise, and local defects.

[0064] Optionally, the color image (RGB format) portion corresponding to the ROI region can be converted to the HSV color space, and thresholding and morphological denoising can be performed on the ROI region based on specific colors (such as the two hue ranges of red) or by contrast (the tray is usually dark green, and the direction indicator is a bright color with high saturation) to obtain multiple candidate contours (i.e. multiple initial candidate regions).

[0065] Optionally, to improve adaptability under different resolutions, shooting distances, and lighting conditions, adaptive thresholds can be used for candidate contour filtering. Specifically, preset area threshold ranges, preset aspect ratio threshold ranges, preset solidity thresholds, and preset roundness thresholds can be dynamically adjusted based on the ROI region's resolution, target size ratio, and contour quality indicators, thereby achieving adaptive filtering of the target region.

[0066] Optionally, when the resolution of the ROI region is lower than a preset resolution threshold, it indicates that the target edge of the ROI region is blurry and the contour points are sparse, making it easy to lose detailed features. In this case, the lower limit of the preset area threshold range is lowered to avoid being mistakenly filtered out due to insufficient pixels; the preset aspect ratio threshold range is widened (e.g., from 1.2~3.0→0.8~4.0) to allow contour deformation caused by low resolution; the preset solidity threshold is lowered to accept some edge breaks or holes; and the preset roundness threshold is lowered to tolerate contour irregularities. When the resolution of the ROI region is equal to or higher than the preset resolution threshold, it indicates that the ROI region contour is fine and noise is increased. In this case, the lower limit of the preset area threshold range is raised to filter out minor noise; the preset aspect ratio threshold range is narrowed to strengthen the preference for standard rectangular structures; the preset solidity threshold is raised to ensure that the target region is a rigid body with a complete structure; and the preset roundness threshold is raised to exclude irregular false targets (such as reflective spots and shadow edges).

[0067] Optionally, the target size ratio refers to the proportion of the area of ​​the smallest bounding rectangle of the target direction marker specified in the preset standard to the total area of ​​the ROI region, reflecting the relative size of the target direction marker in the search area. When the target size ratio is too small (e.g., <5%), the ROI region may contain background noise or small targets at a distance. In this case, the lower limit of the preset area threshold range is increased to suppress minor interference; the preset aspect ratio threshold range is narrowed to retain only targets with compact structures and close to standard rectangles; the preset solidity threshold and preset roundness threshold are increased simultaneously to eliminate loose and non-rigid interference. When the target size ratio is too large (>80%), the target may be truncated or include the background due to severe target occlusion, excessive boundary clipping, or inaccurate detection box expansion. In this case, the upper limit of the preset area threshold range is decreased to avoid misselecting the entire image; the preset aspect ratio threshold range is widened to accept asymmetric deformation caused by occlusion; the preset solidity threshold is moderately lowered to tolerate partial boundary loss; and the preset roundness threshold remains unchanged to avoid completely abandoning shape constraints.

[0068] Optionally, the contour quality index can be determined based on the contour edge gradient and the curvature smoothness of the candidate contour, and is a quantitative indicator used to comprehensively evaluate the geometric credibility of the candidate contour. When the contour quality is low, the lower limit of the preset area threshold range is reduced to ensure that realistic but incomplete candidate contours are retained; the preset aspect ratio threshold range is widened to accommodate severe deformation; the preset solidity threshold is reduced to tolerate holes and breaks; and the preset roundness threshold is reduced to tolerate candidate contours that cannot maintain high roundness due to blurring. When the contour quality is high, the preset area threshold range and the preset aspect ratio threshold range are tightened; and the preset solidity threshold and the preset roundness threshold are increased simultaneously.

[0069] Optionally, the contour quality index is a quantitative indicator used to comprehensively evaluate the geometric reliability of candidate contours. Specifically, the contour edge gradient and curvature smoothness of the candidate contour are obtained; further, the edge gradient evaluation value of the candidate contour is obtained based on the contour edge gradient, and the curvature smoothness evaluation value of the candidate contour is obtained based on the curvature smoothness; the edge gradient evaluation value and the curvature smoothness evaluation value are weighted to obtain the contour quality index. A larger contour edge gradient indicates a more significant image change between the candidate contour edge and adjacent regions, resulting in a clearer edge and a higher corresponding edge gradient evaluation value; a larger curvature smoothness indicates a more continuous curve change and fewer abnormal curvature fluctuations in the candidate contour, resulting in a higher corresponding curvature smoothness evaluation value.

[0070] Optionally, contour quality indicators The following methods can be used to determine this:

[0071]

[0072] in, This is the evaluation value of the edge gradient. This is the evaluation value for curvature smoothness. These are the weighting coefficients for the marginal gradient evaluation values. is the weighting coefficient for the curvature smoothness evaluation value, and .

[0073] Optionally, the average of the local edge gradients corresponding to multiple contour points on the candidate contour can be used to characterize the contour edge gradient of the candidate contour. The following methods can be used to determine this:

[0074]

[0075]

[0076] in, Let m be the local edge gradient of the m-th contour point on the candidate contour. Let be the pixel coordinates of the m-th contour point on the candidate contour. For the preset sampling distance, Let be the unit normal vector corresponding to the m-th contour point on the candidate contour. Let be the pixel coordinates of the sampling point in the positive direction of the normal corresponding to the m-th contour point on the candidate contour. Let be the pixel coordinates of the sampling point in the negative direction of the normal corresponding to the m-th contour point on the candidate contour. Let M be the grayscale value of the corresponding sampling point, and M be the number of contour points contained in the candidate contour. Further analysis of the contour edge gradient... After normalization, the edge gradient evaluation values ​​of the candidate contours are obtained. .

[0077] Optionally, for the curvature smoothness of the candidate contour, the discrete curvature (i.e., the ratio of the tangential change to the arc length) corresponding to the contour point is determined based on the tangential change between the contour point and its adjacent contour points, and the arc length between the contour point and its adjacent contour points. Further, the average discrete curvature is obtained by averaging the discrete curvatures corresponding to multiple contour points of the candidate contour. Next, the variance of the discrete curvature is determined based on the discrete curvatures corresponding to multiple contour points and the average discrete curvature, and this variance is used as the curvature smoothness of the candidate contour.

[0078] Mean Discrete Curvature The following methods can be used to determine this:

[0079]

[0080] in, Let be the discrete curvature of the m-th contour point on the candidate contour.

[0081] Curvature smoothness of candidate profiles characterized by variance of discrete curvature The following methods can be used to determine this:

[0082]

[0083] Curvature smoothness evaluation value The following methods can be used to determine this:

[0084]

[0085] Optionally, the contour edge gradient is obtained by averaging the gradient magnitudes of each pixel on the contour and is used to characterize the sharpness of the contour. The higher the contour edge gradient value, the sharper the contour edge and the higher the contour quality. The curvature smoothness of the contour curve is determined based on the standard deviation of curvature of each pixel on the contour and is used to characterize the geometric regularity of the contour. The higher the curvature smoothness value, the smoother and more continuous the contour edge and the higher the contour quality.

[0086] Optionally, morphological analysis is performed on the target region, dividing it into a main structure and a directional structure. The main structure is preferably rectangular, approximately square, or elongated; the directional structure is preferably a triangular apex structure. The directional structure can also be represented as a rounded tip, a conical end, or other local structures that disrupt symmetry. The orientation recognition process for the target direction marker can be as follows:

[0087] 1. Based on the outline pixel area of ​​the candidate outline, the aspect ratio of the first minimum bounding rectangle of the candidate outline, the solidity of the candidate outline, and the roundness of the candidate outline, combined with the preset area threshold range, the preset aspect ratio threshold range, the preset solidity threshold, and the preset roundness threshold, the target region is selected from multiple candidate outlines.

[0088] 2. Determine the principal axis of the main structure based on the smallest circumscribed rectangle of the main structure (i.e., the second smallest circumscribed rectangle of the main structure);

[0089] 3. Perform polygon approximation on the outline of the target area to find local sharp corners or vertices with the smallest interior angles to determine the location of the pointing structure;

[0090] 4. When the sharp point structure is not obvious, select the point in the outline of the target area that is farthest from the center point of the target area and whose symmetrical point is not on the outline of the target area as an auxiliary local sharp point.

[0091] 5. Determine the orientation of the target direction marker based on the combination relationship between the main axis of the main structure and the position of the pointing structure.

[0092] Using the above methods, stable target orientation identification results can still be obtained even under conditions of low resolution, partial occlusion, or blurred edges.

[0093] In one optional embodiment, determining the target orientation of the target tray based on the target region includes: performing morphological analysis on the target region to divide the target region into a main structure and a directional structure, wherein the main structure is used to characterize the main body of the target orientation marker, and the directional structure is a sharp-angled local structure in the target orientation marker; determining the principal axis of the main structure based on the second minimum circumscribed rectangle of the main structure; approximating the target region with polygons to determine local sharp corner points used to characterize the tip position of the directional structure; determining the orientation of the target orientation marker based on the principal axis and the local sharp corner points; and determining the target orientation based on the orientation.

[0094] It is understandable that the target area is morphologically analyzed and divided into a main structure and a directional structure. The main structure is the rectangular or approximately rectangular base portion of the target direction marker, representing the overall supporting structure of the target direction marker. Its shape is typically a regular quadrilateral with an aspect ratio greater than a preset threshold (e.g., 1.2). The directional structure is a sharp-angled local geometric structure extending from one end of the main structure. Its shape is a triangle, a cone, or an asymmetrical protrusion with an interior angle less than a first preset angle threshold (e.g., 60°). Its function is to disrupt the central symmetry of the main structure, thereby eliminating ambiguity in the 180° direction. For the outline of the main structure, its second minimum circumscribed rotation rectangle is calculated. This rectangle is the rotation rectangle with the smallest area enclosing the outline of the main structure, and the line containing its long side is parallel to the principal axis of the main structure. The aforementioned principal axis only represents the direction of the symmetry axis of the target direction marker and does not possess directionality itself, thus it cannot directly provide a unique orientation. The contour of the target region is approximated by a polygon to obtain a vertex sequence of the approximating polygon. Then, in the vertex sequence, all interior angles are traversed, and the vertex corresponding to the smallest interior angle less than a second preset angle threshold (e.g., 70°) is selected as the local sharp corner point. If the smallest interior angle is not unique, the vertex corresponding to the smallest interior angle farthest from the center point of the main structure is selected as the effective local sharp corner point to enhance noise resistance. The orientation angle can be used to represent the orientation of the target direction indicator. The direction from the center point of the main structure to the effective local sharp corner point is defined as the actual pointing arrow line of the target direction indicator. In order to achieve a unique, continuous orientation output that conforms to the clock rotation law, an absolute reference arrow line parallel to the vertical axis of the image coordinate system and pointing upward is introduced with the center point of the main structure as the reference. The geometric angle θ between the actual pointing arrow line and the absolute reference arrow line is calculated, where θ is the smaller of the two geometric angles (i.e., θ∈[0, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 1 ... 180 The orientation of a target direction marker can be represented by an angle, i.e., an orientation angle. By analyzing the left-right spatial relationship between the actual pointing arrow and the absolute reference arrow, the 180° ambiguity of the orientation angle is eliminated and mapped to a full circumference angle (i.e., 360°). If the actual pointing arrow is located to the right of the absolute reference arrow (corresponding to the 1 o'clock to 6 o'clock position on a clock), then the final orientation angle θ final =θ; If the actual pointing arrow is to the left of the absolute reference arrow (corresponding to the 6 o'clock to 12 o'clock position on a clock), then the geometric angle is compensated in reverse using the full circumference angle to ensure that the obtained orientation angle increases continuously clockwise. At this time, the final orientation angle θ final =360 -θ. Specifically, when the actual pointing arrow coincides with the absolute reference arrow and points upwards, the final pointing angle θ... final =0 When both coincide and point downwards, the final orientation angle θ final =180. Therefore, the final orientation angle θ final The overall calculation logic is completely consistent with the degree mapping of clockwise rotation, at 0 360 It possesses both uniqueness and continuity within its range. Based on the aforementioned orientation angle θ... final The final target orientation of the target pallet is determined. Through the above geometric structure analysis, left and right quadrant determination based on the absolute reference arrow line, and full-circumference angle reverse compensation, the 180° ambiguity of the target orientation of the target pallet is completely eliminated, and a unique, continuous target orientation that can be directly used for robot control is output, which significantly improves the positioning accuracy and anti-interference ability of the target pallet.

[0095] Optionally, based on the orientation angle of the identified target orientation marker, the target orientation of the target pallet, represented by an angle, is obtained, i.e., the target orientation angle. Given the periodicity of angle data, to avoid introducing spurious fluctuations at the 0° / 360° transition boundary, the following three-stage periodic smoothing strategy can be used sequentially to smooth the target direction angle. Perform smoothing:

[0096] Circular mean smoothing, for target orientation angles across multiple consecutive frames. The sequence is weighted by a cyclic mean algorithm based on trigonometric functions to achieve unambiguous smoothing of periodic data;

[0097] Single-frame jump limit, if the current frame's If the absolute value of the difference between the current frame and the previous frame exceeds a preset threshold (e.g., ±45°), then the current frame will be... Forced to limit changes within the legal range of variation, suppressing sudden jumps caused by misidentification or noise;

[0098] Low-pass filtering is applied to the target orientation angle sequence after the above processing. A first-order IIR (Infinite Impulse Response) or FIR (Finite Impulse Response) low-pass filter is used to suppress high-frequency jitter, resulting in a smooth, continuous, and abruptly-free stable target orientation angle output. , as the input signal of the robot control system.

[0099] In one optional embodiment, a polygon approximation is applied to the target region to determine local sharp corner points used to characterize the tip position of the pointing structure. This includes: applying a polygon approximation to the target region to determine multiple interior angles; if the smallest interior angle among the multiple interior angles is less than a preset angle threshold, the vertex corresponding to the smallest interior angle is determined as a local sharp corner point; or, if the smallest interior angle is greater than or equal to the preset angle threshold, points on the contour of the target region that meet preset conditions are determined as local sharp corner points, wherein the preset conditions are that the point is farthest from the center of the target region and the symmetrical point is not on the contour.

[0100] The process involves performing polygon approximation on the contour of the target region to obtain a sequence of approximate vertices. For each vertex, the interior angles formed by its two adjacent sides are calculated, and all interior angle values ​​are recorded to determine multiple interior angles. If the smallest interior angle among these angles is less than a preset angle threshold (e.g., 70°), the vertex corresponding to the smallest interior angle is considered the natural tip of the pointing structure and is directly used as a local sharp corner. If the smallest interior angle is greater than or equal to the preset angle threshold, the target orientation marker is determined to have edge wear, reflective blur, low-resolution blur, or partial occlusion. In this case, the center point of the target region is calculated, and all pixels on the contour of the target region are traversed to select the set of candidate points farthest from the center point. For each candidate point, it is determined whether its symmetrical point about the center point is located on the contour of the target region. The candidate point that is farthest away and whose symmetrical point is not on the contour of the target region is selected as the local sharp corner. By selecting the most asymmetric feature points on the contour of the target region, the pointing structure can still be stably located even when the target orientation marker has wear, blur, or occlusion, improving the accuracy and rationality of the target orientation determination results, thereby enhancing the stability and continuity of robot operations.

[0101] Step S108: Based on the target depth value and the initial center position, determine the target center position of the target tray, and determine the target direction and the target center position as the pose recognition result of the target tray, wherein the target center position is the three-dimensional coordinate in the camera coordinate system;

[0102] It is understandable that, based on the target depth value and the initial center position, the precise three-dimensional target center position of the target tray in the camera coordinate system is calculated. Combined with the target orientation, drift-free, highly stable, and directly usable robot control pose output is achieved, improving the positioning accuracy and grasping control continuity of the target tray under complex working conditions.

[0103] In one optional embodiment, determining the target center position of the target tray based on the target depth value and the initial center position includes: determining the projection center position of the target tray based on the target depth value, the initial center position and camera intrinsic parameters, wherein the projection center position is a three-dimensional coordinate in the camera coordinate system; and performing a correction process on the projection center position to obtain the target center position, wherein the correction process includes jump limitation and filtering.

[0104] It is understandable that, based on the target depth value and initial center position, back projection is performed using camera intrinsic parameters to obtain the projected center position of the target tray in the camera coordinate system. The projected center position is then corrected. First, error analysis is performed between the projected center position and the historical center position determined in the previous frame corresponding to the target color image and target depth image, yielding the center position error. This error is then compared to a preset single-frame jump threshold. If the error does not exceed the threshold, the projected center position is filtered, and the target center position is output. If the error exceeds the threshold, the historical center position of the previous frame is used as the target center position, and the filter state is frozen without updating the filter's internal parameters. By triggering the filter state freeze mechanism through jump detection, the stable filtered result of the previous frame is directly output under abnormal interference, blocking new input updates. This achieves zero-abrupt, oscillation-free, and continuous stability in pose output, improving the safety and operational continuity of the robot control system.

[0105] Optionally, all valid depth values ​​(i.e., depth readings of non-hole, non-invalid pixels in the depth image) can be extracted within the segmented mask area of ​​the target tray. Subsequently, median filtering or other robust statistics (such as truncated mean, M-estimator, etc.) are used to aggregate these valid depth values ​​to obtain a representative depth value (i.e., the target depth value) characterizing the surface of the target tray. Based on the representative depth value and the initial center position, and combined with the camera intrinsic parameters (of the depth camera), back-projection is performed to obtain the coordinates (X, Y, Z) of the projected center position of the target tray.

[0106] Optionally, to address the issue of significant differences in noise characteristics of depth cameras along different spatial axes (i.e., noise in the planar direction (X, Y axes) is mainly high-frequency jitter, while noise in the depth direction (Z axis) is mainly low-frequency drift and outliers), an anisotropic temporal filtering strategy can be adopted. Independent filtering parameters, including jump threshold, filtering window length, and response coefficient, are set for the X, Y, and Z coordinate components of the projection center position coordinates to achieve targeted suppression of different noise characteristics. Specifically, for the X and Y axes (planar directions), a smaller filtering window (e.g., 3-5 frames) and a more lenient jump limit (e.g., a preset single-frame jump threshold of ±50mm) are used to ensure a fast response to the rapid movement of the target tray and avoid control lag due to excessive smoothing. For the Z axis (depth direction), a larger filtering window (e.g., 8-15 frames) and a strict jump limit (e.g., a preset single-frame jump threshold of ±10 mm) are used, combined with a robust statistical filtering mechanism (e.g., median filtering or truncated mean) to effectively suppress abnormal fluctuations in the depth sensor caused by changes in reflectivity, occlusion, or measurement drift, thereby improving the stability of depth estimation.

[0107] Optionally, when the target tray is briefly occluded, the target tray detection fails, the pose recognition confidence is lower than a preset threshold, or the pose output result shows an abnormal jump, the invalid or abnormal value may not be output immediately. Instead, the pose output holding mechanism is automatically triggered: that is, the valid pose result of the previous frame is held as the output of the current frame to avoid sudden control of the downstream robotic arm or robot chassis.

[0108] Optionally, the output hold mechanism can be further combined with confidence criteria and duration criteria: when no valid pose recognition result is output for several consecutive frames, it enters the re-search state; when a stable target is re-identified, normal pose update is resumed.

[0109] Step S110: Based on the pose recognition result, generate the robot's grasping execution command and control the robot to grasp the target tray according to the grasping execution command.

[0110] It is understandable that generating precise grasping execution instructions based on highly stable, interference-resistant, and mutation-free pose recognition results enables robots to achieve zero-impact, high-repeatability pallet grasping under complex working conditions, thereby improving the success rate and stability of automated robot grasping operations.

[0111] Through the above steps S102 to S110, the target depth value can be determined by acquiring the target depth image of the target tray, and the target center position of the target tray can be determined by combining the initial center position determined based on the target color image of the target tray. Based on the target center position and the target direction determined based on the target color image, the pose recognition result of the target tray is determined together, and the robot's grasping execution command is generated based on the pose recognition result to control the robot to grasp the target tray. This achieves the technical effect of improving the accuracy of the pose recognition result of the target tray, thereby improving the success rate and stability of the robot when grasping the target tray. This solves the technical problem in related technologies where the robot's grasping effect on the tray is not ideal due to the inaccurate pose recognition result of the tray.

[0112] Based on the above embodiments and optional embodiments, this application proposes an optional implementation method for pallet pose recognition based on robot grasping. This optional implementation method can be understood as a pallet pose recognition and tracking method based on local orientation mark geometric analysis and split-axis temporal filtering.

[0113] Figure 2 This is a flowchart of an optional pallet pose recognition method based on robot grasping, according to an embodiment of this application. Figure 2 As shown, the steps of the pallet pose recognition and tracking method based on local orientation identifier geometric analysis and axis-sequential filtering include:

[0114] Step S1: Data acquisition and preprocessing.

[0115] The initial color image and initial depth image of the target tray are acquired and then synchronized temporally and spatially to obtain the target color image and target depth image. Subsequently, spatial filtering, temporal filtering, and hole filling are applied to the aligned target depth image to improve the continuity and stability of the depth data. Based on the aligned target color image, an instance segmentation model (such as YOLOv5-Seg) is used to extract the target region, obtaining the detection region (i.e., detection box) and segmentation mask of the target tray, as well as the corresponding detection confidence score. The detection confidence score characterizes the reliability of the determination results of the detection region (i.e., detection box) and segmentation mask.

[0116] The initial depth image can be filtered first, and then aligned with the initial color image; alternatively, depending on the sensor output format, filtering can be performed after alignment. This processing order can be adjusted according to the hardware interface and calibration procedure to ultimately obtain a target depth image aligned with the target color image.

[0117] To improve the stability of coarse localization (i.e., initial center position), instead of directly using the center of the detection box as the initial center position, the centroid of the segmentation mask is used as the initial center position (the centroid of the segmentation mask is calculated based on the actual foreground pixel distribution, which can truly reflect the geometric center of gravity of the target tray and avoid boundary drift caused by edge occlusion or background interference of the detection box), thereby reducing the impact of edge occlusion, local missing data and background interference on position estimation.

[0118] Step S2: Determine the ROI region.

[0119] Based on the detection bounding box of the target tray, a local Region of Interest (ROI) is generated. The target orientation marker is searched only within the ROI, thereby narrowing the search range and reducing computational cost. The target expansion amount of the ROI can be adaptively determined based on the target tray size, orientation marker size, occlusion status, and historical recognition results.

[0120] The target expansion of the ROI region can be determined by the following factors:

[0121] 1. Dimensions of the target tray detection frame (e.g., short side length of pixels and area of ​​the first pixel); 2. Typical installation position of the target orientation marker relative to the target tray body (represented by a preset relative position coefficient), and the approximate ratio of the target orientation marker relative to the target tray (represented by a preset ratio); 3. Whether the edge of the target tray is occluded in the current frame (represented by mask integrity); 4. Stable position of the target orientation marker in historical frames (represented by a weighted average position); 5. Target detection confidence and mask integrity.

[0122] The ROI region must ensure that the target direction marker falls completely within the ROI region, while keeping the area as small as possible to reduce computational load.

[0123] The color image (RGB format) portion corresponding to the ROI region is converted to the HSV color space. Based on specific colors (such as the two hue ranges of red) or by contrast (the tray is usually dark green, and the direction indicator is a bright color with high saturation), the ROI region is thresholded and morphologically denoised to obtain multiple candidate contours (i.e. multiple initial candidate regions).

[0124] Step S3: Orientation recognition of the direction marker.

[0125] To improve adaptability under different resolutions, shooting distances, and lighting conditions, the candidate contour selection criteria employ adaptive thresholds. Specifically, the preset area threshold range, preset aspect ratio threshold range, preset solidity threshold, and preset roundness threshold can be dynamically adjusted based on the ROI region's resolution, target size ratio, and contour quality indicators, thereby achieving adaptive selection of the target region.

[0126] Morphological analysis is performed on the target region, dividing it into a main structure and directional structures. The main structure is preferably rectangular, approximately square, or elongated; the directional structures are preferably triangular apex structures. Directional structures can also be represented as rounded tips, conical ends, or other local structures that disrupt symmetry.

[0127] The orientation recognition process for the target orientation marker is as follows:

[0128] 1. Based on the outline pixel area of ​​the candidate outline, the aspect ratio of the first minimum bounding rectangle of the candidate outline, the solidity of the candidate outline, and the roundness of the candidate outline, combined with the preset area threshold range, the preset aspect ratio threshold range, the preset solidity threshold, and the preset roundness threshold, the target region is selected from multiple candidate outlines.

[0129] 2. Determine the principal axis of the main structure based on the smallest circumscribed rectangle of the main structure (i.e., the second smallest circumscribed rectangle of the main structure);

[0130] 3. Perform polygon approximation on the outline of the target area to find local sharp corners or vertices with the smallest interior angles to determine the location of the pointing structure;

[0131] 4. When the sharp point structure is not obvious, select the point in the outline of the target area that is farthest from the center point of the target area and whose symmetrical point is not on the outline of the target area as an auxiliary local sharp point.

[0132] 5. Determine the orientation of the target direction marker based on the combination relationship between the main axis of the main structure and the position of the pointing structure.

[0133] Using the above methods, stable target orientation identification results can still be obtained even under conditions of low resolution, partial occlusion, or blurred edges.

[0134] Step S4: Pose calculation and timing filtering.

[0135] After determining the orientation of the target direction marker and the initial center position of the target pallet, the target direction and three-dimensional spatial position (i.e., the target center position) of the target pallet are output.

[0136] Based on the orientation of the identified target direction marker, the target direction of the target pallet, represented by an angle, is obtained, i.e., the target direction angle. Given the periodic nature of angle data, to avoid introducing spurious fluctuations at the 0° / 360° transition boundary, the following three-stage periodic smoothing strategy is adopted sequentially for the target direction angle. Perform smoothing:

[0137] Circular mean smoothing, for target orientation angles across multiple consecutive frames. The sequence is weighted by a cyclic mean algorithm based on trigonometric functions to achieve unambiguous smoothing of periodic data;

[0138] Single-frame jump limit, if the current frame's If the absolute value of the difference between the current frame and the previous frame exceeds a preset threshold (e.g., ±45°), then the current frame will be... Forced to limit changes within the legal range of variation, suppressing sudden jumps caused by misidentification or noise;

[0139] Low-pass filtering is applied to the target orientation angle sequence after the above processing. A first-order IIR or FIR low-pass filter is used to suppress high-frequency jitter, resulting in a smooth, continuous, and stable target orientation angle output without abrupt changes. , as the input signal of the robot control system.

[0140] Within the segmented mask area of ​​the target tray, all valid depth values ​​(i.e., depth readings of non-hole, non-invalid pixels in the depth image) are extracted. Subsequently, median filtering or other robust statistics (such as truncated mean, M-estimator, etc.) are used to aggregate these valid depth values ​​to obtain a representative depth value (i.e., the target depth value) characterizing the surface of the target tray. Based on the representative depth value and the initial center position, and combined with the camera intrinsic parameters (from the depth camera), backprojection is performed to obtain the coordinates (X, Y, Z) of the projected center position of the target tray.

[0141] To address the significant differences in noise characteristics of depth cameras across different spatial axes (i.e., noise in the planar direction (X, Y axes) is mainly characterized by high-frequency jitter, while noise in the depth direction (Z axis) is mainly characterized by low-frequency drift and outliers), an anisotropic temporal filtering strategy is adopted. Independent filtering parameters are set for the X, Y, and Z coordinate components of the projection center position, including jump threshold, filter window length, and response coefficient, to achieve targeted suppression of different noise characteristics. Specifically, for the X and Y axes (planar directions), a smaller filtering window (e.g., 3-5 frames) and a more lenient jump limit (e.g., a preset single-frame jump threshold of ±50mm) are used to ensure a fast response to the rapid movement of the target tray and avoid control lag due to excessive smoothing. For the Z axis (depth direction), a larger filtering window (e.g., 8-15 frames) and a strict jump limit (e.g., a preset single-frame jump threshold of ±10mm) are used, combined with a robust statistical filtering mechanism (e.g., median filtering or truncated mean) to effectively suppress abnormal fluctuations in the depth sensor caused by changes in reflectivity, occlusion, or measurement drift, thereby improving the stability of depth estimation.

[0142] Step S5: Abnormal state maintenance and output control.

[0143] When the target tray is briefly occluded, the target tray detection fails, the pose recognition confidence is lower than the preset threshold, or the pose output result shows an abnormal jump, instead of immediately outputting invalid or abnormal values, the pose output holding mechanism is automatically triggered: that is, the valid pose result of the previous frame is held as the output of the current frame to avoid sudden control of the downstream robotic arm or robot chassis.

[0144] The output hold mechanism can be further combined with confidence criteria and duration criteria: when no valid pose recognition result is output for several consecutive frames, it enters the re-search state; when a stable target is re-identified, normal pose update is resumed.

[0145] The above optional implementation methods achieve at least the following effects: A local ROI region is generated based on the detection box of the target tray; target direction markers are searched only in the vicinity of the target tray, reducing false detections caused by similar colors or shapes in the background and reducing computational load; the target direction marker is identified as a combination of the main structure and the pointing structure, and the unique orientation of the target direction marker is determined by the position of the pointing structure relative to the main structure, avoiding 180° ambiguity caused by relying solely on the main axis direction; periodic smoothing is applied to the direction recognition results, and X, Y, and Z-axis filtering is applied to the coordinates of the projection center position; when short-term occlusion, recognition failure, or abnormal jumps occur, the valid pose output of the previous frame is maintained to avoid sudden changes in downstream robot control.

[0146] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0147] This embodiment also provides a tray pose recognition device based on robot grasping. This device is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the terms "module" and "device" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0148] According to an embodiment of this application, an apparatus embodiment for implementing a pallet pose recognition method based on robot grasping is also provided. Figure 3 This is a schematic diagram of a tray pose recognition device based on robot grasping, according to an embodiment of this application. Figure 3 As shown, the above-mentioned pallet pose recognition device based on robot grasping includes a data acquisition module 302, a first determination module 304, a second determination module 306, a third determination module 308, and a grasping module 310. The device will be described below.

[0149] The data acquisition module 302 is used to acquire a target color image and a target depth image of the target tray, wherein the target tray includes a target orientation marker with an asymmetrical structure for indicating the orientation of the target tray;

[0150] The first determining module 304 is connected to the data acquisition module 302 and is used to determine the target depth value based on the target depth image, wherein the target depth value refers to the depth value of the target tray in the camera coordinate system;

[0151] The second determining module 306 is connected to the first determining module 304 and is used to determine the initial center position and target orientation of the target tray based on the target color image, wherein the initial center position is a two-dimensional coordinate in the image coordinate system;

[0152] The third determining module 308, connected to the second determining module 306, is used to determine the target center position of the target tray based on the target depth value and the initial center position, and to determine the target direction and the target center position as the pose recognition result of the target tray, wherein the target center position is a three-dimensional coordinate in the camera coordinate system;

[0153] The grasping module 310, connected to the third determining module 308, is used to generate grasping execution instructions for the robot based on the pose recognition results, and control the robot to grasp the target tray according to the grasping execution instructions.

[0154] This application provides a pallet pose recognition device based on robot grasping. By setting up a data acquisition module 302, a first determination module 304, a second determination module 306, a third determination module 308, and a grasping module 310, the device achieves the following: determining the target depth value by acquiring the target depth image of the target pallet; determining the target center position of the target pallet by combining the initial center position determined based on the target color image of the target pallet; determining the pose recognition result of the target pallet based on the target center position and the target direction determined based on the target color image; and generating a robot grasping execution command based on the pose recognition result to control the robot to grasp the target pallet. This improves the accuracy of the target pallet pose recognition result, thereby improving the success rate and stability of the robot's grasping of the target pallet. This solves the technical problem in related technologies where the robot's grasping effect on the pallet is not ideal due to inaccurate pallet pose recognition results.

[0155] It should be noted that the above modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following ways: the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0156] It should be noted that the data acquisition module 302, the first determining module 304, the second determining module 306, the third determining module 308, and the capture module 310 mentioned above correspond to steps S102 to S110 in the embodiments. The instances and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as part of the device, can run in a computer terminal.

[0157] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant descriptions in the embodiments, and will not be repeated here.

[0158] The aforementioned pallet pose recognition device based on robot grasping may also include a processor and a memory. The data acquisition module 302, the first determination module 304, the second determination module 306, the third determination module 308, the grasping module 310, etc. are all stored in the memory as program units, and the processor executes the aforementioned program units stored in the memory to realize the corresponding functions.

[0159] The processor contains a core that retrieves the corresponding program unit from memory. One or more cores may be configured. Memory may include non-persistent memory in computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory includes at least one memory chip.

[0160] This application provides a non-volatile storage medium storing a program that, when executed by a processor, implements a pallet pose recognition method based on robot grasping.

[0161] This application provides an electronic device. Figure 4 This is a structural diagram of an electronic device provided according to an embodiment of this application. For example... Figure 4 As shown, the electronic device may include: one or more ( Figure 4(Only one is shown) a processor 402, a memory 404, a memory controller, and a peripheral interface, wherein the peripheral interface is connected to an RF module, an audio module, and a display. The electronic device includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it performs the following steps: acquiring a target color image and a target depth image of a target tray, wherein the target tray includes a target orientation marker with an asymmetrical structure indicating the orientation of the target tray; determining a target depth value based on the target depth image, wherein the target depth value refers to the depth value of the target tray in the camera coordinate system; determining the initial center position and target orientation of the target tray based on the target color image, wherein the initial center position is a two-dimensional coordinate in the image coordinate system; determining the target center position of the target tray based on the target depth value and the initial center position, and determining the target orientation and target center position as the pose recognition result of the target tray, wherein the target center position is a three-dimensional coordinate in the camera coordinate system; generating a robot grasping execution command based on the pose recognition result, and controlling the robot to grasp the target tray according to the grasping execution command. The device in this document can be a server, PC, etc.

[0162] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing an initialization program having the following method steps: acquiring a target color image and a target depth image of a target tray, wherein the target tray includes a target orientation marker with an asymmetrical structure for indicating the orientation of the target tray; determining a target depth value based on the target depth image, wherein the target depth value refers to the depth value of the target tray in the camera coordinate system; determining an initial center position and a target orientation of the target tray based on the target color image, wherein the initial center position is a two-dimensional coordinate in the image coordinate system; determining a target center position of the target tray based on the target depth value and the initial center position, and determining the target orientation and the target center position as the pose recognition result of the target tray, wherein the target center position is a three-dimensional coordinate in the camera coordinate system; generating a grasping execution command for the robot based on the pose recognition result, and controlling the robot to grasp the target tray according to the grasping execution command.

[0163] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0164] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0165] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0166] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0167] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0168] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0169] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0170] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0171] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0172] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for pallet pose recognition based on robot grasping, characterized in that, include: Acquire a target color image and a target depth image of the target tray, wherein the target tray includes a target orientation marker with an asymmetrical structure for indicating the orientation of the target tray; Based on the target depth image, a target depth value is determined, wherein the target depth value refers to the depth value of the target tray in the camera coordinate system; Based on the target color image, the initial center position and target orientation of the target tray are determined, wherein the initial center position is a two-dimensional coordinate in the image coordinate system; Based on the target depth value and the initial center position, the target center position of the target tray is determined, and the target direction and the target center position are determined as the pose recognition result of the target tray, wherein the target center position is a three-dimensional coordinate in the camera coordinate system; Based on the pose recognition results, a grasping execution command is generated for the robot, and the robot is controlled to grasp the target tray according to the grasping execution command.

2. The method according to claim 1, characterized in that, Determining the initial center position and target orientation of the target tray based on the target color image includes: Based on the target color image, determine the detection box and segmentation mask for the target tray; The centroid of the segmentation mask is determined as the initial center position; Based on the detection frame and the segmentation mask, the ROI region of the target tray is determined, wherein the ROI region is used to limit the search range of the target direction identifier; Based on the ROI region, determine the target region to which the target direction identifier belongs; Based on the target area, the target orientation of the target tray is determined.

3. The method according to claim 2, characterized in that, The step of determining the ROI region of the target tray based on the detection frame and the segmentation mask includes: Determine the pixel short side length and first pixel area of ​​the detection box, and the second pixel area of ​​the segmentation mask; The initial expansion amount is determined based on the short side length of the pixel, the area of ​​the first pixel, and the area of ​​the second pixel; Based on the initial center position and the preset relative position coefficient, the theoretical center position of the ROI region is determined, wherein the preset relative position coefficient is used to characterize the relative positional relationship between the target direction marker and the target tray, and the theoretical center position is a two-dimensional coordinate in the image coordinate system; Based on the theoretical center position and the weighted average position, the historical offset is determined, wherein the weighted average position is obtained by determining the center position of the target orientation marker in the image coordinate system based on the historical color images preceding the target color image; A scaling factor is determined based on the target detection confidence score, the area of ​​the first pixel, and the area of ​​the second pixel, wherein the target detection confidence score is used to characterize the reliability of the detection box and the segmentation mask determination results; Based on the scaling factor, the initial expansion amount, and the historical offset, the target expansion amount of the ROI region is determined; The ROI region is determined based on the theoretical center location and the target expansion amount.

4. The method according to claim 3, characterized in that, The determination of the initial expansion amount based on the short side length of the pixel, the area of ​​the first pixel, and the area of ​​the second pixel includes: The mask integrity of the segmentation mask is determined based on the ratio between the area of ​​the second pixel and the area of ​​the first pixel. Based on the pixel short side length and a preset ratio, the expected pixel length of the target direction marker is determined, wherein the preset ratio is used to characterize the relative size relationship between the target direction marker and the target tray; The initial expansion amount is determined based on the mask integrity and the expected pixel length.

5. The method according to claim 2, characterized in that, The step of determining the target region to which the target direction identifier belongs based on the ROI region includes: The image of the ROI region is converted to the HSV color space to obtain the three-channel image data corresponding to the ROI region. The three-channel image data includes three channels: hue channel, saturation channel, and brightness channel. Based on the three-channel image data, the preset hue threshold range, the preset saturation threshold, and the preset brightness threshold, the ROI region is segmented to obtain the first region; Determine the local contrast of each of the multiple pixels included in the ROI region; Based on the local contrast corresponding to the multiple pixels and a preset contrast threshold, the ROI region is segmented to obtain a second region; Based on the intersection of the first region and the second region, multiple initial candidate regions are obtained; The target region is determined from the plurality of initial candidate regions based on the contour pixel area corresponding to each of the plurality of initial candidate regions, the aspect ratio of the first minimum circumscribed rotating rectangle corresponding to each of the plurality of initial candidate regions, the solidity corresponding to each of the plurality of initial candidate regions, the roundness corresponding to each of the plurality of initial candidate regions, a preset area threshold range, a preset aspect ratio threshold range, a preset solidity threshold, and a preset roundness threshold.

6. The method according to claim 2, characterized in that, Determining the target orientation of the target tray based on the target area includes: Morphological analysis is performed on the target region to divide it into a main structure and a directional structure. The main structure is used to characterize the main body of the target directional marker, and the directional structure is a sharp-angled local structure in the target directional marker. The principal axis of the main structure is determined based on the second minimum circumscribed rectangle of the main structure. The target region is approximated using polygons to determine local sharp corner points used to characterize the tip position of the pointing structure; Based on the main axis and the local sharp corner points, determine the orientation of the target direction marker; Based on the orientation, the target direction is determined.

7. The method according to claim 6, characterized in that, The step of approximating the target region with polygons to determine local sharp corner points used to characterize the tip position of the pointing structure includes: The target region is approximated using polygons to determine multiple interior angles; If the smallest interior angle among the plurality of interior angles is less than a preset angle threshold, the vertex corresponding to the smallest interior angle is determined as the local apex point; or, When the minimum interior angle is greater than or equal to a preset angle threshold, the points on the contour of the target area that meet the preset conditions are determined as the local sharp corner points, wherein the preset conditions are that the points are farthest from the center of the target area and that the symmetrical points are not on the contour.

8. The method according to any one of claims 1 to 7, characterized in that, Determining the target center position of the target pallet based on the target depth value and the initial center position includes: Based on the target depth value, the initial center position, and the camera intrinsic parameters, the projection center position of the target tray is determined, wherein the projection center position is a three-dimensional coordinate in the camera coordinate system; The projection center position is corrected to obtain the target center position, wherein the correction process includes jump restriction and filtering.

9. A tray pose recognition device based on robot grasping, characterized in that, include: The data acquisition module is used to acquire a target color image and a target depth image of the target tray, wherein the target tray includes a target orientation marker with an asymmetrical structure for indicating the orientation of the target tray; The first determining module is used to determine a target depth value based on the target depth image, wherein the target depth value refers to the depth value of the target tray in the camera coordinate system; The second determining module is used to determine the initial center position and target orientation of the target tray based on the target color image, wherein the initial center position is a two-dimensional coordinate in the image coordinate system; The third determining module is used to determine the target center position of the target tray based on the target depth value and the initial center position, and to determine the target direction and the target center position as the pose recognition result of the target tray, wherein the target center position is a three-dimensional coordinate in the camera coordinate system; The grasping module is used to generate grasping execution instructions for the robot based on the pose recognition results, and control the robot to grasp the target tray according to the grasping execution instructions.

10. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores multiple instructions, which are adapted to be loaded by a processor and executed by the pallet pose recognition method based on robot grasping as described in any one of claims 1 to 8.