Pallet position detection method, apparatus, equipment, and storage medium

By performing region extraction and convolution processing on point cloud images, combined with convolution kernels and plane fitting, the problems of accuracy and speed in pallet pose recognition were solved, and efficient pallet pose detection for unmanned forklifts was achieved.

CN116188577BActive Publication Date: 2026-04-07VISIONNAV ROBOTICS SHENZHEN LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-08
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

During the warehouse inbound and outbound process, the uncertainty of the pallet's position and orientation makes it impossible for unmanned forklifts to quickly and accurately identify the pallet's pose. Existing 3D point cloud feature detection methods are affected by the quality of the point cloud, making it difficult to accurately and stably extract physical features and requiring a long computation time.

Method used

By extracting regions from point cloud images, performing convolution processing and plane fitting using pre-determined convolution kernels, coarse pose extraction is first performed, followed by precise position extraction. The pose of the tray is determined by combining the target feature map and the plane equation.

Benefits of technology

It improves the accuracy and speed of pallet pose detection, reduces the amount of point cloud computing, reduces the impact of background interference, and achieves fast and accurate pallet pose recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116188577B_ABST
    Figure CN116188577B_ABST
Patent Text Reader

Abstract

This invention discloses a pallet pose detection method, apparatus, device, and storage medium. The method includes: extracting regions from a point cloud image of an acquired target region to obtain an initial pallet image; performing convolution processing on the initial pallet image according to a first target convolution kernel to obtain a first target feature map, wherein the first target convolution kernel is determined based on a pre-determined surface area of ​​each pier hole, the total surface area of ​​the pallet, and a first resolution; performing planar fitting processing on the first target feature map to obtain a corresponding pallet planar image and a target plane equation; performing convolution processing on the pallet planar image according to a second target convolution kernel to obtain a second target feature map, wherein the second target convolution kernel is determined based on a pre-determined surface area of ​​each pier hole, the total surface area of ​​the pallet, and a second resolution, wherein the second resolution is greater than the first resolution; and determining the target pose of the pallet based on the second target feature map and the target plane equation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to information processing technology, and relate to but are not limited to a tray pose detection method and device, equipment, and a storage medium. BACKGROUND

[0002] In the warehouse in-out process, due to the influence of operation process, equipment precision, manual operation and other factors, the position and attitude of the tray storage are uncertain, which will cause the unmanned forklift to be unable to quickly and accurately identify the pose of the tray, thereby failing to meet the efficient operation demand of the warehousing logistics industry.

[0003] And in the prior art, the features of three-dimensional point cloud are often used to describe and detect the tray pose, such as using region growing segmentation, RANSAC model segmentation, corner detection and clustering algorithms, but in the implementation process of these algorithms, the extraction of geometric features and feature points is usually affected by the quality of point cloud, and it is difficult to accurately and stably extract the physical features of the tray. At the same time, the amount of point cloud data is large, and it usually takes a long time to process, which is difficult to realize fast detection on limited hardware resources. SUMMARY

[0004] Therefore, the tray pose detection method and device, equipment and storage medium provided by the embodiments of the present application can improve the detection accuracy and speed of the tray pose. The tray pose detection method and device, equipment and storage medium provided by the embodiments of the present application are implemented as follows:

[0005] The tray pose detection method provided by the embodiments of the present application comprises:

[0006] The point cloud image of the target region obtained is subjected to region extraction to obtain an initial tray image, the target region is a region containing a tray, and the initial tray image is an image after filtering out the background region in the point cloud image;

[0007] The initial tray image is subjected to convolution processing according to a first target convolution kernel to obtain a first target feature map, the first target convolution kernel is determined according to a predetermined surface area of each pail hole, a total surface area of the tray and a first resolution;

[0008] The first target feature map is subjected to plane fitting processing to obtain a corresponding tray plane image and a target plane equation;

[0009] The tray plane image is subjected to convolution processing according to a second target convolution kernel to obtain a second target feature map, the second target convolution kernel is determined according to a predetermined surface area of each pail hole, a total surface area of the tray and a second resolution, and the second resolution is greater than the first resolution;

[0010] According to the second target feature map and the target plane equation, a target pose of the tray is determined.

[0011] In some embodiments, the initial tray image is subjected to convolution processing according to a predetermined first target convolution kernel to obtain a first target feature map, including:

[0012] In the first resolution, the initial tray image is subjected to mapping processing to obtain a first projection image of the initial tray image on a projection plane; and the first projection image is subjected to convolution processing according to a predetermined first target convolution kernel to obtain a first target feature map, the first target convolution kernel being determined according to a predetermined surface area of each pail hole, a total surface area of the tray and the first resolution.

[0013] In some embodiments, the first target feature map is subjected to plane fitting processing to obtain a corresponding tray plane image and a target plane equation, including:

[0014] According to the predetermined total surface area of the tray, a candidate tray image is extracted from the first target feature map; and the candidate tray image is subjected to plane fitting processing to obtain a corresponding tray plane image and a target plane equation.

[0015] In some embodiments, the tray plane image is subjected to convolution processing according to a predetermined second target convolution kernel to obtain a second target feature map, including:

[0016] In the second resolution, the tray plane image is subjected to mapping processing to obtain a second projection image of the tray plane image on the projection plane, the second resolution being greater than the first resolution;

[0017] The second projection image is subjected to convolution processing according to a predetermined second target convolution kernel to obtain a second target feature map, the second target convolution kernel being determined according to a predetermined surface area of each pail hole, a total surface area of the tray and the second resolution.

[0018] In some embodiments, the target resolution includes the first resolution and the second resolution, the target projection image includes the first projection image and the second projection image, and the corresponding target convolution kernel is determined according to a predetermined surface area of each pail hole, a total surface area of the tray and the target resolution, including:

[0019] According to the target resolution, a size of the target convolution kernel of the target projection image corresponding to the target resolution is determined;

[0020] The sum of the surface areas of the plurality of pail holes is determined as a first surface area, and the difference between the total surface area of the tray and the first surface area is determined as a second surface area;

[0021] The ratio of the first numerical value to the first surface area is determined as the weight value of each point in the target convolution kernel corresponding to the pier hole region, and the ratio of the second numerical value to the second surface area is determined as the weight value of each point in the target convolution kernel corresponding to the other region in the tray except the pier hole region.

[0022] In some embodiments, the method further comprises:

[0023] The ratio of the first numerical value to the first surface area is determined as the weight value of each point in the target convolution kernel corresponding to the pier hole region, and the ratio of the second numerical value to the first surface area is determined as the weight value of each point in the target convolution kernel corresponding to the other region in the tray except the pier hole region, the first numerical value and the second numerical value being opposite numbers of each other; or

[0024] The weight values of each point in the target convolution kernel corresponding to the pier face region of each pier meet a two-dimensional Gaussian distribution, and the weight values of each point in the target convolution kernel corresponding to the pier hole region and the other region in the tray except the pier hole region are opposite numbers of each other.

[0025] In some embodiments, the target pose of the tray is determined according to the second target feature map and the target plane equation, comprising:

[0026] The depth information corresponding to each point cloud in the second target feature map is determined according to the target plane equation;

[0027] The three-dimensional target pose of the tray is obtained by integrating the two-dimensional coordinates of each point cloud in the second target feature map and the depth information corresponding to each point cloud.

[0028] The tray pose detection device provided by the embodiments of the present application comprises:

[0029] The extraction module is configured to perform region extraction on the acquired point cloud image of the target region to obtain an initial tray image, the target region being a region containing the tray, and the initial tray image being an image after filtering out the background region in the point cloud image;

[0030] The convolution module is configured to perform convolution processing on the initial tray image according to a predetermined first target convolution kernel to obtain a first target feature map, the first target convolution kernel being determined according to the surface area of each pier hole, the total surface area of the tray and a first resolution;

[0031] The fitting module is configured to perform plane fitting processing on the first target feature map to obtain a corresponding tray plane image and a target plane equation;

[0032] The convolution module is further configured to perform convolution processing on the tray planar image according to a predetermined second target convolution kernel to obtain a second target feature map, the second target convolution kernel being determined according to a surface area of each hole, a total surface area of the tray, and a second resolution, the second resolution being greater than the first resolution.

[0033] The determination module is configured to determine the target pose of the tray according to the second target feature map and the target planar equation.

[0034] The computer device provided in the embodiments of the present application includes a memory and a processor, the memory stores a computer program capable of running on the processor, and the processor implements the method provided in the embodiments of the present application when executing the program.

[0035] The computer readable storage medium provided in the embodiments of the present application stores a computer program, and the computer program is executed by a processor to implement the method provided in the embodiments of the present application.

[0036] The tray pose detection method, device, computer device, and computer readable storage medium provided in the embodiments of the present application obtain an initial tray image by performing region extraction on the acquired point cloud image of the target region, perform convolution processing on the initial tray image according to a predetermined first target convolution kernel to obtain a first target feature map, the first target convolution kernel being determined according to a surface area of each hole, a total surface area of the tray, and a first resolution, perform planar fitting processing on the first target feature map to obtain a corresponding tray planar image and a target planar equation, perform convolution processing on the tray planar image according to a predetermined second target convolution kernel to obtain a second target feature map, the second target convolution kernel being determined according to a surface area of each hole, a total surface area of the tray, and a second resolution, the second resolution being greater than the first resolution, and finally determine the target pose of the tray according to the second target feature map and the target planar equation. In this way, the rough pose of the tray image is extracted first, and then the accurate position is extracted based on the extracted rough pose, which can improve the detection accuracy of the tray pose. In addition, by filtering out the interference point cloud in the point cloud image, the detection speed of the tray pose can also be improved, thereby solving the technical problems proposed in the background art. BRIEF DESCRIPTION OF DRAWINGS

[0037] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and serve to explain the technical solutions of the present application together with the specification.

[0038] Figure 1 An application scenario diagram of the tray pose detection method provided in the embodiments of the present application is shown in the following figure.

[0039] Figure 2A structural schematic diagram of a tray provided for an embodiment of the present application is shown in FIG. 1.

[0040] Figure 3 An implementation flowchart of a tray pose detection method provided for an embodiment of the present application is shown in FIG. 2.

[0041] Figure 4 An implementation flowchart of a method for obtaining an initial tray image provided for an embodiment of the present application is shown in FIG. 3.

[0042] Figure 5 An implementation flowchart of another tray pose detection method provided for an embodiment of the present application is shown in FIG. 4.

[0043] Figure 6 An implementation flowchart of a method for determining a target convolution kernel provided for an embodiment of the present application is shown in FIG. 5.

[0044] Figure 7 A structural schematic diagram of a target convolution kernel provided for an embodiment of the present application is shown in FIG. 6.

[0045] Figure 8 A structural schematic diagram of a candidate tray image provided for an embodiment of the present application is shown in FIG. 7.

[0046] Figure 9 An implementation flowchart of a method for detecting a missing tray pose provided for an embodiment of the present application is shown in FIG. 8.

[0047] Figure 10 A structural schematic diagram of a tray pose detection device provided for an embodiment of the present application is shown in FIG. 9.

[0048] Figure 11 A structural schematic diagram of a computer device provided for an embodiment of the present application is shown in FIG. 10. DETAILED DESCRIPTION

[0049] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following will further describe the specific technical solutions of the present application with reference to the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.

[0051] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0052] It should be noted that the terms "first, second, third" used in the embodiments of this application are used to distinguish similar or different objects and do not represent a specific order of objects. It can be understood that "first, second, third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0053] In recent years, with the rise of intelligent logistics and unmanned warehouses, the technology of automated handling and sorting of goods by unmanned forklifts following programmed instructions has become a research and application hotspot. However, in the complex environment of unstructured warehouses, the position and posture of pallets are subject to considerable uncertainty due to factors such as work processes, equipment precision, and manual operations. This makes it difficult for unmanned forklifts to quickly and accurately identify the position and posture of pallets, thus failing to meet the high-efficiency operational needs of the warehousing and logistics industry.

[0054] Furthermore, existing pallet detection technologies often use 3D point cloud features to describe and detect pallet pose. However, methods using geometric features and feature points are usually affected by the quality of the point cloud, making it difficult to accurately and stably extract the physical features of the pallet. Additionally, due to the large volume of point cloud data, processing is typically time-consuming, making it difficult to achieve rapid detection with limited hardware resources. Therefore, there is an urgent need to develop a highly accurate and fast method for unmanned forklift pallet pose detection.

[0055] In view of this, embodiments of this application provide a pallet pose detection method, apparatus, device, and storage medium, which can improve the accuracy and speed of pallet pose detection. This method is applied to a detection device, which can be various types of devices with detection functions and information processing capabilities. For example, the detection device may include a depth camera, LiDAR, infrared camera, infrared scanner, etc. The functions implemented by this method can be achieved by a processor in the detection device calling program code. Of course, the program code can be stored in a computer storage medium. Therefore, the detection device includes at least a processor and a storage medium.

[0056] The following will be described in detail with reference to the accompanying drawings.

[0057] like Figure 1 As shown, Figure 1 This is a schematic diagram of an application scenario for a pallet pose detection method disclosed in an embodiment of this application. In this application scenario, the pallet pose detection system includes an unmanned forklift 11 and a detection device 12 installed on the unmanned forklift.

[0058] The unmanned forklift 11 can be a lurking AGV (Automated Guided Vehicle), a counterbalance AGV, a forklift AGV, etc., and this application embodiment does not limit it to this.

[0059] Among them, the unmanned forklift is an intelligent industrial vehicle robot that integrates forklift technology and AGV technology. Compared with ordinary AGVs, it can not only complete point-to-point material handling, but also realize logistics transportation connecting multiple production links. It is not only good at three major scenarios: high-bay warehouses, off-site receiving areas, and production line transfers, but also plays an irreplaceable role in heavy-duty and special handling scenarios. The application of unmanned forklifts can solve the problems of large material flow and high labor intensity of manual handling in industrial production and warehousing logistics operations.

[0060] Based on product characteristics, forklift-type AGVs are currently widely used in scenarios such as high-bay warehouses, off-site receiving areas, and production line transfers. Especially in the diverse material handling operations of manufacturing enterprises, unmanned forklifts can perform multiple functions according to different needs in stages such as entering and leaving the warehouse, production lines, and storage.

[0061] The detection device 12 can be any device with detection and sensing functions, such as a depth camera, LiDAR, infrared camera, or infrared scanner. Furthermore, in this embodiment, the number of detection devices 12 is not limited; there can be one or more.

[0062] In this embodiment, the installation location of the detection device 12 is not limited. In a preferred embodiment, the detection device 12 may be installed at the midpoint between the fork roots of the unmanned forklift 11 to perform pose detection on pallets in the warehouse.

[0063] Before introducing the technical solution of this application, the structure of the tray will first be explained. Figure 2 A schematic diagram of a tray structure is provided, such as... Figure 2 As shown, the pallet includes multiple pallet feet 201 and an upper pallet plate 202, and there are foot holes 203 between adjacent pallet feet 201. Here, the area where the pallet feet 201 are located is referred to as the foot surface area of ​​the foot foot.

[0064] Figure 3 This is a schematic diagram illustrating the implementation process of the pallet pose detection method provided in this application embodiment, which can improve the accuracy and speed of pallet pose detection. Figure 3 As shown, this method is applied to a detection device installed on an unmanned forklift, and the method may include the following steps 301 to 305:

[0065] Step 301: Extract the target area point cloud image to obtain an initial tray image. The target area is the area containing the tray, and the initial tray image is the image after filtering out the background area from the point cloud image.

[0066] In this embodiment, the control device can send movement commands to the unmanned forklift to control it to travel to a detection point, which can be any position between 1.5m and 3m away from the pallet. Once the unmanned forklift reaches the detection point, the detection device installed on it can acquire a point cloud image of the area where the pallet is located. After acquiring the point cloud image, region extraction is performed to filter out background areas, retaining only the pallet image, referred to as the initial pallet image. This effectively avoids the influence of interfering point clouds in the background area on the accuracy of pallet pose determination and reduces the amount of point cloud computing required for subsequent pallet pose calculations, thus improving calculation speed.

[0067] In some embodiments, step 301 can be achieved by performing the following steps 401 to 402:

[0068] Step 401: Obtain the initial pose of the tray.

[0069] Once the unmanned forklift reaches the inspection point, the initial pose of the pallet can be acquired through the inspection equipment installed on it. This initial pose is the pose of the pallet relative to the inspection equipment, which can be represented by (x, y, z, theta).

[0070] Where x, y, z refer to the lateral distance, longitudinal distance, and depth distance of the pallet relative to the center point of the testing equipment, based on the center point of the testing equipment coordinate system, and theta refers to the angle between the center point of the pallet and the center point of the testing equipment.

[0071] Step 402: Based on the initial pose of the tray and the preset region of interest threshold, region extraction is performed on the point cloud image to obtain the initial tray image, which is the image after filtering out the background region in the point cloud image.

[0072] The region of interest (ROI) is a region of the image that needs to be processed, delineated using shapes such as rectangles, circles, ellipses, or irregular polygons. This region is the focus of attention when analyzing the image.

[0073] In this embodiment of the application, after obtaining the initial pose (coarse pose) of the tray, the tray area can be extracted as the region of interest from the obtained point cloud image based on the preset region of interest threshold and the initial pose of the tray, thereby filtering out the interference of the background area in the point cloud image.

[0074] This effectively avoids the impact of interfering point clouds in the background area on the accuracy of pallet pose determination, and reduces the amount of point cloud computing required for subsequent pallet pose calculations, thereby improving calculation speed and accuracy.

[0075] Step 302: Perform convolution processing on the initial pallet image according to the predetermined first target convolution kernel to obtain the first target feature map. The first target convolution kernel is determined according to the predetermined surface area of ​​each hole, the total surface area of ​​the pallet, and the first resolution.

[0076] In this embodiment of the application, the size information of the pallet is pre-input into the detection device, such as... Figure 2 As shown, the pallet's dimensions include at least the width of the pallet footrest 201, the height of the pallet footrest 201, the height of the upper pallet plate 202, the width of the footrest opening 203, and the overall height of the pallet. In some embodiments, if the pallet has a lower pallet plate, the pallet's dimensions may also include the height of the lower pallet plate.

[0077] After obtaining the pallet's dimensions, the inspection equipment can calculate the surface area of ​​each part of the pallet, such as determining the surface area of ​​each hole and the total surface area of ​​the pallet.

[0078] Understandably, convolution processing of an image is essentially filtering the pixels in the image to extract the information of interest and usefulness. Depending on the type of convolution kernel, the feature map obtained after convolution processing will contain different features with varying emphases.

[0079] Furthermore, to determine the target convolution kernel corresponding to the image, it is necessary to determine the size of the target convolution kernel and the weight of each pixel at that size.

[0080] Therefore, to better meet practical needs and improve detection accuracy, in this embodiment, after extracting the initial tray image, the first target convolution kernel corresponding to the initial tray image can be jointly determined based on the surface area of ​​each hole, the total surface area of ​​the tray, and the first resolution. Specifically, the size of the first target convolution kernel can be determined based on the first resolution (i.e., the resolution of the initial tray image), and the weight of each pixel in the first target convolution kernel can be determined based on the surface area of ​​each hole and the total surface area of ​​the tray.

[0081] In some embodiments, step 302 can be implemented by performing steps 503 to 504 in the following embodiments.

[0082] Step 303: Perform plane fitting processing on the first target feature map to obtain the corresponding tray plane image and target plane equation.

[0083] In this embodiment, the method of fitting the first target feature map is not limited. For example, the Random Sample Consensus (RANSAC) algorithm can be used to perform planar fitting on the first target feature map to obtain the tray planar image corresponding to the first target feature map. In this way, the influence of other interfering objects and interfering point clouds in the first target feature map can be filtered out, improving the accuracy and speed of tray pose detection.

[0084] In addition, when fitting the first target feature map, three point cloud data (with known point cloud coordinates) can be randomly selected from the first target feature map and substituted into the target plane equation shown in Formula 1 to calculate the specific values ​​of a, b, c, and d, thereby determining the composition of the target plane equation.

[0085] ax+by+cz+d=0 (Formula 1);

[0086] Where a is the coefficient of the horizontal axis coordinate of the point cloud, b is the coefficient of the vertical axis coordinate of the point cloud, c is the coefficient of the vertical axis coordinate of the point cloud, and d is the intercept.

[0087] Step 304: Perform convolution processing on the pallet planar image according to the predetermined second target convolution kernel to obtain the second target feature map. The second target convolution kernel is determined according to the predetermined surface area of ​​each hole, the total surface area of ​​the pallet, and the second resolution. The second resolution is greater than the first resolution.

[0088] Here, the method for determining the second target convolutional kernel is the same as that for determining the first target convolutional kernel in step 302, and will not be repeated here. The difference is that the size of the second target convolutional kernel is determined based on the second resolution, while the size of the first target convolutional kernel is determined based on the first resolution.

[0089] In this embodiment, the second resolution is greater than the first resolution. Since the resolution corresponding to the initial tray image is the first resolution and the resolution corresponding to the tray plane image is the second resolution, the clarity of the tray plane image is much greater than that of the initial tray image. Of course, the tray plane image also contains a larger amount of point cloud data.

[0090] In this application, the specific method of convolution processing is not limited. For example, in some embodiments, since global convolution would consume a lot of computing resources when the amount of point cloud data is large, sliding window convolution can be used to perform convolution processing on the second target feature map. Sliding window convolution refers to using the second target convolution kernel to perform convolution processing on the tray planar image at a certain stride interval. That is to say, not every point cloud in the tray planar image is convolved, but a portion of the point cloud is convolved to obtain the accurate planar pose of the tray.

[0091] Of course, there are no restrictions on the value of the step size here; it can be set according to actual needs.

[0092] In some embodiments, step 304 can be achieved by performing steps 505 to 509 in the following embodiments.

[0093] Step 305: Determine the target pose of the tray based on the second target feature map and the target plane equation.

[0094] Understandably, after obtaining the second target feature map, the precise planar pose of the tray (two-dimensional coordinates (x0, y0)) can be obtained. However, in order to determine the target pose of the tray, it is also necessary to know the vertical axis distance of the tray relative to the center point of the detection device, that is, the vertical axis coordinate z0.

[0095] Therefore, in this embodiment of the application, the vertical axis coordinates corresponding to each point cloud can be determined based on the known horizontal and vertical coordinates of each point cloud in the second target feature map and the target plane equation, thereby obtaining the three-dimensional target pose of the tray.

[0096] In some embodiments, step 305 can be achieved by performing steps 510 to 511 as follows.

[0097] In this embodiment, an initial pallet image is obtained by extracting regions from the point cloud image of the target area. The initial pallet image is then convolved with a pre-determined first target convolution kernel to obtain a first target feature map. The first target convolution kernel is determined based on the pre-determined surface area of ​​each pier hole, the total surface area of ​​the pallet, and a first resolution. The first target feature map is then subjected to planar fitting to obtain a corresponding pallet planar image and a target plane equation. Subsequently, the pallet planar image is convolved with a pre-determined second target convolution kernel to obtain a second target feature map. The second target convolution kernel is determined based on the pre-determined surface area of ​​each pier hole, the total surface area of ​​the pallet, and a second resolution, where the second resolution is greater than the first resolution. Finally, the target pose of the pallet is determined based on the second target feature map and the target plane equation. This approach, by first extracting a coarse pose from the pallet image and then performing precise position extraction based on the extracted coarse pose, improves the accuracy of pallet pose detection. Furthermore, filtering out interfering point clouds in the point cloud image also improves the detection speed of the pallet pose.

[0098] This application provides another method for detecting the position and posture of a tray. Figure 5 This is a schematic diagram illustrating the implementation process of the pallet pose detection method provided in the embodiments of this application, as follows: Figure 5 As shown, the method may include the following steps 501 to 511:

[0099] Step 501: Extract the target area point cloud image to obtain an initial tray image. The target area is the area containing the tray, and the initial tray image is the image after filtering out the background area from the point cloud image.

[0100] Step 502: At the first resolution, the initial tray image is mapped to obtain the first projection image of the initial tray image on the projection plane.

[0101] Here, the initial tray image is mapped by projecting it along the z-axis at a first resolution, resulting in a two-dimensional planar image. In other words, the original three-dimensional initial tray image is converted into a two-dimensional planar image, effectively reducing the amount of data and improving data processing speed.

[0102] The resolution of a projected image refers to the number of point clouds per unit surface area. Different resolutions result in different levels of sharpness. In some embodiments, the first resolution is a lower resolution, resulting in a lower-resolution first projected image. In other words, the first projected image has lower sharpness.

[0103] In some embodiments, the projection plane is the XOY plane. Of course, the direction of projection of the initial tray image is not limited in the embodiments of this application; the initial tray image can also be projected along the x-axis or y-axis, and can be set according to specific needs.

[0104] In the embodiments of this application, the specific method of performing planar projection on the initial tray image is not limited. For example, the initial tray image can be orthogonally projected along the z-axis.

[0105] Step 503: Determine the first target convolution kernel based on the pre-determined surface area of ​​each hole, the total surface area of ​​the tray, and the first resolution.

[0106] Depending on the convolution kernel, the feature map obtained by convolutional processing of an image contains different feature emphases. Therefore, to better meet practical needs and improve detection accuracy, in this embodiment, the target convolution kernel to be used is determined jointly based on the pre-determined surface area of ​​each hole, the total surface area of ​​the tray, and the resolution of the projected image to be convolved.

[0107] Furthermore, in convolution operations, the weight distribution of the convolution kernel represents the penalty and gain on the current region. For example, based on the characteristics of the pier hole region, the non-pier hole region, and the pier surface region, different penalties and gains can be set for the pier hole region, the non-pier hole region, and the pier surface region, that is, different convolution kernels can be set for the pier hole region, the non-pier hole region, and the pier surface region.

[0108] Based on this, in some embodiments, to determine the first target convolution kernel corresponding to different regions in the tray contained in the first projected image, the following steps 601 to 603 can be performed:

[0109] Step 601: Determine the size of the first target convolution kernel corresponding to the first projected image based on the resolution of the first projected image.

[0110] The resolution of a projected image refers to the number of points per unit surface area. Different resolutions result in different levels of clarity. When determining the specific structure of a convolutional kernel, the size of the target convolutional kernel used in the current convolutional processing must first be determined. To better suit practical needs, the size of the target convolutional kernel corresponding to the projected image can be determined based on the image's resolution.

[0111] In this application embodiment, the specific method for determining the size of the first target convolution kernel corresponding to the first projected image based on the resolution of the first projected image is not limited.

[0112] In some embodiments, the size of the first target convolution kernel corresponding to the platform surface area of ​​the pallet pier in the first projected image can be determined to be proportional to the width and height of the pier. Here, the form of the first target convolution kernel is not limited; it can be square or rectangular. When the first target convolution kernel is square, it can be represented as follows: Figure 7 The 3x3 convolution kernel shown.

[0113] Step 602: Determine the sum of the surface areas of the multiple pier holes as the first surface area, and determine the difference between the total surface area of ​​the pallet and the first surface area as the second surface area.

[0114] like Figure 2 As shown, the tray includes multiple support holes 203. The sum of the surface areas of these multiple support holes 203 is determined as the first surface area, and the surface area of ​​the other areas in the tray excluding the support holes 203 is taken as the second surface area.

[0115] The sum of the surface areas of multiple holes (i.e., the first surface area) is represented by size_hole, the total surface area of ​​the pallet is represented by size_all, and the second surface area is represented by size_all-size_hole.

[0116] Step 603: Under the size of the first target convolution kernel, determine the target convolution kernel corresponding to each region in the tray based on the first surface area and the second surface area.

[0117] In this embodiment of the application, the method for determining the first target convolutional kernel may include the following:

[0118] (1) Determine the ratio of the first value to the first surface area as the weight of each point in the first target convolution kernel corresponding to the hole area, and determine the ratio of the second value to the second surface area as the weight of each point in the first target convolution kernel corresponding to other areas in the tray except the hole area.

[0119] In this embodiment, the specific values ​​of the first and second values ​​are not limited. For example, the first value can be set to 1 and the second value to -1. Then, the weight of each point in the first target convolution kernel corresponding to the hole area is 1 / size_hole, and the weight of each point in the first target convolution kernel corresponding to other areas in the tray besides the hole area is -1 / (size_all-size_hole).

[0120] Of course, in other embodiments, based on the same setting logic of the first target convolution kernel, the ratio of the first value to the surface area of ​​a single pier hole region can also be determined as the weight of each point in the first target convolution kernel corresponding to each pier hole region, and the ratio of the second value to the third surface area can be determined as the weight of each point in the first target convolution kernel corresponding to other regions in the tray except for the pier hole region, and the third surface area is the difference between the total surface area of ​​the tray and the surface area of ​​a single pier hole region.

[0121] (2) The ratio of the first value to the first surface area is determined to be the weight of each point in the first target convolution kernel corresponding to the hole area, and the ratio of the second value to the first surface area is determined to be the weight of each point in the first target convolution kernel corresponding to other areas in the tray except the hole area. The first value and the second value are opposites of each other.

[0122] If the first value is set to 1 and the second value is set to -1, then the weight of each point in the first target convolution kernel corresponding to the hole area is 1 / size_hole, and the weight of each point in the first target convolution kernel corresponding to other areas in the tray besides the hole area is -1 / size_hole.

[0123] Of course, in other embodiments, based on the same setting logic of the first target convolution kernel, the ratio of the first value to the surface area of ​​a single pier hole region can also be determined as the weight of each point in the first target convolution kernel corresponding to each pier hole region, and the ratio of the second value to the surface area of ​​a single pier hole region can be determined as the weight of each point in the first target convolution kernel corresponding to other regions in the tray except for the pier hole region, with the first value and the second value being opposite numbers to each other.

[0124] (3) Determine that the weights of each point in the first target convolution kernel corresponding to the pier surface area of ​​each foot pier conform to a two-dimensional Gaussian distribution, and that the weights of each point in the first target convolution kernel corresponding to each pier hole area are opposites of the weights of each point in the first target convolution kernel corresponding to other areas in the tray except for the pier hole area.

[0125] It should be noted that by making the weights of each point in the first target convolution kernel corresponding to the pier surface area of ​​each foot block conform to a two-dimensional Gaussian distribution, the point cloud constituting the pier surface area can have stronger noise resistance in the x and y directions, thus making the method applicable to the detection of pallet pose in the presence of noise.

[0126] The weights of each point in the first target convolution kernel corresponding to each pier hole area are set to be opposites of the weights of each point in the first target convolution kernel corresponding to other areas in the tray, excluding the pier hole area. That is, the ratio of the weights of each point in the first target convolution kernel corresponding to each pier hole area to the weights of each point in the first target convolution kernel corresponding to other areas in the tray is set to 1:-1.

[0127] Based on the various methods for determining the first target convolution kernel provided above, one method can be selected according to actual needs during practical use, making the convolution method more flexible and more in line with actual production needs, thereby improving the accuracy of pallet pose detection.

[0128] Step 504: Perform convolution processing on the first projected image according to the first target convolution kernel to obtain the first target feature map.

[0129] It should be noted that when using the target convolution kernel to perform convolution processing on the projected image, it means performing convolution processing on each of the aforementioned regions according to the first target convolution kernels corresponding to the pier hole region, pier surface region, and other regions in the tray besides the aforementioned regions in the pre-determined first projected image, thereby obtaining the first target feature map corresponding to the first projected image. This first target feature map contains the approximate pose of the tray on the XOY plane.

[0130] Step 505: Based on the predetermined total surface area of ​​the tray, extract the candidate tray image from the first target feature map.

[0131] Understandably, as described in step 302 above, before performing position detection on the pallet, the pallet's size information will be input into the detection device. For example, the width of the pallet foot block 201, the height of the pallet foot block 201, the height of the upper pallet plate 202, the width of the foot block hole 203, and the overall height of the pallet will be input into the detection device. In this way, the detection device can know the total surface area of ​​the pallet.

[0132] After obtaining the target feature map corresponding to the projected image, candidate pallet images can be extracted again from the target feature map based on the known total surface area of ​​the pallet. These candidate pallet images further filter out interfering point clouds from the target feature map and reduce the amount of point cloud processing data, thereby improving detection accuracy and speed when performing pallet pose detection based on the candidate pallet images.

[0133] like Figure 8 As shown, a schematic diagram of the construction of a candidate tray image is given, which is composed of multiple point clouds.

[0134] Step 506: Perform plane fitting processing on the candidate tray image to obtain the corresponding tray plane image and target plane equation.

[0135] Step 507: At the second resolution, the tray plane image is mapped to obtain a second projection image of the tray plane image on the projection plane, where the second resolution is greater than the first resolution.

[0136] Here, the mapping process for the pallet plane image is the same as that for the initial pallet image in step 502, and will not be repeated here. The difference is that the resolution of the second projected image is the second resolution, while the resolution of the first projected image is the first resolution. Therefore, the resolution of the second projected image is higher than that of the first projected image, meaning that the number of point clouds in the second projected image is much greater than the number of point clouds in the first projected image.

[0137] Step 508: Determine the second target convolution kernel based on the pre-determined surface area of ​​each hole, the total surface area of ​​the tray, and the second resolution.

[0138] Here, the method for determining the second target convolutional kernel is the same as that for determining the first target convolutional kernel in step 302, and will not be repeated here. The difference is that the size of the second target convolutional kernel is determined based on the second resolution, while the size of the first target convolutional kernel is determined based on the first resolution.

[0139] Step 509: Perform convolution processing on the second projected image according to the second target convolution kernel to obtain the second target feature map.

[0140] In this application, the specific method of convolution processing on the second projected image is not limited. For example, in some embodiments, since global convolution would consume significant computational resources when dealing with large amounts of point cloud data, sliding window convolution can be used to convolve the second target feature map. Sliding window convolution refers to using the second target convolution kernel to convolve the tray planar image at a certain stride interval. That is, not every point cloud in the tray planar image is convolved, but only a portion of the point cloud is convolved to obtain the accurate planar pose of the tray.

[0141] It should be noted that since the resolution corresponding to the second projected image is the second resolution, which is a higher resolution, the second target feature map obtained after convolution processing contains the precise pose of the tray on the XOY plane.

[0142] Step 510: Determine the depth information corresponding to each point cloud in the second target feature map based on the target plane equation.

[0143] In this embodiment of the application, the depth information corresponding to each point cloud in the second target feature map is determined according to the target plane equation, which can be achieved by the following formula 2:

[0144]

[0145] Where z0 is the depth information corresponding to each point cloud, a is the coefficient of the horizontal axis coordinate of the point cloud, b is the coefficient of the vertical axis coordinate of the point cloud, c is the coefficient of the vertical axis coordinate of the point cloud, d is the intercept, x0 is the horizontal coordinate of the point cloud, and y0 is the vertical coordinate of the point cloud.

[0146] Step 511: Combine the two-dimensional coordinates of each point cloud in the second target feature map with the corresponding depth information to obtain the three-dimensional target pose of the tray.

[0147] After calculating the depth information z0 corresponding to each point cloud, the horizontal and vertical coordinates of each point cloud contained in the second target feature map can be combined to obtain the three-dimensional coordinates of each point cloud, thereby obtaining the three-dimensional target pose of the tray.

[0148] In this embodiment, an initial pallet image is obtained by extracting regions from the point cloud image of the target area; at a first resolution, the initial pallet image is mapped to obtain a first projection image of the initial pallet image on the projection plane; a first target convolution kernel is determined based on the pre-determined surface area of ​​each pier hole, the total surface area of ​​the pallet, and the first resolution; the first projection image is convolved based on the first target convolution kernel to obtain a first target feature map; candidate pallet images are extracted from the first target feature map based on the pre-determined total surface area of ​​the pallet; and the candidate pallet images are subjected to plane fitting to obtain the corresponding pallet plane image and target plane equation; a second target convolution kernel is then determined based on the pre-determined surface area of ​​each pier hole, the total surface area of ​​the pallet, and the second resolution; and the second projection image is convolved based on the second target convolution kernel to obtain a second target feature map; finally, the depth information corresponding to each point cloud in the second target feature map is determined based on the target plane equation, thereby combining the two-dimensional coordinates of each point cloud in the second target feature map and the depth information corresponding to each point cloud to obtain the three-dimensional target pose of the pallet. This effectively reduces the amount of point cloud data in the pallet pose detection process and uses convolution kernels that are more in line with actual needs to perform convolution processing on the acquired images, making the extracted pallet features more accurate, thereby effectively improving the accuracy and speed of pallet pose detection.

[0149] The following describes an exemplary application of the embodiments of this application in a real-world application scenario.

[0150] Figure 9 The overall flow of the pallet pose detection method provided in the embodiments of this application is as follows. Figure 9 As shown, the method includes the following steps 901 to 907:

[0151] Step 901: Install a ToF camera (i.e., detection device) at the midpoint between the roots of the forks of the unmanned forklift. Alternatively, it can be installed in other locations or other sensors can be used depending on the actual conditions.

[0152] Step 902: Set up a configuration file in the ToF camera. Enter the pallet information in the configuration file, including the width of the pallet feet, the height of the pallet feet, the distance between two pallet feet, the pallet height, and the height of the upper and lower pallets.

[0153] Step 903: Control the unmanned forklift to drive to the detection point (ToF camera 1.5-3.0m away from the pallet), and send the initial pose ideal_pose (i.e. x, y, z, theta) to the ToF camera.

[0154] Step 904: The ToF camera extracts the ROI from the acquired depth image (i.e., point cloud image) based on the initial pose ideal_pose and the preset ROI threshold.

[0155] Step 905: In the camera coordinate system, the ToF camera re-orthogonally projects the point cloud in the ROI region (i.e., the initial tray image) of the depth image along the z-direction according to the input low resolution, to obtain a new low-resolution projected image (low_resolution_img) (i.e., the first projected image).

[0156] Step 906: Based on the prior information of the tray in step 902, generate rectangular convolution kernels that are proportional to the front of the tray at both high and low resolutions (the convolution kernels have the same resolution as the projected image).

[0157] Specifically, the weight distribution of the convolution kernel in the convolution operation represents the penalty and gain for the pier hole area and the non-pier hole area. Let the area of ​​the front of the pallet (XOZ in the camera coordinate system) be size_all and the area of ​​the pier hole on the front of the pallet be size_hole. The constructed kernel type is an equal-weight convolution kernel, that is, the weight of the corresponding hole area of ​​the pallet in the convolution kernel is set to -1 / size_hole, and the weight of the other areas (including the foot pier and the upper and lower plates) is set to 1 / (size_all-size_hole).

[0158] Step 907, one-stage low-resolution convolution: use the low-resolution equal-weight convolution kernel (i.e. the first target convolution kernel) from step 906 to convolve the low_resolution_img image to obtain the coarse pose of the tray in the XOY plane.

[0159] Step 908: Based on the coarse pose of the XOY plane and the prior tray size information, the candidate tray points (i.e., candidate tray images) are extracted after widening along the XOY plane.

[0160] Step 909: Using the RANSAC plane search algorithm, extract the pallet face point cloud from the candidate pallet points and calculate the plane equation (i.e., the target plane equation), as shown in Formula 1 in the above embodiment.

[0161] Step 910: Re-orthogonally project the data points on the tray surface along the z-direction according to the input high resolution to obtain a new high-resolution projected image (high_resolution_img) (i.e., the second projected image).

[0162] Step 911, two-stage high-resolution convolution: Since global convolution at high resolution consumes a lot of computational resources, the high-resolution equal-weight convolution kernel (i.e., the second target convolution kernel) from step 906 is used to perform sliding window convolution on the high_resolution_img image with a limited range centered on the coarse pose, so as to obtain the accurate pose (x0, y0) of the XOY plane.

[0163] Step 912: Combining the precise pose (x0, y0) and the pallet surface equation information, calculate the final pose (i.e., target pose) of the object in three-dimensional space according to Formula 2 in the above embodiment.

[0164] It should be understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0165] Based on the foregoing embodiments, this application provides a pallet pose detection device, which includes various modules and units included in each module, and can be implemented by a processor; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP) or field programmable gate array (FPGA), etc.

[0166] Figure 10 This is a schematic diagram of the pallet pose detection device provided in the embodiments of this application, as shown below. Figure 10 As shown, the device 1000 includes an extraction module 1001, a convolution module 1002, a fitting module 1003, and a determination module 1004, wherein:

[0167] Extraction module 1001 is used to extract regions from the point cloud image of the acquired target region to obtain an initial tray image. The target region is the region containing the tray, and the initial tray image is the image after filtering out the background region in the point cloud image.

[0168] The convolution module 1002 is used to perform convolution processing on the initial pallet image according to a predetermined first target convolution kernel to obtain a first target feature map. The first target convolution kernel is determined according to the predetermined surface area of ​​each pier hole, the total surface area of ​​the pallet, and a first resolution.

[0169] The fitting module 1003 is used to perform planar fitting processing on the first target feature map to obtain the corresponding tray plane image and target plane equation;

[0170] The convolution module 1002 is further configured to perform convolution processing on the pallet planar image according to a predetermined second target convolution kernel to obtain a second target feature map. The second target convolution kernel is determined according to the predetermined surface area of ​​each pier hole, the total surface area of ​​the pallet, and the second resolution, where the second resolution is greater than the first resolution.

[0171] The determination module 1004 is used to determine the target pose of the tray based on the second target feature map and the target plane equation.

[0172] In some embodiments, the convolution module 1002 is specifically used to perform mapping processing on the initial pallet image at a first resolution to obtain a first projection image of the initial pallet image on the projection plane; and to perform convolution processing on the first projection image according to a predetermined first target convolution kernel to obtain a first target feature map. The first target convolution kernel is determined according to the predetermined surface area of ​​each hole, the total surface area of ​​the pallet, and the first resolution.

[0173] In some embodiments, the fitting module 1003 is specifically used to extract candidate pallet images from the first target feature map based on the predetermined total surface area of ​​the pallet; and to perform planar fitting processing on the candidate pallet images to obtain the corresponding pallet planar image and target plane equation.

[0174] In some embodiments, the convolution module 1002 is further configured to perform mapping processing on the pallet plane image at a second resolution to obtain a second projection image of the pallet plane image on the projection plane, wherein the second resolution is greater than the first resolution; and to perform convolution processing on the second projection image according to a predetermined second target convolution kernel to obtain a second target feature map, wherein the second target convolution kernel is determined according to the predetermined surface area of ​​each pier hole, the total surface area of ​​the pallet, and the second resolution.

[0175] In some embodiments, the determining module 1004 is further configured to: determine the size of the target convolution kernel of the target projection image corresponding to the target resolution based on the target resolution; determine the sum of the surface areas of multiple pier holes as the first surface area, and determine the difference between the total surface area of ​​the tray and the first surface area as the second surface area; and, given the size of the target convolution kernel, determine the ratio of the first value to the first surface area as the weight of each point in the target convolution kernel corresponding to the pier hole area, and determine the ratio of the second value to the second surface area as the weight of each point in the target convolution kernel corresponding to other areas in the tray besides the pier hole area.

[0176] In some embodiments, the determining module 1004 is further configured to, given the size of the target convolution kernel, determine the ratio of a first value to a first surface area as the weight of each point within the target convolution kernel corresponding to the pier hole region, and determine the ratio of a second value to the first surface area as the weight of each point within the target convolution kernel corresponding to other regions in the tray besides the pier hole region, wherein the first value and the second value are opposites of each other; or determine that the weight of each point within the target convolution kernel corresponding to the pier surface region of each pier conforms to a two-dimensional Gaussian distribution, and that the weight of each point within the target convolution kernel corresponding to each pier hole region is opposite to the weight of each point within the target convolution kernel corresponding to other regions in the tray besides the pier hole region.

[0177] In some embodiments, the determining module 1004 is specifically used to determine the depth information corresponding to each point cloud in the second target feature map according to the target plane equation; and to obtain the three-dimensional target pose of the tray by combining the two-dimensional coordinates of each point cloud in the second target feature map and the depth information corresponding to each point cloud.

[0178] The description of the above device embodiments is similar to that of the above method embodiments, and has similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0179] It should be noted that, in the embodiments of this application... Figure 10 The module division of the pallet pose detection device shown is illustrative and represents only one logical functional division; in actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, exist as separate physical units, or be integrated into one unit from two or more other units. The integrated units can be implemented in hardware, as software functional units, or a combination of both.

[0180] It should be noted that, in the embodiments of this application, if the above-described methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0181] This application provides a computer device, which may be a server, and its internal structure diagram may be as follows: Figure 11 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a pallet pose detection method.

[0182] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method provided in the above embodiments.

[0183] This application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps in the method provided in the above-described method embodiments.

[0184] Those skilled in the art will understand that Figure 11 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0185] In one embodiment, the pallet pose detection device provided in this application can be implemented as a computer program, and the computer program can be implemented in the form of... Figure 11It runs on the computer device shown. The computer device's memory can store the various program modules that make up the sampling device, for example, Figure 10 The extraction module, convolution module, fitting module, and determination module are shown. The computer program, comprised of these modules, causes the processor to execute the steps of the pallet pose detection methods of the various embodiments of this application described in this specification.

[0186] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium, storage medium, and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0187] It should be understood that the phrases "one embodiment," "an embodiment," or "some embodiments" mentioned throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment," "in one embodiment," or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments; their similarities or commonalities can be referred to mutually, and for the sake of brevity, they will not be repeated here.

[0188] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three kinds of relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.

[0189] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0190] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or modules can be electrical, mechanical, or other forms.

[0191] The modules described above as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.

[0192] In addition, each functional module in the various embodiments of this application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the integrated modules can be implemented in hardware or in the form of hardware plus software functional units.

[0193] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0194] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0195] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0196] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0197] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0198] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for detecting the pose of a tray, characterized in that, The pallet includes multiple pallet feet, and there are foot holes between adjacent pallet feet. The method includes: Region extraction is performed on the point cloud image of the acquired target area to obtain an initial tray image. The target area is the area containing the tray, and the initial tray image is the image after filtering out the background area from the point cloud image. The initial pallet image is convolved according to a predetermined first target convolution kernel to obtain a first target feature map. The first target convolution kernel is determined based on the surface area of ​​each of the pier holes, the total surface area of ​​the pallet, and a first resolution. The first target feature map is subjected to plane fitting processing to obtain the corresponding tray plane image and target plane equation; The pallet planar image is convolved according to a predetermined second target convolution kernel to obtain a second target feature map. The second target convolution kernel is determined based on the surface area of ​​each of the pier holes, the total surface area of ​​the pallet, and a second resolution, wherein the second resolution is greater than the first resolution. The target pose of the tray is determined based on the second target feature map and the target plane equation.

2. The method according to claim 1, characterized in that, The step of convolving the initial tray image according to a predetermined first target convolution kernel to obtain a first target feature map includes: At the first resolution, the initial tray image is mapped to obtain a first projected image of the initial tray image on the projection plane; The first projected image is convolved according to a predetermined first target convolution kernel to obtain a first target feature map. The first target convolution kernel is determined based on the surface area of ​​each of the pier holes, the total surface area of ​​the tray, and the first resolution.

3. The method according to claim 1, characterized in that, The step of performing planar fitting processing on the first target feature map to obtain the corresponding tray plane image and target plane equation includes: Candidate tray images are extracted from the first target feature map based on the predetermined total surface area of ​​the tray; The candidate tray image is subjected to plane fitting processing to obtain the corresponding tray plane image and the target plane equation.

4. The method according to claim 2, characterized in that, The step of convolving the tray planar image according to a predetermined second target convolution kernel to obtain a second target feature map includes: At the second resolution, the tray plane image is mapped to obtain a second projection image of the tray plane image on the projection plane, where the second resolution is greater than the first resolution. The second projected image is convolved according to a predetermined second target convolution kernel to obtain a second target feature map. The second target convolution kernel is determined based on the surface area of ​​each of the pier holes, the total surface area of ​​the tray, and the second resolution.

5. The method according to claim 4, characterized in that, The target resolution includes a first resolution and a second resolution, and the target projected image includes the first projected image and the second projected image. The corresponding target convolution kernel is determined based on the pre-determined surface area of ​​each of the aforementioned holes, the total surface area of ​​the tray, and the target resolution, including: Based on the target resolution, determine the size of the target convolution kernel of the target projection image corresponding to the target resolution; The sum of the surface areas of the plurality of said holes is determined as the first surface area, and the difference between the total surface area of ​​the pallet and the first surface area is determined as the second surface area; Given the size of the target convolution kernel, the ratio of a first value to the first surface area is determined as the weight of each point in the target convolution kernel corresponding to the hole region, and the ratio of a second value to the second surface area is determined as the weight of each point in the target convolution kernel corresponding to other regions in the tray besides the hole region.

6. The method according to claim 5, characterized in that, The method further includes: Given the size of the target convolutional kernel, the ratio of a first value to the first surface area is determined as the weight of each point within the target convolutional kernel corresponding to the hole region, and the ratio of a second value to the first surface area is determined as the weight of each point within the target convolutional kernel corresponding to other regions in the tray besides the hole region, wherein the first value and the second value are opposites of each other; or The weights of each point in the target convolution kernel corresponding to the pier surface area of ​​each foot pier conform to a two-dimensional Gaussian distribution, and the weights of each point in the target convolution kernel corresponding to each pier hole area are opposite to the weights of each point in the target convolution kernel corresponding to other areas in the tray except for the pier hole area.

7. The method according to claim 1, characterized in that, Determining the target pose of the tray based on the second target feature map and the target plane equation includes: Based on the target plane equation, determine the depth information corresponding to each point cloud in the second target feature map; By combining the two-dimensional coordinates of each point cloud in the second target feature map and the corresponding depth information, the three-dimensional target pose of the tray is obtained.

8. A pallet position detection device, characterized in that, include: The extraction module is used to extract regions from the point cloud image of the acquired target region to obtain an initial tray image. The target region is the region containing the tray, and the initial tray image is the image after filtering out the background region from the point cloud image. A convolution module is used to perform convolution processing on the initial pallet image according to a predetermined first target convolution kernel to obtain a first target feature map. The first target convolution kernel is determined according to the predetermined surface area of ​​each pier hole, the total surface area of ​​the pallet, and a first resolution. The fitting module is used to perform planar fitting processing on the first target feature map to obtain the corresponding tray plane image and target plane equation; The convolution module is further configured to perform convolution processing on the pallet planar image according to a predetermined second target convolution kernel to obtain a second target feature map. The second target convolution kernel is determined according to the predetermined surface area of ​​each of the pier holes, the total surface area of ​​the pallet, and a second resolution, wherein the second resolution is greater than the first resolution. The determination module is used to determine the target pose of the tray based on the second target feature map and the target plane equation.

9. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Target detection method and device

    CN108446694A

  • Tray pose detection method and device, automated guided vehicle and goods transportation system

    CN111533051A