A tray pose recognition method and device
Patent Information
- Application Number
- CN202311599810.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2043-11-24
AI Technical Summary
[0003]本申请实施例提供一种托盘位姿识别方法及装置,用以解决托盘摆放位置误差过大,造成叉取货物失败,甚至一些安全隐患的问题
[0032]本申请提供的一种托盘位姿识别方法及装置,通过深度相机获取托盘应用场景的深度图像数据,并处理多帧深度图像数据,滤波融合出剔除了无效数据的深度图像数据;基于深度相机的成像原理、场景信息和托盘的结构信息对托盘进行识别;结合深度相机的相机标定参数,托盘的结构信息和识别后的托盘区域,确定所述托盘的位姿信息,提高了托盘位姿识别方法的准确性和鲁棒性。
Smart Images

Figure CN117576424B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of unmanned forklift technology, and in particular to a pallet pose recognition method and device. Background Technology
[0002] Many factories and warehouses are now using automated forklifts (AGTs) for picking and moving goods. When using AGTs to pick up goods, the pallet's position must be precisely positioned; otherwise, picking may fail or even pose safety hazards. Therefore, pallet posture recognition methods are needed to calculate the pallet's posture angle and socket position deviation to accurately identify the pallet's position and posture information, enabling calibration during picking and improving the accuracy of AGT end-effector operations. Summary of the Invention
[0003] This application provides a pallet position recognition method and device to solve the problem of excessive pallet placement error causing failure to pick up goods and even some safety hazards.
[0004] Firstly, this application proposes a pallet pose recognition method, including:
[0005] The multi-frame depth images of the tray in the scene are filtered and fused to obtain a tray depth image after invalid data is removed. The multi-frame depth images are acquired by a depth camera.
[0006] Based on the imaging principle of the depth camera, scene information, and tray structure information, the tray depth image after invalid data has been removed is identified to obtain tray candidate regions.
[0007] Based on the camera calibration parameters of the depth camera and the structural information of the tray, similarity matching is performed on the candidate regions of the tray to obtain the tray region;
[0008] Data calculation and fitting are performed on the tray area to determine the tray's pose information in the scene.
[0009] Optionally, the filtering and fusion of multiple frames of depth images of the tray in the scene includes:
[0010] Statistically calculate the mean and standard deviation of the data for each pixel location in the depth image of the tray in the scene;
[0011] Determine whether the residual between each pixel location data and the corresponding mean is greater than a preset location threshold; if so, it is considered invalid data.
[0012] After removing invalid data, the mean is recalculated as the pixel position data after removing invalid data. The pixel position data after removing invalid data constitutes the tray depth image after removing invalid data.
[0013] Optionally, the step of identifying pallet depth images after removing invalid data to obtain pallet candidate regions includes:
[0014] The spatial relative relationship between the tray and the optical center of the depth camera is calculated based on the information of the scene and the camera calibration parameters of the depth camera.
[0015] Based on the spatial relative relationship between the tray and the optical center of the camera, candidate imaging data of the tray in the depth camera are calculated according to the imaging principle of the depth camera.
[0016] The pixel position data of the tray depth image after removing invalid data is matched sequentially with the candidate imaging data. Pixel position data of the depth image that is greater than the preset matching threshold is represented as non-tray pixels, otherwise it is retained as tray pixels, thus filtering out the tray candidate region.
[0017] Optionally, the similarity matching of the candidate tray regions includes:
[0018] The similarity matching of pallet imaging templates in the pallet candidate area is performed using a sliding window method. The pallet imaging templates are obtained by calculation and fitting using the height and width of the pallet forklift surface and the height and width of the pallet legs.
[0019] Determine if the candidate pallet region is greater than the preset similarity threshold. If so, select the candidate pallet region with the highest similarity value as the pallet region. If not, it means there is no pallet.
[0020] Optionally, the step of performing data calculation and fitting on the tray area to determine the tray's pose information in the scene includes:
[0021] Based on the depth image data of the pallet foot areas on both sides of the pallet area, N sets of data are sampled from the center of the pallet foot area, and the attitude angles of N pallets and the mean and standard deviation of the N sets of data are calculated.
[0022] Calculate whether the residual between each attitude angle and the mean is greater than the preset attitude angle threshold. If so, it is considered an outlier. After removing the outliers, the mean is recalculated as the final attitude angle.
[0023] The pixel at the center of the tray is calculated based on the structural information of the tray, and then the offset of the pixel at the center of the tray is calculated by combining the final attitude angle. The pose information of the tray in the scene is the final attitude angle and offset.
[0024] Secondly, this application proposes a tray pose recognition device, comprising:
[0025] The depth image filtering and fusion module is used to filter and fuse multiple frames of depth images of the tray in the scene to obtain a tray depth image after invalid data has been removed. The multiple frames of depth images are acquired by a depth camera.
[0026] The tray candidate region acquisition module is used to identify the tray depth image after invalid data has been removed, based on the imaging principle of the depth camera, scene information and tray structure information, to obtain the tray candidate region.
[0027] The tray region recognition module is used to perform similarity matching on the candidate regions of the tray based on the camera calibration parameters of the depth camera and the structural information of the tray to obtain the tray region;
[0028] The pose recognition module is used to perform data calculation and fitting on the tray area to determine the pose information of the tray in the scene.
[0029] Thirdly, this application proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the tray pose recognition method described in the first aspect.
[0030] Fourthly, this application proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the tray pose recognition method described in the first aspect.
[0031] The beneficial effects of the technical solutions provided in some embodiments of this application include at least the following:
[0032] This application provides a pallet pose recognition method and apparatus, which acquires depth image data of a pallet application scenario using a depth camera, processes multiple frames of depth image data, and filters and fuses them to obtain depth image data that has eliminated invalid data; the pallet is recognized based on the imaging principle of the depth camera, scene information, and pallet structural information; and the pallet pose information is determined by combining the camera calibration parameters of the depth camera, the structural information of the pallet, and the recognized pallet area, thereby improving the accuracy and robustness of the pallet pose recognition method.
[0033] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. The objectives and other advantages of this application are realized and obtained through the structures particularly pointed out in the description, claims and drawings.
[0034] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings.
[0035] The advantages of this application in terms of its additional aspects will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0036] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is a schematic flowchart of the pallet pose recognition method shown in the embodiments of this application;
[0038] Figure 2 This is a schematic diagram illustrating the tray pose recognition process in an embodiment of this application;
[0039] Figure 3 This is a schematic block diagram of the tray recognition device shown in the embodiment of this application;
[0040] Figure 4 This is a schematic diagram of the structure of the electronic device shown in the embodiment of this application. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0042] In the following description, when referring to the accompanying drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of methods and systems consistent with some aspects of this application as detailed in the appended claims.
[0043] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0044] The following is combined with Figures 1-4The pallet pose recognition method and device provided in this application will be described in detail through specific embodiments and application scenarios.
[0045] This application provides a pallet pose recognition method, such as... Figure 1 As shown, it includes the following steps:
[0046] Step S1: Filter and fuse the multi-frame depth images of the tray in the scene to obtain a tray depth image after removing invalid data. The multi-frame depth images are acquired by a depth camera.
[0047] Specifically, based on the statistical characteristics of depth image data, multi-frame filtering and fusion are performed to obtain tray depth image data in the scene after removing noise and some interference data;
[0048] In this embodiment, the depth camera must be able to stably acquire depth image data over a certain distance, and is not limited to structured light depth cameras, time-of-flight (TOF) depth cameras, binocular stereo depth cameras, etc.
[0049] It should be noted that when depth cameras are installed on equipment such as unmanned forklifts, due to the wide variety of equipment types and application scenarios, it is necessary to calibrate the external parameters of the depth camera in combination with the equipment type and application scenario.
[0050] For example, the extrinsic parameters of the camera can be calibrated according to the installation of the depth camera. The calibration parameters involved include pitch, roll, yaw, dx, dy, and dz. In this embodiment, the extrinsic parameters of the camera are calibrated using the following formula.
[0051]
[0052]
[0053]
[0054]
[0055]
[0056] Where angle_yaw represents the yaw angle of the depth camera mount, angle_pitch represents the pitch angle of the depth camera mount, d1 and d2 represent the pixel depth values at two known calibration locations (distances between the calibration locations and the camera plane), and p x1 and p x2 p represents the horizontal pixel coordinates of the two known calibration positions. y1 and p y2fx and fy represent the vertical pixel coordinates of the two known calibration positions, respectively; fx and fy represent the horizontal and vertical pixel focal lengths of the depth camera, respectively; c x and c y These represent the coordinates of the depth camera's optical center in the pixel coordinate system;
[0057] dx, dy, and dz represent the three offsets between the known calibration point and the optical center of the depth camera in three-dimensional coordinates (world coordinate system), d represents the pixel depth value of the known calibration point (the distance from the calibration point to the camera plane), and p x p represents the horizontal pixel coordinates of a known calibration point. y The pixel coordinates in the vertical direction representing the known calibration point.
[0058] It should be noted that the roll angle is installed at zero degrees by default (horizontal installation), which does not affect subsequent calculations.
[0059] In the specific implementation process, multi-frame filtering and fusion requires first collecting 5 to 20 frames of depth image data, and then statistically analyzing and calculating the mean and standard deviation of the data at each pixel position in the depth image of the tray in the scene.
[0060] Determine whether the residual between each pixel location data and the corresponding mean is greater than a preset location threshold. If so, it is considered invalid data. In this embodiment, invalid data is noise data.
[0061] After removing invalid data, the mean is recalculated as the pixel position data after removing invalid data. The pixel position data after removing invalid data constitutes the tray depth image after removing invalid data.
[0062] Step S2: Based on the imaging principle of the depth camera, the scene information, and the structural information of the tray, the tray depth image after invalid data has been removed is identified to obtain the tray candidate region;
[0063] Specifically, based on the imaging principle of the depth camera and scene information, the imaging data of the tray in the depth camera is calculated and inferred. The calculated imaging data is then used to match and filter candidate regions of the tray from pixels in all depth images; this includes the following steps:
[0064] Step S2.1: Calculate the spatial relative relationship between the tray and the optical center of the depth camera based on the scene information and the camera calibration parameters of the depth camera;
[0065] In the specific implementation process, the required scene information includes the tray height and maximum horizontal deviation, and the depth camera calibration parameters include the calibration parameters angle_pitch, angle_yaw, dx, dy, and dz. These parameters are used to calculate the spatial relative relationship between the tray and the optical center of the depth camera, i.e., the tray candidate position information. The tray candidate position information is the p-value of all candidate position pixels. x and p y ;
[0066] Step S2.2: Based on the spatial relative relationship between the tray and the optical center of the camera, calculate the candidate imaging data of the tray in the depth camera according to the imaging principle of the depth camera;
[0067] Specifically, by substituting the candidate position information of the tray obtained in step S2.1 into the formula in step S1 for reverse calculation, the candidate imaging data of the tray in the pixel coordinate system can be obtained. The candidate imaging data is the depth value of the pixel point at the candidate position of the tray.
[0068] Step S2.3: The pixel position data of the tray depth image after removing invalid data is matched with the candidate imaging data in sequence. The pixel position data of the depth image that is greater than the preset matching threshold is represented as non-tray pixels, otherwise it is retained as tray pixels, that is, the tray candidate region is selected.
[0069] Step S3: Based on the camera calibration parameters of the depth camera and the structural information of the tray, perform similarity matching on the candidate regions of the tray to obtain the tray region;
[0070] Specifically, in the candidate tray region, a sliding window method is used to perform similarity matching with all tray imaging templates. If a similarity greater than a preset similarity threshold exists, the region with the highest similarity is selected as the tray region; otherwise, it indicates that there is no tray. The steps include the following:
[0071] Step S3.1: Perform similarity matching on the pallet imaging templates in the pallet candidate area using a sliding window method. The pallet imaging templates are obtained by calculating and fitting using the height and width of the pallet forklift surface and the height and width of the pallet legs.
[0072] In the specific implementation process, based on the known height and width of the pallet forklift surface and the height and width of the pallet legs, the pallet imaging template at different depths and different attitude angles is calculated sequentially.
[0073] The location information that needs to be substituted is calculated from the pallet candidate area. For example, the depth range can be quantized based on the depth range d_range of the pallet candidate area, with a quantization accuracy of 10 mm and a quantization level of d_range / 10; the location information can be quantized based on the pixel location range pix_range of the pallet candidate area, with a quantization accuracy of 2 pixels and a quantization level of pix_range / 2; the attitude angle is quantized based on the range of the actual application scenario. Taking (-30 degrees, 30 degrees) as an example, with a quantization accuracy of 1 degree and a quantization level of 61, the number of pallet imaging templates is 61×d_range / 10×pix_range / 2.
[0074] Step S3.2: Determine whether the candidate pallet region is greater than the preset similarity threshold. If so, select the candidate pallet region with the highest similarity value as the pallet region. If not, it means there is no pallet.
[0075] In this step, the matching method compares the pixel values (depth values) of the corresponding regions one by one. If the error between the pixel value and the pixel value fitted to the template is less than the accuracy of the depth camera, it means that the match is successful. This is accumulated and finally divided by the total number of pixels to obtain the similarity. The higher the similarity, the closer it is to the real tray area.
[0076] In the specific implementation process, the matching strategy is to divide the template into six parts according to the structural information of the tray, and match the left tray leg area, left tray hole area, middle tray leg area, right tray hole area, right tray leg area and the whole tray area from left to right. The similarity of each area must be greater than the preset similarity threshold. If any area is less than or equal to the preset similarity threshold, the matching stops and the next round of matching begins. If the similarity of each area is greater than the preset similarity threshold, the matching is successful.
[0077] It should be noted that after matching, the obtained attitude angles and position information can basically realize the identification and localization of the tray area, but the pose of the tray still needs to take into account the errors generated during the matching process and the accuracy errors of the depth camera itself, so further precise calculations are required.
[0078] Step S4: Perform data calculation and fitting on the tray area to determine the pose information of the tray in the scene.
[0079] To further refine the calculation of the tray pose information, a multi-sample data statistical fitting method is employed to reduce errors generated during the matching process and the accuracy errors inherent in the depth camera itself. The specific steps are as follows:
[0080] Step S4.1: Based on the depth image data of the pallet foot areas on both sides of the pallet area, sample N sets of data at the center of the pallet foot area, and calculate the attitude angles of N pallets and the mean and standard deviation of the N sets of data;
[0081] Specifically, in this embodiment, N is set to 25. Based on the depth image data of the pallet foot areas on both sides of the pallet area, 5x5 depth image data are sampled from the center of each of the left and right pallet foot areas. The attitude angles of the 25 pallets, as well as the mean and standard deviation of the data, are calculated one by one. The formula for calculating the attitude angle is as follows.
[0082]
[0083] Step S4.2: Calculate whether the residual between each attitude angle and the mean is greater than the preset attitude angle threshold. If so, it is considered an outlier. After removing the outlier, the mean is recalculated as the final attitude angle.
[0084] Step S4.3: Calculate the pixel at the center of the tray based on the structural information of the tray, that is, take the center pixel of the middle tray foot area, and then calculate the offset deltax and deltay of the pixel at the center of the tray based on the final attitude angle. The calculation formula is as follows.
[0085]
[0086]
[0087] Wherein, d represents the pixel depth value of the pixel at the center of the tray (distance from the center of the tray to the camera plane), TrayMiddleLeg represents the width of the middle leg of the tray, TrayHeight represents the height of the tray, hc represents the height of the camera optical center relative to the tray when the transport equipment picks up the tray (preset by the transport equipment), and other parameters are consistent with the formula in step S1.
[0088] Obtain the pose information of the tray in the scene: the final pose angle, offset deltax, and deltay.
[0089] It should be noted that this pallet pose recognition method can be applied to different types of pallet scenarios. It only requires configuring the structural parameters of different types of pallets, generating templates for different types of pallets in sequence, and traversing them one by one. The pallet with the highest similarity is then selected as the forklift target.
[0090] It should be noted that this pallet pose recognition method can handle multiple pallets in a unified scene, and the target pallet to be picked can be selected based on the scene information.
[0091] The following describes a specific example of pallet recognition, such as... Figure 2As shown, the diagrams from top to bottom correspond to the processing results of steps S1 to S3. The first image is the fused depth image data after removing noise and some interference data; the second image is the tray candidate region selected by matching and filtering the imaging data calculated and inferred based on the imaging principle of the depth camera and scene information; the third image is the tray candidate region detected and identified by similarity matching based on the structural information of the tray.
[0092] The pallet pose recognition method of this disclosure is applied to a pallet pose recognition device, which is mainly used in logistics and warehousing scenarios such as unmanned forklifts to transport goods.
[0093] This application utilizes depth cameras deployed on unmanned forklifts to identify pallet poses, enabling precise pallet picking, while improving the end-of-line efficiency of unmanned forklifts and reducing deployment costs.
[0094] Example 2
[0095] This application provides a tray pose recognition device. For example... Figure 3 As shown, it includes:
[0096] The depth image filtering and fusion module is used to filter and fuse multiple frames of depth images of the tray in the scene to obtain a tray depth image after invalid data has been removed. The multiple frames of depth images are acquired by a depth camera.
[0097] The tray candidate region acquisition module is used to identify the tray depth image after invalid data has been removed, based on the imaging principle of the depth camera, scene information and tray structure information, to obtain the tray candidate region.
[0098] The tray region recognition module is used to perform similarity matching on the candidate regions of the tray based on the camera calibration parameters of the depth camera and the structural information of the tray to obtain the tray region;
[0099] The pose recognition module is used to perform data calculation and fitting on the tray area to determine the pose information of the tray in the scene.
[0100] According to the pallet pose recognition device provided in the embodiments of this application, depth image data of the application scene is acquired by a depth camera, and multiple frames of depth image data are processed. The depth image data is filtered and fused to remove noise and some interference data. Based on the imaging principle of the depth camera, scene information and pallet structural information, the pallet is identified. Combining the camera calibration parameters of the depth camera, the structural information of the target pallet and the recognition results of the pallet area, the pose information of the target pallet is determined, which improves the accuracy and robustness of the pallet pose recognition method.
[0101] Example 3
[0102] This application provides an electronic device, such as... Figure 4 As shown, an electronic device 400 includes a power interface 401, a data interface 402, a communication interface 403, a memory 404, a processor 405, and a computer program that can run on the processor 405. When the program is executed by the processor 405, it implements the various processes of the above-described pallet pose recognition method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0103] It is understood that the electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a laptop, a handheld computer, an in-vehicle electronic device, a mobile internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), an ATM, or a self-service machine, etc. The embodiments of this application do not specifically limit the scope.
[0104] The electronic device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0105] Example 4
[0106] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described pallet pose recognition method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0107] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0108] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0109] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
[0110] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0111] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0112] The applicant has provided a detailed description of the implementation examples of this application in conjunction with the accompanying drawings. However, those skilled in the art should understand that the above implementation examples are merely preferred embodiments of this application. The detailed description is only intended to help readers better understand the spirit of this application and is not intended to limit the scope of protection of this application. On the contrary, any improvements or modifications made based on the inventive spirit of this application should fall within the scope of protection of this application.
Claims
1. A method for pallet pose recognition, characterized in that, include: The multi-frame depth images of the tray in the scene are filtered and fused to obtain a tray depth image after invalid data is removed. The multi-frame depth images are acquired by a depth camera. The multi-frame depth images include 5 to 20 frames of depth images; the filtering and fusion of the multi-frame depth images of the tray in the scene includes: for each pixel position in the 5 to 20 frames of depth images, statistically calculating the mean and standard deviation of the depth data at that pixel position in each frame; determining whether the residual between each depth data at that pixel position and the corresponding mean is greater than a preset position threshold, if so, indicating invalid data; after removing invalid data, recalculating the mean of the depth data at that pixel position as the pixel position data after removing invalid data, and the tray depth image after removing invalid data is composed of the pixel position data after removing invalid data at each pixel position; Based on the imaging principle of the depth camera, scene information, and tray structure information, the tray depth image after invalid data has been removed is identified to obtain tray candidate regions. The step of identifying the tray depth image after removing invalid data to obtain tray candidate regions includes: calculating the spatial relative relationship between the tray and the optical center of the depth camera based on the scene information and the camera calibration parameters of the depth camera; calculating candidate imaging data of the tray in the depth camera based on the spatial relative relationship between the tray and the optical center of the depth camera and the imaging principle of the depth camera; sequentially matching the pixel position data of the tray depth image after removing invalid data with the candidate imaging data, and representing the pixel position data of the depth image that is greater than a preset matching threshold as non-tray pixels, otherwise retaining them as tray pixels, so as to filter out the tray candidate regions; The scene information includes the tray height and maximum horizontal deviation; the camera calibration parameters include the pitch angle and yaw angle of the depth camera installation, as well as three offsets between the known calibration point and the optical center of the depth camera in the world coordinate system; the spatial relative relationship includes the horizontal and vertical pixel coordinates of all candidate position pixels; the candidate imaging data is the depth value of the candidate position pixels on the tray. Based on the camera calibration parameters of the depth camera and the structural information of the tray, similarity matching is performed on the candidate regions of the tray to obtain the tray region; Data calculation and fitting are performed on the tray area to determine the tray's pose information in the scene.
2. The pallet pose recognition method according to claim 1, characterized in that, The similarity matching of the candidate regions of the tray includes: The similarity matching of pallet imaging templates in the pallet candidate area is performed using a sliding window method. The pallet imaging templates are obtained by calculation and fitting using the height and width of the pallet forklift surface and the height and width of the pallet legs. Determine whether the similarity of the candidate pallet regions is greater than the preset similarity threshold. If so, select the candidate pallet region with the highest similarity value as the pallet region. If not, it means there is no pallet. The similarity calculation includes: comparing the pixel depth value of the corresponding region with the pixel depth value fitted to the tray imaging template one by one, recording the pixels with an error less than the depth camera accuracy as successfully matched, and dividing the number of successfully matched pixels by the total number of pixels as the similarity. The tray imaging template is divided into a left tray leg region, a left tray hole region, a middle tray leg region, a right tray hole region, a right tray leg region, and a full tray region according to the tray's structural information. The left tray leg region, left tray hole region, middle tray leg region, right tray hole region, right tray leg region, and full tray region are matched sequentially. When the similarity of any region is less than or equal to the preset similarity threshold, the current round of matching is stopped and the next round of matching begins. When the similarity of all regions is greater than the preset similarity threshold, the matching is considered successful.
3. The tray pose recognition method according to claim 1, characterized in that, The step of performing data calculation and fitting on the tray area to determine the tray's pose information in the scene includes: Based on the depth image data of the pallet foot areas on both sides of the pallet area, N sets of data are sampled from the center of the pallet foot area, and the attitude angles of N pallets and the mean and standard deviation of the N sets of data are calculated. Calculate whether the residual between each attitude angle and the mean is greater than the preset attitude angle threshold. If so, it is considered an outlier. After removing the outliers, the mean is recalculated as the final attitude angle. The pixel at the center of the tray is calculated based on the structural information of the tray, and then the offset of the pixel at the center of the tray is calculated by combining the final attitude angle. The pose information of the tray in the scene is the final attitude angle and offset.
4. A tray position recognition device, characterized in that, include: The depth image filtering and fusion module is used to filter and fuse multiple frames of depth images of the tray in the scene to obtain a tray depth image after invalid data has been removed. The multiple frames of depth images are acquired by a depth camera. The multi-frame depth image includes 5 to 20 frames of depth images. The depth image filtering and fusion module is specifically used for: for each pixel position in the 5 to 20 frames of depth images, statistically calculating the mean and standard deviation of the depth data at that pixel position in each frame; determining whether the residual between each depth data at that pixel position and the corresponding mean is greater than a preset position threshold, if so, indicating invalid data; after removing invalid data, recalculating the mean of the depth data at that pixel position as the pixel position data after removing invalid data, and constructing the tray depth image after removing invalid data from each pixel position. The tray candidate region acquisition module is used to identify the tray depth image after invalid data has been removed, based on the imaging principle of the depth camera, scene information and tray structure information, to obtain the tray candidate region. Specifically, the tray candidate region acquisition module is used to: calculate the spatial relative relationship between the tray and the optical center of the depth camera based on the scene information and the camera calibration parameters of the depth camera; calculate candidate imaging data of the tray in the depth camera based on the spatial relative relationship between the tray and the optical center of the depth camera and the imaging principle of the depth camera; sequentially match the pixel position data of the tray depth image after removing invalid data with the candidate imaging data, and represent the pixel position data of the depth image that is greater than a preset matching threshold as non-tray pixels, otherwise retain them as tray pixels, so as to filter out the tray candidate region; The scene information includes the tray height and maximum horizontal deviation; the camera calibration parameters include the pitch angle and yaw angle of the depth camera installation, as well as three offsets between the known calibration point and the optical center of the depth camera in the world coordinate system; the spatial relative relationship includes the horizontal and vertical pixel coordinates of all candidate position pixels; the candidate imaging data is the depth value of the candidate position pixels on the tray. The tray region recognition module is used to perform similarity matching on the candidate regions of the tray based on the camera calibration parameters of the depth camera and the structural information of the tray to obtain the tray region; The pose recognition module is used to perform data calculation and fitting on the tray area to determine the pose information of the tray in the scene.
5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the tray pose recognition method as described in any one of claims 1-3.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the tray pose recognition method as described in any one of claims 1-3.
Citation Information
Patent Citations
Data processing method and system for TOF depth camera
CN112446836A
RGB-D-based tray pose estimation method, system and device
CN112907666A