A positioning method, device, storage medium and electronic device for a forklift pallet
By combining the image processing technology of RGB-D sensors and cameras, we can identify and locate the pallets, and solve the problems of large data processing volume and low recognition efficiency in the prior art, and achieve efficient pallet positioning.
Patent Information
- Application Number
- CN202010896387.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-31
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-08-31
AI Technical Summary
In the prior art, there is a problem of large data processing volume and low recognition efficiency in the pallet recognition process. Especially when the pallet is not modified, the recognition efficiency of planar lidar is low and costly.
The depth image is obtained by using the RGB-D sensor, and combined with the planar image acquired by the camera, the area of the tray in the depth image is recognized through the planar image, plane segmentation and template matching are performed, and the position of the tray is determined.
By combining planar images and depth images, the image area that needs to be processed is reduced, the amount of data is reduced, the recognition efficiency is improved, the number of template matching is reduced, and the image processing speed is improved.
Smart Images

Figure CN114202548B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image processing and pattern recognition, and particularly to a positioning method, device, storage medium and electronic device for a forklift tray. Background Art
[0002] In automated and semi-automated warehousing systems, the identification and positioning of trays play an important role. A tray is a horizontal platform device for placing goods and products during the processes of containerization, stacking, handling and transportation, and is widely used in fields such as production, circulation and warehousing. The identification of a tray refers to installing a sensor on a forklift to detect and identify the tray in a warehouse, and the positioning refers to calculating the three-dimensional coordinate information of the tray relative to the forklift based on the sensor information on the basis of identification.
[0003] Current detection technologies are divided into two types according to whether the tray needs to be modified:
[0004] 1) Method of modifying the tray: Stick identifiers on the end face of the tray. For example, stick artificial marks at different parts of the end face, such as sticking black-and-white concentric circles on both sides and in the middle of the end face of the tray, or sticking a highly reflective reflective tape on the entire end face. Use relevant technologies of pattern recognition in the sensor data to identify these artificial marks to complete the identification and positioning of the tray.
[0005] 2) Method of not modifying the tray: Such methods use the existing features of the tray itself to complete identification. For example, detecting two notches on the end face of the tray is a common method.
[0006] Disadvantages of the prior art:
[0007] The method of sticking identifiers on the end face of the tray limits the circulation of the tray and is prone to wear during the use of the tray. In addition, considering economic costs, the application and popularization involve the modification of a large number of existing trays, with high labor costs and time costs.
[0008] The method of not modifying the tray is almost entirely based on a horizontally installed planar lidar. However, the planar lidar has a limited field of view in the vertical direction, and the identification of the tray requires the movement of the forklift for compensation, resulting in low efficiency. In addition, considering economic costs, the current cost of lidar is relatively high, which is not conducive to application and popularization.
[0009] In view of the above problems, Chinese application CN105976375A discloses a method for pallet recognition and positioning based on an RGB-D sensor, including: obtaining a depth image through the sensor; performing plane segmentation on the point cloud of the depth image to obtain one or more planes, constituting a plane set; determining relevant planes that may contain the pallet from the plane set; and matching the relevant planes according to a preset pallet template to identify and position the pallet in the relevant planes. Although the above solution has low requirements for the shooting illumination conditions, during the recognition process, it is necessary to perform noise reduction, eliminate invalid points, window judgment, and plane set processing on the entire depth image, and when performing template matching, it is necessary to match the plane set with all templates, resulting in a large amount of data processing and a complex processing process during the whole process, and the recognition efficiency is relatively low. Summary of the Invention
[0010] Therefore, the technical problem to be solved by the present invention is to overcome the defects of large data processing amount and low recognition efficiency in the prior art during the pallet recognition process, so as to provide a positioning method, device, storage medium, and electronic device for a forklift pallet.
[0011] To achieve the above object, the present invention provides the following solutions:
[0012] In a first aspect, an embodiment of the present invention provides a method for positioning a forklift pallet, including the following steps:
[0013] Obtain a depth image obtained based on an RGB-D sensor;
[0014] Obtain a plane image collected by a camera, where the camera has the same shooting angle as the RGB-D sensor, and the size of the collected plane image is the same as that of the depth image;
[0015] Use the plane image to identify the forklift pallet and determine the area where the forklift pallet is located in the depth image;
[0016] Perform plane segmentation processing on the area where the forklift pallet is located in the depth image to determine a plane set containing the forklift pallet;
[0017] Match the plane image with a pre-prepared pallet template to obtain a matching target pallet template;
[0018] Use the determined plane set containing the forklift pallet to match with the target pallet template, and mark the position of the target pallet template in the second target area;
[0019] Convert the position of the target pallet template in the second target area into coordinates in three-dimensional space to obtain the position of the forklift pallet.
[0020] In one embodiment, the method of using the planar image to identify the forklift tray and determining the area where the forklift tray is located in the depth image includes:
[0021] Performing image recognition on the planar image to determine a first target area where the forklift tray is located in the planar image;
[0022] Obtaining the coordinate interval of the first target area in the planar image;
[0023] Using the coordinate interval to demarcate a second target area where the forklift tray is located in the depth image as the area where the forklift tray is located in the depth image.
[0024] In one embodiment, matching the planar image with a pre-prepared tray template to obtain a matching target tray template includes:
[0025] Extracting a first image feature in the planar image and a second image feature of the tray template;
[0026] Comparing the first image feature and the second image feature one by one and calculating the similarity between each corresponding feature;
[0027] Performing weighted summation on all the calculated similarities to obtain the similarity between the forklift tray in the planar image and the tray template;
[0028] When the similarity reaches a preset similarity threshold, determining that the tray template is the target tray template.
[0029] In one embodiment, using the determined planar set containing the forklift tray to match with the target tray template and marking the position of the target tray template in the second target area includes:
[0030] Calculating the matching degree between the target tray template and the planar set;
[0031] When the matching degree between the target tray template and the planar set reaches a preset matching threshold, determining that the target tray template matches the forklift tray and marking the position of the target tray template in the second target area.
[0032] In a second aspect, an embodiment of the present invention provides a positioning device for a forklift tray, including:
[0033] A first acquisition module, configured to acquire a depth image acquired based on an RGB-D sensor;
[0034] A second acquisition module, configured to acquire a planar image collected by a camera, wherein the camera has the same shooting angle as the RGB-D sensor, and the size of the acquired planar image is the same as that of the depth image;
[0035] A determination module, configured to identify a forklift tray by using the planar image and determine the area where the forklift tray is located in the depth image;
[0036] A segmentation module, configured to perform planar segmentation processing on the area where the forklift tray is located in the depth image to determine a planar set including the forklift tray;
[0037] A matching module, configured to match the planar image with a pre-prepared tray template to obtain a matched target tray template;
[0038] A marking module, configured to match the determined planar set including the forklift tray with the target tray template and mark the position of the target tray template in the second target area;
[0039] A positioning module, configured to convert the position of the target tray template in the second target area into coordinates in a three-dimensional space to obtain the position of the forklift tray.
[0040] In one embodiment, the determination module includes:
[0041] An identification unit, configured to perform image recognition on the planar image to determine a first target area where the forklift tray is located in the planar image;
[0042] An acquisition unit, configured to acquire the coordinate range of the first target area in the planar image;
[0043] A circumscribing unit, configured to circumscribe a second target area where the forklift tray is located in the depth image by using the coordinate range as the area where the forklift tray is located in the depth image.
[0044] In one embodiment, the matching module includes:
[0045] An extraction unit, configured to extract a first image feature in the planar image and a second image feature of the tray template;
[0046] A comparison unit, configured to compare the first image feature and the second image feature one by one and calculate the similarity between each corresponding feature;
[0047] A calculation unit, configured to perform weighted summation on all calculated similarities to obtain the similarity between the forklift tray in the planar image and the tray template;
[0048] A confirmation unit, when the similarity reaches a preset similarity threshold, determines that the pallet template is the target pallet template.
[0049] In one embodiment, the marking module includes:
[0050] A calculation unit that calculates the matching degree between the target pallet template and the plane set;
[0051] A marking unit, when the matching between the target pallet template and the plane set reaches a preset matching threshold, determines that the target pallet template matches the forklift pallet, and marks the position of the target pallet template in the second target area.
[0052] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores computer instructions, and when the computer instructions are executed by a processor, the pallet recognition and positioning method according to any one of the embodiments in the first aspect is implemented.
[0053] In a fourth aspect, an embodiment of the present invention provides an electronic device, including:
[0054] A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the pallet recognition and positioning method according to any one of the embodiments in the first aspect.
[0055] The technical solution of the present invention has the following advantages: By using a camera to collect a planar image, identifying the forklift pallet in the image, matching it with the depth image obtained by the RGB-D sensor, determining the area of the forklift pallet in the depth image, then performing planar segmentation on the depth image area of the matched forklift pallet to determine a plane set containing the forklift pallet, matching the planar image with a pre-prepared pallet template to obtain a target pallet template, and finally using the determined plane set containing the forklift pallet to match with the target pallet template, marking the position of the target pallet template in the second target area, and converting the position into three-dimensional space coordinates to obtain the position of the forklift pallet. In the embodiment of the present invention, by combining the planar image and the depth image, the area where the pallet is located is matched, and image processing is performed on this area, reducing the image area that needs to be processed, greatly reducing the amount of data for image processing. At the same time, by using the planar image to perform pallet template matching, compared with the method of matching the binary image after depth image processing with all templates, the processing speed is faster, reducing the number of depth image matching templates and improving the image processing efficiency. Description of the Drawings
[0056] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0057] Figure 1 Flowchart of a method for positioning a forklift pallet provided by an embodiment of the present invention;
[0058] Figure 2 Flowchart of a method for identifying a forklift pallet using a planar image and determining the area where the forklift pallet is located in a depth image provided by an embodiment of the present invention;
[0059] Figure 3 Flowchart of a method for depth image segmentation based on random Hough transform provided by an embodiment of the present invention;
[0060] Figure 4 Schematic diagram of preparing a pallet end face template offline provided by an embodiment of the present invention;
[0061] Figure 5 Flowchart of a method for matching a planar image with a pre-prepared pallet template to obtain a matching target pallet template provided by an embodiment of the present invention;
[0062] Figure 6 Flowchart of a method for matching the determined plane set containing the forklift pallet with the target pallet template and marking the position of the target pallet template in the second target area provided by an embodiment of the present invention;
[0063] Figure 7 Schematic diagram of the structure of a positioning device for a forklift pallet provided by an embodiment of the present invention;
[0064] Figure 8 Schematic diagram of the structure of the determination module of a positioning device for a forklift pallet provided by an embodiment of the present invention;
[0065] Figure 9 Schematic diagram of the structure of the matching module of a positioning device for a forklift pallet provided by an embodiment of the present invention;
[0066] Figure 10 Schematic diagram of the structure of the marking module of a positioning device for a forklift pallet provided by an embodiment of the present invention;
[0067] Figure 11 Composition diagram of a specific example of an electronic device provided by an embodiment of the present invention. Specific embodiments
[0068] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0069] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0070] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can also be the internal connection of two components. It can be a wireless connection or a wired connection. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0071] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0072] Embodiment 1
[0073] The embodiment of the present invention provides a positioning method for a forklift tray, as Figure 1 shown, including the following steps:
[0074] Step S101, obtaining a depth image acquired based on an RGB-D sensor;
[0075] Depth image = ordinary RGB three-channel color image + Depth Map. The RGB color model is a color standard in the industry. It obtains various colors through the changes of the three color channels of red (R), green (G), and blue (B) and their superposition with each other. RGB represents the colors of the three channels of red, green, and blue. This standard almost includes all the colors that the human vision can perceive and is one of the most widely used color systems at present; in 3D computer graphics, a DepthMap (depth map) is an image or image channel that contains information related to the distance of the surface of the scene object from the viewpoint. Among them, the Depth Map is similar to a grayscale image, except that each pixel value of it is the actual distance of the sensor from the object. Usually, the RGB image and the Depth image are registered, so there is a one-to-one correspondence between the pixel points.
[0076] In the embodiments of the present invention, it is necessary to convert the depth image into a point cloud, that is, to convert the image coordinate system into the world coordinate system using the formula. The formula is:
[0077]
[0078] Where x, y, z are the point cloud coordinate system, x', y' are the image coordinate system, and D is the depth value. In the embodiments of the present invention, by using an RGB-D sensor to obtain the depth image, it is beneficial to intuitively distinguish the distance relationship between each pixel point and the sensor, so as to judge whether the pixel point belongs to the same plane as other pixel points.
[0079] Step S102: Obtain the planar image collected by the camera. Among them, the camera has the same shooting angle as the RGB-D sensor, and the size of the collected planar image is the same as that of the depth image;
[0080] After the camera collects and obtains the planar image, perform binary processing on it so that each pixel on the image has only two possible value or gray level states. Generally, black and white, B&W, and monochrome images are used to represent binary images.
[0081] The RGB-D sensor and the camera are arranged in parallel and the distance is infinitely close. At this time, it can be considered that the planar image collected by the camera and the image collected by the RGB-D sensor are two images taken from the same perspective and have the same size. At this time, the image does not need to be corrected; of course, it can also be considered that the shooting angles of these two images are approximately the same, then the image taken by the camera can be slightly corrected according to the image obtained by the RGB-D sensor to obtain two images with exactly the same shooting angle and size. In practical applications, it is set according to the situation whether the image needs to be corrected, and the present invention does not make a limitation here.
[0082] In the embodiment of the present invention, on the basis of obtaining the RGB-D depth image, a camera is used to collect a planar image, and two photos with the same shooting angle and the same size are obtained, which is beneficial to determining the area of the required image by using the mapping relationship in the subsequent operation process.
[0083] Step S103: Identify the forklift tray by using the planar image, and determine the area where the forklift tray is located in the depth image;
[0084] Since the feature recognition technology of the planar image is relatively mature, and the inventor found that in the working environment of the forklift, there are usually certain requirements for light, and the features of the forklift tray on the planar image captured by cameras such as CCD / CMOS can still be clearly recognized. Therefore, in the embodiment of the present invention, the area recognition can be first performed by using the planar image to circle the approximate area of the forklift tray to obtain a detection frame. Since the shooting angles of the planar image and the depth image are the same and the image sizes are the same, the detection frame can be directly mapped to the depth image, so that the area where the forklift tray is located in the depth image can be determined.
[0085] After determining the depth image area of the forklift tray, the depth image of this area is generated into a point cloud and fused into the final point cloud for output. The actual depth image collected by any depth image acquisition device inevitably has noise. Therefore, it is necessary to perform smoothing processing on the image. First, use a 3x3 window to perform median filtering on the image to filter out random pulse noise, and then use a Gaussian smoothing filter to filter out other random noise and quantization noise to obtain a smoothed image.
[0086] As an alternative embodiment, as Figure 2 shown, identifying the forklift tray by using the planar image and determining the area where the forklift tray is located in the depth image includes the following steps:
[0087] Step S1031: Perform image recognition on the planar image to determine the first target area where the forklift tray is located in the planar image;
[0088] Step S1032: Obtain the coordinate interval of the first target area in the planar image;
[0089] Since the shooting angles and sizes of the planar image and the depth image are approximately the same, it can be considered that the pixel points in the planar image can correspond one by one in the depth image. Therefore, first identify and confirm the position of the first target area of the forklift tray in the planar image, and determine the coordinate interval where the first target area is located through coordinates.
[0090] Step S1033: Use the coordinate interval to delineate the second target area where the forklift tray is located in the depth image, which is used as the area where the forklift tray is located in the depth image.
[0091] Through the mapping relationship, determine the second target area where it is located in the depth image from the coordinate interval of the forklift tray in the planar image.
[0092] In the embodiment of the present invention, by using the position of the forklift tray in the planar image to confirm the position area of the forklift tray in the depth image, the recognition process of the position of the forklift tray in the depth image is greatly simplified, and the recognized position area of the forklift tray is accurate and has a high precision.
[0093] Step S104: Perform planar segmentation processing on the area where the forklift tray is located in the depth image to determine a set of planes containing the forklift tray;
[0094] The planar segmentation methods for depth images are generally divided into three categories:
[0095] The first category: Edge-based methods. The basic idea is to use a suitable convolution operator to convolve the image to obtain the corresponding gradient image of the image. Since the edges of the image often have large differences in image pixels and large gradients, the gradient image of the image is obtained through a suitable convolution kernel, that is, the edge image of the image is obtained. Advantages: For traditional operator gradient detection, only convolution with a suitable convolution kernel is required to quickly obtain the corresponding edge image. Disadvantages: The image edges may not be accurate. The gradients of complex images may not only appear at the image edges but may also appear in the colors and textures inside the image.
[0096] The second category: Region-based methods. Commonly used ones include traditional algorithms combined with genetic algorithms, region growing algorithms, region splitting and merging, watershed algorithms, etc. Among them, the deep learning segmentation algorithm based on regions and semantics is the main direction with more achievements and research in current image segmentation.
[0097] The third category: Graph-based segmentation algorithms, which are simple to implement, relatively fast in speed, and have high precision. Of course, other segmentation algorithms can also be used, such as: depth image segmentation based on random Hough transform, depth image segmentation based on normal component edge fusion, etc.
[0098] In the embodiment of the present invention, taking the depth image segmentation based on random Hough transform as an example, as Figure 3 shown, it includes the steps:
[0099] Step S201: Quantize the entire parameter space into multiple sub-regions;
[0100] Step S202: Randomly select three non-collinear points in the image;
[0101] Step S203, if the distances between any two of the three points are between a pre-set maximum distance threshold and a minimum distance threshold, calculate the plane parameters determined by these three points, obtain the sub-region where the determined plane parameters are located, and increment the cumulative plane number of this sub-region by 1;
[0102] Step S204, if one of the distances between any two of the three points is not between the pre-set maximum distance threshold and the minimum distance threshold, re-select three points;
[0103] Step S205, if there is a sub-region whose cumulative plane number exceeds a pre-specified threshold T, or the number of iterations exceeds P, then the parameters of the sub-region with the most cumulative plane numbers are the found plane F; otherwise, return to Step S202.
[0104] The prior art already has relatively mature technologies for depth image segmentation. The embodiments of the present invention are only used as examples here and are not limited thereto.
[0105] The embodiments of the present invention segment the smoothed depth image to confirm the plane of the forklift tray in the depth image area, which is beneficial to subsequent positioning processing of the tray.
[0106] Step S105, match the plane image with a pre-prepared tray template to obtain a matching target tray template;
[0107] The pre-prepared tray template and the plane image can both be binary images. When there is a new template, only the binary image of the new tray template needs to be stored in the system template library. It is also possible to process the two into binary images during template matching and then perform the matching. In the embodiments of the present invention, in order to better connect the tray template matched by using the planar graph and the matching of the plane set of the depth image with the tray template, a series of tray templates can be prepared in advance using the plane image for planar image matching to determine the target tray template; then a series of tray templates corresponding to the foregoing series of templates are prepared using the depth image. After determining the target tray template, the template prepared using the depth image data corresponding to the target tray template is used for positioning matching. Since the planar image matching algorithm is relatively mature and fast, after determining the template tray template using it, it is not necessary to match the plane set corresponding to the depth image with all the tray templates, which improves the efficiency of positioning and recognition.
[0108] Such as Figure 4As shown in the figure, measure the lengths of various parts of the end face of the measurement template, and discretize it into grids. The side length of the grid is determined according to the resolution of the sensor and the accuracy required by the system. For example, the side length used in the figure is 1 cm. The size of the binary image is the number of grids contained in the minimum outer rectangle of the tray. The assignment rule for the pixel points of the binary image is as follows: If the grid corresponds to the end face of the tray (the shaded part in the figure), the corresponding pixel point of the binary image is assigned a value of 1; conversely, if the grid corresponds to the notch of the tray, the corresponding pixel point of the binary image is assigned a value of 0.
[0109] As an alternative implementation, as Figure 5 shown, the step of matching the planar image with a pre-prepared tray template to obtain a matching target tray template includes the following steps:
[0110] Step S1051, extract the first image feature in the planar image and the second image feature of the tray template;
[0111] Since the images of the planar image and the tray template are both binary images after processing, the image features of the binary image include texture change features, boundary features, gray change features, etc.
[0112] Step S1052, compare the first image feature and the second image feature one by one, and calculate the similarity between each corresponding feature;
[0113] Use the formula where θ is the ratio of the number of pixels with the same pixel value in the current position of the tray template and its corresponding planar image to the total number of pixels of the tray template, and μ is the ratio of the number of pixels with a pixel value of 1 in the template to its total number of pixels.
[0114] Step S1053, perform a weighted sum of all the calculated similarities to obtain the similarity between the forklift tray in the planar image and the tray template;
[0115] Step S1054, when the similarity reaches a preset similarity threshold, determine that this tray template is the target tray template.
[0116] The similarity threshold can be a value set manually or a value obtained by the system through a learning algorithm. The present invention does not make a limitation here. Mark the type and position in the planar image of the tray template whose similarity reaches the preset similarity threshold.
[0117] In the embodiment of the present invention, matching the planar image of the forklift tray captured by the camera with a pre-prepared tray template is beneficial to determining the type and model of the forklift tray, and facilitating subsequent positioning in the depth image.
[0118] Step S106: Use the determined set of planes containing the forklift tray to match with the target tray template, and mark the position of the target tray template in the second target area.
[0119] As an alternative implementation, as Figure 6 shown, the step of using the determined set of planes containing the forklift tray to match with the target tray template and mark the position of the target tray template in the second target area includes the steps of:
[0120] Step S1061: Calculate the matching degree between the target tray template and the set of planes.
[0121] Since both the binary image of the target tray template and the point cloud image of the set of planes contain data such as gray levels, etc., the matching degree between the target tray template and the set of planes containing the forklift tray can be comprehensively calculated based on numerical values such as edge contours, area sizes, and gray states.
[0122] Step S1062: When the matching between the target tray template and the set of planes reaches a preset matching threshold, determine that the target tray template matches the forklift tray, and mark the position of the target tray template in the second target area.
[0123] In the embodiment of the present invention, by matching the binary image of the target tray template with the set of planes containing the forklift tray, the regional position of the target tray template in the depth image can be directly marked, and the point cloud of the forklift tray in the set of planes can be corrected using the target tray template, which is beneficial to obtaining more accurate point cloud data of the forklift tray.
[0124] Step S107: Convert the position of the target tray template in the second target area into coordinates in three-dimensional space to obtain the position of the forklift tray.
[0125] The RGB picture in the RGB-D image provides the x and y coordinates in the pixel coordinate system, while the depth map directly provides the Z coordinate in the camera coordinate system, that is, the distance between the camera and the point.
[0126] According to the information of the RGB-D image and the internal parameters of the camera, the coordinates of any pixel point in the camera coordinate system can be calculated.
[0127] According to the information of the RGB-D image and the internal and external parameters of the camera, the coordinates of any pixel point in the world coordinate system can be calculated.
[0128] Within the camera's field of view, the coordinates of the obstacle points in the camera coordinate system are the point cloud sensor data, which is also the point cloud data in the camera coordinate system. The point cloud sensor data can be calculated based on the coordinates provided by the RGB-D image and the camera's internal parameters.
[0129] The coordinates of all obstacle points in the world coordinate system are the point cloud map data, which is also the point cloud data in the world coordinate system. The point cloud map data can be calculated based on the coordinates provided by the RGB-D image and the camera's internal and external parameters.
[0130] The point cloud data in the camera coordinate system can be used to find the X and Y coordinate values in the camera coordinate system based on the x and y coordinates (i.e., u and v in the formula) in the pixel coordinate system provided by the RGB image and the camera's internal parameters. At the same time, the depth map directly provides the Z coordinate value in the camera coordinate system. Thus, the coordinates in the camera coordinate system can be obtained. The coordinates of the obstacle points in the camera coordinate system are the point cloud sensor data, which is also the point cloud data in the camera coordinate system.
[0131] According to the relationship formula between the camera coordinate system P and the coordinates of the points in the pixel coordinate system P uv :
[0132]
[0133] The coordinate transformation formula from the world coordinate system to the pixel coordinate system:
[0134]
[0135] Obtain the coordinate relationship between the world coordinate system P w and the pixel coordinate system P uv . Therefore, using the above formula, after inputting the pixel coordinates [u, v] and the depth Z, the coordinates of the points in the world coordinate system P w can be obtained. The coordinates of all obstacle points in the world coordinate system are the point cloud map data. According to the internal parameter formula, the three-dimensional coordinates of this point in the camera coordinate system can be calculated as P = [X, Y, Z]. Then, based on the homogeneous transformation matrix T of the camera or the rotation matrix and translation vector R, t, the three-dimensional coordinates of this point in the world coordinate system P w = [X w , Y w , Z w can be obtained. P w = [X w , Y w , Z w is the point cloud calibrated in the world coordinate system.
[0136] In an embodiment of the present invention, by using the information of an RGB-D image and the internal and external parameters of a camera, the coordinate data of the depth image is transformed, and the coordinate position of the target tray template in the camera coordinate system in the depth image area is converted into the coordinate in the world coordinate system, so as to obtain the position of the forklift tray, which is beneficial to complete the positioning of the forklift tray.
[0137] Embodiment 2
[0138] An embodiment of the present invention provides a positioning device for a forklift tray, as Figure 7 shown, including:
[0139] A first acquisition module 301, configured to acquire a depth image acquired based on an RGB-D sensor;
[0140] A second acquisition module 302, configured to acquire a planar image collected by a camera, wherein the camera has the same shooting angle as the RGB-D sensor, and the size of the collected planar image is the same as that of the depth image;
[0141] A determination module 303, configured to identify a forklift tray by using the planar image and determine the area where the forklift tray is located in the depth image;
[0142] A segmentation module 304, configured to perform planar segmentation processing on the area where the forklift tray is located in the depth image to determine a planar set containing the forklift tray;
[0143] A matching module 305, configured to match the planar image with a pre-prepared tray template to obtain a matched target tray template;
[0144] A marking module 306, configured to match the determined planar set containing the forklift tray with the target tray template and mark the position of the target tray template in the second target area;
[0145] A positioning module 307, configured to convert the position of the target tray template in the second target area into coordinates in a three-dimensional space to obtain the position of the forklift tray.
[0146] In an embodiment of the present invention, a planar image is collected by using a camera, and a forklift tray in the image is recognized and matched with a depth image obtained by an RGB-D sensor to determine the area of the forklift tray in the depth image. Then, the depth image area of the forklift tray obtained by matching is subjected to planar segmentation to determine a planar set containing the forklift tray. The planar image is matched with a pre-prepared tray template to obtain a target tray template. Finally, the determined planar set containing the forklift tray is matched with the target tray template to mark the position of the target tray template in the second target area, and the position is converted into three-dimensional space coordinates to obtain the position of the forklift tray. In the embodiment of the present invention, by combining the planar image and the depth image, the area where the tray is located is matched, and image processing is performed on this area, reducing the image area to be processed and greatly reducing the amount of data for image processing. At the same time, by using the planar image to match the tray template, compared with the method of matching the binary image after depth image processing with all templates, the processing speed is faster, the number of depth image matching templates is reduced, and the image processing efficiency is improved.
[0147] Embodiment 3
[0148] An embodiment of the present invention provides a determination module for a positioning device of a forklift tray, as Figure 8 shown, including:
[0149] An identification unit 3031, configured to perform image recognition on the planar image to determine a first target area where the forklift tray is located in the planar image;
[0150] An acquisition unit 3032, configured to acquire a coordinate interval of the first target area in the planar image;
[0151] A circumscribing unit 3033, configured to use the coordinate interval to circumscribe a second target area where the forklift tray is located in the depth image as the area where the forklift tray is located in the depth image.
[0152] In the embodiment of the present invention, by using the position of the forklift tray in the planar image to confirm the position area of the forklift tray in the depth image, the recognition process of the position of the forklift tray in the depth image is greatly simplified, and the recognized position area of the forklift tray is accurate and has a high precision.
[0153] For the specific description of the device part, reference may be made to the above method embodiment, which will not be elaborated here.
[0154] Embodiment 4
[0155] An embodiment of the present invention provides a matching module for a positioning device of a forklift tray, as Figure 9 shown, including:
[0156] An extraction unit 3051 extracts a first image feature in the planar image and a second image feature of the pallet template;
[0157] A comparison unit 3052 compares the first image feature and the second image feature one by one, and calculates the similarity between each corresponding feature;
[0158] A first calculation unit 3053 performs a weighted sum of all the calculated similarities to obtain the similarity between the forklift pallet in the planar image and the pallet template;
[0159] A confirmation unit 3054 determines that the pallet template is the target pallet template when the similarity reaches a preset similarity threshold.
[0160] In the embodiment of the present invention, the planar image of the forklift pallet captured by the camera is matched with the pre-prepared pallet template, which is beneficial to determining the type and model of the forklift pallet and facilitating subsequent positioning in the depth image.
[0161] For the specific description of the device part, reference may be made to the above method embodiment, which will not be elaborated here.
[0162] Embodiment 5
[0163] An embodiment of the present invention provides a marking module for a positioning device of a forklift pallet, as Figure 10 shown, including:
[0164] A second calculation unit 3061 calculates the matching degree between the target pallet template and the planar set;
[0165] A marking unit 3062 determines that the target pallet template matches the forklift pallet when the matching between the target pallet template and the planar set reaches a preset matching threshold, and marks the position of the target pallet template in the second target area.
[0166] In the embodiment of the present invention, the binary image of the target pallet template is matched with the planar set containing the forklift pallet, and the regional position of the target pallet template in the depth image can be directly marked, and the point cloud of the forklift pallet in the planar set can be corrected by using the target pallet template, which is beneficial to obtaining more accurate forklift pallet point cloud data.
[0167] For the specific description of the device part, reference may be made to the above method embodiment, which will not be elaborated here.
[0168] Embodiment 6
[0169] In an embodiment of the present invention, an electronic device is further provided. The electronic device may be the background server in the above embodiment, and its internal structure diagram may be as Figure 11As shown. The electronic device includes a processor, a memory, and a network interface connected via a system bus, and may further include a display screen and an input device. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with external electronic devices via a network connection. When the computer program is executed by the processor, it realizes a method for forklift pallet positioning. The electronic device may further include a display screen and an input device. Its display screen may be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device may be a touch layer covered on the display screen, or a button, a trackball or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad or mouse, etc.
[0170] On the other hand, the electronic device may not include a display screen and an input device. Those skilled in the art can understand that Figure 11 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0171] In one embodiment, an electronic device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented: obtaining a depth image acquired based on an RGB-D sensor; obtaining a planar image acquired by a camera, wherein the camera has the same shooting angle as the RGB-D sensor, and the acquired planar image has the same size as the depth image; using the planar image to identify a forklift pallet and determining the area where the forklift pallet is located in the depth image; performing planar segmentation processing on the area where the forklift pallet is located in the depth image to determine a planar set containing the forklift pallet; matching the planar image with a pre-prepared pallet template to obtain a matching target pallet template; using the determined planar set containing the forklift pallet to match with the target pallet template to mark the position of the target pallet template in the second target area; converting the position of the target pallet template in the second target area into coordinates in a three-dimensional space to obtain the position of the forklift pallet.
[0172] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining a depth image acquired based on an RGB-D sensor; obtaining a planar image acquired by a camera, wherein the camera has the same shooting perspective as the RGB-D sensor, and the size of the acquired planar image is the same as that of the depth image; using the planar image to identify a forklift tray and determining the area where the forklift tray is located in the depth image; performing planar segmentation processing on the area where the forklift tray is located in the depth image to determine a planar set containing the forklift tray; matching the planar image with a pre-prepared tray template to obtain a matching target tray template; using the determined planar set containing the forklift tray to match with the target tray template and marking the position of the target tray template in the second target area; converting the position of the target tray template in the second target area into coordinates in a three-dimensional space to obtain the position of the forklift tray. Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0173] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
[0174] Obviously, the above embodiments are merely examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. And the obvious changes or modifications derived therefrom still fall within the protection scope of the present invention.
Claims
1. A positioning method for a forklift pallet, characterized in that, Including the following steps: Obtain a depth image acquired based on an RGB-D sensor; Obtain a planar image collected by a camera, wherein the camera has the same shooting angle as the RGB-D sensor, and the size of the collected planar image is the same as that of the depth image; the RGB-D sensor and the camera are arranged side by side and the distance therebetween is infinitely close; Identify a forklift tray using the planar image, and determine the area where the forklift tray is located in the depth image; Perform planar segmentation processing on the area where the forklift tray is located in the depth image to determine a planar set containing the forklift tray; Match the planar image with a pre-prepared tray template to obtain a matched target tray template; Match the determined planar set containing the forklift tray with the target tray template, and mark the position of the target tray template in a second target area; Convert the position of the target tray template in the second target area into coordinates in a three-dimensional space to obtain the position of the forklift tray.
2. The method according to claim 1, characterized in that, The step of identifying a forklift tray using the planar image and determining the area where the forklift tray is located in the depth image includes: Perform image recognition on the planar image to determine a first target area where the forklift tray is located in the planar image; Obtain the coordinate interval of the first target area in the planar image; Use the coordinate interval to demarcate a second target area where the forklift tray is located in the depth image as the area where the forklift tray is located in the depth image.
3. The method according to claim 1, characterized in that The step of matching the planar image with a pre-prepared tray template to obtain a matched target tray template includes: Extract a first image feature in the planar image and a second image feature of the tray template; Compare the first image feature and the second image feature one by one, and calculate the similarity between each corresponding feature; Perform weighted summation on all the calculated similarities to obtain the similarity between the forklift tray in the planar image and the tray template; When the similarity reaches a preset similarity threshold, determine that this tray template is the target tray template.
4. The method according to claim 1, wherein The step of matching the determined planar set containing the forklift tray with the target tray template and marking the position of the target tray template in the second target area includes: Calculate the matching degree between the target tray template and the planar set; When the matching degree between the target tray template and the planar set reaches a preset matching threshold, determine that the target tray template matches the forklift tray, and mark the position of the target tray template in the second target area.
5. A positioning device for a forklift pallet, characterized in that, Including: A first acquisition module for obtaining a depth image acquired based on an RGB-D sensor; A second acquisition module for obtaining a planar image collected by a camera, wherein the camera has the same shooting angle as the RGB-D sensor, and the size of the collected planar image is the same as that of the depth image; the RGB-D sensor and the camera are arranged side by side and the distance therebetween is infinitely close; A determination module, configured to identify a forklift tray by using the planar image and determine the area where the forklift tray is located in the depth image; A segmentation module, configured to perform planar segmentation processing on the area where the forklift tray is located in the depth image to determine a planar set containing the forklift tray; A matching module, configured to match the planar image with a pre-prepared tray template to obtain a matching target tray template; A marking module, configured to use the determined planar set containing the forklift tray to match with the target tray template and mark the position of the target tray template in the second target area; A positioning module, configured to convert the position of the target tray template in the second target area into coordinates in a three-dimensional space to obtain the position of the forklift tray.
6. The device according to claim 5, characterized in that The determination module includes: An identification unit, configured to perform image recognition on the planar image to determine a first target area where the forklift tray is located in the planar image; An acquisition unit, configured to acquire the coordinate interval of the first target area in the planar image; An enclosure unit, configured to use the coordinate interval to enclose a second target area where the forklift tray is located in the depth image as the area where the forklift tray is located in the depth image.
7. The device according to claim 5, characterized in that, The matching module includes: An extraction unit, which extracts a first image feature in the planar image and a second image feature of the tray template; A comparison unit, which compares the first image feature and the second image feature one by one and calculates the similarity between each corresponding feature; A first calculation unit, which performs weighted summation on all the calculated similarities to obtain the similarity between the forklift tray in the planar image and the tray template; A confirmation unit, which determines that the tray template is the target tray template when the similarity reaches a preset similarity threshold.
8. The device according to claim 5, characterized in that, The marking module includes: A second calculation unit, which calculates the matching degree between the target tray template and the planar set; A marking unit, which determines that the target tray template matches the forklift tray and marks the position of the target tray template in the second target area when the matching degree between the target tray template and the planar set reaches a preset matching threshold.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the positioning method of the forklift tray as described in any one of claims 1-4 is implemented.
10. An electronic device, characterized in that, It includes: A memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the positioning method of the forklift tray as described in any one of claims 1-4.
Citation Information
Patent Citations
RGB-D-type sensor based tray identifying and positioning method
CN105976375A
Digital printing method and device
CN106778881A
Robot autonomous classification grabbing method based on YOLOv3
CN111080693A