A target object carrying method, device, equipment and medium
By acquiring the scene point cloud of the target object and identifying its category, and using a rotation and translation matrix to adjust the forklift posture information, the problem of intelligent forklifts being unable to accurately determine the forklift posture is solved, thus achieving efficient handling of the target object.
Patent Information
- Application Number
- CN202310717791.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-16
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2043-06-16
AI Technical Summary
Intelligent forklifts have difficulty accurately determining the forking direction and center point when picking up a target object, making it difficult to complete material handling tasks accurately and efficiently.
By acquiring the scene point cloud of the target object, identifying the category of the target object and obtaining its relative position information, and using a preset relative position determination algorithm and rotation and translation matrix, feature points are mapped onto the model point cloud. The optimal rotation and translation matrix is selected, and the pose information of the fork is adjusted to achieve accurate transport.
It enables quick and accurate determination of forklift position information, improving the precision and efficiency of handling tasks. It is applicable to different types of target objects and has versatility and scalability.
Smart Images

Figure CN116862981B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of logistics, and particularly relates to a target object carrying method and device, equipment and a medium. BACKGROUND
[0002] Automated logistics is a development trend of future intelligent manufacturing, and an intelligent forklift which can quickly and efficiently carry goods plays an important role in automated logistics.
[0003] In addition to the requirement of autonomous navigation, the intelligent forklift also needs to identify various attributes of the target object such as the category, position and posture of the target object, so as to carry the target object based on the attribute information of the target object and complete the automatic carrying task.
[0004] However, in the related art, the intelligent forklift usually cannot accurately determine the forking pose information such as the forking direction and the forking center point when forking the target object, which causes difficulty in accurately and efficiently completing the carrying task. How to quickly and accurately determine the forking pose information when forking the target object is a technical problem to be solved at present. SUMMARY
[0005] The present application provides a target object carrying method, device, equipment and medium for quickly and accurately determining the forking pose information when forking the target object.
[0006] In a first aspect, the present application provides a target object carrying method, which comprises:
[0007] obtaining a scene point cloud of a target object, obtaining feature points of the target object contained in the scene point cloud, determining relative position information between each two feature points according to a preset relative position determination algorithm, identifying a target category of the target object, and obtaining a target reference relative position information set corresponding to the target category which is pre-stored;
[0008] For each relative position information, if it is determined that there is a target reference relative position information matching the relative position information, a rotation and translation matrix corresponding to the relative position information is determined based on the relative position information and the target reference relative position information;
[0009] The feature points of the target object contained in the scene point cloud are respectively mapped into a model point cloud corresponding to the target category which is pre-stored based on each rotation and translation matrix, a feature point fusion ratio corresponding to each rotation and translation matrix is determined, and an optimal rotation and translation matrix is selected based on the feature point fusion ratio corresponding to each rotation and translation matrix;
[0010] According to the preferred rotation translation matrix, preset reference fork taking pose information corresponding to the target category is adjusted, and the target object is carried based on the adjusted preset reference fork taking pose information.
[0011] In a second aspect, the present application provides a target object carrying device, which comprises:
[0012] An obtaining module is configured to obtain a scene point cloud of a target object, obtain feature points of the target object contained in the scene point cloud, and determine relative position information between each two feature points according to a preset relative position determination algorithm.
[0013] A determining module is configured to, for each relative position information, if it is determined that there is target reference relative position information matching the relative position information, determine a rotation translation matrix corresponding to the relative position information based on the relative position information and the target reference relative position information.
[0014] A selecting module is configured to respectively map the feature points of the target object contained in the scene point cloud to a model point cloud corresponding to the target category which is pre-stored, based on each rotation translation matrix, determine a feature point fusion ratio corresponding to each rotation translation matrix, and select a preferred rotation translation matrix based on the feature point fusion ratio corresponding to each rotation translation matrix.
[0015] A carrying module is configured to adjust preset reference fork taking pose information corresponding to the target category which is pre-stored, according to the preferred rotation translation matrix, and carry the target object based on the adjusted preset reference fork taking pose information.
[0016] In a third aspect, the present application provides an electronic device, which comprises at least a processor and a memory, and the processor is configured to implement the steps of any of the above target object carrying methods when executing a computer program stored in the memory.
[0017] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program is configured to implement the steps of any of the above target object carrying methods when executed by a processor.
[0018] The application can obtain a scene point cloud of a target object, obtain feature points of the target object contained in the scene point cloud, determine relative position information between each two feature points according to a preset relative position determination algorithm, identify a target category of the target object, and obtain a target reference relative position information set corresponding to the target category which is pre-stored; for each relative position information, if it is determined that there is target reference relative position information matching the relative position information, a rotation and translation matrix corresponding to the relative position information is determined based on the relative position information and the target reference relative position information; the feature points of the target object contained in the scene point cloud are mapped into a model point cloud corresponding to the target category which is pre-stored based on each rotation and translation matrix respectively, and a feature point fusion ratio corresponding to each rotation and translation matrix is determined; an optimal rotation and translation matrix is selected based on the feature point fusion ratio corresponding to each rotation and translation matrix; and then the preset reference fork taking pose information corresponding to the target category which is pre-stored is adjusted according to the optimal rotation and translation matrix, and the target object is carried based on the adjusted preset reference fork taking pose information. Since the application can select an optimal rotation and translation matrix based on the feature point fusion ratio corresponding to each rotation and translation matrix, the optimal rotation and translation matrix can be considered as a rotation and translation matrix most suitable for the current pose of the target object, the preset reference fork taking pose information which is pre-stored is adjusted based on the optimal rotation and translation matrix, and the target object is carried based on the adjusted preset reference fork taking pose information, so that the fork taking pose information when the target object is forked can be determined quickly and accurately, and the carrying task can be completed accurately and efficiently to the greatest extent. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the application or the implementation manners in the related art, the drawings needed to be used in the embodiments or related art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can also be obtained by those skilled in the art according to these drawings.
[0020] Figure 1 A first target object carrying process schematic diagram provided by some embodiments is shown;
[0021] Figure 2 A schematic diagram of determining the relative position relationship between each two feature points provided by some embodiments is shown;
[0022] Figure 3 A second target object carrying process schematic diagram provided by some embodiments is shown;
[0023] Figure 4 A third target object carrying process schematic diagram provided by some embodiments is shown;
[0024] Figure 5Fig. 6 shows a schematic diagram of a fourth target object carrying process according to some embodiments;
[0025] Figure 6 Fig. 7 shows a schematic diagram of a fifth target object carrying process according to some embodiments;
[0026] Figure 7 Fig. 8 shows a schematic diagram of a sixth target object carrying process according to some embodiments;
[0027] Figure 8 Fig. 9 shows a schematic diagram of a target object carrying device according to some embodiments;
[0028] Figure 9 Fig. 10 shows a schematic diagram of an electronic device according to some embodiments. DETAILED DESCRIPTION
[0029] For the purpose of clarity, technical and scientific terms used in the present application may be similarly used in the art. Unless specifically stated otherwise, as apparent from the following discussions, it is appreciated that throughout the present application, technical and scientific terms may be used in accordance with convention and common usage of the terms by one of ordinary skill in the art to which the present application pertains.
[0030] It is to be noted that the explanations of the terms in the present application are only for the convenience of understanding the following described embodiments, and are not intended to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and common meanings.
[0031] The terms "first", "second", "third", and the like in the description and in the claims of the present application and the above drawings are used for distinguishing between similar or identical objects or entities, and do not necessarily mean a specified order or sequence, unless otherwise specified. It will be understood that the terms so used are interchangeable under appropriate circumstances.
[0032] The terms "comprise", "comprising", "include", "including", and "has", "having", and any variations thereof, are intended to cover a non-exclusive inclusion, such that a product or process that comprises several components or steps does not include only those components or steps that are clearly listed, but can include other components or steps not expressly listed or inherent to such product or process.
[0033] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that can perform the functions described with regard to that element.
[0034] It should be finally pointed out that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the above embodiments, or make equivalent replacement for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0035] All the implementation manners of the embodiments of the present application comply with the relevant provisions of the national laws and regulations in data acquisition, storage, use, processing, etc.
[0036] In order to quickly and accurately determine the forking pose information when forking the target object, the present application provides a target object carrying method, device, equipment and medium.
[0037] Embodiment 1:
[0038] Figure 1 A first target object carrying process schematic diagram provided by some embodiments is shown, as shown in the figure, the process includes the following steps: Figure 1 As shown in the figure, the process includes the following steps:
[0039] S101: Obtain the scene point cloud of the target object, obtain the feature points of the target object contained in the scene point cloud, determine the relative position information between each two feature points according to a preset relative position determination algorithm; identify the target category of the target object, and obtain the pre-stored target reference relative position information set corresponding to the target category.
[0040] In a possible implementation manner, the target object carrying method provided by the embodiments of the present application can be applied to forking equipment such as intelligent forklift, and can also be applied to electronic equipment, which can be PC, mobile terminal and other equipment, or server and the like. When the target object carrying method provided by the embodiments of the present application is applied to electronic equipment, the electronic equipment can be connected with forking equipment such as forklift, and the electronic equipment can send the finally determined forking pose information and the like to the forking equipment such as forklift, and the forking equipment such as forklift can receive the forking pose information and carry the target object based on the forking pose information. For the convenience of understanding, the target object carrying method provided by the embodiments of the present application is applied to intelligent forklift (for the convenience of description, referred to as forklift) as an example.
[0041] In a possible implementation, when it is intended to carry the target object, the intelligent forklift can first acquire point cloud information of the target object in an actual scene such as a factory (for the convenience of description, the point cloud information of the target object in the actual scene is referred to as scene point cloud). The specific category of the target object is not specifically limited in the present application. For example, the target object includes but is not limited to a three-legged pallet with goods, a three-legged pallet without goods, a two-legged pallet with goods, a three-legged pallet without goods, a goods shelf, a two-wheeled trolley with goods, a two-wheeled trolley without goods, a round barrel, a square barrel, and the like.
[0042] Optionally, a laser radar module can be installed in the intelligent forklift, and the scene point cloud of the target object can be acquired based on the laser radar module. In addition, the intelligent forklift can also receive the scene point cloud of the target object sent by other laser radar acquisition devices and the like. The manner of acquiring the scene point cloud is not specifically limited in the present application, and can be flexibly set according to requirements.
[0043] In a possible implementation, after the scene point cloud of the target object is acquired, the feature points of the target object contained in the scene point cloud can be obtained. In a possible implementation, when the feature points of the target object contained in the scene point cloud are obtained, the point attributes of each point contained in the scene point cloud can be determined based on a set scene point cloud edge extraction algorithm, wherein the point attributes can include an inner point or an edge point, that is, whether each point contained in the scene point cloud is an inner point or an edge point can be determined based on the set scene point cloud edge extraction algorithm. It can be understood that the edge point is a point located at the edge of the target object, and the inner point is a point inside the target object (also referred to as a non-edge point). Optionally, when whether each point contained in the scene point cloud is an inner point or an edge point is determined based on the set scene point cloud edge extraction algorithm, the points that simultaneously satisfy a preset condition 1 (the distance between adjacent points is less than a given threshold) and a preset condition 2 (the included angle between the normal vectors of adjacent points is less than a given threshold) can be clustered together as the inner points of the target object based on a region growing method, and the points that do not satisfy the preset condition 1 or do not satisfy the preset condition 2 are regarded as the edge points of the target object. The determination of the point attributes of each point contained in the scene point cloud can adopt the prior art, and will not be described here.
[0044] In a possible implementation, to improve the accuracy of the determined fork pose information, the categories of the feature points selected for different object categories can be different. For example, for a non-planar object such as a barrel, the interior points can be selected as the feature points, and the relative position information between the interior points can be calculated subsequently. For a planar object such as a pallet or a shelf, the interior points and the edge points can be selected as the feature points, and the relative position information between the interior points and the edge points can be calculated subsequently. In addition, considering that the interior points and the edge points are selected as the feature points, and the relative position information between the interior points and the edge points is calculated, which is time-consuming, to improve the efficiency, only the edge points can be selected as the feature points, and the relative position information between the edge points can be calculated subsequently. That is, the calculation of the relative position information between the interior points and the edge points is suitable for a scenario with high accuracy requirement and low time requirement, and the calculation of the relative position information between the edge points is suitable for a scenario with low accuracy requirement and high time requirement.
[0045] In a possible implementation, the target feature point attribute corresponding to the target category of the target object can be determined according to the correspondence between the object categories and the feature point attributes that are pre-stored, and then the points in the scene point cloud whose point attributes are the target feature point attribute are determined as the feature points. The relative position information between each two feature points is calculated based on these feature points. The object categories and the feature point attributes corresponding to the object categories are not limited in the present application, and can be flexibly set according to requirements. Optionally, when the target category of the target object is determined, the image of the target object can be obtained based on the image acquisition module of the camera, the image can be input into the target object recognition model that is pre-trained, and the target category of the target object can be obtained (recognized) based on the output result of the target object recognition model, which will not be described herein again.
[0046] In a possible implementation, when the relative position information between each two feature points is determined according to the preset relative position determination algorithm, the point attribute of each two feature points used for calculating the relative position information corresponding to the target category of the target object can be determined according to the correspondence between the object categories and the point attributes of each two feature points used for calculating the relative position information that are pre-stored, and the relative position information between each two feature points satisfying the point attribute can be determined (also referred to as point pair feature) based on the preset relative position determination algorithm. For example, refer to Figure 2 , Figure 2The relative position relationship between each two feature points is determined, and the point attributes of each two feature points are taken as an example. The length of the connecting line vector d between any inner point (referred to as point 1 in the figure for convenience of description) and any edge point (referred to as point 2 in the figure for convenience of description), the included angle 1 between the normal vector n1 of point 1 and the connecting line vector d, the included angle 2 between the normal vector n2 of point 2 and the connecting line vector d, and the included angle 3 between the normal vector n1 and the normal vector n2 are calculated. That is, the relative position information between each two feature points (point pair feature) includes four parameters: the length of the connecting line vector d, the included angle 1, the included angle 2, and the included angle 3. The normal vector n1 of point 1 and the normal vector n2 of point 2 can be determined by using the prior art, for example, the normal vector n1 of point 1 can be determined based on the plane formed by point 1 and the points around point 1, which will not be described herein.
[0047] In a possible implementation, the reference relative position information set corresponding to different object categories can be pre-stored in an offline or the like scenario. When it is desired to carry the target object, the target category of the target object can be identified, and then the target reference relative position information set corresponding to the target category of the target object is obtained. The manner of identifying the target category of the target object is the same as that in the above-described embodiments, which will not be described herein. The process of determining the reference relative position information set corresponding to each object category can be as follows.
[0048] The model point cloud of all objects to be carried in the carrying environment of a factory or the like can be collected by using an RGBD camera or the like installed at the front of a forklift. Optionally, in order to improve the success rate of subsequent point cloud matching as much as possible, the placement position of the object to be carried can be ensured to be a common placement position when the model point cloud is collected, and the forklift is directly opposite the object to be carried at the position to be forked.
[0049] Optionally, for any object to be carried (for the convenience of description, the subsequent examples are taken as the above-mentioned target object), after the model point cloud of the target object is collected, the edge points and the inner points of the target object in the model point cloud can be obtained based on the region growing method, etc. Optionally, the technical personnel, etc. can use the cloudcompare third-party software, etc. to crop the points contained in the model point cloud based on the target object edge composed of the edge points, so as to obtain the required model point cloud. In a possible implementation, the size, model, etc. of the forklift can also be used to set a point cloud effective range suitable for the forklift, such as the point cloud effective range can be a range composed of point clouds within 1 meter from the center point of the point cloud, etc. which can be flexibly set according to the requirements, and the present application does not make specific limitations thereto. Optionally, the forklift can crop the point cloud located outside the point cloud effective range based on the pre-configured point cloud effective range, and retain the point cloud located within the point cloud effective range as the required model point cloud.
[0050] In a possible implementation, the image of the target object can also be obtained based on the RGBD camera, etc. The image is input into the pre-trained target object recognition model, so as to obtain the center point, the rectangular frame, the category, etc. of the target object. The training process of the target object recognition model will be introduced later, and will not be described here. In a possible implementation, the same as the calculation of the relative position information between each two feature points in the scene point cloud, when the reference relative position information set corresponding to the target category of the target object is calculated, the point attribute of each two model points used for calculating the relative position information corresponding to the target category of the target object can be obtained, and then based on the pre-set relative position determination algorithm, the relative position information between each two model points satisfying the point attribute is determined, and each obtained relative position information is taken as the reference relative position information set corresponding to the target category. Optionally, the reference relative position information set corresponding to each object category obtained can be saved in the model Point Pair Features (PPF) point pair hash table.
[0051] The training process of the target object recognition model will be briefly introduced below.
[0052] The technical personnel or the like can label the set position points (such as the center point, the upper left corner, the upper right corner, the lower left corner, the lower right corner, and the like. For convenience of description, the center point is taken as an example for subsequent illustration) of the object contained in the image, the rectangular frame in which the object is located, and the category of the object, thereby composing a sample set for training the original target object recognition model. When the target object recognition model is trained, any sample image containing an object in the sample set can be obtained, and the sample image corresponds to a sample category label, a sample rectangular frame position label, and a sample center point position label. The obtained any sample image can be input into the original target object recognition model, and through the original target object recognition model, the recognition category information of the object corresponding to the sample image and the corresponding recognition rectangular frame position information and recognition center point position information are obtained.
[0053] In a specific implementation, after the recognition category information of the input sample image and the corresponding recognition rectangular frame position information and recognition center point position information are determined, because the sample category label, the sample rectangular frame position label, and the sample center point position label of the sample image are pre-stored, whether the recognition result of the target object recognition model is accurate can be determined according to whether the sample category label is consistent with the recognition category information, whether the sample rectangular frame position label is consistent with the recognition rectangular frame position information, and whether the sample center point position label is consistent with the recognition center point position information. In a specific implementation, if the above conditions are not met, it indicates that the recognition result of the target object recognition model is inaccurate, and the parameters of the target object recognition model need to be adjusted, thereby training the target object recognition model.
[0054] In a specific implementation, when the parameters in the target object recognition model are adjusted, the gradient descent algorithm can be used to perform back propagation on the gradient of the parameters of the target object recognition model, thereby training the target object recognition model. In a possible implementation, the above operation can be performed on each sample image in the sample set, and when a preset convergence condition is met, it is determined that the training of the target object recognition model is completed.
[0055] The preset convergence condition can be that the number of sample images in the sample set that are correctly recognized by the original target object recognition model is greater than a set number, or the number of iterations for training the target object recognition model reaches a set maximum number of iterations, and the like. In a specific implementation, the preset convergence condition can be flexibly set, and is not limited herein.
[0056] In a possible implementation, when the original target object recognition model is trained, the sample images in the sample set can be divided into training sample images and test sample images, and the original target object recognition model can be trained based on the training sample images first, and then the reliability of the trained target object recognition model can be verified based on the test sample images. Details are not described herein again.
[0057] S102: For each relative position information, if it is determined that there is target reference relative position information matching the relative position information, a rotation and translation matrix corresponding to the relative position information is determined based on the relative position information and the target reference relative position information.
[0058] In a possible implementation, for each relative position information, it can be determined whether there is target reference relative position information matching the relative position information in the set of target reference relative position information. Optionally, the similarity between each reference relative position information and the relative position information can be calculated, and any reference relative position information with a similarity exceeding a set similarity threshold is determined as the target reference relative position information matching the relative position information. For example, the reference relative position information with the highest similarity exceeding the set similarity threshold can be determined as the target reference relative position information matching the relative position information. The similarity threshold can be flexibly set according to requirements, which is not limited in the present application. The similarity between two relative position information can be determined by using existing technologies, which is not described herein again.
[0059] In a possible implementation, after determining the target reference relative position information matching any relative position information, a rotation and translation matrix corresponding to the relative position information can be determined based on the relative position information and the target reference relative position information. For example, the actual position of the relative position information and the actual position of the target reference relative position information can be obtained, and if the relative position information is the same as the actual position of the target reference relative position information after being rotated and translated according to a certain rotation and translation matrix, the rotation and translation matrix can be determined as the rotation and translation matrix corresponding to the relative position information. The specific value of the rotation and translation matrix can be flexibly set according to requirements, which is not limited in the present application. For example, the rotation and translation matrix can include the translation distance along the X-axis, the Y-axis, and the Z-axis, and the rotation angle (Euler angle, such as roll angle, yaw angle, and pitch angle) around the X-axis, the Y-axis, and the Z-axis. Thus, a series of candidate rotation and translation matrices can be obtained.
[0060] S103: The feature points of the target object included in the scene point cloud are mapped into the model point cloud corresponding to the target category pre-stored based on each rotation and translation matrix, respectively, to determine the feature point fusion ratio corresponding to each rotation and translation matrix; and the preferred rotation and translation matrix is selected based on the feature point fusion ratio corresponding to each rotation and translation matrix.
[0061] In a possible implementation, a preferred rotation and translation matrix can be selected from the series of candidate rotation and translation matrices obtained in S102. Optionally, when selecting the preferred rotation and translation matrix, the feature points of the target object contained in the scene point cloud can be respectively mapped to the model point cloud corresponding to the target category of the target object pre-stored in the database based on each rotation and translation matrix. For each feature point, if the distance between any model point in the model point cloud and the feature point is less than a set distance threshold, the feature point can be considered as a fusion feature point. Then, the ratio of the number of fusion feature points to the total number of feature points can be determined as the feature point fusion ratio. The set distance threshold can be flexibly set according to requirements, which is not limited in the present application. In the present application, the rotation and translation operation of the points in the scene point cloud in the opposite direction of the rotation and translation coefficients contained in the rotation and translation matrix is referred to as inverse rotation and translation. For example, if the rotation and translation coefficients contained in the rotation and translation matrix are moving A distance in the positive direction of the X axis, the inverse rotation and translation can be moving A distance in the negative direction of the X axis, which is not described herein again.
[0062] In a possible implementation, in order to accurately determine the feature point fusion ratio, the process of determining the feature point fusion ratio corresponding to any rotation and translation matrix can be as follows:
[0063] For each feature point, the model point in the model point cloud closest to the feature point is determined. If the distance between the feature point and the model point is less than a set distance threshold, the feature point is determined as a fusion feature point.
[0064] Based on the number of fusion feature points and the total number of feature points, the feature point fusion ratio corresponding to any rotation and translation matrix is determined.
[0065] Specifically, when determining the feature point fusion ratio corresponding to a certain rotation and translation matrix, for example, after the feature points of the target object contained in the scene point cloud are respectively mapped to the model point cloud corresponding to the target category of the target object pre-stored in the database based on the rotation and translation matrix, for each feature point, the model point in the model point cloud closest to the feature point can be determined, and then the distance between the feature point and the model point is determined. If the distance is less than a set distance threshold, it can be considered that the feature point and the model point are successfully paired, and the feature point can be determined as a fusion feature point. Whether each feature point is a fusion feature point can be determined based on this method. After determining whether each feature point is a fusion feature point, the feature point fusion ratio corresponding to the rotation and translation matrix can be determined based on the number of fusion feature points finally determined and the total number of feature points. For example, the ratio of the number of fusion feature points to the total number of feature points can be determined as the feature point fusion ratio corresponding to the rotation and translation matrix.
[0066] After determining the feature point fusion ratio corresponding to each rotation translation matrix, an optimal rotation translation matrix can be selected based on the feature point fusion ratio corresponding to each rotation translation matrix. For example, the rotation translation matrix with the highest feature point fusion ratio can be determined as the optimal rotation translation matrix, or one rotation translation matrix can be randomly selected from the rotation translation matrices with a feature point fusion ratio higher than a set fusion ratio threshold as the optimal rotation translation matrix.
[0067] S104: Adjust the pre-stored preset reference forklift pose information corresponding to the target category according to the optimal rotation translation matrix, and carry the target object based on the adjusted preset reference forklift pose information.
[0068] In one possible implementation, after the optimal rotation translation matrix is determined, the pre-stored preset reference forklift pose information corresponding to the target category of the target object can be adjusted according to the optimal rotation translation matrix. For example, the preset reference forklift pose information can include forklift pose information such as a forklift center point and a forklift direction, and the preset reference forklift position information can be rotated and translated according to the optimal rotation translation matrix. After adjusting the forklift pose information such as the forklift center point and the forklift direction, the forklift can fork the target object based on the adjusted forklift pose information such as the forklift center point and the forklift direction, thereby achieving efficient and accurate carrying of the target object. The forklift pose information can be flexibly set according to requirements, and the application does not make specific limitations.
[0069] The application can obtain a scene point cloud of a target object, obtain feature points of the target object contained in the scene point cloud, determine relative position information between each two feature points according to a preset relative position determination algorithm, identify a target category of the target object, and obtain a target reference relative position information set corresponding to the target category which is pre-stored; for each relative position information, if it is determined that there is target reference relative position information matching the relative position information, a rotation and translation matrix corresponding to the relative position information is determined based on the relative position information and the target reference relative position information; the feature points of the target object contained in the scene point cloud are mapped into a model point cloud corresponding to the target category which is pre-stored based on each rotation and translation matrix respectively, and a feature point fusion ratio corresponding to each rotation and translation matrix is determined; an optimal rotation and translation matrix is selected based on the feature point fusion ratio corresponding to each rotation and translation matrix; and then the preset reference fork taking pose information corresponding to the target category which is pre-stored is adjusted according to the optimal rotation and translation matrix, and the target object is carried based on the adjusted preset reference fork taking pose information. Since the optimal rotation and translation matrix can be selected based on the feature point fusion ratio corresponding to each rotation and translation matrix, the optimal rotation and translation matrix can be considered as a rotation and translation matrix most suitable for the current pose of the target object, the preset reference fork taking pose information which is pre-stored is adjusted based on the optimal rotation and translation matrix, and the target object is carried based on the adjusted preset reference fork taking pose information, so that the fork taking pose information when the target object is forked can be determined quickly and accurately to the greatest extent, and the carrying task can be completed accurately and efficiently to the greatest extent.
[0070] In addition, the target object carrying method provided by the application has universality and is suitable for target objects of different categories, without the need to customize and develop codes for specific target objects. If it is necessary to expand the category of the object to be carried in the future, the target object recognition model only needs to be optimized and trained, and the model point cloud of the object to be carried needs to be collected, so that the carrying task can be completed quickly and conveniently.
[0071] Embodiment 2
[0072] Considering that the number of feature points contained in the scene point cloud is relatively large, it takes a long time to map each feature point contained in the scene point cloud into a corresponding model point cloud and determine the feature point fusion ratio corresponding to the rotation and translation matrix. In order to improve efficiency, in the embodiment of the application, before the feature points of the target object contained in the scene point cloud are mapped into the model point cloud corresponding to the target category which is pre-stored based on each rotation and translation matrix after the rotation and translation matrix corresponding to the relative position information is determined, the method further comprises:
[0073] input the obtained image of the target object into a pre-trained target object recognition model, and obtain a reference point at a set position of a rectangular frame in which the target object is located based on an output result of the target object recognition model;
[0074] map the reference point of the rectangular frame into a scene point cloud of the target object to obtain first position information of the reference point in the scene point cloud, and for each rotation and translation matrix, map a preset position model point of a pre-stored model point cloud corresponding to the target category into the scene point cloud based on the rotation and translation matrix to obtain second position information of the preset position model point corresponding to the rotation and translation matrix;
[0075] For each rotation and translation matrix, if the distance between the first position information and the second position information corresponding to the rotation and translation matrix is less than a set distance threshold, subsequent steps are performed based on the rotation and translation matrix.
[0076] In a possible implementation, in order to improve efficiency, before the feature points of the target object contained in the scene point cloud are mapped into the pre-stored model point cloud corresponding to the target category, the obtained image of the target object can be input into a pre-trained target object recognition model, and then based on the output result of the target object recognition model, a rectangular frame in which the target object is located, a target category, and a point at a set position (such as a center point, a top-left corner, a top-right corner, a bottom-left corner, a bottom-right corner, etc.) of the rectangular frame in which the target object is located are obtained. The point at the set position (such as the center point) of the rectangular frame is taken as the reference point. For convenience of description, the following will be described by taking the set position as the center point as an example. The reference point (such as the center point) of the rectangular frame in which the target object is located in the image can be mapped into the scene point cloud of the target object to obtain first position information of the reference point in the scene point cloud.
[0077] Optionally, the model point cloud corresponding to the target category of the target object can also be obtained according to the pre-stored correspondence between the object category and the model point cloud, and a model point at a preset position of the model point is obtained. The preset position can be a center point position, a top-left corner, a top-right corner, a bottom-left corner, a bottom-right corner, etc. corresponding to the set position of the rectangular frame, which can be flexibly set according to requirements, and the present application does not make specific limitation thereon. Optionally, for the convenience of understanding, the following will be described by taking the set position of the rectangular frame as the center point and the preset position of the model point as the center point as an example. For each rotation and translation matrix, the model center point (pre-set position model point) corresponding to the target category can be mapped into the scene point cloud based on the rotation and translation matrix, so as to obtain second position information of the model center point. For the convenience of description, the second position information is referred to as the second position information corresponding to the rotation and translation matrix.
[0078] In one possible implementation, for each rotation and translation matrix, it can be determined whether the distance between the first position information and the second position information corresponding to the rotation and translation matrix is less than a set distance threshold. If it is less, it can be preliminarily considered that the center point of the image and the center point of the model point cloud can be matched, and the rotation and translation matrix is relatively reasonable. Based on the rotation and translation matrix, the step of mapping each feature point of the target object contained in the scene point cloud to the corresponding model point cloud and determining the feature point fusion ratio corresponding to the rotation and translation matrix can be performed. This will not be elaborated here.
[0079] Optionally, if the distance between the first position information and the second position information corresponding to the rotation and translation matrix is not less than a set distance threshold, it can be considered that the center point of the image does not match the center point of the model point cloud, and the rotation and translation matrix is unreasonable. The rotation and translation matrix can be removed (deleted) without proceeding to the subsequent step of determining the feature point fusion ratio of the rotation and translation matrix, thus achieving the initial screening of the rotation and translation matrix.
[0080] Since this application can quickly perform preliminary screening of rotation and translation matrices based on the set position reference points in the image and the preset position model points in the model point cloud, it can improve efficiency.
[0081] Example 3:
[0082] To improve efficiency, based on the above embodiments, in this embodiment, after determining the rotation and translation matrix corresponding to the relative position information, and before mapping the feature points of the target object contained in the scene point cloud to the pre-saved model point cloud corresponding to the target category based on each rotation and translation matrix, the method further includes:
[0083] For each rotation and translation matrix, determine whether each rotation and translation coefficient contained in the rotation and translation matrix is not greater than the corresponding set rotation and translation coefficient threshold. If so, proceed with subsequent steps based on the rotation and translation matrix.
[0084] In addition to the manner of improving efficiency in Embodiment 2, the present application also provides a manner of improving efficiency. Specifically, considering that the target object is usually placed on the ground and cannot be separated from the ground, the rotation and translation coefficients (also referred to as diagonal element coefficients) in the rotation and translation matrix, such as the translation height, the roll angle, the yaw angle, the pitch angle, etc., cannot be too large. Therefore, for each rotation and translation matrix, before mapping the feature points of the target object contained in the scene point cloud to the corresponding model point cloud based on the rotation and translation matrix, it is also determined whether each rotation and translation coefficient contained in the rotation and translation matrix is not greater than a corresponding set rotation and translation coefficient threshold. If any rotation and translation coefficient contained in the rotation and translation matrix is greater than the corresponding set rotation and translation coefficient threshold, it is considered that the rotation and translation matrix is unreasonable, and the rotation and translation matrix can be excluded (deleted) and the subsequent step of determining the feature point fusion ratio of the rotation and translation matrix is not performed, thereby realizing preliminary screening of the rotation and translation matrix.
[0085] Optionally, if each rotation and translation coefficient contained in the rotation and translation matrix is not greater than the corresponding set rotation and translation coefficient threshold, it is preliminarily considered that the rotation and translation matrix is reasonable, and the step of mapping each feature point of the target object contained in the scene point cloud to the corresponding model point cloud based on the rotation and translation matrix and determining the feature point fusion ratio corresponding to the rotation and translation matrix based on the rotation and translation matrix can be performed again, which will not be described herein.
[0086] For the convenience of understanding, the target object carrying process provided by the present application is exemplified by a specific embodiment below. Referring to Figure 3 , Figure 3 a second target object carrying process provided by some embodiments is shown, which includes the following steps:
[0087] S301: Obtain the scene point cloud of the target object, obtain the feature points of the target object contained in the scene point cloud, determine the relative position information between each two feature points according to a preset relative position determination algorithm, identify the target category of the target object, and obtain the target reference relative position information set corresponding to the target category which is pre-stored.
[0088] S302: For each relative position information, if it is determined that there is a target reference relative position information matching the relative position information, determine the rotation and translation matrix corresponding to the relative position information based on the relative position information and the target reference relative position information.
[0089] S303: For each rotation and translation matrix, it is determined whether each rotation and translation coefficient contained in the rotation and translation matrix is not greater than a corresponding set rotation and translation coefficient threshold. If yes, S304 is performed based on the rotation and translation matrix.
[0090] S304: input the obtained image of the target object into the target object recognition model pre-trained, obtain the reference point located at the set position of the rectangular frame where the target object is located based on the output result of the target object recognition model; map the reference point of the rectangular frame to the scene point cloud of the target object to obtain the first position information of the reference point in the scene point cloud; and for each rotation translation matrix, map the pre-stored preset position model point of the model point cloud corresponding to the target category to the scene point cloud based on the rotation translation matrix to obtain the second position information of the preset position model point corresponding to the rotation translation matrix; for each rotation translation matrix, if the distance between the first position information and the second position information corresponding to the rotation translation matrix is less than the set distance threshold, S305 is performed based on the rotation translation matrix.
[0091] S305: respectively based on each rotation translation matrix, map the feature points of the target object contained in the scene point cloud to the model point cloud corresponding to the target category pre-stored, determine the feature point fusion ratio corresponding to each rotation translation matrix; based on the feature point fusion ratio corresponding to each rotation translation matrix, select the preferred rotation translation matrix.
[0092] S306: according to the preferred rotation translation matrix, adjust the pre-stored preset reference fork taking pose information corresponding to the target category, and based on the adjusted preset reference fork taking pose information, carry out the target object.
[0093] Wherein, the execution order between S303 and S304 is not limited in the present application, that is, S303 can be executed first and then S304, or S304 can be executed first and then S303. Figure 3 The execution order of S303 and S304 is not limited in the present application, that is, S303 can be executed first and then S304, or S304 can be executed first and then S303.
[0094] Embodiment 4:
[0095] In order to improve the accuracy of the determined fork taking pose information, on the basis of the above embodiments, in the present application, after determining the rotation translation matrix corresponding to the relative position information, before respectively mapping the feature points of the target object contained in the scene point cloud to the model point cloud corresponding to the target category pre-stored, the method further comprises:
[0096] For each rotation translation matrix, the scene point cloud is mapped into the model point cloud based on the rotation translation matrix, scene point clouds in the model point cloud bounding box are determined based on the model point cloud bounding box, and subsequent steps are performed based on the scene point clouds in the model point cloud bounding box.
[0097] In a possible implementation, similar to the rectangular frame in which the target object in the image is located, the point cloud in the point cloud bounding box in the model point cloud can be considered as the point cloud corresponding to the target object. In order to improve the accuracy of the determined forklift pose information, for each rotation translation matrix, in addition to mapping the feature points of the target object contained in the scene point cloud into the model point cloud corresponding to the target category that is pre-stored, the scene point cloud is also mapped into the model point cloud based on the rotation translation matrix, the scene point clouds in the point cloud bounding box are determined based on the point cloud bounding box in the model point cloud, the scene point clouds in the point cloud bounding box are considered as the point cloud corresponding to the target object, and the scene point clouds outside the point cloud bounding box can be considered as noise or point clouds that have little relationship with the target object. In order to avoid noise interference and improve the accuracy of the determined forklift pose information, optionally, the scene point clouds outside the point cloud bounding box can be removed, the feature points contained in the scene point clouds in the point cloud bounding box are mapped into the model point cloud corresponding to the target category that is pre-stored based on the scene point clouds in the point cloud bounding box, the feature point fusion ratio corresponding to each rotation translation matrix is determined, the optimal rotation translation matrix is selected based on the feature point fusion ratio corresponding to each rotation translation matrix, and the pre-stored preset reference forklift pose information corresponding to the target category is adjusted according to the optimal rotation translation matrix. The steps of carrying the target object based on the adjusted preset reference forklift pose information are not described herein again.
[0098] Since the application can determine the feature point fusion ratio based on the scene point clouds in the point cloud bounding box, the accuracy of the determined feature point fusion ratio can be improved, and the accuracy of the determined forklift pose information can be further improved.
[0099] Embodiment 5:
[0100] In order to verify the rotation translation matrix and improve the accuracy of the determined optimal rotation translation matrix, based on the above embodiments, in the embodiments of the application, the selection of the optimal rotation translation matrix based on the feature point fusion ratio corresponding to each rotation translation matrix comprises:
[0101] Each rotation translation matrix is sorted in descending order according to the feature point fusion ratio corresponding to each rotation translation matrix, and one rotation translation matrix is selected in sequence according to the sorted order, and the following operations are performed on the selected rotation translation matrix.
[0102] Based on the selected rotation translation matrix, each model point included in the model point cloud is mapped into the scene point cloud, and based on distances between each model point and points in the scene point cloud and a set distance threshold, a matching success percentage of the rotation translation matrix is determined;
[0103] If the matching success percentage is greater than a set matching success percentage threshold, it is determined that the verification passes, the selected rotation translation matrix is determined as the preferred rotation translation matrix, and the selection process is stopped; otherwise, it is determined that the verification fails, and the next rotation translation matrix is selected according to the ordered sequence.
[0104] In a possible implementation, after the feature point fusion ratio corresponding to each rotation translation matrix is obtained, each rotation translation matrix can be sorted in descending order according to the feature point fusion ratio corresponding to the rotation translation matrix, that is, the greater the feature point fusion ratio corresponding to the rotation translation matrix, the higher the sequence of the rotation translation matrix, and vice versa. Optionally, a rotation translation matrix can be selected in turn according to the descending order, and then the selected rotation translation matrix is subjected to the following verification process:
[0105] Optionally, based on the selected rotation translation matrix, each model point included in the model point cloud can be mapped into the scene point cloud, and then distances between each model point and points in the scene point cloud are determined, and based on the distances and a set distance threshold, a matching success percentage is determined, where the matching success percentage is referred to as the matching success percentage of the rotation translation matrix for ease of description. Optionally, when the matching success percentage is determined, for each model point, if a distance between any point in the scene point cloud and the model point is less than the set distance threshold, the model point is considered as a matching success model point; if distances between the model point and each point in the scene point cloud are all not less than the set distance threshold, the model point is considered as a non-matching success model point. Then, a ratio of a number of the matching success model points to a total number of the model points can be determined as the matching success percentage.
[0106] In a possible implementation, the matching success percentage can also be determined in the following manner:
[0107] For each model point, a point in the scene point cloud closest to the model point is determined, if a distance between the point and the model point is not less than a set distance threshold, the model point is determined as a non-matching success model point; otherwise, the model point is determined as a matching success model point.
[0108] Based on a number of the matching success model points and a total number of the model points, a matching success percentage of any rotation translation matrix is determined.
[0109] Optionally, in order to accurately determine the matching success percentage, for each model point, the distance between each point in the scene point cloud and the model point can be determined, the point closest to the model point is obtained, it is judged whether the distance between the point and the model point is less than the set distance threshold, if the distance between the point and the model point is less than the set distance threshold, it can be considered that the point matches the model point, the model point can be determined as a matching successful model point, otherwise, if the distance between the point and the model point is not less than the set distance threshold, it can be considered that the point does not match the model point, and the model point can be determined as an unmatching successful model point.
[0110] After determining whether each model point is a matching successful model point or an unmatching successful model point, the number of matching successful model points can be obtained, and the matching success percentage can be determined based on the number of matching successful model points and the total number of model points, for example, the ratio of the number of matching successful model points to the total number of model points can be determined as the matching success percentage.
[0111] In a possible implementation, after the matching success percentage of the selected rotation and translation matrix is determined, it is judged whether the matching success percentage is greater than a set matching success percentage threshold, if the matching success percentage is greater than the set matching success percentage threshold, it is determined that the verification passes, at this time, the process of selecting the next rotation and translation matrix can be no longer performed, the currently selected rotation and translation matrix can be determined as a preferred rotation and translation matrix, and then the pre-stored preset reference forking pose information corresponding to the target category is adjusted according to the preferred rotation and translation matrix, and the target object is carried based on the adjusted preset reference forking pose information, which will not be described herein.
[0112] In a possible implementation, if the matching success percentage of the currently selected rotation and translation matrix is not greater than the set matching success percentage threshold, it is determined that the verification fails, at this time, the next rotation and translation matrix can be selected in turn according to the sorted order, and the above verification process is performed on the next rotation and translation matrix, which will not be described herein.
[0113] In a possible implementation, after the rotation translation matrices are sorted in descending order, a comparison can be made on whether the difference between the two rotation translation matrices adjacent in the sorting is less than a set difference threshold value. If yes, one of the rotation translation matrices can be deleted and only one of the rotation translation matrices can be reserved. That is, the non-maximum suppression method can be used to eliminate very similar rotation translation matrices. For example, assuming that the rotation translation matrices are sorted in descending order as follows: rotation translation matrix A, rotation translation matrix B, rotation translation matrix C, and rotation translation matrix D, and the difference between the rotation translation matrix B and the rotation translation matrix C is less than the set difference threshold value, it can be considered that the rotation translation matrix B and the rotation translation matrix C are similar, and the rotation translation matrix B or the rotation translation matrix C can be deleted. For example, the rotation translation matrix C at the rear end in the sorting can be deleted preferentially, and the rotation translation matrices finally reserved in the descending sorting are as follows: the rotation translation matrix A, the rotation translation matrix B, and the rotation translation matrix D. Based on the above verification process, the rotation translation matrix A, the rotation translation matrix B, and the rotation translation matrix D are verified in turn, and details are not described herein.
[0114] Since the rotation translation matrix can be verified in the application, the verified rotation translation matrix can be used as a preferred rotation translation matrix, so that the accuracy of the determined preferred rotation translation matrix can be improved, and the accuracy of the determined fork pose information can be improved, thereby ensuring that the carrying task is accurately and efficiently completed.
[0115] Embodiment 6:
[0116] In view of the fact that in an actual factory carrying scene, a shelf or the like can cause a certain occlusion to a target object, in order to exclude the interference of the occlusion on the matching success percentage, on the basis of the above embodiments, in the embodiments of the application, after the matching success percentage of the rotation translation matrix is determined, if the matching success percentage of the rotation translation matrix is not greater than a set matching success percentage threshold value, before it is determined that the verification fails, the method further includes:
[0117] For each unmatched model point, it is determined whether there is a point in a set direction of the unmatched model point in the scene point cloud. If yes, the unmatched model point is updated to a matched model point.
[0118] Based on the number of the updated matched model points, the matching success percentage of the rotation translation matrix is updated.
[0119] In a possible implementation, if it is determined that the matching success percentage of any rotation and translation matrix is not greater than the set matching success percentage, it can be further determined whether each unmatched model point is blocked before determining that the verification fails. Optionally, for each unmatched model point, it can be determined whether there is a point in the scene point cloud located in a set direction of the unmatched model point. The set direction can be flexibly set according to requirements, and the present application does not make a specific limitation thereon. For example, it can be determined whether there is a point in the scene point cloud whose coordinate values of the Y axis and the Z axis are the same as those of the unmatched model point and whose coordinate value of the X axis is greater than that of the unmatched model point (located in front of the unmatched model point). If there is, it can be considered that the unmatched model point is blocked, and the unmatched model point can be updated to a matched model point.
[0120] After it is determined whether each unmatched model point is blocked, the matching success percentage can be updated based on the number of updated matched model points. For example, the ratio of the number of updated matched model points to the total number of model points can be determined as the final updated matching success percentage. If the updated matching success percentage is greater than the set matching success percentage threshold, it can be determined that the verification passes, and the corresponding rotation and translation matrix can be determined as the preferred rotation and translation matrix. If the updated matching success percentage is still not greater than the set matching success percentage threshold, it can be determined that the verification fails, and the next rotation and translation matrix is selected in turn according to the sorted order to perform the above verification process, which will not be described herein again.
[0121] Since the present application can exclude the interference of the matching success percentage caused by the blocking problem, the accuracy of the selected preferred rotation and translation matrix can be improved.
[0122] Embodiment 7:
[0123] Considering that the importance of the inliers and the edge points in the point cloud corresponding to the target object varies with the category of the target object, in order to improve the verification efficiency, the rotation and translation matrix can be verified based on only the important points when the rotation and translation matrix is verified, and the rotation and translation matrix can not be verified based on each point. Optionally, in order to improve the verification efficiency, in the embodiments of the present application, before the model points in the model point cloud are mapped into the scene point cloud, the method further includes:
[0124] According to the pre-stored correspondence between the object category and the model point attribute, the target model point attribute corresponding to the target category is determined, wherein the model point attribute includes at least one of an inlier and an edge point; and each target model point with the target model point attribute is selected from each model point included in the model point cloud.
[0125] The mapping of each model point included in the model point cloud to the scene point cloud includes:
[0126] According to the target model points, the model points are updated, and the updated model points are mapped to the scene point cloud.
[0127] In a possible implementation, in order to determine the points that are more important for the target object, the target model point attribute corresponding to the target category of the target object can be determined according to the pre-stored correspondence between the object category and the model point attribute. The model point attribute can include at least one of an inner point and an edge point. The model point attribute corresponding to the object category is not specifically limited in the present application, and can be flexibly set according to requirements. For example, for an object with unclear edges, a small number of point clouds, and no high requirement on time consumption, the model point attribute corresponding to the object category can be set as an inner point, and the rotation and translation matrix is verified based on the inner point (for convenience of description, it can also be referred to as inner point verification. For the convenience of understanding, the inner point verification will be illustrated again later). For an object with clear edges, a large number of point clouds, and a high requirement on time consumption, the model point attribute corresponding to the object category can be set as an edge point, and the rotation and translation matrix is verified based on the edge point (for convenience of description, it can also be referred to as edge point verification. For the convenience of understanding, the edge point verification will be illustrated again later). The inner point verification has higher accuracy but longer time consumption; the edge point verification has lower accuracy but shorter time consumption.
[0128] In a possible implementation, after the target model point attribute is determined, target model points with the target model point attribute can be selected from the model points included in the model point cloud, and then the model points are updated according to the target model points, and the updated model points are mapped to the scene point cloud to perform the above verification process of the rotation and translation matrix. For the convenience of understanding, the target object carrying process based on the inner point verification will be illustrated first. Referring to Figure 4 , Figure 4 A third target object carrying process provided by some embodiments is shown in a schematic diagram, which includes the following steps:
[0129] S401: Obtain the scene point cloud of the target object, obtain the feature points of the target object included in the scene point cloud, determine the relative position information between each two feature points according to a pre-set relative position determination algorithm; identify the target category of the target object, and obtain the pre-stored target reference relative position information set corresponding to the target category.
[0130] S402: For each relative position information, if it is determined that there is target reference relative position information matching the relative position information, a rotation and translation matrix corresponding to the relative position information is determined based on the relative position information and the target reference relative position information.
[0131] S403: Based on each rotation and translation matrix, respectively, the feature points of the target object contained in the scene point cloud are mapped into the model point cloud corresponding to the target category pre-stored, and the feature point fusion ratio corresponding to each rotation and translation matrix is determined.
[0132] S404: According to the correspondence between the pre-stored object category and the model point attribute, the target model point attribute corresponding to the target category is determined, and taking the target model point attribute as an inner point, each target model point (for convenience of description, referred to as inner point model point) with the model point attribute as the inner point can be selected from each model point contained in the model point cloud. And according to the feature point fusion ratio corresponding to each rotation and translation matrix, the rotation and translation matrices are sorted in descending order; according to the order of sorting, a rotation and translation matrix is selected in turn, and the following operations are performed on the selected rotation and translation matrix:
[0133] Based on the selected rotation and translation matrix, each inner point model point contained in the model point cloud is mapped into the scene point cloud, and based on the distance between each inner point model point and the point in the scene point cloud and the set distance threshold, the matching success percentage of the rotation and translation matrix is determined; if the matching success percentage is greater than the set matching success percentage threshold, it is determined that the verification is passed, the selected rotation and translation matrix is determined as the preferred rotation and translation matrix, and the selection process is stopped; otherwise, it is determined that the verification is not passed, and the next rotation and translation matrix is selected according to the order of sorting.
[0134] S405: After the preferred rotation and translation matrix is selected, the pre-stored target category corresponding preset reference fork pose information can be adjusted according to the preferred rotation and translation matrix, and the target object is transported based on the adjusted preset reference fork pose information.
[0135] For convenience of understanding, the following will be illustrated based on the target object transportation process based on edge point verification. Referring to Figure 5 , Figure 5 shows a fourth target object transportation process provided by some embodiments, which includes the following steps:
[0136] S501: Obtain the scene point cloud of the target object, obtain the feature points of the target object contained in the scene point cloud, determine the relative position information between each two feature points according to the pre-set relative position determination algorithm, identify the target category of the target object, and obtain the target reference relative position information set corresponding to the target category pre-stored.
[0137] S502: For each relative position information, if it is determined that there is target reference relative position information matching the relative position information, a rotation and translation matrix corresponding to the relative position information is determined based on the relative position information and the target reference relative position information.
[0138] S503: Based on each rotation and translation matrix respectively, the feature points of the target object contained in the scene point cloud are mapped into the model point cloud corresponding to the target category pre-stored, and a feature point fusion ratio corresponding to each rotation and translation matrix is determined.
[0139] S504: According to the correspondence between the pre-stored object category and the model point attribute, the target model point attribute corresponding to the target category is determined. Taking the target model point attribute as an edge point as an example, each target model point (referred to as an edge point model point for convenience of description) whose model point attribute is an edge point can be selected from each model point contained in the model point cloud. And according to the feature point fusion ratio corresponding to each rotation and translation matrix, each rotation and translation matrix is sorted in descending order; according to the sorted order, a rotation and translation matrix is selected in turn, and the following operations are performed on the selected rotation and translation matrix:
[0140] Based on the selected rotation and translation matrix, each edge point model point contained in the model point cloud is mapped into the scene point cloud, and based on the distance between each edge point model point and the points in the scene point cloud and the set distance threshold, the matching success percentage of the rotation and translation matrix is determined. If the matching success percentage is greater than the set matching success percentage threshold, it is determined that the verification is passed, the selected rotation and translation matrix is determined as the optimal rotation and translation matrix, and the selection process is stopped; otherwise, it is determined that the verification is not passed, and the next rotation and translation matrix is selected according to the sorted order.
[0141] S505: After the optimal rotation and translation matrix is selected, the pre-stored preset reference fork taking pose information corresponding to the target category can be adjusted according to the optimal rotation and translation matrix, and the target object is transported based on the adjusted preset reference fork taking pose information.
[0142] Embodiment 8:
[0143] Considering that the accuracy of the rotation and translation matrix determined based on the relative position information of the point pair may be low, in order to improve the accuracy of the rotation and translation matrix, on the basis of the above embodiments, in the embodiments of the present application, after the rotation and translation matrix is selected according to the sorted order, before the method further comprises:
[0144] According to the correspondence between the pre-stored object category and the point attribute used for fine-tuning, the target fine-tuning point attribute corresponding to the target category is determined.
[0145] The point attribute points used for the target refinement are selected from the scene point cloud, and the point attribute points used for the target refinement are matched with corresponding model points in the model point cloud based on a preset Iterative Closest Point (ICP) matching algorithm, and the selected rotation and translation matrix is refined.
[0146] In a possible implementation, before the rotation and translation matrix is verified, the rotation and translation matrix can also be refined, and the above verification process is performed again based on the refined rotation and translation matrix. After the verification passes, the target object can be carried based on the refined rotation and translation matrix, thereby improving the success rate of carrying. Optionally, when the rotation and translation matrix is refined, the target point attribute used for refinement corresponding to the target category can be determined according to the pre-stored correspondence between the object category and the point attribute used for refinement. The point attribute used for refinement corresponding to the object category can be flexibly set according to requirements, and the present application does not make a specific limitation thereon. For example, for a non-planar object such as a barrel, the rotation and translation matrix can be refined (which can also be referred to as ICP matching) based on only the inner points, and the target point attribute used for refinement can be the inner points. For example, for an object with a planar main body, the rotation and translation matrix can be refined based on the inner points and the edge points, and the corresponding target point attribute used for refinement can be the inner points and the edge points.
[0147] Optionally, the point attribute points (inner points and / or edge points) used for target refinement can be selected from the scene point cloud, and then the point attribute points used for target refinement are matched with corresponding model points in the model point cloud based on a preset Iterative Closest Point (ICP) matching algorithm, so as to refine the selected rotation and translation matrix and obtain a more accurate refined rotation and translation matrix. The refinement of the rotation and translation matrix based on the ICP matching algorithm can use the prior art, and will not be described here.
[0148] For example, for a non-planar object such as a barrel, the point attribute points used for target refinement such as the inner points can be selected from the scene point cloud, and the inner points are matched with corresponding model inner points in the model point cloud based on a preset ICP matching algorithm, so as to obtain a more accurate refined rotation and translation matrix.
[0149] For example, if the object is a planar object, and the point attributes used for target refinement are the interior points and the edge points, the two-step approach can be used to achieve refinement: in the first step, the inverse rotation and translation can be performed on the scene point cloud based on the selected rotation and translation matrix, the edge points in the scene point cloud can be mapped to the model bounding box in the model point cloud, the edge points outside the model bounding box can be removed, the edges of the scene point cloud can be cropped, the edge points in the remaining scene point cloud edges can be matched with the points in the model bounding box (model edges) based on the ICP matching algorithm, and a preliminary refined rotation and translation matrix can be obtained. In the second step, the inverse rotation and translation can be performed on the interior points in the scene point cloud based on the preliminary refined rotation and translation matrix, the interior points in the scene point cloud can be mapped to the model point cloud, the interior points outside the model bounding box can be removed, the interior points can be cropped, the interior points in the remaining scene point cloud can be matched with the interior points in the model point cloud based on the ICP matching algorithm, and a final refined rotation and translation matrix can be obtained.
[0150] The rotation and translation matrix can be refined, and the accuracy of the determined rotation and translation matrix can be improved.
[0151] Embodiment 9:
[0152] In order to improve the accuracy of the determined rotation and translation matrix, considering that the point cloud can contain noise points irrelevant to the target object and affect the accuracy of the determined rotation and translation matrix, in the above embodiments, before obtaining the feature points of the target object in the scene point cloud, the method further includes:
[0153] The scene point cloud is converted to a set of forklift device coordinate system, the scene point cloud outside the effective range of the set of forklift device is cropped according to the pre-saved effective range of the scene point cloud corresponding to the set of forklift device, and the subsequent steps are performed based on the scene point cloud within the effective range of the scene point cloud.
[0154] In a possible implementation, in order to improve the accuracy of the determined rotation translation matrix, the noise points irrelevant to the target object in the point cloud can be pruned first. Optionally, in order to improve the pruning efficiency, the point cloud effective range suitable for the forklift (forklift device) can be pre-configured, the point cloud in the point cloud effective range is considered as the point cloud related to the target object, and the point cloud outside the point cloud effective range is considered as the noise point irrelevant to the target object. Specifically, after obtaining the scene point cloud of the target object, the scene point cloud can be converted to the set forklift device coordinate system suitable for the forklift, and then the scene point cloud outside the scene point cloud effective range corresponding to the set forklift device (forklift) is pruned according to the pre-stored scene point cloud effective range, the scene point cloud outside the scene point cloud effective range is removed, and the above steps of obtaining the feature points of the target object contained in the scene point cloud are performed based on the scene point cloud in the scene point cloud effective range. For easy understanding, refer to Figure 6 , Figure 6 A fifth target object carrying process schematic diagram provided by some embodiments is shown, which includes the following steps:
[0155] S601: Obtain the scene point cloud of the target object.
[0156] S602: Convert the scene point cloud to the set forklift device coordinate system, prune the scene point cloud outside the scene point cloud effective range corresponding to the set forklift device according to the pre-stored scene point cloud effective range, and take the scene point cloud in the scene point cloud effective range as the scene point cloud of the target object.
[0157] S603: Obtain the feature points of the target object contained in the scene point cloud, determine the relative position information between each two feature points according to the pre-set relative position determination algorithm, identify the target category of the target object, and obtain the target reference relative position information set corresponding to the target category pre-stored.
[0158] S604: For each relative position information, if it is determined that there is a target reference relative position information matching the relative position information, the rotation translation matrix corresponding to the relative position information is determined based on the relative position information and the target reference relative position information.
[0159] S605: Map the feature points of the target object contained in the scene point cloud to the model point cloud corresponding to the target category pre-stored based on each rotation translation matrix respectively, determine the feature point fusion ratio corresponding to each rotation translation matrix, and select the optimal rotation translation matrix based on the feature point fusion ratio corresponding to each rotation translation matrix.
[0160] S606: Adjust the pre-stored preset reference forking pose information corresponding to the target category according to the preferred rotation and translation matrix, and carry the target object based on the adjusted preset reference forking pose information.
[0161] Embodiment 10:
[0162] In order to improve efficiency, on the basis of the above embodiments, in the embodiments of the present application, after obtaining the scene point cloud of the target object, before obtaining the feature points of the target object contained in the scene point cloud, the method further comprises:
[0163] inputting the obtained image of the target object into a pre-trained target object recognition model, and obtaining at least one rectangular frame in which the target object is located based on the output result of the target object recognition model;
[0164] For each rectangular frame, map the reference point at the set position of the rectangular frame to the scene point cloud, and determine whether the reference point of the rectangular frame is located outside the effective range of the scene point cloud. If yes, delete the rectangular frame; otherwise, retain the rectangular frame.
[0165] If there is a retained rectangular frame, the subsequent steps are performed.
[0166] In a possible implementation, before obtaining the feature points of the target object contained in the scene point cloud, the image of the target object can also be obtained, the image is input into a pre-trained target object recognition model, and at least one rectangular frame in which the target object is located is obtained based on the output result of the target object recognition model. Wherein, the higher the accuracy of the target object recognition model, the fewer the number of rectangular frames obtained, and in an ideal case, the rectangular frame in which the target object is located can only be one. The lower the accuracy of the target object recognition model, the more the number of rectangular frames obtained, and the number of rectangular frames is not specifically limited in the present application.
[0167] In a possible implementation, for each rectangular frame, the reference point (such as the center point of the rectangular frame) at the set position of the rectangular frame can be mapped to the scene point cloud, it is determined whether the reference point of the rectangular frame is located outside the effective range of the scene point cloud, if it is located outside the effective range of the scene point cloud, the rectangular frame can be deleted, and if it is located within the effective range of the scene point cloud, the rectangular frame can be retained.
[0168] In a possible implementation, if there is a retained rectangular frame, the subsequent steps of obtaining the feature points contained in the scene point cloud and the like can be performed. If there is no retained rectangular frame, it can be considered that there is no target object meeting the carrying condition, and a set prompt information can be output.
[0169] Since the application can first determine whether there is a reserved rectangular frame, if there is, only then the subsequent step of determining the cross pose is performed, which can improve efficiency and save energy consumption.
[0170] Embodiment 11:
[0171] In order to improve the accuracy of the determined rotation translation matrix, on the basis of the above embodiments, in the embodiments of the application, if there is a reserved rectangular frame, before obtaining the feature points of the target object contained in the scene point cloud, the method further comprises:
[0172] According to a preset rectangular frame expansion ratio, the reserved rectangular frame is expanded;
[0173] When the expanded rectangular frame is mapped to the scene point cloud, the scene point cloud contained in the expanded rectangular frame is obtained, and based on the scene point cloud contained in the expanded rectangular frame and a set scene point cloud edge extraction algorithm, the edge contour of the target object is obtained, and the subsequent step is performed based on the scene point cloud contained in the edge contour.
[0174] In a possible implementation, if there is a reserved rectangular frame, considering that the rectangular frame may not contain the edge contour of the target object in its entirety, in order to obtain complete edge contour information of the target object, for each reserved rectangular frame, the rectangular frame can be expanded according to a preset rectangular frame expansion ratio to obtain an expanded rectangular frame, and the expanded rectangular frame can be mapped to the scene point cloud to obtain the scene point cloud contained in the expanded rectangular frame, and then based on the scene point cloud contained in the expanded rectangular frame and a set scene point cloud edge extraction algorithm, the edge contour of the target object is obtained. Optionally, the edge contour of the target object can be located in the annular gap between the rectangular frame and the expanded rectangular frame. The rectangular frame expansion ratio can be flexibly set according to requirements, which is not limited in the application.
[0175] For example, assuming that the minimum values of the scene point cloud contained in the rectangular frame in the X, Y and Z directions are x 0min , y 0min , z 0min , the maximum values of the scene point cloud contained in the rectangular frame in the X, Y and Z directions are x 0max , y 0max , z 0max , and the rectangular frame is represented by Region0{x 0min , x 0max , y 0min , y 0max , z 0min , z 0max}. The minimum values of the scene point cloud contained in the expanded rectangular frame in the X, Y and Z directions are x 0min , y 0min , z 0min , and the maximum values of the scene point cloud contained in the expanded rectangular frame in the X, Y and Z directions are x 0max , y 0max , z 0max , and the expanded rectangular frame is represented by Region1{x 0min , x 0min , y 0min , y 0max , z 0max , z 0max}.1min , y 1min , z 1min , the maximum values of the scene point cloud contained in the rectangular frame in the X, Y, Z directions are respectively: x 1max , y 1max , z 1max , the maximum values of the scene point cloud contained in the rectangular frame in the X, Y, Z directions are respectively: x 1min , x 1max , y 1min , y 1max , z 1min , z 1max , the minimum values of the scene point cloud contained in the edge contour line of the target object in the X, Y, Z directions are respectively: x 2min , y 2min , z2 min , the maximum values of the scene point cloud contained in the edge contour line in the X, Y, Z directions are respectively: x 2max , y 2max , z 2max , the edge contour line is represented by Region2{x 2min , x 2max , y 2min , y 2max , z 2min , z 2max}. Optionally, x 1min x 2min x 0min x 0max x 2max x 1max , y 1min y 2min y 0min y 0max y 2max y 1max , z 1min z 2min z 0min z 0max z 2max z 1max .
[0176] In one possible implementation, the above steps of obtaining feature points of the target object contained in the scene point cloud can be performed based on the scene point cloud contained in the edge contour. For ease of understanding, refer to Figure 7 , Figure 7 Fig. 6 shows a sixth target object carrying process provided by some embodiments, which includes the following steps:
[0177] S701: Obtain the scene point cloud of the target object.
[0178] S702: convert the scene point cloud to the set forking device coordinate system, according to the pre-stored effective range of the scene point cloud corresponding to the set forking device, clip the scene point cloud located outside the effective range of the scene point cloud, and take the scene point cloud located within the effective range of the scene point cloud as the scene point cloud of the target object.
[0179] S703: input the obtained image of the target object into the pre-trained target object recognition model, obtain at least one rectangular frame in which the target object is located based on the output result of the target object recognition model, map the reference point located at a set position of each rectangular frame to the scene point cloud, judge whether the reference point of the rectangular frame is located outside the effective range of the scene point cloud, if yes, delete the rectangular frame, otherwise, retain the rectangular frame, and if there is a retained rectangular frame, proceed to S704.
[0180] S704: for each retained rectangular frame, expand the rectangular frame according to a preset rectangular frame expansion ratio, obtain the scene point cloud contained in the expanded rectangular frame when the expanded rectangular frame is mapped to the scene point cloud, and obtain the edge contour of the target object based on the scene point cloud contained in the expanded rectangular frame and the set scene point cloud edge extraction algorithm.
[0181] S705: for each edge contour, obtain the feature points of the target object contained in the scene point cloud in the edge contour, determine the relative position information between each two feature points according to a preset relative position determination algorithm, identify the target category of the target object, and obtain a set of target reference relative position information corresponding to the target category pre-stored. For each relative position information, if it is determined that there is a target reference relative position information matching the relative position information, the relative position information and the target reference relative position information are determined based on the relative position information and the target reference relative position information, and each rotation and translation matrix corresponding to the edge contour can be obtained. The feature points of the target object contained in the scene point cloud in the edge contour can be mapped into the model point cloud corresponding to the target category pre-stored based on each rotation and translation matrix respectively, and the feature point fusion ratio corresponding to each rotation and translation matrix is determined.
[0182] S706: select an optimal rotation and translation matrix based on the feature point fusion ratios corresponding to the rotation and translation matrices of all edge contours.
[0183] Among them, all the rotation and translation matrices corresponding to the feature point fusion ratios can be sorted in descending order according to the feature point fusion ratios corresponding to the rotation and translation matrices. According to the order of the sorting, a rotation and translation matrix is selected in turn, and the following operations are performed on the selected rotation and translation matrix:
[0184] mapping each model point in the model point cloud to the scene point cloud based on the selected rotation translation matrix, determining a matching success percentage of the rotation translation matrix based on distances between each model point and points in the scene point cloud and a distance threshold value;
[0185] If the matching success percentage is greater than a set matching success percentage threshold value, it is determined that the verification is passed, the selected rotation translation matrix is determined as the preferred rotation translation matrix, and the selection process is stopped; otherwise, it is determined that the verification is not passed, and the next rotation translation matrix is selected according to the ordered sequence.
[0186] S707: Adjusting pre-stored preset reference fork taking pose information corresponding to the target category according to the preferred rotation translation matrix, and carrying the target object based on the adjusted preset reference fork taking pose information.
[0187] Embodiment 12:
[0188] Based on the same technical concept, the present application also provides a target object carrying device. Referring to Figure 8 , Figure 8 shows a schematic diagram of a target object carrying device provided by some embodiments, which comprises:
[0189] The obtaining module 81 is configured to obtain a scene point cloud of a target object, obtain feature points of the target object contained in the scene point cloud, determine relative position information between each two feature points according to a preset relative position determination algorithm, identify a target category of the target object, and obtain a target reference relative position information set corresponding to the target category which is pre-stored.
[0190] The determining module 82 is configured to, for each relative position information, if it is determined that there is target reference relative position information matching the relative position information, determine a rotation translation matrix corresponding to the relative position information based on the relative position information and the target reference relative position information.
[0191] The selecting module 83 is configured to respectively map the feature points of the target object contained in the scene point cloud to a model point cloud corresponding to the target category which is pre-stored based on each rotation translation matrix, determine a feature point fusion ratio corresponding to each rotation translation matrix, and select a preferred rotation translation matrix based on the feature point fusion ratio corresponding to each rotation translation matrix.
[0192] The carrying module 84 is configured to adjust preset reference fork taking pose information corresponding to the target category which is pre-stored according to the preferred rotation translation matrix, and carry the target object based on the adjusted preset reference fork taking pose information.
[0193] In a possible implementation, the selecting module 83 is further configured to:
[0194] input the obtained image of the target object into a pre-trained target object recognition model, and obtain a reference point located at a set position of a rectangular frame in which the target object is located based on an output result of the target object recognition model;
[0195] map the reference point of the rectangular frame to a scene point cloud of the target object to obtain first position information of the reference point in the scene point cloud, and for each rotation and translation matrix, map pre-stored preset position model points of a model point cloud corresponding to the target category to the scene point cloud based on the rotation and translation matrix to obtain second position information of the preset position model points corresponding to the rotation and translation matrix;
[0196] for each rotation and translation matrix, if a distance between the first position information and the second position information corresponding to the rotation and translation matrix is less than a set distance threshold, perform the step of mapping the feature points of the target object included in the scene point cloud to the pre-stored model point cloud corresponding to the target category based on the rotation and translation matrix.
[0197] In a possible implementation, the selecting module 83 is further configured to:
[0198] for each rotation and translation matrix, determine whether each rotation and translation coefficient included in the rotation and translation matrix is not greater than a corresponding set rotation and translation coefficient threshold, and if so, perform the step of mapping the feature points of the target object included in the scene point cloud to the pre-stored model point cloud corresponding to the target category based on the rotation and translation matrix.
[0199] In a possible implementation, the selecting module 83 is further configured to:
[0200] for each rotation and translation matrix, map the scene point cloud to the model point cloud based on the rotation and translation matrix, determine scene point clouds in the scene point cloud located in a point cloud bounding box in the model point cloud based on the point cloud bounding box in the model point cloud, and perform the step of mapping the feature points of the target object included in the scene point cloud to the pre-stored model point cloud corresponding to the target category based on each rotation and translation matrix based on the scene point clouds located in the point cloud bounding box.
[0201] In a possible implementation, the selecting module is specifically configured to:
[0202] for each feature point, determine a model point in the model point cloud closest to the feature point, and if a distance between the feature point and the model point is less than a set distance threshold, determine the feature point as a fusion feature point;
[0203] Determine the feature point fusion ratio corresponding to any rotation translation matrix based on the number of fused feature points and the total number of feature points.
[0204] In a possible implementation, the selecting module 83 is specifically configured to:
[0205] Sort the rotation translation matrices in descending order according to the feature point fusion ratio corresponding to each rotation translation matrix, and select a rotation translation matrix in turn according to the sorted order, and perform the following operations on the selected rotation translation matrix:
[0206] Map each model point included in the model point cloud to the scene point cloud based on the selected rotation translation matrix, and determine the matching success percentage of the rotation translation matrix based on the distance between each model point and the points in the scene point cloud and the set distance threshold.
[0207] If the matching success percentage is greater than the set matching success percentage threshold, it is determined that the verification is passed, the selected rotation translation matrix is determined as the preferred rotation translation matrix, and the selection process is stopped; otherwise, it is determined that the verification is not passed, and the next rotation translation matrix is selected according to the sorted order.
[0208] In a possible implementation, the selecting module is specifically configured to:
[0209] For each model point, determine the point in the scene point cloud closest to the model point, and if the distance between the point and the model point is not less than the set distance threshold, determine the model point as an unmatched successful model point; otherwise, determine the model point as a matched successful model point.
[0210] Determine the matching success percentage of any rotation translation matrix based on the number of matched successful model points and the total number of model points.
[0211] In a possible implementation, the selecting module 83 is specifically configured to:
[0212] For each unmatched successful model point, determine whether there is a point in the scene point cloud located in the set direction of the unmatched successful model point, and if there is, update the unmatched successful model point to a matched successful model point.
[0213] Update the matching success percentage of the rotation translation matrix based on the number of updated matched successful model points.
[0214] In a possible implementation, the selecting module 83 is further configured to:
[0215] determine a target model point attribute corresponding to the target category according to a correspondence relationship between object categories and model point attributes pre-stored, wherein the model point attribute comprises at least one of an inner point or an edge point; and select target model points with the target model point attribute from the model points in the model point cloud;
[0216] update the model points according to the target model points, and map the updated model points to the scene point cloud.
[0217] In a possible implementation, the selecting module 83 is further configured to:
[0218] determine a target model point attribute corresponding to the target category according to a correspondence relationship between object categories and model point attributes pre-stored, wherein the model point attribute comprises at least one of an inner point or an edge point; and select target model points with the target model point attribute from the model points in the model point cloud;
[0219] select points with the target model point attribute from the scene point cloud, match the points with the target model point attribute with corresponding model points in the model point cloud based on a preset iterative closest point (ICP) matching algorithm, and refine the selected rotation and translation matrix.
[0220] In a possible implementation, the obtaining module 81 is further configured to:
[0221] convert the scene point cloud to a set of forklift device coordinates, clip scene point cloud located outside a valid range of the scene point cloud corresponding to the set of forklift devices according to the valid range of the scene point cloud pre-stored, and obtain feature points of the target object included in the scene point cloud located in the valid range of the scene point cloud.
[0222] In a possible implementation, the obtaining module 81 is specifically configured to:
[0223] determine point attributes of points included in the scene point cloud based on a set of scene point cloud edge extraction algorithms, wherein the point attribute comprises an inner point or an edge point;
[0224] determine a target feature point attribute corresponding to the target category according to a correspondence relationship between object categories and feature point attributes pre-stored, and determine points with the target feature point attribute included in the scene point cloud as the feature points.
[0225] obtain respective point attributes of each two feature points for calculating relative position information corresponding to the target category, and determine relative position information between each two feature points satisfying the point attribute based on a preset relative position determination algorithm.
[0226] In a possible implementation, the obtaining module 81 is further configured to:
[0227] input the obtained image of the target object into a pre-trained target object recognition model, and obtain at least one rectangular frame in which the target object is located based on an output result of the target object recognition model;
[0228] For each rectangular frame, a reference point at a set position of the rectangular frame is mapped to the scene point cloud, and it is determined whether the reference point of the rectangular frame is located outside the effective range of the scene point cloud. If yes, the rectangular frame is deleted, otherwise, the rectangular frame is retained.
[0229] If there is a retained rectangular frame, the step of obtaining feature points of the target object included in the scene point cloud is performed.
[0230] In a possible implementation, the obtaining module 81 is further configured to:
[0231] If there is a retained rectangular frame, the retained rectangular frame is expanded according to a preset rectangular frame expansion ratio.
[0232] When the expanded rectangular frame is mapped to the scene point cloud, the edge contour of the target object is obtained based on the scene point cloud included in the expanded rectangular frame and a set scene point cloud edge extraction algorithm, and the feature points of the target object included in the scene point cloud of the edge contour are obtained.
[0233] Embodiment 13
[0234] Based on the same technical concept, the present application provides an electronic device. Figure 9 An electronic device structure schematic diagram provided by some embodiments is shown in FIG. 1, which includes a processor 91, a communication interface 92, a memory 93 and a communication bus 94, wherein the processor 91, the communication interface 92 and the memory 93 complete mutual communication through the communication bus 94. Figure 9 The memory 93 stores a computer program, and when the program is executed by the processor 91, the processor 91 executes the steps of any target object carrying method described above.
[0235] Since the principle of solving the problem of the above-mentioned electronic device is similar to that of the target object carrying method, the implementation of the above-mentioned electronic device can refer to the implementation of the method, and the repeated parts will not be described again.
[0236]
[0237] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0238] The communication interface 92 is used for communication between the above electronic device and other devices.
[0239] The memory can include a Random Access Memory (RAM) and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located away from the aforementioned processor.
[0240] The processor mentioned above can be a general-purpose processor, including a central processing unit, a network processor (NP), etc.; can also be a Digital Signal Processing (DSP), an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc.
[0241] On the basis of the above embodiments, the embodiments of the present application provide a computer readable storage medium, the computer readable storage medium stores a computer program executable by an electronic device, when the program runs on the electronic device, makes the electronic device execute the steps of the target object carrying method as described above.
[0242] The above computer readable storage medium can be any available medium or data storage device accessible by a processor in an electronic device, including but not limited to a magnetic memory such as a floppy disk, a hard disk, a magnetic tape, a magneto-optical disk (MO), etc., an optical memory such as a CD, a DVD, a BD, a HVD, etc., and a semiconductor memory such as a ROM, an EPROM, an EEPROM, a non-volatile memory (NAND FLASH), a solid state disk (SSD), etc.
[0243] Based on the same technical concept, the present application also provides a computer program product, the computer program product includes computer programs / instructions, when the computer programs / instructions are executed by a processor, the target object carrying method as described above is realized.
[0244] The methods in the present application can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented by software, the methods can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer programs or instructions are loaded and executed by a computer, the processes or functions described in the present application are performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, a core network device, an OAM, or other programmable devices.
[0245] The computer programs or instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer programs or instructions can be transmitted from one website site, computer, server, or data center to another website site, computer, server, or data center through a wired or wireless manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium, for example, a floppy disk, a hard disk, a magnetic tape; an optical medium, for example, a digital video disc; and a semiconductor medium, for example, a solid-state disk. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile storage media.
[0246] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0247] The present application is described with reference to flowcharts and / or block diagrams according to the methods, devices (systems), and computer program products of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks
[0248] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks of the block or blocks. Figure 1
[0249] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that are executed on the computer or other programmable apparatus provide steps for implementing the Figure 1 function specified in the flow or flows and / or blocks of the block or blocks. Figure 1
[0250] Obviously, numerous modifications and variations of the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of theappended claims and their equivalents, the application can be practiced otherwise than as specifically described.
Claims
1. A target object carrying method characterized by comprising: The method comprises: acquiring a scene point cloud of a target object, obtaining feature points of the target object contained in the scene point cloud, determining relative position information between each two feature points according to a preset relative position determination algorithm; identifying a target category of the target object, and obtaining a target reference relative position information set corresponding to the target category which is pre-stored; for each relative position information, if it is determined that there is target reference relative position information matching the relative position information, then based on the relative position information and the target reference relative position information, a rotation and translation matrix corresponding to the relative position information is determined; based on each rotation and translation matrix, the feature points of the target object contained in the scene point cloud are mapped into a model point cloud corresponding to the target category which is pre-stored, a feature point fusion ratio corresponding to each rotation and translation matrix is determined, and an optimal rotation and translation matrix is selected based on the feature point fusion ratio corresponding to each rotation and translation matrix; based on the optimal rotation and translation matrix, a preset reference fork taking pose information corresponding to the target category which is pre-stored is adjusted, wherein the preset reference fork taking pose information contains at least one of a fork center point and a fork direction; and based on the adjusted preset reference fork taking pose information, the target object is handled; after the determination of the rotation and translation matrix corresponding to the relative position information, before the mapping of the feature points of the target object contained in the scene point cloud into the model point cloud corresponding to the target category which is pre-stored based on each rotation and translation matrix, the method further comprises: inputting an obtained image of the target object into a target object recognition model which is pre-trained, obtaining a reference point located at a set position of a rectangular frame in which the target object is located based on an output result of the target object recognition model; mapping the reference point of the rectangular frame into the scene point cloud of the target object to obtain first position information of the reference point in the scene point cloud, and for each rotation and translation matrix, mapping a preset position model point of the model point cloud corresponding to the target category which is pre-stored into the scene point cloud based on the rotation and translation matrix to obtain second position information of the preset position model point corresponding to the rotation and translation matrix; for each rotation and translation matrix, if a distance between the first position information and the second position information corresponding to the rotation and translation matrix is less than a set distance threshold, then the subsequent steps are performed based on the rotation and translation matrix.
2. The method of claim 1, wherein, after the determination of the rotation and translation matrix corresponding to the relative position information, before the mapping of the feature points of the target object contained in the scene point cloud into the model point cloud corresponding to the target category which is pre-stored based on each rotation and translation matrix, the method further comprises: for each rotation and translation matrix, it is judged whether each rotation and translation coefficient contained in the rotation and translation matrix is not greater than a corresponding set rotation and translation coefficient threshold, if yes, then the subsequent steps are performed based on the rotation and translation matrix.
3. The method of claim 1, wherein, After the relative position information corresponding to the rotation translation matrix is determined, before the feature points of the target object contained in the scene point cloud are mapped into the model point cloud corresponding to the target category based on each rotation translation matrix respectively, the method further comprises: For each rotation translation matrix, the scene point cloud is mapped into the model point cloud based on the rotation translation matrix, the scene point cloud located in the point cloud bounding box in the model point cloud is determined based on the point cloud bounding box in the model point cloud, and subsequent steps are performed based on the scene point cloud located in the point cloud bounding box.
4. The method according to any of claims 1 to 3, characterized in that, The feature point fusion ratio corresponding to any rotation translation matrix comprises: For each feature point, the model point in the model point cloud closest to the feature point is determined, and if the distance between the feature point and the model point is less than a set distance threshold, the feature point is determined as a fusion feature point. Based on the number of fusion feature points and the total number of feature points, the feature point fusion ratio corresponding to any rotation translation matrix is determined.
5. The method of claim 4, wherein, The preferred rotation translation matrix is selected based on the feature point fusion ratio corresponding to each rotation translation matrix, which comprises: Each rotation translation matrix is sorted in descending order according to the feature point fusion ratio corresponding to each rotation translation matrix, and a rotation translation matrix is selected in turn according to the order of the sorting, and the following operations are performed on the selected rotation translation matrix: Based on the selected rotation translation matrix, each model point contained in the model point cloud is mapped into the scene point cloud, and the matching success percentage of the rotation translation matrix is determined based on the distance between each model point and the points in the scene point cloud and the set distance threshold. If the matching success percentage is greater than a set matching success percentage threshold, it is determined that the verification is passed, the selected rotation translation matrix is determined as the preferred rotation translation matrix, and the selection process is stopped; otherwise, it is determined that the verification is not passed, and the next rotation translation matrix is selected according to the order of the sorting.
6. The method of claim 5, wherein, The matching success percentage of any rotation translation matrix comprises: For each model point, the point in the scene point cloud closest to the model point is determined, and if the distance between the point and the model point is not less than a set distance threshold, the model point is determined as an unmatching successful model point; otherwise, the model point is determined as a matching successful model point. Based on the number of matching successful model points and the total number of model points, the matching success percentage of any rotation translation matrix is determined.
7. The method of claim 6, wherein, After the matching success percentage of the rotation translation matrix is determined, if the matching success percentage of the rotation translation matrix is not greater than a set matching success percentage threshold, before it is determined that the verification is not passed, the method further comprises: For each unmatching successful model point, it is judged whether there is a point in the scene point cloud located in the set direction of the unmatching successful model point, and if there is, the unmatching successful model point is updated as a matching successful model point. The matching success percentage of the rotation translation matrix is updated based on the number of updated matching successful model points.
8. The method of claim 5, wherein, Before the model point in the model point cloud is mapped into the scene point cloud, the method further comprises: According to the correspondence relationship between the object category and the model point attribute pre-stored, a target model point attribute corresponding to the target category is determined, wherein the model point attribute comprises at least one of an inner point and an edge point; and each target model point with the target model point attribute is selected from each model point included in the model point cloud; The mapping of each model point included in the model point cloud to the scene point cloud comprises: According to the target model point, each model point is updated, and the updated each model point is mapped to the scene point cloud.
9. The method according to any one of claims 5-8, characterized in that, After the rotation and translation matrix is selected in the order, before the mapping of each model point included in the model point cloud to the scene point cloud based on the selected rotation and translation matrix, the method further comprises: According to the correspondence relationship between the object category and the point attribute used for refining pre-stored, a target point attribute used for refining corresponding to the target category is determined; The point with the target point attribute used for refining is selected from the scene point cloud, the point with the target point attribute used for refining is matched with the corresponding model point in the model point cloud based on the preset iterative closest point (ICP) matching algorithm, and the selected rotation and translation matrix is refined.
10. The method of claim 1, wherein, After the scene point cloud of the target object is obtained, before the feature point of the target object included in the scene point cloud is obtained, the method further comprises: The scene point cloud is converted to a set forklift device coordinate system, scene point clouds located outside the effective range of the scene point cloud are cropped according to the effective range of the scene point cloud corresponding to the set forklift device pre-stored, and subsequent steps are performed based on the scene point clouds located within the effective range of the scene point cloud.
11. The method according to any of claims 1-3, 5-8, 10, characterized by, The feature point of the target object included in the scene point cloud comprises: Based on the set scene point cloud edge extraction algorithm, the point attribute of each point included in the scene point cloud is determined, and the point attribute comprises an inner point or an edge point; According to the correspondence relationship between the object category and the feature point attribute pre-stored, a target feature point attribute corresponding to the target category is determined, and the point with the target feature point attribute included in the scene point cloud is determined as the feature point; According to the preset relative position determination algorithm, the relative position information between each two feature points is determined, which comprises: The point attribute of each two feature points for calculating the relative position information corresponding to the target category is obtained, and the relative position information between each two feature points satisfying the point attribute is determined based on the preset relative position determination algorithm.
12. The method of claim 10, wherein, After the scene point cloud of the target object is obtained, before the feature point of the target object included in the scene point cloud is obtained, the method further comprises: The obtained image of the target object is input into a pre-trained target object recognition model, and at least one rectangular box in which the target object is located is obtained based on the output result of the target object recognition model; For each rectangular box, a reference point located at a set position of the rectangular box is mapped to the scene point cloud, it is judged whether the reference point of the rectangular box is located outside the effective range of the scene point cloud, if yes, the rectangular box is deleted, otherwise, the rectangular box is retained; If there is a remaining rectangular frame, a subsequent step is performed.
13. The method of claim 12, wherein, If there is a remaining rectangular frame, before obtaining the feature points of the target object contained in the scene point cloud, the method further comprises: For the remaining rectangular frame, the rectangular frame is expanded according to a preset rectangular frame expansion ratio. When the expanded rectangular frame is mapped to the scene point cloud, the scene point cloud contained in the expanded rectangular frame is obtained, the edge contour of the target object is obtained based on the scene point cloud contained in the expanded rectangular frame and a set scene point cloud edge extraction algorithm, and the subsequent step is performed based on the scene point cloud contained in the edge contour.
14. A target carrying device, characterized by comprising: The device comprises: An obtaining module is configured to obtain a scene point cloud of a target object, obtain feature points of the target object contained in the scene point cloud, and determine relative position information between each two feature points according to a preset relative position determination algorithm; identify a target category of the target object, and obtain a set of target reference relative position information corresponding to the target category which is pre-stored; A determining module is configured to, for each relative position information, if it is determined that there is target reference relative position information matching the relative position information, determine a rotation and translation matrix corresponding to the relative position information based on the relative position information and the target reference relative position information; A selecting module is configured to respectively map the feature points of the target object contained in the scene point cloud to a model point cloud corresponding to the target category which is pre-stored based on each rotation and translation matrix, determine a feature point fusion ratio corresponding to each rotation and translation matrix, and select an optimal rotation and translation matrix based on the feature point fusion ratio corresponding to each rotation and translation matrix; A carrying module is configured to adjust preset reference fork taking pose information corresponding to the target category which is pre-stored according to the optimal rotation and translation matrix, wherein the preset reference fork taking pose information comprises at least one of a fork taking center point and a fork taking direction, and carry the target object based on the adjusted preset reference fork taking pose information. The selecting module is further configured to input an obtained image of the target object into a target object recognition model which is pre-trained, obtain a reference point located at a set position of a rectangular frame in which the target object is located based on an output result of the target object recognition model, map the reference point of the rectangular frame to the scene point cloud of the target object to obtain first position information of the reference point in the scene point cloud, and for each rotation and translation matrix, map a preset position model point of the model point cloud corresponding to the target category which is pre-stored to the scene point cloud based on the rotation and translation matrix to obtain second position information of the preset position model point corresponding to the rotation and translation matrix; for each rotation and translation matrix, if a distance between the first position information and the second position information corresponding to the rotation and translation matrix is less than a set distance threshold, the step of mapping the feature points of the target object contained in the scene point cloud to the model point cloud corresponding to the target category which is pre-stored is performed based on the rotation and translation matrix.
15. An electronic device, comprising: The electronic device comprises at least a processor and a memory, the processor being configured to implement the steps of the object handling method according to any one of claims 1 to 13 when executing a computer program stored in the memory.
16. A computer-readable storage medium, characterized in that, The computer program is stored in the memory and is configured to implement the steps of the object handling method according to any one of claims 1 to 13 when executed by the processor.
Citation Information
Patent Citations
Object grabbing method and device, storage medium and electronic device
CN114241286A