Key point detection method, device and equipment
By generating dense point clouds and combining image key point detection models, the problem of insufficient detection accuracy in the prior art is solved, and high-precision and cost-effective operation object detection is achieved.
Patent Information
- Application Number
- CN202510647913.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-29
AI Technical Summary
The prior art has insufficient detection accuracy in the detection of work objects, especially in complex environments, which are prone to missed or missed detection, and high-precision detection equipment is expensive.
By fusing multi-frame point clouds to generate dense point clouds, and combining key point pixel coordinates in the image to determine point cloud coordinates, lidar and cameras are used for detection, and the detection accuracy is improved using a pre-trained key point detection model.
It realizes accurate positioning of key points of the work object, improves detection accuracy and reduces equipment costs, and adapts to the detection needs in complex environments.
Smart Images

Figure CN120563609A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of detection technology, and in particular to a key point detection method, device and equipment. Background Art
[0002] In engineering operations, it's often necessary to inspect the work object to facilitate intelligent operations. For example, in logistics operations, vehicles or containers awaiting loading or unloading can be inspected to facilitate loading and unloading. The accuracy of work object inspection can impact the effectiveness of subsequent intelligent operations, leading to a need for improved accuracy. Summary of the Invention
[0003] A brief overview of the present disclosure is provided below to provide a basic understanding of some aspects of the present disclosure. However, it should be understood that this overview is not an exhaustive overview of the present disclosure. It is not intended to identify key or important parts of the present disclosure, nor is it intended to limit the scope of the present disclosure. Its purpose is simply to present certain concepts of the present disclosure in a simplified form as a prelude to the more detailed description that will be given later.
[0004] One of the objectives of the present disclosure is to provide a key point detection method, apparatus and device.
[0005] According to a first aspect of the present disclosure, a key point detection method is provided, comprising:
[0006] Acquire, by a first acquisition device, n frames of point clouds of the work object at a current moment, and generate a dense point cloud of the work object at the current moment based on the n frames of point clouds, where n is a positive integer greater than 1;
[0007] Acquire an image of the work object through a second acquisition device, and determine pixel coordinates of key points of the work object based on the image and a pre-trained key point detection model; and
[0008] The point cloud coordinates of the key point in the dense point cloud are determined based on the pixel coordinates of the key point.
[0009] In some embodiments, generating a dense point cloud of the operation object at the current moment based on the point clouds of the previous n frames includes:
[0010] Based on the i-th frame point cloud in the first n frames of point clouds, each point in the other frame point clouds is converted into the i-th frame point cloud to generate the dense point cloud, where i is any value from 1 to n.
[0011] In some embodiments, based on the i-th frame point cloud in the first n frames of point clouds, converting each point in other frame point clouds into the i-th frame point cloud to generate the dense point cloud includes:
[0012] For each frame of point cloud in the other frame point clouds, the pose of each point in the current coordinate system is converted into the pose in the reference coordinate system on which the i-th frame point cloud is based.
[0013] In some embodiments, the key point detection method further includes:
[0014] For the k-th frame point cloud in the first n frames of point cloud, correction is performed based on the motion parameters of the k-1-th frame point cloud, and the corrected point cloud is used as the k-th frame point cloud, where k is any value from 1 to n.
[0015] In some embodiments, determining the point cloud coordinates of the key point in the dense point cloud based on the pixel coordinates of the key point includes:
[0016] Determine the coordinate transformation relationship between point cloud coordinates and pixel coordinates; and
[0017] The point cloud coordinates of the key point are obtained according to the coordinate transformation relationship and the pixel coordinates of the key point.
[0018] In some embodiments, the first acquisition device includes a lidar; and / or the second acquisition device includes a camera.
[0019] In some embodiments, the work object includes a polygonal prism structure, and the key point is a vertex of the polygonal prism structure.
[0020] In some embodiments, the key point detection method further includes:
[0021] The work object is loaded or unloaded according to the point cloud coordinates of the key points.
[0022] According to a second aspect of the present disclosure, there is provided a key point detection device, comprising:
[0023] The first acquisition device is configured to obtain the first n frames of point cloud of the work object at the current moment, where n is a positive integer greater than 1;
[0024] a point cloud generating module configured to generate a dense point cloud of the work object at the current moment based on the point clouds of the previous n frames;
[0025] a second acquisition device, configured to acquire an image of the work object;
[0026] A first determination module is configured to determine the pixel coordinates of the key points of the work object based on the image and a pre-trained key point detection model;
[0027] The second determining module is configured to determine the point cloud coordinates of the key point in the dense point cloud based on the pixel coordinates of the key point.
[0028] According to a third aspect of the present disclosure, a key point detection device is provided, comprising:
[0029] processor; and
[0030] A memory, wherein instructions are stored in the memory, and when the instructions are executed by the processor, the steps of the key point detection method described above are implemented.
[0031] According to a fifth aspect of the present disclosure, a loader is provided, wherein the loader includes the key point detection device as described above or the key point detection equipment as described above.
[0032] According to a fifth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein instructions are stored on the non-transitory computer-readable storage medium, and when the instructions are executed by a processor, the steps of the key point detection method as described above are implemented.
[0033] According to a sixth aspect of the present disclosure, a computer program product is provided, wherein the computer program product includes instructions, and when the instructions are executed by a processor, the steps of the key point detection method described above are implemented.
[0034] Other features and advantages of the present disclosure will become more apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0036] The present disclosure can be more clearly understood from the following detailed description with reference to the accompanying drawings, in which:
[0037] Figure 1 A schematic diagram showing a flow chart of a key point detection method according to an exemplary embodiment of the present disclosure is shown;
[0038] Figure 2 A schematic diagram of a process for generating a dense point cloud according to an exemplary embodiment of the present disclosure is shown;
[0039] Figure 3 A schematic diagram showing a loader and a work object according to an exemplary embodiment of the present disclosure is shown;
[0040] Figure 4 A schematic diagram illustrating key points of an object for marking according to an exemplary embodiment of the present disclosure is shown;
[0041] Figure 5 A schematic diagram of an object coordinate system according to an exemplary embodiment of the present disclosure is shown;
[0042] Figure 6 A schematic structural diagram of a key point detection device according to an exemplary embodiment of the present disclosure is shown;
[0043] Figure 7 A schematic structural diagram of a key point detection device according to an exemplary embodiment of the present disclosure is shown;
[0044] Figure 8 A schematic diagram of the structure of a computer system that can implement the embodiments of the present disclosure is shown.
[0045] Note that in the embodiments described below, the same reference numerals are sometimes used in common across different drawings to denote the same parts or parts having the same functions, and their repeated descriptions are omitted. In this specification, similar reference numerals and letters are used to denote similar items. Therefore, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.
[0046] To facilitate understanding, the positions, sizes, and ranges of various structures shown in the drawings and the like may not always represent actual positions, sizes, and ranges. Therefore, the disclosed invention is not limited to the positions, sizes, and ranges disclosed in the drawings and the like. Furthermore, the drawings are not necessarily drawn to scale, and some features may be exaggerated to illustrate details of specific components. DETAILED DESCRIPTION
[0047] Various exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure.
[0048] The following description of at least one exemplary embodiment is merely illustrative and is not intended to limit the present disclosure, its application, or use. In other words, the structures and methods herein are presented in an exemplary manner to illustrate various embodiments of the structures and methods of the present disclosure. However, those skilled in the art will appreciate that these are merely exemplary of the disclosure that may be implemented, and are not exhaustive. Furthermore, the drawings are not necessarily drawn to scale, and some features may be exaggerated to illustrate details of specific components.
[0049] Technologies, methods and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods and equipment should be considered part of the authorization specification.
[0050] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0051] With the rapid development of intelligent technology, more and more engineering operations can be performed unmanned or almost unmanned. Furthermore, the advancement of intelligentization places higher demands on the detection accuracy of operational objects. For example, in intelligent port logistics operations, automated yard cranes and unmanned container trucks can achieve unmanned container transportation. In material loading and unloading, unmanned loaders can perform intelligent loading and unloading by inspecting train carriages and containers.
[0052] Among them, the work object can be detected by a camera or a laser radar. When the work object is far away, the scanning accuracy of the laser radar may be low. In some examples, a point cloud about the work object can be obtained by laser radar scanning, and then the work object is detected based on the point cloud. The point cloud obtained by the laser radar may lose global posture information and may lose part of the information of the work object. In addition, when there are outliers in the point cloud, it may cause erroneous detection results. Alternatively, in some examples, a binocular camera can be used to obtain image and point cloud data to detect the work object. However, the accuracy of the point cloud constructed based on the binocular camera is poor, and it is difficult to ensure the quality of the image and point cloud under complex lighting conditions. Alternatively, in some examples, the heading angle and detection frame of the work object can be directly detected by the target detection model. However, the accuracy of the target detection model directly affects the detection accuracy of the carriage. In complex environments such as dust and strong light, missed detection or false detection may occur, and the detection accuracy is not high.
[0053] In order to solve the above problems, the present disclosure proposes a key point detection method and device, which generates a dense point cloud by fusing multiple frames of point clouds, and determines the point cloud coordinates of the key points in the dense point cloud according to the pixel coordinates of the key points in the image, thereby improving the key point detection accuracy of the work object, helping to achieve accurate positioning of the work object, and improving the work efficiency of subsequent engineering operations.
[0054] According to one aspect of the present disclosure, a key point detection method is provided. Figure 1 As shown, in an exemplary embodiment of the present disclosure, a key point detection method may include:
[0055] Step S110 , obtaining the first n frames of point cloud of the working object at the current moment through the first acquisition device, and generating a dense point cloud of the working object at the current moment based on the first n frames of point cloud.
[0056] Where n is a positive integer greater than 1. A point cloud is a collection of multiple discrete points, each of which may have corresponding coordinate information, i.e., point cloud coordinates. In some embodiments, the first acquisition device may include a laser radar, which scans the work object via the laser radar to obtain point cloud data about the work object. The laser radar may be a 16-beam, 32-beam, or other beam laser radar. In some cases, the acquired point cloud of the work object may be relatively sparse and may not cover the key points of the work object, resulting in the loss of key point features of the work object. Generating a dense point cloud by using the relatively sparse point clouds of the first n frames at the current moment can achieve a more comprehensive scan of the key points of the work object, reduce the possibility of losing key point features of the work object, and improve key point detection accuracy. Furthermore, this solution can obtain the point clouds of the first n frames by scanning with a lower beam laser radar (such as a 16-beam or 32-beam laser radar), and generate a dense point cloud based on the point clouds of the first n frames, without using a high-beam laser radar, which can effectively reduce costs.
[0057] In some embodiments, generating a dense point cloud of the work object at the current moment based on the first n frames of point clouds may include converting each point in the point clouds of the other frames into the point cloud of the first frame of the first n frames of point clouds to generate the dense point cloud. Where i can be any value from 1 to n. For example, based on the point cloud of the first frame of the first n frames of point clouds, each point in the point clouds of the second frame through the nth frame of the first n frames of point clouds may be converted into the point cloud of the first frame of the first n frames of point clouds.
[0058] In this way, the i-th frame point cloud can be used as a benchmark to convert the other n-1 frames of point clouds in the previous n frames of point cloud except the i-th frame point cloud into the i-th frame point cloud, realizing the fusion of n frames of point clouds, and thus constructing a dense point cloud at the current moment.
[0059] In a specific example, for each frame point cloud in other frame point clouds, the pose of each point in the current coordinate system can be converted to the pose in the reference coordinate system based on the i-th frame point cloud. Specifically, for the points in each frame point cloud, the coordinate system based on which each frame point cloud is based can be used to represent them. Considering that the coordinate system based on which each frame point cloud is based may be different, for the above-mentioned first n frames point cloud, the pose of each point in each other frame point cloud can be represented by the reference coordinate system based on which the i-th frame point cloud is based. In this way, the other frame point clouds can be converted to the i-th frame point cloud, and the pose of each point in the dense point cloud in the reference coordinate system can be determined, so as to facilitate the subsequent determination of the pose of the key points, and the pose of the key points can be input into the engineering operation device for operation. Among them, the reference coordinate system based on the i-th frame (such as the first frame) point cloud can be a global coordinate system, then the global pose information of each point in the dense point cloud can be determined, which helps the engineering operation device to work better.
[0060] Considering that the work object may move relative to the first acquisition device, resulting in poor quality of the acquired point cloud data, the acquired point cloud can be corrected to improve its accuracy, especially in motion scenes. In some embodiments, the k-th frame of the point cloud among the first n frames of the point cloud can be corrected based on the motion parameters of the k-1-th frame of the point cloud, and the corrected point cloud can be used as the k-th frame of the point cloud.
[0061] Here, k can be any value between 1 and n. In a specific example, each of the first n frames of the point cloud can be corrected to improve the acquisition accuracy of the point cloud, thereby improving the accuracy of subsequent key point detection. Alternatively, correction can be performed on only some of the first n frames of the point cloud, without limitation. If the scanning interval of the first acquisition device is short and the angular velocity and linear velocity of the workpiece relative to the first acquisition device are constant, the motion parameters of the previous frame can be used to predict the motion of the current frame to correct the point cloud of the current frame.
[0062] The following describes some embodiments of correcting the point cloud of the current frame based on the motion parameters of the point cloud of the previous frame.
[0063] In a specific example, the pose of each point in the current frame can be interpolated according to the motion parameters of the previous frame according to the timestamp to obtain the dedistorted point cloud of the current frame.
[0064] For example, the dedistorted point cloud of the current frame (frame k) It can be expressed as That is, a collection The elements in in, The original point cloud P belonging to the current frame k Points in, for example It can be the point where the mth vertical laser beam and the nth horizontal laser beam intersect in the kth frame point cloud; T k Indicates the pose of the current frame, T k-1 Represents the pose of the previous frame (i.e., the k-1 frame); represents the motion parameters of the previous frame; exp represents the exponential function; N is the total number of scans in a single frame, and δt can represent the time interval between scans. The motion parameters can be used to describe the motion of an object with six degrees of freedom (three translational degrees of freedom and three rotational degrees of freedom) in three-dimensional space, for example, and can be represented by Lie algebra, denoted as The pose T of the current frame k and the pose T of the previous frame k-1 It can be represented by a 4×4 transformation matrix. In this way, a dedistorted point cloud can be obtained based on the uniform velocity model, thereby effectively and quickly reducing or even eliminating the point cloud distortion caused by motion.
[0065] Considering that the point cloud may be distorted due to motion, the pose of the corresponding current frame point cloud may have errors. Therefore, after obtaining the dedistorted point cloud, the optimized pose of the current frame can be determined based on the dedistorted point cloud of the current frame.
[0066] In a specific example, for each frame of the first n frames of point cloud, a corresponding dedistorted point cloud can be obtained, and feature extraction is performed based on the dedistorted point cloud of each frame to obtain the edge features and plane features of the dedistorted point cloud of each frame, and then the dedistorted point clouds of each frame except the dedistorted point cloud of the i-th frame are feature matched with the dedistorted point cloud of the i-th frame, and the optimized pose of the current frame is determined.
[0067] The feature extraction of the point cloud may be based on the following process, for example.
[0068] Taking the acquisition of point cloud data by LiDAR as an example, for each scanning point of each frame of point cloud, for example, for the point where the mth vertical laser beam and the nth horizontal laser beam intersect in the kth frame of point cloud The local smoothness of the point can be calculated by the following formula
[0069] in In the horizontal direction Adjacent points The current point With adjacent points The average Euclidean distance can be used to measure the smoothness of the local surface. The edge feature ε can be determined k and plane feature Sk , for example, the point with the highest local smoothness can be selected as the edge feature ε k , the point with the lowest local smoothness can be selected as the plane feature S k .
[0070] Determining the optimized pose of the current frame may be based on the following process, for example.
[0071] The features of the current frame point cloud are matched with the edges or planes stored based on the KD-Tree for global map matching. The matching process can be expressed as follows:
[0072]
[0073] Among them, “·” can represent dot product operation; “×” can represent cross product operation, p ε It can be expressed based on the above edge feature ε k Determine the edge feature vector; p S It can be expressed based on the above plane feature S k Determine the plane eigenvectors; It can represent the global straight line geometric center, p n can represent a unit vector, in some examples p n Can be omitted; It can represent the global plane geometric center; Can represent the direction vector of a straight line; When feature matching is performed on each other frame point cloud and the dedistorted point cloud of the i-th frame, the global line geometric center, the global plane geometric center, the line direction vector, and the plane normal vector can be determined based on the dedistorted point cloud of the i-th frame.
[0074] By constructing the objective function ∑W(p ε )f ε (p ε )+∑W(p S )f S (p S ), to determine the optimal pose Specifically, the objective function ∑W(p ε )f ε (p ε )+∑W(p S )f S (p S )The smallest pose T k As the optimized pose Among them, W(p ε ) and W(p S ) represent edge weight and plane weight respectively. By iteratively optimizing the above objective function, the optimized pose can be obtained. In a specific example, the optimized pose can be obtained by adding left perturbation Among them, the disturbance increment δ ξ It can be represented by Lie algebra, that is, δ ξ ∈se(3), the point cloud pose T with coordinates p p (where p can be a vector of dimension 3, for example) can be represented by the rotation parameter R and the translation parameter t, T p About the perturbation increment δ ξ The Jacobian matrix That is, the Jacobian matrix In this way, we can use the Jacobian matrix J p , iteratively update the perturbation increment δ ξ , until the maximum number of iterations or the convergence condition is met, thus obtaining the optimized pose
[0075] When the optimized posture is obtained, the optimized motion parameters that can more accurately describe the motion of the current frame can be determined. and the pose T of the previous frame k-1 Determine the optimal motion parameter Δξ * .
[0076] In a specific example, the optimized motion parameter Δξ can be determined based on the following formula: * :
[0077] The optimized motion parameter Δξ * It can be used to reflect the actual motion of the current frame after optimization. In this way, the optimized motion parameter Δξ can be used to reflect the actual motion of the current frame after optimization. * To correct the point cloud of the current frame, the corrected point cloud It can be expressed as:
[0078]
[0079] By optimizing the motion parameter Δξ * , the source point cloud can be distorted with high precision and the corrected point cloud can be updated to the global map. The global map can be initialized based on the Simultaneous Localization and Mapping (SLAM) algorithm.
[0080] The above describes some embodiments of correcting the original point cloud of the current frame, realizing two-stage distortion compensation of the point cloud, wherein the first stage of distortion compensation obtains a dedistorted point cloud, and the second stage of distortion compensation corrects the original point cloud based on the optimized motion parameters obtained from the dedistorted point cloud to obtain a higher quality point cloud. Then, a corresponding dense point cloud can be generated based on the corrected point cloud to improve the detection accuracy of subsequent key points.
[0081] In a specific example, the dense point cloud generation process can be as follows Figure 2 As shown. In step S201, the first n frames of point cloud are acquired by laser radar. In step S202, feature extraction is performed on the first n frames of point cloud. In step S203, the first n frames of point cloud are corrected. In step S204, pose optimization and point cloud fusion are performed. For example, the pose of each frame of point cloud is optimized, and the first n frames of point cloud are fused. In step S205, it is determined whether the requirements are met. In a specific example, the determination of whether the requirements are met may be, for example, determining whether n frames of point cloud have been accumulated and fused, that is, determining whether the accumulated n frames of point cloud have been fused based on the i-th frame of point cloud. If the judgment is yes, step S206 is executed to generate a dense point cloud. If the judgment is no, n frames of point cloud can be reacquired based on the fused and corrected point cloud, that is, step S207 is executed to update the point cloud. After executing step S207, step S201 is executed.
[0082] like Figure 1 As shown, in an exemplary embodiment of the present disclosure, the key point detection method may further include:
[0083] Step S120: Acquire an image of the work object through a second acquisition device, and determine the pixel coordinates of the key points of the work object based on the image and a pre-trained key point detection model.
[0084] In some embodiments, the second acquisition device may include a camera, such as a monocular camera or a binocular camera. The image of the work object may be input into a pre-trained key point detection model to obtain the pixel coordinates of the key points of the work object in the image.
[0085] The key point detection model can be a neural network model, wherein the key point detection model can be trained based on the following process.
[0086] A sample set comprising a preset number of object images annotated with key points is prepared. For example, object images may be collected for a certain period of time in the working environment of the working object, and key points are annotated on each object image, such as the pixel coordinates of the key points, to generate samples.
[0087] Furthermore, in some embodiments, a portion of the object images in the sample set can be determined as a test set, and another portion of the object images can be determined as a training set. The test set and the training set can be determined manually or automatically, and the determination process can be random. Alternatively, the sample set can be used as the training set.
[0088] Furthermore, the keypoint detection model is trained using the training set, and the trained keypoint detection model is tested using the test set to obtain the accuracy of the keypoint detection model. The training process specifically includes adjusting model parameters in the keypoint detection model. By comparing the accuracy of the keypoint detection model with a preset accuracy, it can be determined whether further training is required. Specifically, when the accuracy of the trained keypoint detection model is greater than or equal to the preset accuracy, it can be considered that the accuracy of the keypoint detection model meets the requirements, and training can be terminated. The trained keypoint detection model can then be used to detect keypoints in an image. If the accuracy of the trained keypoint detection model is less than the preset accuracy, it can be considered that the keypoint detection model requires further optimization. In this case, the keypoint detection model can be further trained by increasing the number of samples in the training set, specifically by expanding the sample set and / or the training set, or by increasing the ratio of the number of samples in the training set to the total number of samples in the sample set. Alternatively, the keypoint detection model itself can be adjusted, and the adjusted keypoint detection model can be trained until a keypoint detection model that meets the requirements is obtained.
[0089] In some cases, it may be difficult for the second acquisition device to completely capture the appearance information of the work object, resulting in that the acquired image may not cover all the required key points. For example, only an image containing a part of the key points can be acquired, and accordingly, only the pixel coordinates of a part of the key points may be obtained according to the key point detection model. In some embodiments, the pixel coordinates of the first part of the key points can be determined based on the image and the pre-trained key point detection model. Then, the equation for fitting the edge line of the work object can be determined based on the pixel coordinates of the first part of the key points. Among them, the size information of the work object (such as at least one of the length, width and height) can be determined based on the pixel coordinates of the first part of the key points, and then the pixel coordinates of the second part of the key points can be determined based on the size information of the work object and the equation of the determined edge line. In a specific example, the equation for fitting the edge line can be determined by the coordinates of two key points on the edge. Determining the size information of the work object based on the pixel coordinates of the key points may include, for example: determining the object coordinates of the key points in the object coordinate system based on the pixel coordinates of the key points, and determining the size information of the work object based on the object coordinates of the key points. As Figure 5 As shown, the object coordinate system may be, for example, a coordinate system OXYZ, wherein the origin O may be located at a corner of the work object away from the work surface.
[0090] In this way, when the second acquisition device can only obtain a part of the key points, the coordinates of other key points that the second acquisition device cannot capture can be solved according to the size information of the work object and the fitted equation.
[0091] like Figure 1 As shown, in an exemplary embodiment of the present disclosure, the key point detection method may further include:
[0092] Step S130: determining the coordinates of the key points in the dense point cloud based on the pixel coordinates.
[0093] In some embodiments, the coordinate transformation relationship between the point cloud coordinates and the pixel coordinates can be determined, and then the point cloud coordinates of the key points can be obtained based on the coordinate transformation relationship and the pixel coordinates of the key points. The coordinate transformation relationship between the point cloud coordinates of the first acquisition device and the pixel coordinates of the second acquisition device can be obtained by using a perspective-n-points (PNP) algorithm based on the acquisition parameters of the second acquisition device. For example, the external parameters (such as the rotation matrix and the translation matrix) between the camera and the lidar can be obtained based on the intrinsic parameters of the camera using the PNP algorithm, wherein the external parameters can be used to indicate the coordinate transformation relationship between the point cloud coordinates and the pixel coordinates, and the internal parameters can include, for example, the focal length of the camera, the position of the principal point, the distortion coefficient, etc., and the internal parameters can be calibrated in advance.
[0094] In one specific example, based on the coordinate transformation relationship, the corresponding pixel coordinates in the image for each point cloud coordinate in the dense point cloud can be calculated to construct a corresponding lookup table. Then, the corresponding point cloud coordinates can be found in the lookup table based on the pixel coordinates of the key points in the image. This allows the point cloud coordinates of key points to be quickly determined, improving detection efficiency.
[0095] In some embodiments, the work object may include a polygonal prism structure, wherein the key point may be the vertex of the polygonal prism structure, so that the work object can be easily positioned. For example, the work object may include a quadrangular prism (cuboid) structure. Figure 3 As shown, the operation object may be, for example, at least one of a carriage 310 and a container 320 .
[0096] In some embodiments, the key point detection method may further include: loading or unloading the work object according to the point cloud coordinates of the key points. Figure 3 As shown, the loader 330 can locate the work object (such as the carriage 310 or the container 320) according to the point cloud coordinates of the key points of the work object, and load or unload the work object.
[0097] For example, the cab of a loader 330 can be equipped with a number of cameras 340 and lidars 350 to perceive the port's operating environment. The transformation matrix T1 between the coordinate systems of the cameras 340 and the lidars 350, as well as the transformation matrix T2 between the coordinate systems of the lidars 350 and the loader 330, can be pre-calibrated. Transformation matrices T1 and T2 can then be used to transform the key point detection results of port operating objects into the coordinate system of the loader 330, allowing the loader 330 to perform loading and unloading operations based on the key point detection results of the operating objects.
[0098] Figure 4 The six key points 420 of the operation object 410 are shown as an example. In other specific examples, the operation object 410 may have other numbers of key points 420. The operation object 410 may be, for example, Figure 3 The carriage 301 or container 302 shown. In some embodiments, as Figure 4 As shown, a certain number of key points 420 can be marked on the work surface of the work object 410. In some embodiments, the key points 420 can be used to determine the geometric center of the work object 410, so that the work object 410 can be more accurately positioned based on the geometric center of the work object 410. In other words, the key points 420 of the work object 410 can be points that can determine the geometric center of the work object 410.
[0099] According to another aspect of the present disclosure, a key point detection device is also provided. Figure 6 As shown, the key point detection device may include a first acquisition device 610, a point cloud generation module 620, a second acquisition device 630, a first determination module 640, and a second determination module 650. The first acquisition device 610 may be configured to obtain the first n frames of point cloud of the work object at the current moment. Wherein, n is a positive integer greater than 1. The point cloud generation module 620 may be configured to generate a dense point cloud of the work object at the current moment based on the first n frames of point cloud. The second acquisition device 630 may be configured to acquire an image of the work object. The first determination module 640 may be configured to determine the pixel coordinates of the key points of the work object based on the image and a pre-trained key point detection model. The second determination module 650 may be configured to determine the point cloud coordinates of the key points in the dense point cloud based on the pixel coordinates of the key points.
[0100] In addition, some embodiments of the key point detection device can be described with reference to some embodiments of the above-mentioned key point detection method, which will not be described in detail here.
[0101] According to another aspect of the present disclosure, a key point detection device is also provided. Figure 7As shown, the key point detection device 700 may include a processor 710 and a memory 720. The memory 720 may store instructions, which, when executed by the processor 710, implement the steps of the key point detection method described in any of the aforementioned embodiments of the present disclosure.
[0102] The processor 710 can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component to implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor can be a microprocessor or any conventional processor, etc., and can be an X86 architecture or an ARM architecture, etc.
[0103] The memory 720 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus random access memory (DRRAM). It should be noted that the memory described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0104] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, on which instructions are stored. When the instructions are executed, the steps of the key point detection method described in any of the foregoing embodiments of the present disclosure can be implemented.
[0105] Similarly, the non-transitory computer-readable storage medium in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory. It should be noted that the non-transitory computer-readable storage medium described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0106] According to another aspect of the present disclosure, a loader is further provided, wherein the loader may include the key point detection device 600 or the key point detection equipment 700 as described above.
[0107] According to another aspect of the present disclosure, a computer program product is further provided. The computer program product includes a computer program. When the computer program is executed by a processor, the steps of the key point detection method described in any of the aforementioned embodiments of the present disclosure can be implemented.
[0108] The instructions may be any set of instructions to be executed directly by one or more processors, such as machine code, or any set of instructions to be executed indirectly, such as a script. The terms "instructions," "application," "process," "steps," and "program" (including "computer program") are used interchangeably herein. The instructions may be stored in object code format for direct processing by one or more processors, or as a script or collection of independent source code modules in any other computer language, including those interpreted on demand or compiled in advance. The instructions may include instructions that cause one or more processors to act as the various neural networks herein. The functions, methods, and routines of the instructions are explained in more detail elsewhere herein.
[0109] Figure 8A schematic block diagram of a computer system 800 on which embodiments of the present disclosure may be implemented is shown. The computer system 800 includes a bus 810 or other communication mechanism for transmitting information, and a processing device 820 coupled to the bus 810 for processing information. The computer system 800 also includes a memory 830 coupled to the bus 810 for storing instructions to be executed by the processing device 820. The memory 830 may be a random access memory (RAM) or other dynamic storage device. The memory 830 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by the processing device 820. The computer system 800 may also include a read-only memory (ROM) 840 or other static storage device coupled to the bus 810 for storing static information and instructions for the processing device 820. A storage device 850, such as a magnetic disk or optical disk, is provided and coupled to the bus 810 for storing information and instructions. The computer system 800 can be coupled via bus 810 to an output device 860 for providing output to a user, such as, but not limited to, a display (such as a cathode ray tube (CRT) or liquid crystal display (LCD)), speakers, and the like. Input devices 870, such as a keyboard, mouse, microphone, and the like, are coupled to bus 810 for communicating information and command selections to the processing device 820. The computer system 800 can perform embodiments of the present disclosure. Consistent with certain implementations of the present disclosure, the computer system 800 provides results in response to the processing device 820 executing one or more sequences of one or more instructions contained in memory 830. Such instructions can be read into memory 830 from another computer-readable medium, such as storage device 850. Execution of the sequences of instructions contained in memory 830 causes the processing device 820 to perform the methods described herein. Alternatively, hardwired circuitry can be used in place of or in combination with software instructions to implement the present teachings. Thus, implementations of the present disclosure are not limited to any specific combination of hardware circuitry and software. In various embodiments, the computer system 800 can be connected to one or more other computer systems like the computer system 800 across a network via a network interface 880 to form a networked system. The network can include a private network or a public network such as the Internet. In a networked system, one or more computer systems can store data and supply data to other computer systems. As used herein, the term "computer-readable medium" refers to any medium that participates in providing instructions to the processing device 820 for execution. Such media can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks such as storage device 850. Volatile media include dynamic memory such as memory 830. Transmission media include coaxial cables, copper wire, and optical fibers, including the wiring comprising bus 810.Common forms of computer-readable media or computer program products include, for example, floppy disks, flexible disks, hard disks, magnetic tape, or any other magnetic medium, CD-ROMs, digital video disks (DVDs), Blu-ray discs, any other optical media, thumb drives, memory cards, RAM, PROMs and EPROMs, flash EPROMs, any other memory chips or cartridges, or any other tangible medium from which a computer can read. Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to processing device 820 for execution. For example, the instructions may initially be carried on a disk on a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 800 may receive the data on the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector coupled to bus 810 may receive the data carried in the infrared signal and place the data on bus 810. Bus 810 carries the data to memory 830, from which processing device 820 retrieves the instructions and executes them. Optionally, the instructions received by memory 830 may be stored on storage device 850 either before or after execution by processing device 820 .
[0110] According to various embodiments, instructions configured to be executed by the processing device 820 to perform the method are stored on a computer-readable medium. A computer-readable medium can be a device that stores digital information. For example, a computer-readable medium includes a compact disk read-only memory (CD-ROM) as known in the art for storing software. The computer-readable medium is accessed by a processor adapted to execute the instructions configured to be executed.
[0111] In the technical solution disclosed herein, a dense point cloud is generated by fusing multiple point clouds. A key point detection model is used to obtain the pixel coordinates of key points, which are then used to determine their point cloud coordinates within the dense point cloud. This allows for highly accurate detection of key points on work objects, enabling accurate positioning of the work objects. Certain embodiments of the present disclosure can effectively address the issue of low-beam LiDAR being unable to adequately scan key points on work objects. Furthermore, they can provide a foundation for planning subsequent material loading and unloading by loaders, helping to advance intelligent port operations.
[0112] The words "left," "right," "front," "back," "top," "bottom," "up," "down," "high," "low," and the like, if any, in the specification and claims, are used for descriptive purposes and are not necessarily intended to describe invariant relative positions. It should be understood that the words so used are interchangeable under appropriate circumstances such that the embodiments of the present disclosure described herein, for example, are capable of operation in other orientations than those shown or otherwise described herein. For example, when the device in the figures is turned over, features previously described as "above" other features could now be described as "below" the other features. The device can also be otherwise oriented (rotated 90 degrees or in other orientations) and relative spatial relationships will be interpreted accordingly.
[0113] In the specification and claims, when an element is referred to as being "on," "attached," "connected," "coupled," or "in contact with," etc., another element, the element may be directly on, directly attached, directly connected, directly coupled, or directly in contact with the other element, or one or more intervening elements may be present. In contrast, when an element is referred to as being "directly" "on," "directly attached," "directly connected," "directly coupled," or "in direct contact with" another element, there will be no intervening elements. In the specification and claims, when a feature is arranged "adjacent" to another feature, it may mean that the feature has a portion that overlaps with the adjacent feature or a portion that is located above or below the adjacent feature.
[0114] As used herein, the word "exemplary" means "serving as an example, instance, or illustration," rather than as a "model" to be precisely copied. Any implementation described as exemplary is not necessarily to be construed as preferred or advantageous over other implementations. Furthermore, this disclosure is not to be bound by any expressed or implied theory presented in the technical field, background, summary, or detailed description.
[0115] As used herein, the term "substantially" is intended to encompass any minor variations due to design or manufacturing imperfections, device or component tolerances, environmental influences, and / or other factors. The term "substantially" also allows for deviations from a perfect or ideal condition due to parasitic effects, noise, and other practical considerations that may be present in actual implementations.
[0116] Additionally, terms such as "first," "second," and the like may also be used herein for reference purposes only and are not intended to be limiting. For example, the terms "first," "second," and other numerical terms referring to structures or elements do not imply a sequence or order unless the context clearly indicates otherwise.
[0117] It should also be understood that when the term “include / comprises” is used in this document, it indicates the presence of the specified features, integers, steps, operations, units and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, units and / or components and / or their combinations.
[0118] In this disclosure, the term "provide" is used in a broad sense to cover all ways of obtaining an object, and thus "providing an object" includes but is not limited to "purchasing", "preparing / manufacturing", "arranging / setting up", "installing / assembling", and / or "ordering" an object, etc.
[0119] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the present disclosure. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0120] Those skilled in the art will appreciate that the boundaries between the above-mentioned operations are merely illustrative. Multiple operations can be combined into a single operation, a single operation can be distributed among additional operations, and operations can be performed at least partially overlapping in time. Moreover, alternative embodiments can include multiple instances of a particular operation, and the order of operations can be changed in various other embodiments. However, other modifications, variations, and replacements are also possible. Aspects and elements of all embodiments disclosed above can be combined in any manner and / or in combination with aspects or elements of other embodiments to provide multiple additional embodiments. Therefore, this specification and the accompanying drawings should be regarded as illustrative, not restrictive.
[0121] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. The various embodiments disclosed herein may be combined in any manner without departing from the spirit and scope of the present disclosure. Those skilled in the art will also appreciate that various modifications may be made to the embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A key point detection method, characterized in that: The key point detection method comprises: Acquire, by a first acquisition device, n frames of point clouds of the work object at a current moment, and generate a dense point cloud of the work object at the current moment based on the n frames of point clouds, where n is a positive integer greater than 1; Acquire an image of the work object through a second acquisition device, and determine pixel coordinates of key points of the work object based on the image and a pre-trained key point detection model; and The point cloud coordinates of the key point in the dense point cloud are determined based on the pixel coordinates of the key point.
2. The key point detection method according to claim 1, characterized in that: Generating a dense point cloud of the operation object at the current moment according to the first n frames of point cloud includes: Based on the i-th frame point cloud in the first n frames of point clouds, each point in the other frame point clouds is converted into the i-th frame point cloud to generate the dense point cloud, where i is any value from 1 to n.
3. The key point detection method according to claim 2, characterized in that: Based on the i-th frame point cloud in the first n frames of point clouds, converting each point in the other frame point clouds into the i-th frame point cloud to generate the dense point cloud includes: For each frame of point cloud in the other frame point clouds, the pose of each point in the current coordinate system is converted into the pose in the reference coordinate system on which the i-th frame point cloud is based.
4. The key point detection method according to claim 1, characterized in that: The key point detection method further includes: For the k-th frame point cloud in the first n frames of point cloud, correction is performed based on the motion parameters of the k-1-th frame point cloud, and the corrected point cloud is used as the k-th frame point cloud, where k is any value from 1 to n.
5. The key point detection method according to claim 1, characterized in that: Determining the point cloud coordinates of the key point in the dense point cloud based on the pixel coordinates of the key point includes: Determine the coordinate transformation relationship between point cloud coordinates and pixel coordinates; and The point cloud coordinates of the key point are obtained according to the coordinate transformation relationship and the pixel coordinates of the key point.
6. The key point detection method according to claim 1, characterized in that: The first acquisition device includes a laser radar; and / or the second acquisition device includes a camera.
7. The key point detection method according to claim 1, characterized in that: The working object includes a polygonal prism structure, and the key point is a vertex of the polygonal prism structure.
8. The key point detection method according to claim 1, characterized in that: The key point detection method further includes: The work object is loaded or unloaded according to the point cloud coordinates of the key points.
9. A key point detection device, characterized in that: The key point detection device comprises: The first acquisition device is configured to obtain the first n frames of point cloud of the work object at the current moment, where n is a positive integer greater than 1; a point cloud generating module configured to generate a dense point cloud of the work object at the current moment based on the point clouds of the previous n frames; a second acquisition device, configured to acquire an image of the work object; A first determining module is configured to determine the pixel coordinates of the key points of the work object based on the image and a pre-trained key point detection model; and The second determining module is configured to determine the point cloud coordinates of the key point in the dense point cloud based on the pixel coordinates of the key point.
10. A key point detection device, characterized in that: The key point detection device includes: processor; and A memory having instructions stored thereon, wherein when the instructions are executed by the processor, the steps of the key point detection method according to any one of claims 1 to 8 are implemented.
11. A loader, characterized in that: The loader includes the key point detection device according to claim 9 or the key point detection equipment according to claim 10.
12. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores instructions, and when the instructions are executed by the processor, the steps of the key point detection method according to any one of claims 1 to 8 are implemented.
13. A computer program product, characterized in that The computer program product includes instructions, and when the instructions are executed by a processor, the steps of the key point detection method according to any one of claims 1 to 8 are implemented.