Object real parameter solving method, device and equipment

By generating target object prediction point clouds under standard space and building correspondence relationships, the problem of poor accuracy in labeling symmetrical objects and non-standard structural objects in the prior art is solved, and more accurate object real parameters are obtained and a complete predicted point cloud acquisition is achieved.

CN120107355APending Publication Date: 2025-06-06PASSINI ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510163220.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing indoor scene pose labeling tools have poor accuracy when labeling symmetrical objects and non-standard structural objects. Especially due to visual occlusion and other relationships, it is difficult to obtain real parameter information such as the 6D pose of the object.

Method used

By obtaining the original image in the scene of the target object, the target object predicted point cloud is generated in the standard space, the corresponding relationship between the target object image and the predicted point cloud is constructed, and the real parameters of the object are obtained, including scale information and position information.

Benefits of technology

It realizes a more accurate and comprehensive acquisition of real parameters of target objects, and can obtain a complete object prediction point cloud, which is suitable for subsequent more accurate operations, such as grabbing, etc.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107355A_ABST
    Figure CN120107355A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the technical field of visual identification, and relates to an object real parameter solving method, which comprises the steps of generating a standard target object image of a target object meeting standard space input based on a scene 2D image, and generating a first corresponding relation between an original target object image and the standard target object image; generating a target object prediction point cloud in a standard space based on the standard target object image; constructing a second corresponding relation between the standard target object image and the target object prediction point cloud; extracting a third corresponding relation between the original target object image and the target object depth point cloud based on the original image; constructing a fourth corresponding relation between the target object prediction point cloud and the target object depth point cloud based on the first corresponding relation, the second corresponding relation and the third corresponding relation; and solving real parameters based on the fourth corresponding relation. According to the technical scheme, the real parameters of the object can be obtained based on the target object prediction point cloud in the standard space, so that more accurate and comprehensive real parameters of the object are obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of visual recognition technology, and in particular to a method, device and equipment for obtaining real parameters of an object. Background Art

[0002] 6D pose data represents the coordinate relationship of an object in three-dimensional space, including the rotation angle of the object and the displacement vector of the object. An important task in the field of computer vision is to estimate the pose of an object in three-dimensional space. Currently, it can be applied in many scenarios such as daily life, transportation, medical treatment and industrial production, such as robot grasping objects, robot hand-eye calibration, joint object labeling, three-dimensional scene reconstruction, augmented reality technology, unmanned driving technology and other tasks.

[0003] The existing pose estimation model based on deep learning requires a large amount of accurate pose data. To obtain accurate pose data, the collected data needs to be annotated. Generally, a point cloud combining color images and depth maps is used to annotate the pose of objects. The pose dataset can be used in various scenarios as needed, such as: outdoor scenes are oriented to autonomous driving technology, annotating the pose information of pedestrians and vehicles; indoor scenes are oriented to robot grasping, three-dimensional scene reconstruction and other tasks, annotating the pose information of indoor objects. Different scenarios have different requirements for detailed pose information. For example, for the pose of vehicles and pedestrians, the rotation angles of the x-axis and y-axis are limited, so detailed model information is not required, and the annotation box uses a three-dimensional bounding box; indoor scenes require detailed information on the pose of objects due to downstream task requirements and higher degrees of freedom than outdoor scenes, such as symmetrical objects, non-standard structure objects, joint structure objects, etc.

[0004] Existing indoor scene pose annotation tools, including 6D-PAT, 6DPoseAnnotator, CVAT, Vision6D, etc., are applied to different scenarios. Indoor object annotation generally uses color pictures and three-dimensional models or three-dimensional bounding boxes to annotate objects, but two-dimensional color pictures lack depth information, and the pose annotation accuracy is poor; for symmetrical objects, such as water bottles that consider color texture, the annotation pose needs to be adjusted according to color information during annotation, while non-standard structural objects have unclear shapes and high degrees of freedom in placement pose, such as apples and bananas. Existing pose annotation tools usually perform pose recognition based on the local point cloud of the target object due to visual occlusion and other relationships, so they often cannot obtain accurate object 6D pose and other real parameter information. Summary of the invention

[0005] The purpose of the embodiments of the present application is to propose a method, device and equipment for obtaining the real parameters of an object, so as to obtain the real parameters of the object based on the predicted point cloud of the target object in the standard space, so as to obtain more accurate and comprehensive real parameters of the target object.

[0006] In a first aspect, the present application provides a method for obtaining the real parameters of an object, which adopts the following technical solution:

[0007] A method for obtaining real parameters of an object, the method comprising the following steps:

[0008] Acquire an original image of a scene including a target object; the original image includes a scene 2D image and a corresponding scene depth image;

[0009] Based on the scene 2D image, a standard target image of the target satisfying the standard space input is generated, and a first correspondence relationship between an original target image and the standard target image is generated; wherein the original target image is a target part in the scene 2D image;

[0010] Generate a target object prediction point cloud in the standard space based on the standard target object image; construct a second corresponding relationship between the standard target object image and the target object prediction point cloud;

[0011] Extracting a third corresponding relationship between the original target object image and the target object depth point cloud based on the original image;

[0012] Constructing a fourth correspondence between the target object prediction point cloud and the target object depth point cloud based on the first correspondence, the second correspondence and the third correspondence;

[0013] The real parameter is obtained based on the fourth corresponding relationship; wherein the real parameter includes real scale information of the object and / or posture information of the object.

[0014] Furthermore, the constructing of the second corresponding relationship between the standard target image and the target prediction point cloud specifically includes the following steps:

[0015] Filtering the invisible part of the target object prediction point cloud to obtain a filtered target object prediction point cloud;

[0016] The second corresponding relationship between the standard target object image and the filtered target object prediction point cloud is constructed.

[0017] Furthermore, generating a standard target image of a target object that meets the standard space input based on the scene 2D image specifically includes the following steps:

[0018] Generating a target object mask image based on the scene 2D image;

[0019] Extracting an initial target object image of the target object from the scene 2D image based on the target object mask image;

[0020] The initial target object image is preprocessed to obtain the standard target object image.

[0021] Furthermore, the preprocessing of the initial target object image to obtain the standard target object image specifically includes the following steps:

[0022] Filling pixels on the background of the initial target image;

[0023] Performing size filling on the initial target image after the pixel filling;

[0024] A scaling operation is performed on the size-filled initial target image to obtain the standard target image that meets the standard space input requirement.

[0025] Furthermore, the third corresponding relationship between the original target object image and the target object depth point cloud extracted based on the original image specifically includes the following steps:

[0026] Generating a target object mask image based on the scene 2D image;

[0027] Extracting an original target object image from the scene 2D image using the target object mask image;

[0028] The third corresponding relationship between the original target object image and the depth object point cloud is acquired.

[0029] Furthermore, obtaining the real parameter based on the fourth corresponding relationship specifically includes the following steps:

[0030] Obtaining an IPC error term formula constructed based on the fourth corresponding relationship;

[0031] Based on the error term formula, an optimization objective function of the least squares problem is constructed to find the position parameters and scale parameters of the target object when the sum of squared errors is minimized.

[0032] In a second aspect, an embodiment of the present application provides a device for obtaining real parameters of an object, the device comprising:

[0033] An image acquisition module, used to acquire an original image of a scene including a target object; the original image includes a scene 2D image and a corresponding scene depth image;

[0034] A first generating module is used to generate a standard target image of the target object that meets the standard space input based on the scene 2D image, and generate a first corresponding relationship between the original target image and the standard target image; wherein the original target image is the target part in the scene 2D image;

[0035] A second generating module is used to generate a target object prediction point cloud in the standard space based on the standard target object image; and to construct a second corresponding relationship between the standard target object image and the target object prediction point cloud;

[0036] A third generating module, configured to extract a third corresponding relationship between the original target object image and the target object depth point cloud based on the original image;

[0037] A fourth generating module, configured to construct a fourth corresponding relationship between the target object prediction point cloud and the target object depth point cloud based on the first corresponding relationship, the second corresponding relationship and the third corresponding relationship;

[0038] A real obtaining module is used to obtain the real parameters of the predicted point cloud of the target object based on the fourth corresponding relationship; wherein the real parameters include the real scale information of the object and / or the posture information of the object.

[0039] Furthermore, the second generating module includes:

[0040] A point cloud filtering submodule, used for filtering the invisible part of the point cloud in the target object prediction point cloud to obtain a filtered target object prediction point cloud;

[0041] The relationship construction submodule is used to construct the second corresponding relationship between the standard target object image and the filtered target object prediction point cloud.

[0042] In a third aspect, an embodiment of the present application provides a controller, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method for obtaining the real parameters of the object described in any one of the above items are implemented.

[0043] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for obtaining the real parameters of an object described in any one of the above items are implemented.

[0044] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0045] The embodiment of the present application generates a target image of the target object that meets the standard space input based on the scene 2D image, and generates a first correspondence between the original target image and the standard target image; generates a target prediction point cloud in the standard space based on the target image, and constructs a second correspondence between the target image and the target prediction point cloud; extracts a third correspondence between the original target image of the target object and the target depth point cloud based on the original image; constructs a fourth correspondence between the target prediction point cloud and the target depth point cloud based on the first correspondence, the second correspondence and the third correspondence; and obtains real parameters of the target prediction point cloud (including: real scale information of the object and / or posture information of the object) based on the fourth correspondence, thereby providing reference information for subsequent operations based on the target object.

[0046] In addition, compared with the object point cloud generated in the real scene based on 2D images, etc., since the object point cloud is generated based on the actual captured image, and there are invisible parts in the image, the object point cloud obtained based on this often lacks the invisible part. The method described in the embodiment of the present application can obtain a complete object prediction point cloud, which can facilitate subsequent more accurate capture and other operations on the object prediction point cloud. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the scheme in the present application, a brief introduction is given below to the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0048] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;

[0049] Figure 2 It is a flowchart of an embodiment of a method for obtaining real parameters of an object of the present application;

[0050] Figure 3 It is a schematic diagram of the framework structure of an embodiment of a device for obtaining real parameters of an object of the present application;

[0051] Figure 4 It is a structural diagram of an embodiment of a computer device of the present application. DETAILED DESCRIPTION

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by technicians in the technical field of the present application; the terms used in the specification of the application herein are only for the purpose of describing specific embodiments and are not intended to limit the present application; the terms "including" and "having" and any variations thereof in the specification and claims of the present application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of the present application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0053] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0054] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0055] like Figure 1 As shown, Figure 1 is an exemplary system architecture diagram to which the present application can be applied.

[0056] An embodiment of the present application provides an image capturing system 100 , which includes an image sensor 110 and a controller 120 .

[0057] Image Sensor

[0058] The image sensor is used to collect a scene 2D image and a corresponding scene depth image in the scene of the target object.

[0059] The image sensor 120 may be any type of image sensor that is currently available or will be developed in the future and can capture 2D images and depth images. For example, it may be an RGBD depth camera, or it may be a combination of two independent cameras, a 2D camera and a depth camera. For ease of understanding, the embodiments of the present application are mainly described in detail by taking the image sensor 120 as an RGBD depth camera (referred to as a "camera") as an example.

[0060] Exemplarily, taking an RGBD camera as an example, the camera can capture an RGB image of the scene within the field of view and a corresponding depth image (which may be called a "raw image"), and then send the raw image to a storage device or a server, etc.

[0061] Controller

[0062] The controller 120 is connected to the image sensor 110 in a wired or wireless communication manner.

[0063] For the limitations of the controller, please refer to the limitations of the method for obtaining the real parameters of the object in the following embodiments.

[0064] It should be noted that the above-mentioned wireless connection methods may include but are not limited to 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or developed in the future.

[0065] The controller provided in the embodiments of the present application may be, but is not limited to: a computer terminal (Personal Computer, PC); an industrial control computer terminal (Industrial Personal Computer, IPC); a mobile terminal; a server; a system including a terminal and a server, and implemented through the interaction between the terminal and the server; a programmable logic controller (Programmable Logic Controller, PLC); a field programmable gate array (Field-Programmable Gate Array, FPGA); a digital signal processor (Digital Signal Processer, DSP) or a microcontroller unit (Microcontroller unit, MCU) and other similar controllers. The controller 120 can generate program instructions according to a pre-fixed program in combination with data signals collected by an external image sensor 110, etc. Specifically, the controller may be, for example, Figure 4 Computer equipment shown.

[0066] It should be noted that the controller described in the embodiments of the present application can be provided separately or partially integrated in the image sensor, both of which fall within the scope of protection of the present application.

[0067] Based on the system of the above embodiment, the embodiment of the present application provides a method for obtaining real parameters of an object, which is generally executed by the controller 130.

[0068] like Figure 2 As shown, Figure 2 1 is a flow chart of an embodiment of a method for obtaining the real parameters of an object of the present application; the method for obtaining the real parameters of the object may include the following method steps:

[0069] Step 210 acquires an original image of a scene including a target object; the original image includes a scene 2D image and a corresponding scene depth image;

[0070] Step 220 generates a standard target image of the target object that meets the standard space input based on the scene 2D image, and generates a first correspondence between the original target image and the standard target image; wherein the original target image is the target part in the scene 2D image.

[0071] Step 230 generates a target object prediction point cloud in a standard space based on the target object image; and constructs a second corresponding relationship between the target object image and the target object prediction point cloud;

[0072] Step 240 extracts a third correspondence between the original target object image and the target object depth point cloud based on the original image;

[0073] Step 250 constructs a fourth correspondence between the target object prediction point cloud and the target object depth point cloud based on the first correspondence, the second correspondence, and the third correspondence;

[0074] Step 260 obtains real parameters based on the fourth corresponding relationship; wherein the real parameters include real scale information of the object and / or posture information of the object.

[0075] The embodiment of the present application generates a target image of the target object that meets the standard space input based on the scene 2D image, and generates a first correspondence between the original target image and the target image; generates a target prediction point cloud in the standard space based on the target image, and constructs a second correspondence between the target image and the target prediction point cloud; extracts a third correspondence between the original target image of the target object and the target depth point cloud based on the original image; constructs a fourth correspondence between the target prediction point cloud and the target depth point cloud based on the first correspondence, the second correspondence and the third correspondence; and obtains real parameters of the target prediction point cloud (including: real scale information of the object and / or posture information of the object) based on the fourth correspondence, thereby providing reference information for subsequent operations based on the target object.

[0076] In addition, compared with the object point cloud generated in the real scene based on 2D images, etc., since the object point cloud is generated based on the actual captured image, and there are invisible parts in the image, the object point cloud obtained based on this often lacks the invisible part. The method described in the embodiment of the present application can obtain a complete object prediction point cloud, which can facilitate subsequent more accurate capture and other operations on the object prediction point cloud.

[0077] For ease of understanding, the above method steps are further described in detail below.

[0078] Step 210 acquires an original image of the target scene; the original image includes a scene 2D image and a corresponding scene depth image.

[0079] In an optional embodiment, the controller obtains an original image of a scene including a target object obtained by an RGBD image sensor from a memory or a server according to a preset address, where the original image includes a scene RGB image (ie, a scene 2D image) and a corresponding scene depth image.

[0080] It should be noted that the target object scene in the above original image may include only the target object, or may include other objects besides the target object.

[0081] like Figure 1 As shown, illustratively, taking multiple objects placed on a table as an example, the target object scene of the original image may only include the target object 200 described in the embodiment of the present application (the target object may be any one of the multiple objects), or may also include the table and other multiple objects placed on the table.

[0082] Step 220 generates a standard target image of the target object that meets the standard space input based on the scene 2D image, and generates a first correspondence between the original target image and the standard target image; wherein the original target image is the target part in the scene 2D image.

[0083] It should be noted that, in order to ensure the effective use of multi-source data and reduce the search space of the model, the 3D model generated based on the subsequent step 230 is generally predicted in the canonical space, so the target object image needs to be pre-converted into a format that meets the canonical space input.

[0084] Among them, since the sizes of the target images extracted in different situations are inconsistent, the corresponding camera intrinsic parameters may also be different. For the subsequent target 3D image generation model, it is necessary to set the input size and camera intrinsic parameters to be fixed. The target 3D image prediction under fixed image size and camera intrinsic parameters is called standard space (CannonicalSpace).

[0085] In an optional embodiment, step 220 generates a standard target image of a target object that meets the standard space input based on the scene 2D image, which may specifically include the following method steps:

[0086] Step 221 generates a target object mask image based on the scene 2D image.

[0087] Specifically, the target object mask map can be extracted based on the scene 2D image based on various existing or future developed methods, such as: generating the target object mask map through an instance segmentation neural network model, an open vocabulary detection network model, etc.

[0088] Step 222 extracts an initial target object image of the target object from the scene 2D image based on the target object mask image.

[0089] Step 223 pre-processes the initial target image to obtain a standard target image.

[0090] The embodiment of the present application can convert the scene 2D image into the target object image through the above conversion steps, so that the 2D-2D correspondence between the target object part in the scene 2D image (i.e., the original target object image) and the standard target object image can be found, that is, the matching point pairs in the two images can be found.

[0091] Further, in an optional embodiment, step 223 may include the following method steps:

[0092] Step 2231 performs pixel filling on the background in the initial target image.

[0093] Specifically, the above pixels can be set arbitrarily as needed, for example, the background can be uniformly filled with white.

[0094] Step 2232 fills the size of the initial target image after pixel filling (for example, by padding processing).

[0095] In the embodiment of the present application, the goal of image preprocessing is to scale the target object image to a fixed size. In order to keep the aspect ratio of the original target object unchanged, padding processing is required.

[0096] Step 2233 performs a scaling operation on the size-filled initial target image to obtain a standard target image that meets the standard space input requirements.

[0097] The embodiment of the present application can obtain a target image with a uniform background and uniform size through the above method steps.

[0098] It should be noted that, in addition to the above embodiments, the embodiments of the present application can pre-process the initial target object image based on various existing or future developed methods. As long as a standard target object image that meets the standard space input requirements can be obtained, it falls within the scope of protection of the present application.

[0099] In the embodiment of the present application, an initial target object image of the target object is extracted from the scene 2D image, and then the initial target object image is subjected to a series of preprocessing to obtain a final standard target object image.

[0100] Step 230 generates a target object prediction point cloud in a standard space based on the standard target object image; and constructs a second corresponding relationship between the target object image and the target object prediction point cloud.

[0101] In an embodiment of the present application, a target prediction point cloud in a standard space can be generated based on a standard target image based on various currently existing or future developed methods, such as: three-dimensional Gaussian sputtering (3DGS), diffusion model (DifffusionModel) and neural radiation field (NERF); or, other forms of target 3D images (such as: CAD or MESH) can be first constructed based on the target image, and then other forms of target 3D images can be converted into target prediction point clouds, all of which fall within the scope of protection of the present application.

[0102] In an optional embodiment, the above-mentioned “constructing a first corresponding relationship between the target object image and the target object prediction point cloud” may specifically include the following method steps:

[0103] Step 231 filters the invisible part of the target object prediction point cloud to obtain a filtered target object prediction point cloud.

[0104] Step 232 constructs a second correspondence between the target object image and the filtered target object prediction point cloud.

[0105] In the embodiment of the present application, the invisible part of the object is filtered by projection filtering in combination with the principle of pinhole imaging. The principle is that points p1 and p2 on the same straight line will be projected to the same position on the image. For the point cloud set projected to the same or adjacent pixels, only the point cloud with the closest depth is retained, thereby realizing the function of filtering the point cloud corresponding to the invisible part of the object, so as to obtain a more accurate 2D-3D second correspondence between the target object image and the predicted object point cloud.

[0106] The embodiment of the present application can filter out the invisible parts in the predicted object point cloud through the above method steps, so that the 2D-3D correspondence relationship between the target object image and the predicted object point cloud can be finally constructed more accurately.

[0107] Step 240 extracts a third correspondence between the original target object image and the target object depth point cloud based on the original image.

[0108] In an optional embodiment, step 240 extracts a third correspondence between the original target object image and the target object depth point cloud based on the original image, which may specifically include the following method steps:

[0109] Step 241 generates a target object mask image based on the scene 2D image.

[0110] The method for generating the target object mask image can be found in the previous embodiment and will not be repeated here.

[0111] Step 242 extracts the original target object image from the scene 2D image using the target object mask image.

[0112] Step 243 obtains a third correspondence between the original target object image and the depth object point cloud.

[0113] Based on the previous embodiments, since there is a corresponding relationship between the original scene 2D image and the scene depth image itself, the original target object image is extracted from the scene 2D image through the target object mask map, and the 2D-3D third corresponding relationship between the original target object image and the corresponding depth object point cloud in the depth point cloud can be directly obtained.

[0114] Step 250 constructs a fourth correspondence between the target object prediction point cloud and the target object depth point cloud based on the first correspondence, the second correspondence, and the third correspondence.

[0115] In the embodiment of the present application, by combining the first correspondence (i.e., the 2D-2D correspondence between the original target object image and the standard target object image in the scene 2D image), the second correspondence (i.e., the 2D-3D correspondence between the standard target object image and the target object prediction point cloud), and the third correspondence (i.e., the 2D-3D correspondence between the original target object image and the target object depth point cloud), a 3D-3D correspondence between the matched target object prediction point cloud and the target object depth point cloud can be obtained.

[0116] The embodiment of the present application uses the 2D-2D correspondence between the original target object image and the standard target object image as a matching bridge, which is a fixed conversion relationship. Therefore, the fourth 3D-3D correspondence between the target object prediction point cloud and the target object depth point cloud (i.e., matching point pairs) is finally obtained. This fourth correspondence has high accuracy.

[0117] It should be noted that in the above-mentioned correspondence between the target object prediction point cloud and the target object depth point cloud, due to inevitable errors and other reasons, usually only part of the point sets form matching point pairs, that is, only part of the points in the target object prediction point cloud and the target object depth point cloud can be matched to form the corresponding fourth correspondence.

[0118] Step 260 obtains the real parameter based on the fourth corresponding relationship.

[0119] The real parameters may include the real scale information of the object and / or the posture information of the object. In addition, the real parameters may also include other parameters as needed.

[0120] It should be noted that the real parameters of the object can be obtained based on the fourth correspondence based on various existing or future developed methods, such as: based on the IPC pose solution algorithm and the PNP pose solution method.

[0121] For ease of understanding, the IPC pose solution algorithm is taken as an example for detailed explanation below.

[0122] In an optional embodiment, step 260 may include the following method steps:

[0123] Step 261 obtains the IPC error term formula (1) constructed based on the fourth corresponding relationship.

[0124] Specifically, we can set p i Predict the i-th point in the matching point set A in the point cloud for the target object, p i ′ is the i-th point in the matching point set B in the depth point cloud of the target object, then p i and p i ′ This constitutes the fourth corresponding relationship mentioned above, namely the 3D-3D matching point pair.

[0125] The error term formula based on the i-th point pair (1) is:

[0126] e i =p i -(R×s×p i ′ +t) (1)

[0127] Among them, R and t are the rotation and translation to be solved, s is the scale of the object to be solved, and e i is the error of the i-th point pair.

[0128] Step 262 constructs the optimization objective function (2) of the least squares problem based on the above error term formula (1) to find R, t, s (i.e., the true parameters) that minimize the sum of squared errors.

[0129]

[0130] In one embodiment, the solution can be performed based on the SVD singular value decomposition method, and the specific solution process is described as follows:

[0131] We can first define the centroids of two sets of points (set as point set A and point set B), that is, the average position of all points in each set. Then the above optimization objective function (2) can be simplified to (3)

[0132]

[0133] Among them, p is the centroid corresponding to the point set A, p ′ is the centroid of point set B.

[0134] The right side of the equal sign in the above equation is related to the center of mass. As long as we solve for R, we can set the second term equal to zero to solve for t.

[0135] Then calculate the centroid coordinates of each point to center the point set so that the centroid is at the origin, thereby eliminating the effect of translation; calculate the covariance matrix H between the centralized point sets, which can be transposed and multiplied with point set B; perform singular value decomposition (SVD) on the covariance matrix H to obtain the matrix U, singular value S and matrix V T ; Calculate the rotation matrix R through matrix multiplication, check whether the rows and columns of the rotation matrix are negative. If so, it means that the rotation matrix is ​​inverse, so V needs to be T The last line of is inverted to ensure that the rotation matrix is ​​valid; the scale factor s is calculated by the ratio of the sum of the singular values ​​S and the sum of the squares of the centered point set A of the target object prediction point cloud.

[0136] The embodiment of the present application can obtain the predicted point cloud of the target object (also regarded as an object) in the real scene through the above method steps. The posture R, t (that is, the posture information of the above object), and the real scale factor s of the corresponding object (that is, the real scale information of the above object) can be used as the real parameters of the object.

[0137] In the embodiment of the present application, since the depth point cloud of the target object corresponds to the real value in the real scene, such as: the real posture information and the real scale information in the scene, therefore, in combination with the fourth correspondence between the target object prediction point cloud and the target object depth point cloud, the real parameters of the target object depth point cloud (such as: the real scale information and / or the posture information in the scene) can be obtained to facilitate subsequent other operations on the object prediction point cloud.

[0138] The embodiment of the present application generates a target image of the target object that meets the standard space input based on the scene 2D image, and generates a first correspondence between the original target image and the target image; generates a target prediction point cloud in the standard space based on the target image, and constructs a second correspondence between the target image and the target prediction point cloud; extracts a third correspondence between the original target image of the target object and the target depth point cloud based on the original image; constructs a fourth correspondence between the target prediction point cloud and the target depth point cloud based on the first correspondence, the second correspondence and the third correspondence; and obtains real parameters of the target prediction point cloud (including: real scale information of the object and / or posture information of the object) based on the fourth correspondence, thereby providing reference information for subsequent operations based on the target object.

[0139] In addition, compared with the object point cloud generated in the real scene based on 2D images, etc., since the object point cloud is generated based on the actual captured image, and there are invisible parts in the image, the object point cloud obtained based on this often lacks the invisible part. Therefore, the method described in the embodiment of the present application can obtain a complete object prediction point cloud, which can facilitate subsequent more accurate capture and other operations on the object prediction point cloud.

[0140] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0141] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0142] Further references Figure 3 , as a response to the above Figure 2 The present application provides an embodiment of a device for obtaining the real parameters of an object. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied in a controller.

[0143] like Figure 3 As shown, the real parameter obtaining device 400 of the object described in this embodiment includes:

[0144] The image acquisition module 410 is used to acquire an original image of a scene including a target object; the original image includes a scene 2D image and a corresponding scene depth image;

[0145] A first generating module 420 is used to generate a standard target image of a target object that meets the standard space input based on the scene 2D image, and generate a first correspondence between the original target image and the standard target image; wherein the original target image is a target part in the scene 2D image;

[0146] The second generating module 430 is used to generate a target object prediction point cloud in a standard space based on the standard target object image; and to construct a second corresponding relationship between the standard target object image and the target object prediction point cloud;

[0147] A third generating module 440 is used to extract a third corresponding relationship between the original target object image and the target object depth point cloud based on the original image;

[0148] A fourth generating module 450 is used to construct a fourth corresponding relationship between the target object prediction point cloud and the target object depth point cloud based on the first corresponding relationship, the second corresponding relationship and the third corresponding relationship;

[0149] The real obtaining module 460 is used to obtain the real parameters of the predicted point cloud of the target object based on the fourth corresponding relationship; wherein the real parameters include the real scale information of the object and / or the posture information of the object.

[0150] In an optional embodiment, the second generating module 430 may include:

[0151] The point cloud filtering submodule is used to filter the invisible part of the point cloud in the target object prediction point cloud to obtain the filtered target object prediction point cloud;

[0152] The relationship construction submodule is used to construct a second corresponding relationship between the standard target image and the filtered target prediction point cloud.

[0153] In an optional embodiment, the first generating module 420 may include:

[0154] The mask generation submodule is used to generate a target object mask map based on the scene 2D image;

[0155] An image extraction submodule, used for extracting an initial target object image from a scene 2D image based on a target object mask image;

[0156] The image processing submodule is used to preprocess the initial target object image to obtain a standard target object image.

[0157] In an optional embodiment, the image processing submodule may include:

[0158] A pixel filling unit, used for filling pixels of the background in the initial target image;

[0159] A size filling unit, used for filling the size of the initial target image after pixel filling;

[0160] The scaling operation unit is used to perform a scaling operation on the initial target object image after the size filling, so as to obtain a standard target object image that meets the standard space input requirements.

[0161] In an optional embodiment, the third generating module 440 may include:

[0162] The mask generation submodule is used to generate a target object mask map based on the scene 2D image;

[0163] The target extraction submodule is used to extract the original target image from the scene 2D image through the target mask image;

[0164] The relationship acquisition submodule is used to obtain the third corresponding relationship between the original target object image and the depth object point cloud.

[0165] In an optional embodiment, the real obtaining module 460 may include:

[0166] An error acquisition submodule, used to obtain an IPC error term formula constructed based on the fourth corresponding relationship;

[0167] The parameter calculation submodule is used to construct the optimization objective function of the least squares problem based on the error term formula, and calculate the position and scale parameters of the target object when the sum of squared errors is minimized.

[0168] In order to solve the above technical problems, the present application embodiment also provides a controller, such as: Figure 4 Computer equipment shown.

[0169] The computer device may be a terminal or a server.

[0170] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as basic cloud computing services such as big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application.

[0171] The computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 6 with components 61-63, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (Application Specific Integrated Circuit, ASIC), programmable gate arrays (Field-Programmable Gate Array, FPGA), digital processors (Digital Signal Processor, DSP), embedded devices, etc.

[0172] The memory 61 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 61 can be an internal storage unit of the computer device 6, such as a hard disk or memory of the computer device 6. In other embodiments, the memory 61 can also be an external storage device of the computer device 6, such as a plug-in hard disk equipped on the computer device 6, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), etc. Of course, the memory 61 can also include both the internal storage unit of the computer device 6 and its external storage device. In this embodiment, the memory 61 is generally used to store the operating system and various application software installed on the computer device 6, such as the program code of the method for obtaining the real parameters of the object, etc. In addition, the memory 61 can also be used to temporarily store various types of data that have been output or are to be output.

[0173] The processor 62 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments. The processor 62 is generally used to control the overall operation of the computer device 6. In this embodiment, the processor 62 is used to run the program code stored in the memory 61 or process data, such as running the program code of the method for obtaining the real parameters of the object.

[0174] The network interface 63 may include a wireless network interface or a wired network interface. The network interface 63 is generally used to establish a communication connection between the computer device 6 and other electronic devices.

[0175] The present application also provides another embodiment, namely, providing a computer-readable storage medium, wherein the computer-readable storage medium stores a control program for obtaining the real parameters of an object, and the control program for obtaining the real parameters of the object can be executed by at least one processor so that the at least one processor performs the steps of the method for obtaining the real parameters of the object as described above.

[0176] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0177] Obviously, the embodiments described above are only some embodiments of the present application, rather than all embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application is described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions recorded in the aforementioned specific implementation methods, or to perform equivalent replacement of some of the technical features therein. Any equivalent structure made using the contents of the specification and drawings of this application, directly or indirectly used in other related technical fields, is similarly within the scope of patent protection of this application.

Claims

1. A method for obtaining the real parameters of an object, characterized in that: The method comprises the following steps: Acquire an original image of a scene including a target object; the original image includes a scene 2D image and a corresponding scene depth image; Based on the scene 2D image, a standard target image of the target satisfying the standard space input is generated, and a first correspondence relationship between an original target image and the standard target image is generated; wherein the original target image is a target part in the scene 2D image; Generate a target object prediction point cloud in the standard space based on the standard target object image; construct a second corresponding relationship between the standard target object image and the target object prediction point cloud; Extracting a third corresponding relationship between the original target object image and the target object depth point cloud based on the original image; Constructing a fourth correspondence between the target object prediction point cloud and the target object depth point cloud based on the first correspondence, the second correspondence and the third correspondence; The real parameter is obtained based on the fourth corresponding relationship; wherein the real parameter includes real scale information of the object and / or posture information of the object.

2. The method for obtaining the real parameters of an object according to claim 1 or 2, characterized in that: The step of constructing a second corresponding relationship between the standard target object image and the target object prediction point cloud specifically includes the following steps: Filtering the invisible part of the target object prediction point cloud to obtain a filtered target object prediction point cloud; The second corresponding relationship between the standard target object image and the filtered target object prediction point cloud is constructed.

3. The method for obtaining the real parameters of an object according to claim 1 or 2, characterized in that: The step of generating a standard target object image of a target object satisfying the standard space input based on the scene 2D image specifically comprises the following steps: Generating a target object mask image based on the scene 2D image; Extracting an initial target object image of the target object from the scene 2D image based on the target object mask image; The initial target object image is preprocessed to obtain the standard target object image.

4. The method for obtaining the real parameters of an object according to claim 3, characterized in that: The preprocessing of the initial target image to obtain the standard target image specifically includes the following steps: Filling pixels on the background of the initial target image; Performing size filling on the initial target image after the pixel filling; A scaling operation is performed on the size-filled initial target image to obtain the standard target image that meets the standard space input requirement.

5. The method for obtaining the real parameters of an object according to claim 1 or 2, characterized in that: The third corresponding relationship between the original target object image and the target object depth point cloud extracted based on the original image specifically includes the following steps: Generating a target object mask image based on the scene 2D image; Extracting an original target object image from the scene 2D image using the target object mask image; The third corresponding relationship between the original target object image and the depth object point cloud is acquired.

6. The method for obtaining the real parameters of an object according to claim 1 or 2, characterized in that: The obtaining of the real parameter based on the fourth corresponding relationship specifically comprises the following steps: Obtaining an IPC error term formula constructed based on the fourth corresponding relationship; Based on the error term formula, an optimization objective function of the least squares problem is constructed to find the position parameters and scale parameters of the target object when the sum of squared errors is minimized.

7. A device for obtaining the real parameters of an object, characterized in that: The device comprises: An image acquisition module, used to acquire an original image of a scene including a target object; the original image includes a scene 2D image and a corresponding scene depth image; A first generating module is used to generate a standard target image of the target object that meets the standard space input based on the scene 2D image, and generate a first corresponding relationship between the original target image and the standard target image; wherein the original target image is the target part in the scene 2D image; A second generating module is used to generate a target object prediction point cloud in the standard space based on the standard target object image; and to construct a second corresponding relationship between the standard target object image and the target object prediction point cloud; A third generating module, configured to extract a third corresponding relationship between the original target object image and the target object depth point cloud based on the original image; A fourth generating module, configured to construct a fourth corresponding relationship between the target object prediction point cloud and the target object depth point cloud based on the first corresponding relationship, the second corresponding relationship and the third corresponding relationship; A real obtaining module is used to obtain the real parameters of the predicted point cloud of the target object based on the fourth corresponding relationship; wherein the real parameters include the real scale information of the object and / or the posture information of the object.

8. The device for obtaining the real parameters of an object according to claim 7, characterized in that: The second generation module comprises: A point cloud filtering submodule, used for filtering the invisible part of the point cloud in the target object prediction point cloud to obtain a filtered target object prediction point cloud; The relationship construction submodule is used to construct the second corresponding relationship between the standard target object image and the filtered target object prediction point cloud.

9. A controller comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method for obtaining the real parameters of an object as described in any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for obtaining the real parameters of an object according to any one of claims 1 to 6 are implemented.