Transparent object measurement method and device based on guide generation type point cloud diffusion model

By guiding the generative point cloud diffusion model, utilizing binocular cameras and network computing, the reconstruction failure problem caused by optical properties in transparent object measurement is solved, accurate transparent object surface point clouds are generated, and efficient transparent object measurement and cross-dataset generalization capabilities are achieved.

CN120672826APending Publication Date: 2025-09-19TSINGHUA UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510703228.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional binocular stereo reconstruction methods fail in transparent object measurement and cannot accurately reconstruct the surface of transparent objects. This is because the optical properties of transparent objects cause light path deflection and light penetration, and existing methods cannot effectively calculate the three-dimensional position of transparent objects.

Method used

A guided generative point cloud diffusion model is adopted, and images are collected using a binocular camera. The back-projected point cloud and edge point cloud are calculated through a mask segmentation network and a pose estimation network. The generative point cloud diffusion model is used for iterative denoising to generate the target point cloud of the transparent object.

Benefits of technology

It achieves accurate measurement of transparent objects and generates a complete point cloud of the transparent object surface, which improves the accuracy of measurement and has the ability to generalize across data sets, reducing the training cost in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672826A_ABST
    Figure CN120672826A_ABST
Patent Text Reader

Abstract

The invention provides a transparent object measurement method and device based on a guide generation type point cloud diffusion model, and belongs to the field of three-dimensional measurement. The method comprises the following steps: acquiring a binocular image of a transparent object to be measured by using a binocular camera; inputting the binocular image into a preset mask segmentation network and a pose estimation network to obtain a binocular mask and a seven-dimensional pose estimation result of the object; based on the binocular mask, calculating a back projection point cloud of the object and a binocular edge point cloud, and obtaining a target point cloud of the object through iterative noise reduction by using a preset generative point cloud diffusion model; and converting the target point cloud into a camera coordinate system of the binocular camera based on a seven-dimensional pose estimation result, and completing the measurement of the transparent object to be measured. The method can be suitable for existing binocular camera equipment, and the transparent object is measured by calculating a guide function based on projection, noise addition and back projection, using an edge point cloud fusion method and using the estimated seven-dimensional pose.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of three-dimensional measurement, and in particular relates to a transparent object measurement method and device based on a guided generative point cloud diffusion model. Background Art

[0002] In transparent object measurement tasks, due to the special optical properties of the transparent object surface, such as transmission and refraction, the traditional binocular stereo reconstruction method will fail in the stereo matching process, resulting in reconstruction failure. The principle of binocular stereo matching is to find the pixel coordinates of the same point in the binocular camera, and through triangulation, calculate the absolute position of the measurement point based on the relative position of the binocular camera and the measurement point. For transparent objects, refraction will deflect the light path, making the triangulation method based on the non-deflected light path invalid; while transmission causes only a small amount of light to be reflected on the surface of the transparent object, and most of the light penetrates the surface of the object. The triangulation method is more inclined to reconstruct the surface or scene after penetrating the surface of the transparent object. Therefore, in the transparent object measurement task, a three-dimensional reconstruction method that is not based on traditional stereo matching is needed. Summary of the Invention

[0003] The present invention aims to overcome the shortcomings of existing technologies by proposing a method and apparatus for transparent object measurement based on a guided generative point cloud diffusion model. This method, applicable to existing binocular camera systems, calculates a guidance function based on projection, noise addition, and back-projection, and employs an edge point cloud fusion method to measure transparent objects using an estimated 7-dimensional pose.

[0004] The first embodiment of the present invention provides a transparent object measurement method based on a guided generative point cloud diffusion model, comprising:

[0005] Use a binocular camera to collect binocular images of the transparent object to be measured;

[0006] Inputting the binocular image into a preset mask segmentation network and a pose estimation network respectively to obtain a binocular mask and a 7-dimensional pose estimation result of the transparent object to be measured;

[0007] Based on the binocular mask, calculating the back-projected point cloud of the transparent object to be measured and the binocular edge point cloud;

[0008] Based on the back-projected point cloud and the binocular edge point cloud, a preset generative point cloud diffusion model is used to obtain a target point cloud of the transparent object to be measured through iterative noise reduction;

[0009] Based on the 7-dimensional pose estimation result, the target point cloud is transformed into the camera coordinate system of the binocular camera, and the measurement of the transparent object to be measured is completed.

[0010] In a specific embodiment of the present invention, before using a binocular camera to collect a binocular image of the transparent object to be measured, the method further includes:

[0011] Calibrate the binocular camera to obtain the camera's intrinsic parameter matrix:

[0012]

[0013] Among them, K l , K r Represents the intrinsic parameter matrix of the left and right cameras, f lx 、f rx Represents the x-axis focal length of the left and right cameras, f ly 、f ry Respectively represent the focal length of the left and right camera y axis, u l0 、u r0 Respectively represent the coordinates of the left and right camera principal points on the x-axis, v l0 、v r0 Represent the coordinates of the left and right camera principal points on the y-axis respectively.

[0014] In a specific embodiment of the present invention, the mask segmentation network adopts the Segment Anything model, and the pose estimation network adopts the LaPose pose model.

[0015] In a specific embodiment of the present invention, the step of calculating the back-projected point cloud and the binocular edge point cloud of the transparent object to be measured based on the binocular mask includes:

[0016] 1) The binocular left mask S of the transparent object to be measured 0,l and binocular right mask S 0,r Discretize and back-project to a plane perpendicular to the camera optical axis and passing through the center of the 3D segmentation frame of the transparent object to be measured to obtain the back-projected point cloud X0 b ;

[0017] 2) From S 0,l The leftmost column and S 0,r In the rightmost column, points in the same row are selected to form a pair of matching point pairs. Using all matching point pairs, the binocular edge point cloud X is calculated according to the triangulation method. r .

[0018] In a specific embodiment of the present invention, obtaining the target point cloud of the transparent object to be measured includes:

[0019] 1) Obtain the Gaussian point cloud of the t-step iteration in the denoising process by random sampling, denoted as X t , t﹥1;

[0020] 2) Using the back-projected point cloud, the Gaussian point cloud X of the t-th iteration is transformed t Input the preset generative point cloud diffusion model μ θ (X t ,t), the network outputs the denoised partial target point cloud of the t-1th iteration;

[0021]

[0022] Among them, X t-1 unknown is the denoised partial target point cloud of the t-1th iteration, σ t is the variance constant of the t-th step, S t,l ,S t,r are the binocular left mask and binocular right mask of the transparent object to be measured in the t-th iteration, respectively, F guide is the bootstrap function;

[0023] 3) Adding noise to the binocular edge point cloud and fusing it with the denoised partial target point cloud obtained in step 2) to obtain the denoised target point cloud in step t-1;

[0024] Among them, for X r Add noise to get the noised partial target point cloud of the t-1th iteration:

[0025]

[0026] Then we get the denoised target point cloud of the t-1th iteration:

[0027] X t-1 =concat(X t-1 known ,X t-1 unknown [N r +1:N total ])

[0028] Among them, ε1, ε2 are the first noise and the second noise sampled from two independent normal distributions, N r Point Cloud X t-1 known The number of points, N total Point Cloud X t-1 The number of points;

[0029] 4) Determine t:

[0030] If t>1, set t=t-1 and return to step 2);

[0031] Otherwise, the target point cloud X0 is generated.

[0032] In a specific embodiment of the present invention, in the step t iterative Gaussian point cloud X t Input the preset generative point cloud diffusion model μ θ (X t ,t), also includes:

[0033] training the generative point cloud diffusion model;

[0034] The training of the generative point cloud diffusion model includes:

[0035] 1) constructing a training set, wherein the training samples in the training set include surface point clouds of objects of the same category as the transparent object to be measured and their corresponding labels;

[0036] 2) Constructing a generative point cloud diffusion model μ θ (X t ,t), the model adopts Point-Voxel CNN network;

[0037] 3) Using the training set obtained in step 1), the generative point cloud diffusion model constructed in step 2) is trained to obtain a trained generative point cloud diffusion model.

[0038] In a specific embodiment of the present invention, the F guide The calculation process is as follows:

[0039] 1) Back-project point cloud X0 b Add noise to get the back-projected point cloud after adding noise at the t-th iteration by Render the back-projected point cloud for radius X t b Mask S t ,in α t =1-β t , α t is the variance constant of the iterative noise addition in the t-th step, α i is the variance constant of the noise added in the iterative process of step i, is the cumulative value of the variance constant of the iterative noise addition in the t-th step, β t is the variance constant of the sampling noise in the t-th step; r render,t is the noise rendering radius of the t-th step iteration, and η is the rendering radius scale coefficient;

[0040] 2) Mask S t Discrete into 2D point cloud X t Projected into 2D point cloud Y t , calculate two and Y t Chamfer distance between:

[0041]

[0042] in, Point Cloud The i-th point in y t,j is the point cloud Y t The i-th point in, N is the point cloud X t b The number of points;

[0043] Add the normalization coefficient to get the final guide function value Where s is the adjustment coefficient; z and f are the distance from the camera optical center to the back-projection plane and the camera focal length respectively; z is the three-dimensional coordinate t of the center point of the three-dimensional segmentation frame of the object in the seven-dimensional pose estimation result. O2C The z-axis component value in , f is the focal length of the left camera f lx The value of L CD,l The chamfer distance calculated for the left camera, L CD,r The chamfer distance calculated for the right camera.

[0044] A second embodiment of the present invention provides a transparent object measurement device based on a guided generative point cloud diffusion model, comprising:

[0045] A binocular image acquisition module is used to acquire binocular images of the transparent object to be measured using a binocular camera;

[0046] A mask and pose estimation module is used to input the binocular image into a preset mask segmentation network and a preset pose estimation network respectively to obtain a binocular mask and a 7-dimensional pose estimation result of the transparent object to be measured;

[0047] A back-projection point cloud and edge point cloud generation module, configured to calculate the back-projection point cloud and the binocular edge point cloud of the transparent object to be measured based on the binocular mask;

[0048] A target point cloud generation module is used to obtain the target point cloud of the transparent object to be measured based on the back-projected point cloud and the binocular edge point cloud, using a preset generative point cloud diffusion model and performing iterative noise reduction;

[0049] The measurement module is used to transform the target point cloud into the camera coordinate system of the binocular camera based on the 7-dimensional pose estimation result, and the measurement of the transparent object to be measured is completed.

[0050] A third embodiment of the present invention provides an electronic device, including:

[0051] at least one processor; and a memory communicatively coupled to the at least one processor;

[0052] The memory stores instructions that can be executed by the at least one processor, and the instructions are configured to execute the above-mentioned transparent object measurement method based on the guided generative point cloud diffusion model.

[0053] A fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions for enabling the computer to execute the aforementioned transparent object measurement method based on a guided generative point cloud diffusion model.

[0054] The characteristics and beneficial effects of the present invention are:

[0055] The present invention can be applied to existing binocular camera equipment. By calculating the guidance function based on projection, noise addition, and back-projection, and using the edge point cloud fusion method, the estimated 7-dimensional pose is used to measure transparent objects. The present invention can use inaccurate 7-dimensional pose to generate a complete and accurate transparent object surface point cloud, thereby more accurately measuring transparent objects. In view of the weak generalization characteristics of the transparent object surface reconstruction algorithm based on completion, the present invention realizes cross-dataset generalized surface reconstruction based on the generative diffusion model of category-level object geometry prior. The cross-dataset generalization of the present invention reduces the use cost caused by additional training in actual application scenarios and improves the use efficiency in actual application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 This is an overall flow chart of a transparent object measurement method based on a guided generative point cloud diffusion model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0057] The present invention proposes a transparent object measurement method and device based on a guided generative point cloud diffusion model. The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0058] The first embodiment of the present invention provides a transparent object measurement method based on a guided generative point cloud diffusion model, comprising:

[0059] Use a binocular camera to collect binocular images of the transparent object to be measured;

[0060] Inputting the binocular image into a preset mask segmentation network and a pose estimation network respectively to obtain a binocular mask and a 7-dimensional pose estimation result of the transparent object to be measured;

[0061] Based on the binocular mask, calculating the back-projected point cloud of the transparent object to be measured and the binocular edge point cloud;

[0062] Based on the back-projected point cloud and the binocular edge point cloud, a preset generative point cloud diffusion model is used to obtain a target point cloud of the transparent object to be measured through iterative noise reduction;

[0063] Based on the 7-dimensional pose estimation result, the target point cloud is transformed into the camera coordinate system of the binocular camera, and the measurement of the transparent object to be measured is completed.

[0064] In a specific embodiment of the present invention, the transparent object measurement method based on the guided generative point cloud diffusion model has the following overall process: Figure 1 As shown, the following steps are included:

[0065] 1) Use a binocular camera to collect binocular images of the transparent object to be measured.

[0066] In this embodiment, the collected binocular images are respectively denoted as I l ,I r , where I l Represents the left view of the binocular camera, I r This embodiment has no special requirements on the binocular measurement equipment and the transparent object to be measured.

[0067] In a specific embodiment of the present invention, a RealSense D415 binocular camera is used, and the camera resolution is 1280×960.

[0068] Furthermore, before image acquisition, the binocular camera needs to be calibrated to obtain the camera's intrinsic parameter matrix:

[0069]

[0070] Among them, K l , K r Represents the intrinsic parameter matrix of the left and right cameras respectively, f lx 、f rx Represents the x-axis focal length of the left and right cameras, f ly 、f ry Respectively represent the focal length of the left and right camera y axis, u l0 、u r0 Respectively represent the coordinates of the left and right camera principal points on the x-axis, v l0 、v r0 Represent the coordinates of the left and right camera principal points on the y-axis respectively.

[0071] 2) Inputting the binocular image obtained in step 1) into a preset mask segmentation network to obtain a binocular mask of the transparent object to be measured; inputting the binocular image obtained in step 1) into a preset pose estimation network to obtain a 7-dimensional pose estimation result of the transparent object to be measured.

[0072] In this embodiment, the binocular image I obtained in step 1) is l ,I r Input to the existing mask segmentation network F seg , obtain the binocular left mask S of the transparent object to be measured 0,l and binocular right mask S 0,r ; The binocular image I l ,I r Input to the existing pose estimation network F pose , obtain the 7-dimensional pose estimation result R of the transparent object to be measured O2C ,t O2C and s, where R O2C Represents the three-dimensional rotation matrix of the object, t O2C Represents the 3D coordinates of the center point of the 3D segmentation box of the object, and s represents the scale of the diagonal of the 3D segmentation box.

[0073] In a specific embodiment of the present invention, the prompt points are manually input and the SegmentAnything model is used as the mask segmentation network F seg Segment transparent object mask. Use LaPose pose model as pose estimation network F pose Estimate the 7D pose of transparent objects.

[0074] 3) Using the binocular mask obtained in step 2), calculate the back-projected point cloud of the transparent object to be measured and the binocular edge point cloud; the specific steps are as follows:

[0075] 3-1) The left and right masks S of the transparent object to be measured 0,l ,S 0,r Discretize and back-project to a plane perpendicular to the camera optical axis and passing through the center of the 3D segmentation frame of the transparent object to be measured to obtain the back-projected point cloud X0 b .

[0076] 3-2) From S 0,l The leftmost column and S 0,r In the rightmost column, points in the same row are selected to form a pair of matching point pairs. Using all matching point pairs, the binocular edge point cloud X is calculated according to the triangulation method. r .

[0077] 4) Obtain the Gaussian point cloud of the t-step iteration in the denoising process by random sampling and record it as X t In a specific embodiment of the present invention, the initial value of t is 1000.

[0078] 5) Using the back-projected point cloud obtained in step 3), the Gaussian point cloud X of the t-th iteration is transformed into t Input the preset generative point cloud diffusion model, and the network outputs the denoised partial target point cloud of the t-1th iteration.

[0079] In this embodiment, a generative point cloud diffusion model μ containing category priors is used. θ (X t ,t), the input of the network is X t , and then X t Perform denoising operations and calculate the projection loss function to guide the denoising process to generate part of the target point cloud:

[0080]

[0081] Among them, X t-1 unknown is the denoised partial target point cloud of the t-1th iteration, σ t is the variance constant of the t-th step, which can be any real number. In this embodiment, it is β t , β t S is the variance constant of the sampling noise in the t-th step. t,l ,S t,r are the binocular left mask and binocular right mask of the transparent object to be measured in the t-th iteration, ε2 is the second noise sampled from the normal distribution, and F guide is the bootstrap function.

[0082] Furthermore, in this embodiment, in the Gaussian point cloud X of the t-th iteration t Before inputting the preset generative point cloud diffusion model, the method of this embodiment further includes:

[0083] The generative point cloud diffusion model is trained.

[0084] In a specific embodiment of the present invention, the training of the generative point cloud diffusion model includes:

[0085] According to the category of transparent objects to be measured, data of the same category are selected from the ShapeNet dataset to form a training set. Then, the DDPM generative point cloud diffusion model μ is constructed using the Point-Voxel CNN network structure. θ (X t ,t), and use the training set to train the network. The ShapeNet dataset is a public dataset that contains a total of 3 million data instances, of which 3135 data subsets are divided according to categories. Each category of data subset contains a surface point cloud of an object of a specific category, and a single training data instance contains the surface point cloud of the object. In this embodiment, the data subset of the mug category is used as the training set for reconstructing a transparent mug. This subset contains 497 point cloud instances. This embodiment uses the AdamW optimizer to train the diffusion network μ θ (X t,t), the learning rate is set to 0.001. The batch size is set to 16 during training, and a total of 100,000 steps are trained.

[0086] Furthermore, in this embodiment, the guide function F is calculated guide Guide the denoising process to generate part of the target point cloud, F guide The specific calculation steps are as follows:

[0087] 5-1) Back-project point cloud X0 b Add noise to get the back-projected point cloud after adding noise at the t-th iteration by Render the back-projected point cloud for radius X t b Mask S t ,in α t =1-β t , α t is the variance constant of the iterative noise addition in the t-th step, α i is the variance constant of the noise added in the iterative process of step i, is the cumulative value of the variance constant of the iterative noise addition in the t-th step, β t is the variance constant of the iterative sampling noise in the t-th step. In a specific embodiment of the present invention, β t The value of is 1*10 at t=0 -5 , becomes 0.008 when t=1000, and β at the middle step t The value of is a linear interpolation between the two. render,t is the noise rendering radius of the iterative noise addition in step t, and η is the rendering radius scale coefficient, which is 1.5 in this embodiment.

[0088] 5-2) Mask S t Discrete into 2D point cloud At the same time, X t Projected into 2D point cloud Y t , calculate two and Y t Chamfer distance between:

[0089]

[0090] in, Point Cloud The i-th point in y t,j is the point cloud Y t The i-th point in, N is the point cloud X t b points.

[0091] Add the normalization coefficient to get the final guide function value Where s is the adjustment coefficient, and its value range is 100 to 150. In this embodiment, s = 120. z and f are the distance from the camera optical center to the back-projection plane and the camera focal length, respectively. z is the translation component t of the seven-dimensional pose. O2C The z-axis component value in f is the focal length of the left camera. lx The value of L CD,l The chamfer distance calculated for the left camera, L CD,r The chamfer distance calculated for the right camera.

[0092] 6) Add noise to the binocular edge point cloud obtained in step 3) and fuse it with the denoised partial target point cloud obtained in step 5) to obtain the denoised target point cloud in step t-1. The specific steps are as follows:

[0093] 6-1) To X r Add noise to get the noised partial target point cloud of the t-1th iteration:

[0094]

[0095] 6-2) Fuse the results of step 6-1) and step 5) to obtain the denoised target point cloud of the t-1th iteration:

[0096] X t-1 =concat(X t-1 known ,X t-1 unknown [N r +1:N total ])

[0097] Among them, ε1, ε2 are the first noise and the second noise sampled from two independent normal distributions, N r Point Cloud X t-1 known The number of points, N total Point Cloud X t-1 points.

[0098] 7) Determine t:

[0099] If t>1, set t=t-1 and return to step 5);

[0100] Otherwise, the target point cloud X0 is generated and goes to step 8).

[0101] 8) Based on the estimated 7-dimensional pose estimation result obtained in step 2), the target point cloud obtained in step 6) is transformed into the camera coordinate system to complete the measurement of the transparent object to be measured.

[0102] In this embodiment, the denoised target point cloud X0 of the 0th iteration is transformed into the camera coordinate system through the 7-dimensional pose estimated by LaPose. The transformed point cloud is the measurement result of the transparent object, where the transformation matrix is

[0103] To implement the above embodiment, a second embodiment of the present invention provides a transparent object measurement device based on a guided generative point cloud diffusion model, comprising:

[0104] A binocular image acquisition module is used to acquire binocular images of the transparent object to be measured using a binocular camera;

[0105] A mask and pose estimation module is used to input the binocular image into a preset mask segmentation network and a preset pose estimation network respectively to obtain a binocular mask and a 7-dimensional pose estimation result of the transparent object to be measured;

[0106] A back-projection point cloud and edge point cloud generation module, configured to calculate the back-projection point cloud and the binocular edge point cloud of the transparent object to be measured based on the binocular mask;

[0107] A target point cloud generation module is used to obtain the target point cloud of the transparent object to be measured based on the back-projected point cloud and the binocular edge point cloud, using a preset generative point cloud diffusion model and performing iterative noise reduction;

[0108] The measurement module is used to transform the target point cloud into the camera coordinate system of the binocular camera based on the 7-dimensional pose estimation result, and the measurement of the transparent object to be measured is completed.

[0109] It should be noted that the aforementioned explanation of the embodiment of a transparent object measurement method based on a guided generative point cloud diffusion model is also applicable to a transparent object measurement device based on a guided generative point cloud diffusion model in this embodiment, and will not be repeated here. According to an embodiment of the present invention, a transparent object measurement device based on a guided generative point cloud diffusion model is proposed, which uses a binocular camera to collect a binocular image of the transparent object to be measured; the binocular image is input into a preset mask segmentation network and a pose estimation network respectively to obtain a binocular mask and a 7-dimensional pose estimation result of the transparent object to be measured; based on the binocular mask, the back-projection point cloud and the binocular edge point cloud of the transparent object to be measured are calculated; based on the back-projection point cloud and the binocular edge point cloud, a preset generative point cloud diffusion model is used to iteratively reduce noise to obtain a target point cloud of the transparent object to be measured; based on the 7-dimensional pose estimation result, the target point cloud is transformed into the camera coordinate system of the binocular camera, and the measurement of the transparent object to be measured is completed. This makes it possible to generate a complete and accurate point cloud of the transparent object surface using an inaccurate 7-dimensional pose, thereby making the measurement of transparent objects more accurate.

[0110] In a specific embodiment of the present invention, before using a binocular camera to collect a binocular image of the transparent object to be measured, the method further includes:

[0111] Calibrate the binocular camera to obtain the camera's intrinsic parameter matrix:

[0112]

[0113] Among them, K l , K r Represents the intrinsic parameter matrix of the left and right cameras, f lx 、f rx Represents the x-axis focal length of the left and right cameras, f ly 、f ry Respectively represent the focal length of the left and right camera y axis, u l0 、u r0 Respectively represent the coordinates of the left and right camera principal points on the x-axis, v l0 、v r0 Represent the coordinates of the left and right camera principal points on the y-axis respectively.

[0114] In a specific embodiment of the present invention, the mask segmentation network adopts the SegmentAnything model, and the pose estimation network adopts the LaPose pose model.

[0115] In a specific embodiment of the present invention, the step of calculating the back-projected point cloud and the binocular edge point cloud of the transparent object to be measured based on the binocular mask includes:

[0116] 1) The binocular left mask S of the transparent object to be measured 0,l and binocular right mask S 0,r Discretize and back-project to a plane perpendicular to the camera optical axis and passing through the center of the 3D segmentation frame of the transparent object to be measured to obtain the back-projected point cloud X0 b ;

[0117] 2) From S 0,l The leftmost column and S 0,r In the rightmost column, points in the same row are selected to form a pair of matching point pairs. Using all matching point pairs, the binocular edge point cloud X is calculated according to the triangulation method. r .

[0118] In a specific embodiment of the present invention, obtaining the target point cloud of the transparent object to be measured includes:

[0119] 1) Obtain the Gaussian point cloud of the t-step iteration in the denoising process by random sampling, denoted as X t , t﹥1;

[0120] 2) Using the back-projected point cloud, the Gaussian point cloud X of the t-th iteration is transformed t Input the preset generative point cloud diffusion model μ θ (X t ,t), the network outputs the denoised partial target point cloud of the t-1th iteration;

[0121]

[0122] Among them, X t-1 unknown is the denoised partial target point cloud of the t-1th iteration, σ t is the variance constant of the t-th step, S t,l ,S t,r are the binocular left mask and binocular right mask of the transparent object to be measured in the t-th iteration, respectively, F guide is the bootstrap function;

[0123] 3) Adding noise to the binocular edge point cloud and fusing it with the denoised partial target point cloud obtained in step 2) to obtain the denoised target point cloud in step t-1;

[0124] Among them, for X r Add noise to get the noised partial target point cloud of the t-1th iteration:

[0125]

[0126] Then we get the denoised target point cloud of the t-1th iteration:

[0127] X t-1 =concat(X t-1 known ,X t-1 unknown [N r +1:N total ])

[0128] Among them, ε1, ε2 are the first noise and the second noise sampled from two independent normal distributions, N r Point Cloud X t-1 known The number of points, N total Point Cloud X t-1 The number of points;

[0129] 4) Determine t:

[0130] If t>1, set t=t-1 and return to step 2);

[0131] Otherwise, the target point cloud X0 is generated.

[0132] In a specific embodiment of the present invention, in the step t iterative Gaussian point cloud X t Input the preset generative point cloud diffusion model μ θ (X t ,t), also includes:

[0133] training the generative point cloud diffusion model;

[0134] The training of the generative point cloud diffusion model includes:

[0135] 1) constructing a training set, wherein the training samples in the training set include surface point clouds of objects of the same category as the transparent object to be measured and their corresponding labels;

[0136] 2) Constructing a generative point cloud diffusion model μ θ (X t ,t), the model adopts Point-Voxel CNN network;

[0137] 3) Using the training set obtained in step 1), the generative point cloud diffusion model constructed in step 2) is trained to obtain a trained generative point cloud diffusion model.

[0138] In a specific embodiment of the present invention, the F guide The calculation process is as follows:

[0139] 1) Back-project point cloud X0 b Add noise to get the back-projected point cloud after adding noise at the t-th iteration by Render the back-projected point cloud for radius X t b Mask S t ,in α t =1-β t , α t is the variance constant of the iterative noise addition in the t-th step, α i is the variance constant of the noise added in the iterative process of step i, is the cumulative value of the variance constant of the iterative noise addition in the t-th step, β t is the variance constant of the sampling noise in the t-th step; r render,t is the noise rendering radius of the t-th step iteration, and η is the rendering radius scale coefficient;

[0140] 2) Mask S t Discrete into 2D point cloud X t Projected into 2D point cloud Y t , calculate two and Y t Chamfer distance between:

[0141]

[0142] in, Point Cloud The i-th point in y t,j is the point cloud Y t The i-th point in, N is the point cloud X t b The number of points;

[0143] Add the normalization coefficient to get the final guide function value Where s is the adjustment coefficient; z and f are the distance from the camera optical center to the back-projection plane and the camera focal length respectively; z is the three-dimensional coordinate t of the center point of the three-dimensional segmentation frame of the object in the seven-dimensional pose estimation result. O2C The z-axis component value in , f is the focal length of the left camera f lx The value of L CD,l The chamfer distance calculated for the left camera, L CD,r The chamfer distance calculated for the right camera.

[0144] To implement the above embodiment, a third aspect of the present invention provides an electronic device, including:

[0145] at least one processor; and a memory communicatively coupled to the at least one processor;

[0146] The memory stores instructions that can be executed by the at least one processor, and the instructions are configured to execute the above-mentioned transparent object measurement method based on the guided generative point cloud diffusion model.

[0147] To implement the above embodiment, the fourth aspect of the present invention proposes a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the above-mentioned transparent object measurement method based on the guided generative point cloud diffusion model.

[0148] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0149] The computer-readable medium may be included in the electronic device or may exist independently and not incorporated into the electronic device. The computer-readable medium carries one or more programs. When executed by the electronic device, the one or more programs cause the electronic device to perform the transparent object measurement method based on the guided generative point cloud diffusion model described in the above embodiment.

[0150] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0151] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0152] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0153] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0154] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing it in a suitable manner if necessary, and then storing it in a computer memory.

[0155] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0156] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0157] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0158] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A transparent object measurement method based on a guided generative point cloud diffusion model, characterized in that: include: Use a binocular camera to collect binocular images of the transparent object to be measured; Inputting the binocular image into a preset mask segmentation network and a pose estimation network respectively to obtain a binocular mask and a 7-dimensional pose estimation result of the transparent object to be measured; Based on the binocular mask, calculating the back-projected point cloud of the transparent object to be measured and the binocular edge point cloud; Based on the back-projected point cloud and the binocular edge point cloud, a preset generative point cloud diffusion model is used to obtain a target point cloud of the transparent object to be measured through iterative noise reduction; Based on the 7-dimensional pose estimation result, the target point cloud is transformed into the camera coordinate system of the binocular camera, and the measurement of the transparent object to be measured is completed.

2. The method according to claim 1, characterized in that Before using the binocular camera to collect the binocular image of the transparent object to be measured, the method further includes: Calibrate the binocular camera to obtain the camera's intrinsic parameter matrix: Among them, K l , K r Respectively represent the intrinsic parameter matrices of the left and right cameras, f lx 、f rx Represents the x-axis focal length of the left and right cameras, f ly 、f ry Respectively represent the focal length of the left and right camera y axis, u l0 、u r0 Respectively represent the coordinates of the left and right camera principal points on the x-axis, v l0 、v r0 Represent the coordinates of the left and right camera principal points on the y-axis respectively.

3. The method according to claim 1, characterized in that The mask segmentation network adopts the SegmentAnything model, and the pose estimation network adopts the LaPose pose model.

4. The method according to claim 1, wherein The step of calculating the back-projected point cloud of the transparent object to be measured and the binocular edge point cloud based on the binocular mask includes: 1) The binocular left mask S of the transparent object to be measured 0,l and binocular right mask S 0,r Discretize and back-project to a plane perpendicular to the camera optical axis and passing through the center of the 3D segmentation frame of the transparent object to be measured to obtain the back-projected point cloud X0 b ; 2) From S 0,l The leftmost column and S 0,r In the rightmost column, points in the same row are selected to form a pair of matching point pairs. Using all matching point pairs, the binocular edge point cloud X is calculated according to the triangulation method. r .

5. The method according to claim 4, characterized in that The step of obtaining a target point cloud of the transparent object to be measured includes: 1) Obtain the Gaussian point cloud of the t-step iteration in the denoising process by random sampling, denoted as X t , t﹥1; 2) Using the back-projected point cloud, the Gaussian point cloud X of the t-th iteration is transformed t Input the preset generative point cloud diffusion model μ θ (X t ,t), the network outputs the denoised partial target point cloud of the t-1th iteration; Among them, X t-1 unknown is the denoised partial target point cloud of the t-1th iteration, σ t is the variance constant of the t-th step, S t,l ,S t,r are the binocular left mask and binocular right mask of the transparent object to be measured in the t-th iteration, respectively, F guide is the bootstrap function; 3) Adding noise to the binocular edge point cloud and fusing it with the denoised partial target point cloud obtained in step 2) to obtain the denoised target point cloud in step t-1; Among them, for X r Add noise to get the noised partial target point cloud of the t-1th iteration: Then we get the denoised target point cloud of the t-1th iteration: X t-1 =concat(X t-1 known ,X t-1 unknown [N r +1:N total ]) Among them, ε1, ε2 are the first noise and the second noise sampled from two independent normal distributions, N r Point Cloud X t-1 known The number of points, N total Point Cloud X t-1 The number of points; 4) Determine t: If t>1, set t=t-1 and return to step 2); Otherwise, the target point cloud X0 is generated.

6. The method according to claim 5, characterized in that In the Gaussian point cloud X of the t-th iteration t Input the preset generative point cloud diffusion model μ θ (X t ,t), also includes: training the generative point cloud diffusion model; The training of the generative point cloud diffusion model includes: 1) constructing a training set, wherein the training samples in the training set include surface point clouds of objects of the same category as the transparent object to be measured and their corresponding labels; 2) Constructing a generative point cloud diffusion model μ θ (X t ,t), the model adopts Point-Voxel CNN network; 3) Using the training set obtained in step 1), the generative point cloud diffusion model constructed in step 2) is trained to obtain a trained generative point cloud diffusion model.

7. The method according to claim 5, characterized in that The F guide The calculation process is as follows: 1) Back-project point cloud X0 b Add noise to get the back-projected point cloud after adding noise at the t-th iteration by Render the back-projected point cloud for radius X t b Mask S t ,in α t is the variance constant of the iterative noise addition in the t-th step, α i is the variance constant of the noise added in the iterative process of step i, is the cumulative value of the variance constant of the iterative noise addition in the t-th step, β t is the variance constant of the sampling noise in the t-th step; r render,t is the noise rendering radius of the t-th step iteration, and η is the rendering radius scale coefficient; 2) Mask S t Discrete into 2D point cloud X t Projected into 2D point cloud Y t , calculate two and Y t Chamfer distance between: in, Point Cloud The i-th point in y t,j is the point cloud Y t The i-th point in, N is the point cloud X t b The number of points; Add the normalization coefficient to get the final guide function value Where s is the adjustment coefficient; z and f are the distance from the camera optical center to the back-projection plane and the camera focal length respectively; z is the three-dimensional coordinate t of the center point of the three-dimensional segmentation frame of the object in the seven-dimensional pose estimation result. O2C The z-axis component value in , f is the focal length of the left camera f lx The value of L CD,l The chamfer distance calculated for the left camera, L CD,r The chamfer distance calculated for the right camera.

8. A transparent object measurement device based on a guided generative point cloud diffusion model, characterized in that: include: A binocular image acquisition module is used to acquire binocular images of the transparent object to be measured using a binocular camera; A mask and pose estimation module is used to input the binocular image into a preset mask segmentation network and a preset pose estimation network respectively to obtain a binocular mask and a 7-dimensional pose estimation result of the transparent object to be measured; A back-projection point cloud and edge point cloud generation module, configured to calculate the back-projection point cloud and the binocular edge point cloud of the transparent object to be measured based on the binocular mask; A target point cloud generation module is used to obtain the target point cloud of the transparent object to be measured based on the back-projected point cloud and the binocular edge point cloud, using a preset generative point cloud diffusion model and performing iterative noise reduction; The measurement module is used to transform the target point cloud into the camera coordinate system of the binocular camera based on the 7-dimensional pose estimation result, and the measurement of the transparent object to be measured is completed.

9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are configured to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Object depth and normal estimation method and device, electronic equipment and storage medium

    CN121505505A