Device and method for training model for determining shape of object and method for operating computer-controlled machine depending on shape of object
By converting digital images into point cloud and voxel representations and denoising these representations using diffusion models, the problem of determining object shapes from noisy real-time world digital images is solved, and effective processing of incomplete and noisy point cloud data and accurate 3D shape representation generation is achieved.
Patent Information
- Application Number
- JP2024187015
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-24
- Filing Date
- 2024-10-23
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is difficult to effectively determine the shape of an object from digital images of the noisy real-time world, especially when processing incomplete and noisy point cloud data.
The shape model of the object is generated by converting the digital image into point cloud representations and voxel representations and denoising these representations using a diffusion model. The method also includes using pseudo-real depth images and contour images to further train the model to improve its accuracy.
The ability to effectively determine the shape of an object from a noisy real-time world digital image is realized, able to process incomplete and noisy point cloud data, and generate accurate 3D shape representations.
Smart Images

Figure 2025072343000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an apparatus and method for training a model, in particular a neural network, to determine the shape of an object, and to a method for operating a computer-controlled device in response to the shape of an object. Summary of the Invention [Means for solving the problem]
[0002] A computer-implemented method for training a model, in particular a neural network, to determine a shape of an object includes the steps of: determining a first point cloud representation of the object in response to a first digital image, the first point cloud representation comprising points representing a first view of the object; determining a second point cloud representation of the object in response to a second digital image, the second point cloud representation comprising points representing a second view of the object; and determining a first voxel representation of the object in response to the first point cloud representation, the first voxel representation comprising points representing the first view. The method includes: mapping the first voxel representation to a voxel representation of the shape using a model; and providing a ground truth for training the model in response to the first and second point cloud representations or in response to the first and second voxel representations, where the second voxel representation of the object is determined in response to the second point cloud representation, the second voxel representation comprising voxels representing a second view, and the ground truth is a voxel representation comprising voxels of the first and second voxel representations. This provides a model for determining the shape of an object from a noisy real-world digital image.
[0003] The model may be a diffusion model configured to remove noise from a noisy input of the diffusion model, the diffusion model configured to output a voxel representation of a shape, the noisy input having a plurality of input elements, the noisy input including elements representing a first voxel representation of the object and elements representing noise, in particular noise randomly sampled from a distribution, and the elements representing the first voxel representation are unperturbed, in particular do not include additional noise.
[0004] The method may include determining a depth image of the object according to a voxel representation of the shape of the object, providing a pseudo ground truth depth image, and training a model according to a difference between the depth image and the ground truth depth image, which further improves the model.
[0005] The step of providing a pseudo ground truth depth image may include a step of determining a pseudo ground truth depth image of the object depending on the first digital image and the second digital image, in particular a step of mapping the first digital image and the second digital image to the pseudo ground truth depth image of the object using a first artificial neural network, the first artificial neural network being configured to map the first digital image to the pseudo ground truth depth image.
[0006] Providing the pseudo ground truth depth image may include providing training data points that include the first digital image and the pseudo ground truth depth image.
[0007] The method may include determining a silhouette image of the object according to a voxel representation of the shape of the object, providing a ground truth silhouette image, and training a model according to a difference between the silhouette image and the ground truth silhouette image, which further improves the model.
[0008] Providing a ground truth silhouette image may include providing training data points including the first digital image and the ground truth silhouette image.
[0009] A computer implemented method for operating a computer controlled device, in particular a robot, a vehicle, a home appliance, a power tool, a manufacturing device, a personal assistant or an access control system, comprises the steps of capturing a digital image, in particular using a sensor; determining a point cloud representation of an object in response to the digital image, the point cloud representation comprising points representing a view of the object; determining a voxel representation of the object in response to the point cloud representation, the voxel representation comprising voxels representing the view; mapping the voxel representation to a voxel representation of a shape using a model, in particular a neural network; determining a shape in response to the voxel representation of the shape, in particular using an artificial neural network configured for mapping the voxel representation of the shape to a shape; and operating the computer controlled device in response to the shape.
[0010] An apparatus for operating a computer-controlled apparatus or for training a model, in particular a neural network, for operating a computer-controlled apparatus comprises at least one processor and at least one memory, the at least one processor configured to execute instructions that, when executed by the at least one processor, cause the apparatus to perform a method, and the at least one memory containing instructions.
[0011] The computer program comprises computer readable instructions which, when executed by a computer, cause the computer to perform a method.
[0012] Further embodiments can be derived from the following description and drawings. [Brief description of the drawings]
[0013] [Figure 1] 1 shows a schematic diagram of an apparatus; [Diagram 2] 1 illustrates generally steps in a method for training a neural network for determining the shape of an object. [Diagram 3] 2 illustrates generally steps in a method for operating a computer-controlled device in response to the shape of an object; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] 1 illustrates a schematic diagram of an apparatus 100. The apparatus 100 includes at least one processor 102 and at least one memory 104. The at least one memory 104 may include a non-transitory memory. The at least one memory 104 may include a transitory memory.
[0015] The device 100 is configured to process digital images, which may be captured using or received from a sensor 106 or which may be read from at least one memory 104.
[0016] The sensor 106 may be located outside the device 100. The device 100 may include an interface for the sensor 106. The device 100 according to this example includes the sensor 106. The sensor 106 is configured to capture a digital image. The sensor 106 may be a camera, a radar sensor, a lidar sensor, an ultrasonic sensor, an infrared sensor, or a motion sensor. The digital image may be a visual image, a radar image, a lidar image, an ultrasonic image, an infrared image, or a video image.
[0017] According to one example, the digital image has three color channels, e.g., red, green, and blue, of size N. According to one example, the digital image has pixels associated with a depth value indicating the distance between a sensor for capturing the digital image and the object that the pixel represents in the digital image.
[0018] The device 100 is configured to operate a computer-controlled device 108. The device 100 is configured to operate the computer-controlled device 108 in response to digital images. The sensor 106 is configured to capture digital images of an environment of the computer-controlled device 108.
[0019] The computer controlled device 108 may be a robot, a vehicle, an appliance, a power tool, a manufacturing device, a personal assistant, or an access control system.
[0020] At least one processor 102 is configured to execute a method for operating a computer-controlled device 108 .
[0021] The device 100 may be part of a computer-controlled device 108. The device 100 may be configured to operate an actuator 110 for operating the computer-controlled device 108. The actuator 110 may be located external to the device 100. The device 100 may include an interface to the actuator 110. The device 100 in this example includes the actuator 110.
[0022] The actuator 110 may be configured to operate a robot, a vehicle, a power tool, or a piece of manufacturing equipment. The actuator 110 may also be configured to operate a home appliance, a personal assistant, or a part of an access control system.
[0023] The device 100 may be configured to operate the actuator 110 to operate a robot, a vehicle, a power tool, or manufacturing equipment, or a portion thereof. The device 100 may be configured to operate the actuator 110 to operate a home appliance, a personal assistant, or a portion of an access control system.
[0024] The digital image may include at least a portion of the object.
[0025] The apparatus 100 may be configured to determine a shape of the object in response to the digital image. The apparatus 100 may be configured to determine a target position of the computer-controlled apparatus 108, or a part thereof, in response to the shape of the object. The apparatus 100 may be configured to move the computer-controlled apparatus 108, or a part thereof, to the target position.
[0026] For example, the apparatus 100 may be configured to determine a target position that avoids collision with an object, and move a robot or vehicle to the target position that avoids collision with the object.
[0027] The object may be a traffic participant. The digital image may include a first part of the traffic participant without a second part of the traffic participant. The shape of the object may be a shape of the traffic participant including the second part. The target position may be a position that avoids the position of the second part. This means that the target position is determined to avoid a collision with the second part, which is not visible in the digital image.
[0028] The portion of the robot may be configured to grasp the object. For example, the apparatus 100 may be configured to determine a target position for grasping the object and move the portion of the robot configured to grasp the object to the target position to grasp the object.
[0029] The object may be a cup with a handle. The digital image may include a part of the cup without the handle. The shape of the object may be a shape including a handle. The target position may be the position of the handle. This means that the target position is determined to grasp the cup by the handle, which is not visible in the digital image.
[0030] Given a noisy partial real-world point cloud of an object, the goal of our method is to complete its 3D shape by predicting the geometry of the unseen parts using a model.
[0031] The completion task is formulated as a conditional generation problem, which generates a complete shape given an input partial point cloud.
[0032] Shape completion problems are multi-modal in nature. In this example, the model is a denoising diffusion probabilistic model (DDPM).
[0033] The model allows for generating multiple plausible completions for a single incomplete point cloud, while effectively learning category-specific shape priors from only partial real-world data, without considering the complete shape during the training process.
[0034] To overcome the limitation of dealing with noisy observations without relying on training on synthetic data, the method can exploit additional geometric cues, including depth and silhouette information.
[0035] 2 shows diagrammatically the steps in a method for training a model, in particular a neural network 200, for determining a voxel representation 202 of the shape of an object according to a first digital image 204 comprising a first view 206 of the object and a second digital image 208 comprising a second view 210 of the object. According to an example, the voxel representation 202 of the shape is a three-dimensional reconstruction of the shape. The method is not limited to two views. A certain number V of views can be used, V being 2 or greater.
[0036] The method is used to determine a shape representation, i.e. a voxel representation 202, from a point cloud representation. This means that the method uses a transformation from point cloud to voxel grid, for example with a fixed deterministic algorithm, especially in data pre-processing. According to one example, the shape representation is a shape distribution that is learned from real-world observations, i.e. a first digital image 204 and a second digital image 208 of a real-world object.
[0037] Learning shape representations poses significant challenges, primarily due to the presence of significant noise in digital image data resulting from imprecise sensor measurements, occlusions, and errors involved in segmenting objects from 3D scenes or images. The method does not require synthetic data, i.e., does not require clean synthetic shapes of objects.
[0038] The method for training includes a step 212 .
[0039] Step 212 includes determining a first point cloud representation 214 of the object as a function of the first digital image 204. The first point cloud representation 214 includes points that represent the first view 206. The first digital image 204 is a digital image of the real world, i.e., not a synthetic digital image.
[0040] The method for training includes a step 216 .
[0041] Step 216 includes determining a second point cloud representation 218 of the object as a function of the second digital image 208. The second point cloud representation 218 includes points that represent the second view 210. The second digital image 208 is a digital image of the real world, i.e., not a synthetic digital image.
[0042] The first point cloud representation 214 may be, for example, a first noisy point cloud
number
number
number
number
[0043] The method for training includes a step 220 .
[0044] Step 220 includes determining a first voxel representation 222 of the object as a function of the first point cloud representation 214. The first voxel representation 222 includes voxels representing the first view 206.
[0045] The method for training includes a step 224 .
[0046] Step 224 includes determining a second voxel representation 226 of the object as a function of the second point cloud representation 218. The second voxel representation 226 includes voxels representing the second view 210.
[0047] To simplify 3D data processing, e.g. noisy point clouds
number
number
number
[0048] According to one example, the first voxel representation 222 may be
number
number
[0049] The method for training includes a step 228 .
[0050] Step 228 involves mapping the first voxel representation 222 to the voxel representation of the shape 202 using the model.
[0051] According to one example, the model comprises a diffusion model f configured to remove noise from a noisy input of the diffusion model. θ In this example, the diffusion model f θ is configured to output a voxel representation 202 of the shape.
[0052] The noisy input according to one example has a number of input elements. The noisy input according to one example includes elements representing the first voxel representation 222 of the object and elements representing noise. The noise is, for example, random noise sampled from a distribution. The elements representing the first voxel representation 222 are not perturbed, i.e., no noise is added to the elements representing the first voxel representation 222.
[0053] According to one example, the diffusion model f θ The forward pass of is a distributed schedule β1, ..., β T Unoccupied voxels at T time steps according to
number
number
number
number
[0054] where N is a normal distribution, I is an identity matrix of appropriate dimensions, and the forward process is independent of the adjustment factor c0 that represents the occupied voxels.
[0055] According to one example, the backward process is defined as a Markov chain with learned Gaussian transitions that aims to iteratively remove the noise added in the forward process. The backward process is adjusted for an adjustment factor c0, which in this example represents the occupied voxels.
number
number
[0056] Where:
number
number
[0057] The method for training includes a step 230 .
[0058] Step 230 includes providing a training ground truth 232. The ground truth 232 in this example is a voxel representation, for example an occupancy grid that includes voxels.
[0059] According to one example, the ground truth 232 is determined as a function of the first point cloud representation 214 and the second point cloud representation 218. For example, the first point cloud representation 214, e.g., the first noisy point cloud
number
number
number
[0060] According to one example, the ground truth 232 is determined as a function of the first voxel representation 222 and the second voxel representation 226 .
[0061] According to one example, the voxels of the first voxel representation 222 and the second voxel representation 226 are concatenated into the ground truth 232 voxel representation.
[0062] The method for training includes a step 234 .
[0063] Step 234 involves determining a depth image 236 of the object according to the voxel representation 202 of the object's shape.
[0064] For example, the depth image 236 includes a depth map that is rendered using an occupancy grid of the voxel representation 202 of the shape.
number
[0065] We determine the depth map using a pre-trained Omnidata model, as described, for example, in “Multimodal shape completion via conditional generative adversarial networks” by Rundi Wu, Xuelin Chen, Yixin Zhuang, and Baoquan Chen, Proceedings of the European Conference on Computer Vision (ECCV), 2020.
[0066] Step 234 involves determining a silhouette image 238 of the object according to the voxel representation 202 of the object's shape.
[0067] This means that a two-dimensional silhouette image 238 is determined, for example, using a volume rendering of a three-dimensional reconstruction of the shape.
[0068] Therefore, to render a silhouette pixel, M points are sampled along the ray r. i For each of the occupancy values
number
number
number
number
number
[0069] Since two views v1 and v2 are used to form the ground truth occupancy grid, the silhouettes are also rendered for similar viewpoints.
[0070] The method for training includes a step 240 .
[0071] Step 240 includes providing a pseudo ground truth depth image 242 of the object in response to the first digital image 204. The first digital image 204 may be mapped to the pseudo ground truth depth image 242 of the object using a first artificial neural network. The first artificial neural network may be configured to map the digital image to the depth image.
[0072] Step 240 includes providing a ground truth silhouette image 244 of the object in response to the first digital image 204. The first digital image 204 may be mapped to the ground truth silhouette image 244 of the object using a second artificial neural network. The second artificial neural network may be configured to map the digital image to the silhouette image.
[0073] According to one example, a first digital image 204, a second digital image 208, a simulated ground truth depth image 242, and a ground truth silhouette image 244 are provided as data points of the training data.
[0074] The training data may include multiple data points.
[0075] The method steps for training may be performed for a number of data points, and for each data point, a number of voxel representations 202, a number of depth images 236, and a number of silhouette images 238 associated with the respective data point are determined.
[0076] The method for training includes a step 244 .
[0077] Step 244 involves determining parameters of the model, e.g. weights of the neural network 200, as a function of the loss. The loss in this example depends on the difference between the voxel representation of the shape 202 and the ground truth 232, the difference between the depth image 236 and the simulated ground truth depth image 242, and the difference between the silhouette image 238 and the ground truth silhouette image 244, determined for at least one pair of the first digital image 204 and the second digital image 208.
[0078] Diffusion Model
number
number
number
number
[0079] In this example, the object shape silhouette is constrained to match the ground truth silhouette.
[0080] Diffusion Model
number
number
number
number
[0081] Although individual silhouette images may lack pertinent information, the presence of multiple silhouette images belonging to distinct instances of similar shapes represents an advantage in learning object shapes.
[0082] Learning of object shape is further improved by depth images 236, e.g., depth maps, and pseudo ground truth depth images 242.
[0083] Diffusion Model
number
number
number
number
number
number
[0084] The parameters w and q can be solved using least squares optimization.
[0085] Diffusion Model
number
[0086] The combining loss may include a weighting factor for a loss term in the combining loss.
[0087] FIG. 3 illustrates generally steps in a method for operating a computer-controlled device 108 in response to the shape of an object.
[0088] The method for operating a computer controlled device 108 includes step 302 .
[0089] Step 302 involves capturing a digital image 304 using, for example, the sensor 106 .
[0090] The method for operating a computer controlled device 108 includes a step 306 .
[0091] Step 306 includes determining a point cloud representation 308 of the object according to the digital image 304. The point cloud representation 308 includes points that represent a view 310 of the object.
[0092] The method of operating the computer controlled device 108 includes step 312 .
[0093] Step 312 includes determining a voxel representation 314 of the object according to the point cloud representation 308. The voxel representation 314 includes voxels that represent the view 310.
[0094] The method for determining the shape includes a step 316 .
[0095] Step 316 involves mapping the voxel representation 314 to a voxel representation 318 of the shape using a model, for example the neural network 200 .
[0096] The method for determining the shape includes a step 320 .
[0097] Step 320 involves determining a shape as a function of the voxel representation of the shape 318. Figure 3 shows a perspective view of the shape 322. The shape is determined, for example, using an artificial neural network configured to map the voxel representation to a shape.
[0098] The method for determining the shape includes a step 324 .
[0099] Step 324 includes operating the computer controlled device 108 in response to the shape.
[0100] For example, a target position for avoiding collision with an object is determined, and the robot or vehicle is moved so as to avoid the target position for avoiding collision with the object.
[0101] For example, the object is a traffic participant. The digital image 304 includes a first part of the traffic participant without a second part of the traffic participant. The shape of the object is the shape of the traffic participant including the second part. The target position may be a position that avoids the position of the second part. This means that the target position is determined to avoid a collision with the second part, which is not visible in the digital image.
[0102] A portion of the robot may be configured to grasp an object, for example, a target position for grasping the object is determined, and the portion of the robot configured to grasp the object is moved to the target position to grasp the object.
[0103] The object may be a cup with a handle. The digital image 304 may include a part of a cup without the handle. The shape of the object may be a shape including a handle. The target position may be the position of the handle. This means that according to the method the target position is determined to grasp the cup by the handle, which is not visible in the digital image.
[0104] According to one example, the at least one memory 104 includes instructions executable by the at least one processor 102 and that, when executed by the at least one processor 102, cause the device 100 to perform a method for training a model, e.g., the neural network 200, and / or a method for operating a computer-controlled device 108.
[0105] The method for operating the computer controlled device 108 may include training a model, such as a neural network, according to a method for training a model.
Claims
1. A computer-implemented method for training a model, in particular a neural network (200), for determining the shape of an object, comprising: The method comprises: determining (212) a first point cloud representation (214) of the object in response to a first digital image (204), the first point cloud representation (214) including points representing a first view (206) of the object; determining (216) a second point cloud representation (218) of the object in response to a second digital image (208), the second point cloud representation (218) including points representing a second view (210) of the object; determining (220) a first voxel representation (222) of the object in response to the first point cloud representation (214), the first voxel representation (222) including voxels representing the first view (206); mapping (228) the first voxel representation (222) to a voxel representation (202) of the shape using the model; providing (230) a ground truth (232) for training the model in response to the first point cloud representation (214) and the second point cloud representation (218) or in response to the first voxel representation (222) and the second voxel representation (226), wherein the second voxel representation (226) of the object is determined in response to the second point cloud representation (218), the second voxel representation (226) comprising voxels representing the second view (210), and the ground truth (232) is a voxel representation comprising voxels of the first voxel representation (222) and the second voxel representation (226); The method of claim 1, further comprising:
2. the model is a diffusion model configured to remove noise from a noisy input of the diffusion model; the diffusion model is configured to output the voxel representation (202) of the shape; the noisy input having a plurality of input elements; the noisy input comprises an element representing the first voxel representation (222) of the object and an element representing noise, in particular noise randomly sampled from a distribution; the elements representing the first voxel representation (222) are undisturbed, in particular do not contain additional noise; 2. The method according to claim 1 .
3. The method comprises: determining (232) a depth image (236) of the object as a function of the voxel representation (202) of the shape of the object; Providing (240) a pseudo ground truth depth image (242); training the model as a function of differences between the depth image (236) and the pseudo ground truth depth image (242); Including, 3. The method according to claim 1 or 2.
4. The step of providing (240) a pseudo ground truth depth image (242) comprises: determining the pseudo ground truth depth image (242) of the object as a function of the first digital image (204), in particular mapping the first digital image (204) to the pseudo ground truth depth image (242) of the object using a first artificial neural network, the first artificial neural network is configured to map the first digital image (204) to the pseudo ground truth depth image (242); 4. The method according to claim 3 .
5. The step of providing (240) a pseudo ground truth depth image (242) comprises: providing training data points including the first digital image (204) and the pseudo ground truth depth image (242); 4. The method according to claim 3 .
6. The method comprises: determining (232) a silhouette image (238) of the object in response to the voxel representation (202) of the shape of the object; Providing (240) a ground truth silhouette image (244); training the model as a function of differences between the silhouette image (238) and the ground truth silhouette image (244); Including, 6. The method according to claim 1, wherein the first and second electrodes are connected to a first electrode.
7. The step of providing (240) a ground truth silhouette image (244) comprises: providing training data points including the first digital image (204) and the ground truth silhouette image (244); 7. The method according to claim 6 .
8. 1. A computer-implemented method for operating a computer-controlled device, in particular a robot, a vehicle, a home appliance, a power tool, a manufacturing device, a personal assistant, or an access control system, comprising: In particular, capturing (302) a digital image (304) using a sensor (106); determining (306) a point cloud representation (308) of an object in response to the digital image (304), the point cloud representation (308) including points representing a view (310) of the object; determining (312) a voxel representation (314) of the object in response to the point cloud representation (308), the voxel representation (314) including voxels representing the view (310); Mapping (316) said voxel representation (314) onto a voxel representation (318) of said shape using a model, in particular a neural network (200); determining (320) the shape in response to the voxel representation (318) of the shape using an artificial neural network configured in particular to map the voxel representation (318) of the shape onto the shape; operating (322) the computer controlled device (108) in response to the shape; A method comprising:
9. An apparatus (100) for operating a computer-controlled device (108) or for training a model, in particular a neural network (200), for operating a computer-controlled device (108), comprising: The device (100) comprises: At least one processor (102); At least one memory (104); Equipped with The at least one processor (102) is configured to execute instructions that, when executed by the at least one processor (102), cause the apparatus (100) to perform a method according to any one of claims 1 to 8; The at least one memory (104) includes the instructions. An apparatus comprising:
10. A computer program comprising computer readable instructions which, when executed by a computer, causes the computer to carry out a method according to any one of claims 1 to 8.