Systems and methods for 3D object reconstruction using neural networks
Patent Information
- Application Number
- CN202180051105.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-25
- Filing Date
- 2021-08-13
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2041-08-13
Smart Images

Figure CN115885315B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to three-dimensional scanning technology, and more specifically, to three-dimensional scanning technology utilizing neural networks. Background Technology
[0002] 3D scanning technology can construct 3D models of the surfaces of physical objects. Applications of 3D scanning span many fields, including industrial design and manufacturing, computer animation, science, education, medicine, art, and design. Summary of the Invention
[0003] This disclosure relates to 3D scanning technology. One method of 3D scanning uses so-called "structured light," in which a projector projects a known pattern of light onto the surface of an object (hereinafter referred to as the "projected pattern"). For example, light from the projector can be guided through a glass slide with a printed pattern. The shape of the object's surface is inferred from distortions in the light pattern captured by a camera. One or more cameras can be used to obtain a reflected image of the pattern on the object. By measuring the position of the pattern in the image (e.g., measuring the distortion of the pattern), a computer system can determine the position on the object's surface using simple geometric calculations (such as, for example, triangulation algorithms).
[0004] To determine the position on an object's surface, a computer system needs to know which point within the projected pattern corresponds to which point in the image. According to some embodiments, a trained neural network can be used to infer the correspondence between image pixels and the coordinates of the projected pattern.
[0005] According to some embodiments, a method is provided for resolving ambiguity of imaging elements in a structured light 3D scanning method. The method includes obtaining an image of an object. The image includes a plurality of imaging elements of an imaging pattern. The imaging pattern corresponds to a projection pattern projected onto a surface of the object, and the projection pattern includes a plurality of projection elements. The method further includes using a neural network to output a correspondence between the plurality of imaging elements and the plurality of projection elements. The method further includes using the correspondence between the plurality of imaging elements and the plurality of projection elements to reconstruct the shape of the surface of the object.
[0006] According to some embodiments, a method is provided for determining the correspondence between a projected pattern and an image of the projected pattern incident on a surface of an object. The method includes acquiring an image of the object while projecting the pattern onto its surface. The method further includes using a neural network to output the correspondence between corresponding pixels in the image and the coordinates of the projected pattern. The method also includes using the correspondence between the corresponding pixels in the image and the coordinates of the projected pattern to reconstruct the shape of the object's surface.
[0007] According to some embodiments, a method for training a neural network is provided. The neural network is trained using simulated data comprising multiple simulated images of a projection pattern projected onto the surface of a simulated object. The projection pattern includes multiple projection elements, and each of the simulated images includes a simulated pattern comprising multiple simulated elements. The multiple simulated elements correspond to corresponding projection elements of the projection pattern. The simulated data also includes data indicating the shape of the corresponding simulated object and data indicating the correspondence between the simulated elements and the corresponding projection elements. Using the simulated data, the neural network is trained to determine the correspondence between the multiple projection elements of the projection pattern and the multiple simulated elements of the simulated pattern. The trained neural network is stored for subsequent image reconstruction.
[0008] According to some embodiments, another method for training a neural network is provided. The method includes generating simulated data comprising: multiple simulated images of a projection pattern projected onto the surface of a corresponding simulated object; data indicating the shape of the corresponding simulated object; and data indicating the correspondence between corresponding pixels in the simulated images and coordinates of the projection pattern. The method also includes using the simulated data to train the neural network to determine the correspondence between the images and the projection pattern. The method further includes storing the trained neural network for subsequent use in reconstructing images.
[0009] According to some embodiments, a computer system is provided. The computer system includes one or more processors and a memory storing instructions for performing any of the methods described herein.
[0010] According to some embodiments, a non-transitory computer-readable storage medium is provided that stores instructions. The non-transitory computer-readable storage medium includes instructions that, when executed by a computer system, cause the computer system to perform any of the methods described herein. Attached Figure Description
[0011] To better understand the various embodiments described, reference should be made to the following description of embodiments, in conjunction with the accompanying drawings, wherein similar reference numerals refer to corresponding parts throughout the drawings.
[0012] Figure 1A –1B illustrates an imaging system according to some embodiments.
[0013] Figure 1C A projection pattern according to some embodiments is shown.
[0014] Figure 1D An imaging pattern according to some embodiments is shown.
[0015] Figure 2 This is a block diagram of an imaging system according to some embodiments.
[0016] Figure 3 This is a block diagram of a remote device according to some embodiments.
[0017] Figure 4 The inputs and outputs of a neural network according to some embodiments are shown.
[0018] Figure 5A –5C shows a flowchart of a 3D reconstruction method according to some embodiments.
[0019] Figure 6A –6B shows a flowchart of a method for training a neural network according to some embodiments.
[0020] Figure 7 A flowchart of another 3D reconstruction method according to some embodiments is shown. Detailed Implementation
[0021] Examples of the embodiments are now illustrated in the accompanying drawings. Numerous specific details are set forth in the following description to provide a thorough understanding of the various embodiments described. However, it will be apparent to those skilled in the art that the corresponding embodiments described can be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure the inventive aspects of the embodiments.
[0022] Figure 1A –1B illustrates a three-dimensional (“3D”) imaging environment 100 according to an embodiment of the present invention, which includes a projector 110 and one or more cameras 112 (e.g., sensors). Note that in various embodiments, more than one projector and / or more than one camera may be used. Figure 1A As shown, projector 110 is configured to project a projection pattern (sometimes referred to as "structured illumination") onto an object 120 to be imaged. For this purpose, in some embodiments, light from the projector is projected through a glass slide printed with the projection pattern. The projection pattern includes multiple projection elements. Non-limiting examples of the projection pattern include a series of contrast lines (e.g., black and white lines), a series of contrast zigzag lines, and a grid pattern of dots. Further examples of the projection pattern are described in U.S. Application No. 11 / 846,494, which is hereby incorporated by reference in its entirety.
[0023] Rays 190-1 to 190-4 each correspond to a corresponding projection element of the projection pattern (e.g., a different line in the projection pattern). For example, ray 190-1 represents a first projection element in the projection pattern projected from projector 110 onto surface 121 of object 120, while ray 190-2 represents another projection element projected from projector 110 onto surface 121 of object 120. Rays 190 are reflected at surface 121 of object 120 (as reflected rays 192-1 to 192-4, each corresponding to ray 190-1 to 190-4 respectively). At least a portion of the light is captured by one or more cameras 112.
[0024] In some embodiments, while the surface 121 of object 120 is illuminated with a projection pattern, camera 112 captures multiple images of object 120. In some embodiments, the projection pattern is strobe-illuminated onto the surface of object 120, and multiple images are captured each time the projection pattern is illuminated onto the surface of object 120. As used herein, the term "strobe" refers to repeating at a fixed rate (e.g., 15 frames per second).
[0025] It should be noted that although the projector 110 and the camera 112 are in Figure 1A-1B The images are shown separately, but in some embodiments, the projector 110 and camera 112 are integrated into a single housing as a 3D scanner 200. Figure 2 Users of the 3D scanner 200 can scan an object and collect data by moving the 3D scanner 200 relative to the object 120. Therefore, in some embodiments, images of the object 120 are captured by the camera 112 from different angles or positions.
[0026] Each of the multiple images shows an imaging pattern corresponding to the projection pattern, which is distorted due to the surface of object 120. Therefore, the imaging pattern includes multiple imaging elements, each corresponding to a corresponding projection element in the projection pattern.
[0027] Figure 1B Another example of the operation of the 3D imaging environment 100 is shown. For example... Figure 1B As shown, the projector 110 will be with Figure 1A The same projection pattern is projected toward object 120. In this example, another object 122 is positioned in the path of ray 190-3, so ray 190-3 is reflected at the surface of object 122. Note that object 122 can be a component of object 120 (e.g., the handle of a drinking cup) or a separate object. Camera 112 captures multiple images of objects 120 and 122. As shown, the presence of object 122 causes ray 192-3 (corresponding to ray 190-3) to reflect onto the surface of object 122. Figure 1AThe light is incident on camera 112 at different positions. Therefore, multiple images captured by camera 112 will display the imaging elements corresponding to rays 190-3 and 192-3 at different positions in the imaging pattern (and possibly in different orders) relative to other imaging elements in the imaging pattern.
[0028] Figure 1C An example of a projected pattern 130 emitted from projector 110 is shown, and Figure 1D An example of an imaging pattern 132 captured by camera 112 is shown.
[0029] exist Figure 1C In the example shown, the projection pattern 130 includes a plurality of projection elements 140. For example, Figure 1C The projected pattern 130 shown can be projected onto objects 120 and 122 by projector 110, as follows: Figure 1B As shown. In this case, projection element 140 is a repeating line. In some embodiments, as discussed in U.S. Application 11 / 846,494, these lines have alternating thick and thin regions. In some embodiments, the alternating thick and thin regions between lines are different, but these lines are still considered non-coded elements.
[0030] Figure 1D This is an example of an image pattern 132 captured by a camera 112 while the projector 110 projects the projection pattern 130 toward objects 120 and 122, such as... Figure 1B As shown. The image includes multiple imaging elements 142. In this case, the imaging elements 142 are distorted lines. Each of the imaging elements 142 (e.g., distorted lines) corresponds to a corresponding projection element 140 (e.g., a corresponding line) in the projection pattern 130. In this example, imaging element 142-1 in imaging pattern 132 corresponds to projection element 140-1 in projection pattern 130; imaging element 142-2 in imaging pattern 132 corresponds to projection element 140-2 in projection pattern 130; imaging element 142-3 in imaging pattern 132 corresponds to projection element 140-3 in projection pattern 130; and imaging element 142-4 in imaging pattern 132 corresponds to projection element 140-4 in projection pattern 130. Note that due to the geometry of the scanned object (as shown in the reference...) Figure 1B As shown, the imaging elements can be transposed relative to the positions of the corresponding projection elements 140 in the projection pattern 130. For example, the relative order of imaging elements 142-3 and 142-4 is the reverse of that of projection elements 140-3 and 140-4.
[0031] To construct a model of an object's surface using structured light methods, a computer system needs to know the correspondence between an image and a projection pattern (e.g., the coordinates of the projection pattern corresponding to each pixel in the image and / or the correspondence between imaging elements and projection elements). Two general approaches can resolve this ambiguity: one utilizes a pattern with coded elements, and the other relies on a pattern with non-coded elements. In a pattern with coded elements, the elements in the pattern possess unique identifying features that allow the computer system to identify each imaging element. In a pattern with non-coded elements, the elements in the pattern (e.g., lines or repeating elements) lack the individual unique features that allow the specific elements of the pattern to be identified in the captured image. For non-coded elements (e.g., lines), additional methods are needed to determine the correspondence between the image and the projection pattern.
[0032] In some specific embodiments, the projection pattern is a non-coded light pattern, such that the projection elements of the projection pattern are non-coded elements. In some embodiments, as will be described in detail below, a neural network is used to determine the correspondence between the projection pattern and an image of an object having the projection pattern projected onto it. In some embodiments, the non-coded light pattern includes structured light patterns, such as lines or other repeating elements.
[0033] although Figure 1C and Figure 1D The illustration shows instances where multiple projection elements 140 in the projection pattern 130 are lines, but it will be understood that projection elements 140 and imaging elements 142 can take any shape. For example, projection elements 140 can be bars, zigzag elements, dots, squares, a series of three small bars, etc. Figures A1 and A2 in Appendix A of U.S. Provisional Application No. 63 / 070,066 provide non-limiting examples of images showing an imaging pattern formed by an object illuminated by a projection pattern, which is incorporated herein by reference in its entirety.
[0034] Figure 2 This is a block diagram of a 3D scanner 200 according to some embodiments. Figure 3 A computer system for a 200 D scanner, or a 3D scanner 200, typically includes a memory 204, one or more processors 202, a power supply 206, a user input / output (I / O) subsystem 208, and one or more sensors 203 (e.g., including a camera 112). Figure 1A-1B The processor 202 executes modules, programs, and / or instructions stored in memory 204 and thereby performs processing operations.
[0035] In some embodiments, processor 202 includes at least one central processing unit. In some embodiments, processor 202 includes at least one graphics processing unit. In some embodiments, processor 302 includes at least one neural processing unit (NPU) for executing the neural networks described herein. In some embodiments, processor 202 includes at least one field-programmable gate array (FPGA).
[0036] In some embodiments, memory 204 stores one or more programs (e.g., instruction sets) and / or data structures. In some embodiments, memory 204 or a non-transitory computer-readable storage medium of memory 204 stores the following programs, modules, and data structures, or subsets or supersets thereof:
[0037] ● Operating system 212, which includes programs for handling various basic system services and for performing hardware-related tasks;
[0038] ● Network communication module 218, which is used to connect the 3D scanner to other computer systems (e.g., remote device 236) via one or more communication networks 250;
[0039] ● User interface module 220, which receives commands and / or inputs from the user through user input / output (I / O) subsystem 208, and provides outputs for presentation and / or display on user input / output (I / O) subsystem 208;
[0040] ● Data processing module 224, which is used to process or preprocess data from sensor 203, including optionally performing method 500 ( Figures 5A-5C ) and / or method 700 ( Figure 7 Any or all of the operations described herein. Alternatively, in various embodiments, any or all data processing may be performed by remote device 236, to which 3D scanner 200 is coupled via network 250;
[0041] ● Data acquisition module 226, which controls the readout of the camera, projector, and sensors; and
[0042] ● Storage device 230, which includes buffers, RAM, ROM and / or other memory for storing data used and generated by 3D scanner 200.
[0043] The modules identified above (e.g., data structures and / or programs, including instruction sets) do not need to be implemented as separate software programs, processes, or modules, and therefore various subsets of these modules can be combined or otherwise rearranged in various embodiments. In some embodiments, memory 204 stores a subset of the modules identified above. Furthermore, memory 204 may store additional modules not described above. In some embodiments, the modules stored in memory 204, or the non-transitory computer-readable storage medium of memory 204, provide instructions for performing corresponding operations in the methods described below. In some embodiments, some or all of these modules can be implemented using dedicated hardware circuitry (e.g., an FPGA) containing some or all of the module functionality. The elements identified above (one or more) can be executed by one or more processors 202.
[0044] In some embodiments, the user input / output (I / O) subsystem 208 communicatively couples the 3D scanner 200 to one or more devices, such as one or more remote devices 236, via a communication network 250 and / or via a wired and / or wireless connection. In some embodiments, the communication network 250 is the Internet. In some embodiments, the user input / output (I / O) subsystem 208 may communicatively couple the 3D scanner 200 to one or more integrated or peripheral devices, such as a touch-sensitive display.
[0045] In some embodiments, the projector 110 includes one or more lasers. In some embodiments, the one or more lasers include vertical-cavity surface-emitting lasers (VCSELs). In some embodiments, the projector 110 also includes an array of light-emitting diodes (LEDs) that generate visible light. In some embodiments, instead of lasers, the projector 110 includes a flash lamp or some other light source.
[0046] The communication bus 210 optionally includes a circuit system (sometimes referred to as a chipset) that interconnects and controls communication between system components.
[0047] Figure 3 This is a block diagram of a remote device 236 coupled to a 3D scanner 200 via a network 250 according to some embodiments. The remote device 236 typically includes a memory 304, one or more processors 302, a power supply 306, a user input / output (I / O) subsystem 308, and a communication bus 310 for interconnecting these components. The processor 302 executes modules, programs, and / or instructions stored in the memory 304 and thereby performs processing operations.
[0048] In some embodiments, processor 302 includes at least one central processing unit. In some embodiments, processor 302 includes at least one graphics processing unit. In some embodiments, processor 302 includes at least one neural processing unit (NPU) for executing the neural networks described herein. In some embodiments, processor 302 includes at least one field-programmable gate array (FPGA).
[0049] In some embodiments, memory 304 stores one or more programs (e.g., instruction sets) and / or data structures. In some embodiments, memory 304 or a non-transitory computer-readable storage medium of memory 304 stores the following programs, modules, and data structures, or subsets or supersets thereof:
[0050] ● Operating system 312, which includes programs for handling various basic system services and for performing hardware-related tasks;
[0051] ● Network communication module 318, which is used to connect remote device 236 to other computer systems (e.g., 3D scanner 200) via one or more communication networks 250;
[0052] ● User interface module 320, which receives commands and / or inputs from the user through user input / output (I / O) subsystem 308, and provides outputs for presentation and / or display on user input / output (I / O) subsystem 308;
[0053] ● Data processing module 324, which processes data from 3D scanner 200, including optionally performing method 500 ( Figures 5A-5C ) and / or method 700 ( Figure 7 Any or all operations described herein. In some embodiments, the data processing module 324 includes a neural network module 340 (for determining the correspondence between an image and a projected pattern) and a triangulation module 344 for executing a triangulation algorithm to determine the spatial coordinates of an object using the correspondence determined by the neural network module 340. In some embodiments, the neural network module 340 includes instructions for executing neural networks 340-a and 340-b, which are referred to below. Figure 4 Describe it;
[0054] ● Neural network training module 328, which is used to train a neural network to determine the correspondence between imaging elements in an imaging pattern and projection elements in a projection pattern, including optionally performing operations related to method 600 ( Figures 6A-6B (or any or all operations described in the alternative methods mentioned herein.)
[0055] ● Storage device 330, which includes buffers, RAM, ROM and / or other memory for storing data used and generated by remote device 236.
[0056] The modules identified above (e.g., data structures and / or programs, including instruction sets) do not need to be implemented as separate software programs, processes, or modules, and therefore various subsets of these modules can be combined or otherwise rearranged in various embodiments. In some embodiments, memory 304 stores a subset of the modules identified above. Furthermore, memory 304 may store additional modules not described above. In some embodiments, the modules stored in memory 304, or the non-transitory computer-readable storage medium of memory 304, provide instructions for performing corresponding operations in the methods described below. In some embodiments, some or all of these modules can be implemented using dedicated hardware circuitry (e.g., an FPGA) containing some or all of the module functionality. One or more elements identified above may be executed by one or more processors 302.
[0057] In some embodiments, the user input / output (I / O) subsystem 308 communicatively couples the remote device 236 to one or more devices, such as one or more 3D scanners 200 or external displays, via a communication network 250 and / or via wired and / or wireless connections. In some embodiments, the communication network 250 is the Internet. In some embodiments, the user input / output (I / O) subsystem 308 may communicatively couple the remote device 236 to one or more integrated or peripheral devices, such as touch-sensitive displays.
[0058] The communication bus 310 optionally includes a circuit system (sometimes referred to as a chipset) that interconnects and controls communication between system components.
[0059] Figure 4 The inputs and outputs of a neural network 340-a according to some embodiments are shown. As described above, one problem arising in the context of structured light 3D scanning methods is the ambiguity of elements imaged on the surface of an object (e.g., the need to know which element in the image corresponds to which element on the slide). According to some embodiments, this problem is solved by using a neural network trained to determine the correspondence between projected elements and elements imaged on the surface of the object (or more generally, the correspondence between image pixels and their corresponding coordinates within the projection pattern). For this purpose, the imaging pattern 132 (refer to...) Figure 1CThe image (132) is provided as input to the neural network 340 as a projection pattern. In some embodiments, the neural network 340 outputs a correspondence 402. In some embodiments, the correspondence 402 is an "image" having the same number of pixels as the image 132. For example, in some embodiments, the input to the neural network 340 is a 9-megapixel photograph of the object's surface while the projection pattern is projected onto it, and the output of the neural network is a 9-megapixel output "image" where each pixel corresponds one-to-one with the input image. That is, each pixel in the "image" of correspondence 402 is a value representing the correspondence between the image imaged in the image 132 and the projection pattern. For example, a value "4" in correspondence 402 indicates that those pixels in the image 132 correspond to coordinates (or lines) with a value of "4" in the projection pattern (note that the range of values within the projection pattern coordinate system is arbitrary and can be 0 to 1, 0 to 10, or have other ranges).
[0060] In some embodiments, the neural network 340-a receives additional input. For example, the neural network receives information about the projected pattern.
[0061] In some embodiments, neural network 340-a outputs a coarse value for the correspondence, and neural network 340-b outputs a fine value for the correspondence. In some embodiments, 340-b operates in a manner similar to neural network 340-a, except that neural network 340-b receives an image of the surface of an object as input and the output of neural network 340-a (e.g., neural networks 340-a and 340-b are cascaded). Note that any number of neural networks in any arrangement can be used. For example, in some embodiments, three or four neural networks are used, some arranged in a cascaded manner and some arranged independently.
[0062] Figure 5A–5C illustrates a flowchart of a method 500 for providing 3D reconstruction from a 3D imaging environment 100 according to some embodiments. Method 500 is executed at a computing device having a processor and memory storing a program configured to be executed by the processor. In some embodiments, method 500 is executed at a computing device communicating with a 3D scanner 200. In some embodiments, certain operations of method 500 are performed by a computing device different from the 3D scanner 200 (e.g., a computer system receiving and / or transmitting data from the 3D scanner 200). In some embodiments, certain operations of method 500 are performed by a computing device storing a neural network 340, such that the trained neural network 340 is available as part of a 3D reconstruction based on images captured by the 3D scanner 200. Some operations in method 500 may optionally be combined and / or the order of some operations may optionally be changed.
[0063] In various embodiments, method 500 may include any features or operations of method 700, as described below, provided that such features or operations are not inconsistent with the described method 500. For the sake of brevity, some details described with reference to method 700 will not be repeated here.
[0064] Method 500 includes obtaining (510) an object (e.g., Figure 1A and 1B The image shows an object 120. The image includes a plurality of imaging elements 142 of an imaging pattern 132. The imaging pattern 132 corresponds to a projection pattern 130 that is projected onto the surface of the object 120. The projection pattern 130 includes a plurality of projection elements 140. Method 500 further includes using a neural network (520) (e.g., neural network 340-a) to output a correspondence between the plurality of imaging elements 142 and the plurality of projection elements 140. In some embodiments, the neural network directly outputs the correspondence (e.g., at least a plurality of nodes in the output layer of the neural network have a one-to-one correspondence with at least a plurality of nodes in the input layer). In some embodiments, the neural network is a classification neural network (e.g., classifying each pixel of the input image by the correspondence between the input image and the projection pattern).
[0065] It is important to note that conventional neural networks are trained to recognize different instances of the same object. For example, an instance of human-written characters can be used to train a neural network to recognize human-written characters. Conversely, according to the embodiments described herein, it has been found that neural networks can be trained to determine correspondences between projected elements and elements imaged on the surface of an object, even if the training data does not include another instance of that object. For example, by training a neural network on data from objects with various features, the neural network can be used to determine element correspondences when scanning the skull of a previously undiscovered extinct species of whale, even if the training data does not include skulls of that species.
[0066] The complex geometry of objects (e.g., narrow features, sharp edges, deep grooves, etc.) complicates the determination of element correspondences. Here, the inventors also discovered that using a trained neural network can improve image resolution and integrity, especially when "sharp" features are present in the object.
[0067] In some embodiments, method 500 includes inputting (522) the value of each corresponding pixel of an image of object 120 to the corresponding node in the input layer of a neural network (e.g., neural network 304-a).
[0068] In some embodiments, each corresponding pixel of the image of object 120 corresponds to a corresponding node in the output layer of a neural network (524). The value of the corresponding node in the output layer of the neural network represents the correspondence between the corresponding pixel and a plurality of projection elements 140 of the projection pattern 130 (e.g., the value represents the coordinates on the projection pattern).
[0069] In some embodiments, the output layer of the neural network has (526) the same size as the image of object 120 (e.g., the neural network outputs an "image" with the same number of pixels as the input image, as referenced). Figure 4 The above).
[0070] In some embodiments, the output layer of the neural network is smaller than the size of the image (528). In some embodiments, the output layer of the neural network is larger than the size of the image.
[0071] In some embodiments, method 500 includes inputting (530) information about the projection pattern 130 into the input layer of a neural network.
[0072] In some embodiments, a plurality of projection elements 140 of a projection pattern 130 projected onto the surface of object 120 include (532) non-coded elements (e.g., any projection pattern having non-coded elements described herein). In some embodiments, a plurality of projection elements 140 of a projection pattern 130 projected onto the surface of object 12 include lines.
[0073] In some embodiments, the neural network is trained using simulated data (535). The simulated data includes a plurality of simulated images, and each of the plurality of simulated images includes a simulated pattern comprising a plurality of simulated elements. Each of the plurality of simulated elements corresponds to a corresponding projection element among a plurality of projection elements projected onto the surface of a corresponding simulated object. Each of the plurality of simulated images also includes correspondence data indicating the correspondence between the plurality of simulated elements of the simulated image and the plurality of projection elements of the projection pattern.
[0074] In some embodiments, each of the plurality of simulated images includes (536) texture information about the corresponding simulated object.
[0075] In some embodiments, the texture information of the corresponding simulated object is (538) texture information other than the natural texture of the corresponding simulated object.
[0076] In some embodiments, the texture information of the corresponding simulated object includes (540) features of a plurality of projection elements 140 similar to a projection pattern.
[0077] In some embodiments, the texture information of the corresponding simulated object includes (542) text.
[0078] In some embodiments, the texture information of the corresponding simulated object includes (544) lines.
[0079] Operations 534-544 are discussed below regarding method 600. Figures 6A-6B (This will be described in more detail. That is, in some embodiments, the neural network used in method 500 is trained using method 600.)
[0080] In some embodiments, multiple neural networks are used. The neural networks may be cascaded or operate independently of each other. As a non-limiting example of a cascaded network, in some embodiments, the neural network is a first neural network (e.g., neural network 340-a), and method 500 further includes using (550) a second neural network (e.g., neural network 340-b) to output offsets (e.g., thinning) of the correspondences between the plurality of imaging elements 142 and the plurality of projection elements 140 determined by the first neural network. Thus, in some embodiments, the resolution of the 3D reconstruction is enhanced by using two neural networks: (i) a first neural network that identifies the correspondences between projection elements and elements imaged on the surface of an object, and (2) a second neural network that identifies the offsets of the identified correspondences. In some embodiments, the second neural network directly outputs the offsets (e.g., at least a plurality of nodes in the output layer of the second neural network have a one-to-one correspondence with each pixel in the input image). The inventors have found that this two-stage approach produces a significant and unexpected improvement in the resolution of the resulting image.
[0081] Figure 6A –6B illustrates a flowchart of a method 600 for training a neural network 340 according to some embodiments. Method 600 is executed (601) at a computing device having a processor and memory storing a program configured to be executed by the processor. In some embodiments, method 600 is executed at a computing device communicating with a 3D scanner 200. In some embodiments, certain operations of method 600 are performed by a computing device different from the 3D scanner 200 (e.g., a computer system receiving and / or transmitting data from the 3D scanner 200). In some embodiments, certain operations of method 600 are performed by a computing device that trains and stores the neural network 340, such that the trained neural network 340 is available as part of a 3D reconstruction based on images captured by the 3D scanner 200. Some operations in method 600 may optionally be combined and / or the order of some operations may be (optionally) changed.
[0082] According to some embodiments, method 600 uses simulated (also called “synthetic”) data, where the spatial relationships between the projector, camera, and object are known for each training image. The challenge in training a neural network to determine the correspondence between projected elements and elements imaged on the surface of an object lies in the difficulty of obtaining the “foundational facts” for training. Typically, hundreds of thousands of elements are projected onto the surface of an object. Existing algorithms for determining line correspondences are hampered by the problems addressed by the neural networks of this disclosure. Therefore, existing algorithms cannot be used to provide the foundational facts for training such neural networks. Furthermore, unlike image analysis, character recognition, and similar applications, manual labeling is impractical in 3D scanning / reconstruction applications and is as prone to error as existing algorithms. These problems are addressed by training a neural network using simulated data, where the exact correspondences and geometry of the image acquisition are known. In this way, training data can be generated for countless different object shapes and geometries, as well as the geometry of the camera and projector relative to the object.
[0083] Method 600 includes generating (610) simulation data. The simulation data includes i) a plurality of simulation images (e.g., as shown in Figures A1-A2 of Appendix A to U.S. Provisional Application 63 / 070,066), ii) object data (e.g., as shown in Figures A3-A6 of Appendix A to U.S. Provisional Application 63 / 070,066), and iii) correspondence data. The plurality of simulation images are images of a known projection pattern projected onto the surface of a simulated object. The projection pattern includes a plurality of elements, and each of the images includes an imaging element of the simulation pattern. Each of the imaging elements corresponds to a corresponding element among the plurality of elements of the known projection pattern. The object data includes data indicating the shape of the corresponding simulated object, and the correspondence data includes data indicating the correspondence between the imaging elements in the simulation images and the corresponding elements among the plurality of elements of the known projection pattern. Method 600 also includes using the simulation data to train (620) a neural network to determine the correspondence between the corresponding elements among the plurality of elements of the known projection pattern and the imaging elements of the known projection pattern projected onto a real object and / or surface (e.g., imaging element 142 of imaging pattern 132). Method 600 also includes storing (630) a trained neural network 340 for subsequent use in reconstructing images (e.g., as used in method 500). Figures 5A-5C ).
[0084] In some embodiments, the simulated data further includes (611) information about the texture (e.g., color) of the simulated object. In some embodiments, multiple simulated images further include texture information about the respective simulated objects. The challenge in training a neural network to determine the correspondence between projected elements and elements imaged on the surface of an object is that the object itself has color, and the color may vary with the image of the object (e.g., due to changes in the object's own color, or due to lighting, shadows, etc.). This makes it difficult to distinguish patterns from the texture of the object itself. This problem is addressed by using simulated training data with a variety of textures and reflectivities (effectively making the problem more challenging during the training phase, so that the neural network becomes more effective once trained). In particular, the inventors have found that texturing simulated objects to include text, patterns, or other abrupt (high-contrast) texture features is particularly effective in teaching a neural network to distinguish between object textures and projected elements (e.g., because text involves high-contrast variations between light and dark, so do projected patterns).
[0085] In some embodiments, the texture information of the corresponding simulated object is (612) texture information other than the natural texture of the corresponding simulated object.
[0086] In some embodiments, the texture information of the corresponding simulated object includes (613) features of a plurality of elements similar to a known projection pattern.
[0087] In some embodiments, the texture information of the corresponding simulated object includes (614) text.
[0088] In some embodiments, the texture information of the corresponding simulated object includes (615) lines.
[0089] In some embodiments, the corresponding simulated object includes (616) one or more sharp features.
[0090] In some embodiments, an alternative method for training a neural network is provided, which includes generating simulated data comprising: a plurality of simulated images of a projection pattern projected onto a surface of a corresponding simulated object; data indicating the shape of the corresponding simulated object; and data indicating the correspondence between corresponding pixels in the simulated images and coordinates of the projection pattern. The alternative method also includes using the simulated data to train the neural network to determine the correspondence between the images and the projection pattern. The alternative method further includes storing the trained neural network for subsequent use in reconstructing images. Note that in some embodiments, the alternative method for training the neural network may share any features or operations of the method 600 described above, provided that such features or operations are not inconsistent with the alternative method.
[0091] Figure 7A flowchart of a method 700 for providing 3D reconstruction from a 3D imaging environment 100, according to some embodiments, is shown. Method 700 is executed at a computing device having a processor and memory storing programs configured to be executed by the processor. In some embodiments, method 700 is executed at a computing device communicating with a 3D scanner 200. In some embodiments, certain operations of method 700 are performed by a computing device different from the 3D scanner 200 (e.g., a computer system receiving and / or transmitting data from the 3D scanner 200). In some embodiments, certain operations of method 700 are performed by a computing device storing a neural network 340, such that the trained neural network 340 is available as part of a 3D reconstruction based on images captured by the 3D scanner 200. Some operations in method 700 may optionally be combined and / or the order of some operations may optionally be changed.
[0092] In various embodiments, method 700 may include any features or operations of method 500 described above, as long as such features or operations are not inconsistent with the described method 700. For the sake of brevity, some details described with reference to method 500 will not be repeated here.
[0093] Method 700 includes obtaining (702) an image of an object while projecting a pattern onto the surface of the object. In some embodiments, the projection pattern is generated by passing light through a glass slide. In some embodiments, a coordinate system is associated with the projection pattern. The coordinate system describes the positioning of each location of the projection pattern on the glass slide.
[0094] Method 700 further includes using a neural network (704) to output the correspondence between corresponding pixels in the image and coordinates of the projection pattern (e.g., relative to a coordinate system). For this purpose, in some embodiments, the image is provided to the input layer of the neural network (e.g., neural network 340-a). In some embodiments, the output layer of the neural network directly generates the corresponding coordinates for each pixel within the projection pattern. For example, the neural network outputs an output image with the same number of pixels as the input image, wherein each pixel of the output image has a one-to-one correspondence with a pixel in the input image, and the value of the input pixel coordinates is stored on the projection pattern. In this way, the output image is spatially correlated with the input image.
[0095] In some embodiments, the neural network is trained using the method 600 described above or an alternative method.
[0096] Note that in some embodiments, the neural network outputs two coordinates (e.g., x and y coordinates on a slide pattern) for each pixel of the input image. Alternatively, in some embodiments, the neural network outputs only a single coordinate for each pixel of the input image. In such embodiments, the other coordinates are known or can be inferred from the polar geometry of the scanner 200.
[0097] In some embodiments, multiple neural networks can be used to determine the projected pattern coordinates of each pixel in an input image. For example, in some embodiments, a first neural network determines coarse coordinates, while a second neural network determines fine coordinates (e.g., refines the coordinates of the first neural network). In various embodiments, the first and second neural networks can be arranged in a cascaded manner (e.g., such that the output of the first neural network is input into the second neural network), or the two neural networks can operate independently, with their outputs combined. In various embodiments, more than two neural networks can be used (e.g., four neural networks).
[0098] In some embodiments, the input image is a multi-channel image. As a non-limiting example, the input image may include 240 x 320 pixels, but may store more than one value per pixel (e.g., three values in the case of an RGB image). In some embodiments, additional channels are provided to feed additional information into the neural network. Continuing with the non-limiting example, the input image is then 240 x 320 x n, where n is the number of channels. For example, in some embodiments, information about the projection pattern is fed into the neural network as an additional "channel" for each image. In some embodiments, one or more channels include information obtained when the projection pattern is not illuminating the surface of an object. For example, a grayscale image of the projection pattern illuminating the surface of an object may be stacked with an RGB image and obtained at a time close to the grayscale image (e.g., within 200 milliseconds), where the RGB image was obtained without the projection pattern illuminating the surface of the object (recall that in some embodiments, the projection pattern illuminating the surface of the object in a stroboscopic manner).
[0099] In some embodiments, the output image is a multi-channel image. In some embodiments, one channel of the multi-channel output image provides a correspondence, as described above. Continuing with the above non-limiting example, each channel of the output image may include 240 x 320 pixels. The output then has a size of 240 x 320 x m, where m is the number of channels. One of the channels stores the value of the correspondence (e.g., the value of one or more coordinates on the projected pattern). In some embodiments, another channel in the output image stores a confidence value for each correspondence value for each pixel. The confidence value of the correspondence value for each pixel can be used for reconstruction (e.g., by weighting the data in a different way or discarding data with too low a confidence value). In some embodiments, the output image may also include a channel describing the curvature of an object, a channel describing the texture of an object, or any other information spatially related to the input image.
[0100] Those skilled in the art will understand that the input and output images can be of any size. For example, in addition to the 240x320 pixel image described in the above non-limiting examples, in some embodiments, a 9-megapixel image (or any other size image) may be used.
[0101] It is important to note that conventional neural networks are trained to recognize different instances of the same object. For example, an instance of human-written characters can be used to train a neural network to recognize human-written characters. Conversely, according to the embodiments described herein, it has been found that neural networks can be trained to determine the correspondence between corresponding pixels in an image and coordinates of a projected pattern, even if the training data does not include another instance of the object. For example, by training a neural network on data from objects with various features, the neural network can be used to determine correspondences when scanning the skulls of previously undiscovered extinct whale species, even if the training data does not include skulls of that species.
[0102] The complex geometry of objects (e.g., narrow features, sharp edges, deep grooves, etc.) exacerbates the difficulty of determining correspondences. Here, the inventors also discovered that using a trained neural network can improve image resolution and integrity, especially in the presence of “sharp” features in the object. Figures A7–A18 of Appendix A to U.S. Provisional Application 63 / 070,066 provide examples of 3D image reconstruction using conventional methods and neural networks according to the invention (note that Figures A7–A9 are single-image reconstructions, while Figures A10–A18 are multi-image reconstructions). These reconstructed images demonstrate significantly better quality images reconstructed according to the invention, including better image resolution and integrity.
[0103] Method 700 also includes using the correspondence between corresponding pixels in the (706) image and the coordinates of the projected pattern to reconstruct the shape of the object's surface (e.g., using a triangulation algorithm).
[0104] It should be understood that Figure 5A -5B and Figure 6A The specific order of operations described in –6B is merely an example and is not intended to suggest that the described order is the only possible order of operations. One skilled in the art will recognize various methods for reordering the operations described herein.
[0105] For purposes of explanation, the foregoing description has been illustrated with reference to specific embodiments. However, the illustrative discussion above is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments described have been selected and described to best explain the principles of the invention and its practical application, thereby enabling others skilled in the art to best utilize the invention and the corresponding described embodiments with various modifications suitable for the contemplated particular uses.
[0106] It should also be understood that although the terms first, second, etc., may be used in some cases to describe various elements herein, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, without departing from the scope of this disclosure, a first neural network may be referred to as a second neural network, and similarly, a second neural network may be referred to as a first neural network. Unless the context clearly indicates otherwise, both the first neural network and the second neural network are neural networks, but they are not the same neural network.
[0107] The terminology used in the description of the corresponding embodiments described herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the description of the various embodiments described and in the appended claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms, unless the context clearly indicates otherwise. It should also be understood that, as used herein, the term “and / or” refers to and covers any and all possible combinations of one or more of the associated listed items. It should be further understood that, when used in this specification, the terms “includes” and “comprising” specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0108] As used herein, depending on the context, the term "if" is optionally interpreted as meaning "when," "upon," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrases "if determined" or "if detected" are optionally interpreted as meaning "when determined," "in response to determination," "when detected," or "in response to detection."
Claims
1. A method for reconstruction, comprising: Obtain an image of an object, wherein the image includes multiple imaging elements of an imaging pattern, wherein: The imaging pattern corresponds to a projection pattern, which is projected onto the surface of the object via a glass slide; and The projection pattern includes multiple projection elements; An image of the object is provided as input to the neural network; The neural network is used to output the correspondence between the corresponding pixels of the plurality of imaging elements and the coordinates of the plurality of projection elements, wherein the coordinates describe the position of the corresponding projection element of the projection pattern on the slide; as well as The shape of the object's surface is reconstructed using the correspondence between the corresponding pixels of the plurality of imaging elements and the coordinates of the plurality of projection elements.
2. The method of claim 1, wherein the neural network is a first neural network; and The method further includes: A second neural network is used to output the offset of the correspondence between the corresponding pixels of the plurality of imaging elements and the coordinates of the plurality of projection elements, as determined by the first neural network.
3. The method of claim 1, wherein using the neural network to output the correspondence between the corresponding pixels of the plurality of imaging elements and the coordinates of the plurality of projection elements comprises inputting the value of each corresponding pixel of the image of the object to a corresponding node in the input layer of the neural network.
4. The method according to claim 1, wherein: Each corresponding pixel of the image of the object corresponds to a corresponding node in the output layer of the neural network; and The value of the corresponding node in the output layer of the neural network represents the correspondence between the corresponding pixel and the coordinates of the plurality of projection elements of the projection pattern.
5. The method of claim 4, wherein the output layer of the neural network has the same size as the image of the object.
6. The method of claim 4, wherein the output layer of the neural network is smaller than the size of the image.
7. The method of claim 6, wherein using the neural network to output the correspondence between the corresponding pixels of the plurality of imaging elements and the coordinates of the plurality of projection elements comprises inputting information about the projection pattern into the input layer of the neural network.
8. The method of claim 1, wherein the plurality of projection elements of the projection pattern projected onto the surface of the object include uncoded elements.
9. The method of claim 8, wherein the plurality of projection elements of the projection pattern projected onto the surface of the object comprises lines.
10. The method according to claim 1, wherein: The neural network is trained using simulated data, which includes multiple simulated images, each of which contains: A simulated imaging pattern comprising multiple simulated elements, wherein each of the multiple simulated elements corresponds to a corresponding projection element among the multiple projection elements projected onto the surface of a corresponding simulated object; as well as Data indicating the correspondence between multiple simulated elements of the simulated imaging pattern and multiple projected elements of the projected pattern.
11. The method of claim 10, wherein for each of the plurality of simulated images, the corresponding simulated object includes texture information.
12. The method of claim 11, wherein the texture information of the corresponding simulated object is texture information other than the natural texture of the corresponding simulated object.
13. The method of claim 12, wherein the texture information of the corresponding simulated object includes features of the plurality of projection elements similar to the projection pattern.
14. The method of claim 11, wherein the texture information of the corresponding simulated object includes text.
15. The method of claim 11, wherein the texture information of the corresponding simulated object includes lines.
16. A computer system comprising one or more processors and a memory, the memory storing instructions for performing the method according to any one of claims 1 to 15.
17. A non-transitory computer-readable storage medium storing instructions that, when executed by a computer system, cause the computer system to perform the method according to any one of claims 1 to 15.
Citation Information
Patent Citations
System and method for three-dimensional measurement of the shape of material objects
US7768656B2
Phase unwrapping method based on deep learning
CN111523618A
3D object scanning method using structured light
US20190371053A1