An image reconstruction method, apparatus and device
By acquiring multiple frames of images through a speckle projector and a binocular camera system, generating coded maps and matching key points, the problem of poor reconstruction results in existing 3D imaging systems is solved, achieving high-precision and stable 3D reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU HIKROBOT TECH CO LTD
- Filing Date
- 2023-02-23
- Publication Date
- 2026-04-21
AI Technical Summary
Existing 3D imaging systems, when acquiring multi-line structured light images based on cameras, exhibit poor 3D reconstruction results, long reconstruction times, and poor stability.
A speckle projector is used to project speckles, and n frames of images are acquired by a first camera and a second camera to generate first and second coded images. A three-dimensional reconstructed image is generated based on key point pairs.
It achieves high-precision and robust 3D reconstruction, improving the accuracy and stability of 3D reconstruction.
Smart Images

Figure CN116228980B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an image reconstruction method, apparatus and device. Background Technology
[0002] A 3D imaging system can consist of a laser and a camera. The laser projects line-structured light onto the surface of the object being measured (i.e., the target object), and the camera captures an image of the object, resulting in a line-structured light image. After obtaining the line-structured light image, the center line of the light stripes can be acquired and transformed according to pre-calibrated sensor parameters to obtain the spatial coordinates (i.e., 3D coordinates) of the object at its current position. Based on these spatial coordinates, 3D reconstruction of the object can then be achieved.
[0003] To achieve 3D reconstruction of the object under test, it is necessary to acquire line structured light images of different locations on the object. This involves a laser projecting line structured light onto different locations on the object, with each location corresponding to a separate line structured light image. The camera then acquires only one line structured light image at a time. However, when 3D reconstruction is performed based on multiple line structured light images acquired by the camera, problems such as poor 3D reconstruction results exist. Summary of the Invention
[0004] This application provides an image reconstruction method applied to a three-dimensional imaging system, the three-dimensional imaging system including a first camera, a second camera, and a speckle projector, the method comprising:
[0005] Acquire n frames of first image and n frames of second image, where n is a positive integer greater than 1; wherein, each time the speckle projector projects speckles onto the object under test, the first image of the object under test is acquired by the first camera and the second image of the object under test is acquired by the second camera;
[0006] A first coded image is generated based on the n first images, and a second coded image is generated based on the n second images; wherein, for each pixel in the first coded image, the coded value corresponding to the pixel in the first coded image is determined based on the pixel value corresponding to the pixel in each first image; for each pixel in the second coded image, the coded value corresponding to the pixel in the second coded image is determined based on the pixel value corresponding to the pixel in each second image.
[0007] Based on the first coded image and the second coded image, a key point pair corresponding to each pixel in the first coded image is determined; wherein, for each key point pair, the key point pair includes a first pixel in the first coded image and a second pixel in the second coded image, and the first pixel and the second pixel are pixels corresponding to the same position point on the object under test;
[0008] A 3D reconstructed image of the object under test is generated based on key point pairs.
[0009] This application provides an image reconstruction apparatus for use in a three-dimensional imaging system, the three-dimensional imaging system including a first camera, a second camera, and a speckle projector, the apparatus comprising:
[0010] The acquisition module is used to acquire n frames of first images and n frames of second images, where n is a positive integer greater than 1; wherein, each time the speckle projector projects speckles onto the object under test, the first image of the object under test is acquired by the first camera and the second image of the object under test is acquired by the second camera;
[0011] A generation module is configured to generate a first coded image based on the n frames of first images and a second coded image based on the n frames of second images; wherein, for each pixel in the first coded image, the coded value corresponding to the pixel in the first coded image is determined based on the pixel value corresponding to the pixel in each frame of the first image; and for each pixel in the second coded image, the coded value corresponding to the pixel in the second coded image is determined based on the pixel value corresponding to the pixel in each frame of the second image.
[0012] The matching module is used to determine key point pairs corresponding to each pixel in the first coded image based on the first coded image and the second coded image; wherein, for each key point pair, the key point pair includes a first pixel in the first coded image and a second pixel in the second coded image, the first pixel and the second pixel are pixels corresponding to the same position point on the object under test, and a three-dimensional reconstructed image corresponding to the object under test is generated based on the key point pair.
[0013] This application provides an electronic device, including: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; wherein, the processor is configured to execute the machine-executable instructions to implement the image reconstruction method of the example above.
[0014] As can be seen from the above technical solutions, in this embodiment, speckle projectors project speckles onto the object under test, and a first image of the object under test is acquired by a first camera, and a second image of the object under test is acquired by a second camera. When the speckle projector projects speckles onto the object under test n times, n frames of the first image and n frames of the second image can be obtained. A first coded image is generated based on the n frames of the first image, and a second coded image is generated based on the n frames of the second image. A three-dimensional reconstructed image corresponding to the object under test can be generated based on the first and second coded images, thus achieving three-dimensional reconstruction. The three-dimensional reconstruction effect is excellent, and an accurate and reliable three-dimensional reconstructed image can be obtained. For example, since the coded values of the same location point on the object under test are highly similar in the first and second coded images, accurate and reliable key point pairs can be found from the first and second coded images. When generating a three-dimensional reconstructed image based on these key point pairs, an accurate and reliable three-dimensional reconstructed image can be obtained. The reconstruction accuracy of the three-dimensional reconstructed image is higher, and its robustness is better. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings of the embodiments of this application.
[0016] Figure 1 This is a flowchart illustrating an image reconstruction method according to one embodiment of this application;
[0017] Figure 2 This is a schematic diagram of the structure of a three-dimensional imaging system according to one embodiment of this application;
[0018] Figure 3 This is a schematic diagram of the image acquisition process in one embodiment of this application;
[0019] Figure 4 This is a flowchart illustrating an image reconstruction method according to one embodiment of this application;
[0020] Figure 5 This is a schematic diagram illustrating the determination of the encoded value of a pixel in one embodiment of this application;
[0021] Figure 6 This is a schematic diagram of the structure of an image reconstruction apparatus according to one embodiment of this application;
[0022] Figure 7 This is a hardware structure diagram of an electronic device according to one embodiment of this application. Detailed Implementation
[0023] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms “a,” “the,” and “the” as used in this application and claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to any and all possible combinations comprising one or more of the associated listed items.
[0024] It should be understood that although the terms first, second, third, etc., may be used to describe various information in embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" may also be interpreted as "when," "when," or "in response to a determination."
[0025] This application proposes an image reconstruction method that can be applied to a three-dimensional imaging system. The three-dimensional imaging system may include a first camera, a second camera, and a speckle projector. For example, the first camera, the second camera, and the speckle projector may be deployed on the same device, i.e., the three-dimensional imaging system may consist of one device. Alternatively, the second camera and the speckle projector may be deployed on different devices, i.e., the three-dimensional imaging system may consist of multiple devices. There is no limitation on this.
[0026] See Figure 1 The diagram shown illustrates the process of this image reconstruction method, which may include:
[0027] Step 101: Acquire n frames of first image and n frames of second image, where n can be a positive integer greater than 1; wherein, each time the speckle projector projects speckles onto the object under test, the first image of the object under test is acquired by the first camera and the second image of the object under test is acquired by the second camera.
[0028] Step 102: Generate a first coded image based on n frames of the first image, and generate a second coded image based on n frames of the second image; wherein, when generating the first coded image, for each pixel in the first coded image, the coded value corresponding to the pixel in the first coded image is determined based on the pixel value corresponding to the pixel in each frame of the first image; when generating the second coded image, for each pixel in the second coded image, the coded value corresponding to the pixel in the second coded image is determined based on the pixel value corresponding to the pixel in each frame of the second image.
[0029] For example, based on n frames of a first image, a first coded image can be generated using a first encoding strategy, which indicates how the encoded value of each pixel in the first coded image is generated; based on n frames of a second image, a second coded image can be generated using a second encoding strategy, which indicates how the encoded value of each pixel in the second coded image is generated; wherein, the encoded value generation method of the pixel indicated by the first encoding strategy and the encoded value generation method of the pixel indicated by the second encoding strategy are the same encoded value generation method, that is, the first encoding strategy and the second encoding strategy can be the same.
[0030] In one possible implementation, for each pixel in the first coded image, the coded value corresponding to the pixel in the first coded image is determined based on the pixel value corresponding to the pixel in each frame of the first image. This may include, but is not limited to: determining an intermediate value based on n pixel values corresponding to the pixel in n frames of the first image; determining n reference values corresponding to the n pixel values based on the n pixel values and the intermediate value; and determining the coded value corresponding to the pixel in the first coded image based on the n reference values. Wherein, the intermediate value is the average of the n pixel values; the reference value is the comparison result between the pixel value and the intermediate value, or the reference value is the difference between the pixel value and the intermediate value; if the pixel value is greater than the intermediate value, the comparison result between the pixel value and the intermediate value is the first value; if the pixel value is not greater than the intermediate value, the comparison result between the pixel value and the intermediate value is the second value.
[0031] For each pixel in the second encoded image, the encoded value corresponding to that pixel in the second encoded image is determined based on the pixel value corresponding to that pixel in each frame of the second image. This can include, but is not limited to: determining an intermediate value based on n pixel values corresponding to that pixel in n frames of the second image; determining n reference values corresponding to the n pixel values based on the n pixel values and the intermediate value; and determining the encoded value corresponding to that pixel in the second encoded image based on the n reference values. The intermediate value can be the average of the n pixel values; the reference value can be the comparison result between the pixel value and the intermediate value, or the reference value can be the difference between the pixel value and the intermediate value; if the pixel value is greater than the intermediate value, the comparison result between the pixel value and the intermediate value is the first value; if the pixel value is not greater than the intermediate value, the comparison result between the pixel value and the intermediate value is the second value.
[0032] In another possible implementation, for each pixel in the first encoding map, the encoding value corresponding to the pixel in the first encoding map is determined based on the pixel value corresponding to the pixel in each frame of the first image. This may include, but is not limited to: determining the neighboring pixels corresponding to the pixel from the neighborhood window of the pixel; and determining the encoding value corresponding to the pixel in the first encoding map based on the pixel value corresponding to the pixel in each frame of the first image and the pixel values corresponding to the neighboring pixels in each frame of the first image. For example, for each frame of the first image, a reference value for the pixel in the first image is determined based on the pixel value corresponding to the pixel in the first image and the pixel values corresponding to adjacent pixels in the first image. Based on the reference values corresponding to the pixel in each of the n frames of the first image, the encoded value corresponding to the pixel in the first encoded image is determined. The reference value can be a comparison result between the pixel value corresponding to the pixel and the pixel values corresponding to adjacent pixels, or the reference value can be the difference between the pixel value corresponding to the pixel and the pixel values corresponding to adjacent pixels. If the pixel value corresponding to an adjacent pixel is greater than the pixel value corresponding to the pixel, the comparison result can be the first value; if the pixel value corresponding to an adjacent pixel is not greater than the pixel value corresponding to the pixel, the comparison result can be the second value.
[0033] For each pixel in the second encoding image, determining the encoding value of the pixel in the second encoding image based on the pixel value corresponding to the pixel in each frame of the second image may include: determining the neighboring pixels corresponding to the pixel from the neighborhood window of the pixel; and determining the encoding value of the pixel in the second encoding image based on the pixel value corresponding to the pixel in each frame of the second image and the pixel values corresponding to the neighboring pixels in each frame of the second image. For each frame of the second image, a reference value is determined based on the pixel value corresponding to the pixel in that frame and the pixel values corresponding to adjacent pixels in that frame. Based on the reference values corresponding to the pixel in each of the n frames of the second image, the encoded value is determined in the second encoded image. The reference value is either the comparison result between the pixel value corresponding to the pixel and the pixel values corresponding to adjacent pixels, or the difference between the pixel value corresponding to the pixel and the pixel values corresponding to adjacent pixels. If the pixel value corresponding to an adjacent pixel is greater than the pixel value corresponding to the pixel, the comparison result is the first value; if the pixel value corresponding to an adjacent pixel is not greater than the pixel value corresponding to the pixel, the comparison result is the second value.
[0034] Step 103: Determine the key point pair corresponding to each pixel in the first encoding map based on the first encoding map and the second encoding map; wherein, for each key point pair, the key point pair may include the first pixel in the first encoding map and the second pixel in the second encoding map, the first pixel and the second pixel are pixels corresponding to the same position point on the object being measured.
[0035] For example, determining the keypoint pair corresponding to each pixel in the first encoded image based on the first encoded image and the second encoded image may include, but is not limited to: for a first pixel in the first encoded image, where the first pixel is each pixel in the first encoded image (i.e., each pixel in the first encoded image is taken as the first pixel in turn), determining multiple candidate pixels corresponding to the first pixel in the second encoded image; for each candidate pixel, determining the similarity between the first pixel and the candidate pixel based on the encoded value corresponding to the first pixel in the first encoded image and the encoded value corresponding to the candidate pixel in the second encoded image; based on the similarity between the first pixel and each candidate pixel, selecting a candidate pixel that matches the first pixel from the multiple candidate pixels as the second pixel; and generating a keypoint pair based on the first pixel and the second pixel.
[0036] For example, generating a first coded image based on n frames of first images and generating a second coded image based on n frames of second images may include, but is not limited to: performing binocular correction on the n frames of first images and n frames of second images to obtain corrected n frames of first images and corrected n frames of second images; wherein, binocular correction is used to ensure that the same position point on the measured object has the same pixel height in the corrected first image and the corrected second image; generating the first coded image based on the corrected n frames of first images and generating the second coded image based on the corrected n frames of second images. For a first pixel in the first coded image, determining multiple candidate pixels corresponding to the first pixel in the second coded image may include, but is not limited to: for a first pixel in the first coded image, determining multiple candidate pixels corresponding to the first pixel in the second coded image based on the pixel height of the first pixel in the first coded image; wherein, the pixel height of each candidate pixel in the second coded image is the same as the pixel height of the first pixel in the first coded image.
[0037] Step 104: Generate a 3D reconstructed image of the object under test based on key point pairs.
[0038] For example, a 3D reconstructed image of the object under test can be generated based on all key point pairs, or a 3D reconstructed image of the object under test can be generated based on some key point pairs; there are no restrictions on this.
[0039] As can be seen from the above technical solutions, in this embodiment, speckle projectors project speckles onto the object under test, and a first image of the object under test is acquired by a first camera, and a second image of the object under test is acquired by a second camera. When the speckle projector projects speckles onto the object under test n times, n frames of the first image and n frames of the second image can be obtained. A first coded image is generated based on the n frames of the first image, and a second coded image is generated based on the n frames of the second image. A three-dimensional reconstructed image corresponding to the object under test can be generated based on the first and second coded images, thus achieving three-dimensional reconstruction. The three-dimensional reconstruction effect is excellent, and an accurate and reliable three-dimensional reconstructed image can be obtained. For example, since the coded values of the same location point on the object under test are highly similar in the first and second coded images, accurate and reliable key point pairs can be found from the first and second coded images. When generating a three-dimensional reconstructed image based on these key point pairs, an accurate and reliable three-dimensional reconstructed image can be obtained. The reconstruction accuracy of the three-dimensional reconstructed image is higher, and its robustness is better.
[0040] The image reconstruction method of this application embodiment will be described below in conjunction with specific application scenarios.
[0041] To acquire 3D reconstructed images, various techniques can be employed, including surface projection structured light, binocular speckle scanning, Time-of-Flight (TOF), and single-line laser contour scanning. When using surface projection structured light, DLP (Digital Light Processing) or LCD (Liquid Crystal Display) projection technologies can be employed, using LED (Light Emitting Diode) light sources. However, this results in large projection volumes and energy dispersion, making it unsuitable for 3D positioning applications due to its large size and high power consumption, especially at long distances and wide fields of view. Binocular speckle scanning combines binocular parallax with laser speckle matching, but this method suffers from low detection accuracy and poor edge contours, hindering contour scanning and 3D positioning applications. TOF, limited by camera resolution, offers only centimeter-level accuracy, insufficient for automated high-precision positioning applications. Single-line laser contour scanning utilizes a single laser to scan the object's depth information; however, its slow scanning speed and poor stability fail to meet positioning requirements.
[0042] Taking the single-line laser contour scanning method as an example, a laser can be used to project line structured light onto the surface of the object being measured, and a camera can be used to capture images of the object, resulting in a line structured light image. After obtaining the line structured light image, the spatial coordinates (i.e., 3D coordinates) of the object at its current position can be obtained based on the line structured light image, thus achieving 3D reconstruction of the object. However, to achieve 3D reconstruction of the object, it is necessary to acquire line structured light images at different positions of the object. That is, the laser projects line structured light onto different positions of the object, with each position corresponding to a line structured light image. Since the camera only acquires the line structured light image corresponding to one position at a time, the camera needs to acquire multiple line structured light images to complete the 3D reconstruction. This results in a long 3D reconstruction time, slow scanning speed, and poor stability, failing to meet the positioning requirements of 3D reconstruction.
[0043] In response to the above findings, this application proposes an image reconstruction method, which is a binocular matching method based on speckle projector (speckle projector can also be called speckle sensor) projection. The speckle projector projects speckles onto the object under test, and a first image of the object under test is acquired by a first camera and a second image of the object under test is acquired by a second camera. When the speckle projector projects speckles onto the object under test n times, n frames of first image and n frames of second image can be obtained. Based on the n frames of first image and n frames of second image, an accurate and reliable three-dimensional reconstructed image can be obtained. The three-dimensional reconstructed image has higher reconstruction accuracy and better robustness, thereby enabling efficient acquisition of high-precision depth map data and point cloud data.
[0044] This application proposes an image reconstruction method based on speckle projector projection, which can be applied to a three-dimensional imaging system. The type of three-dimensional imaging system is not limited; it can be any system with three-dimensional imaging capabilities, such as any system in the field of machine vision or industrial automation.
[0045] See Figure 2 The diagram shows a schematic of a 3D imaging system. This system can employ a binocular multi-speckle sensor system and may include, but is not limited to, a first camera, a second camera, a speckle projector, and a processor. The first camera can be the left-side camera, located on the left side of the 3D imaging system, also referred to as the left imaging unit. The second camera can be the right-side camera, located on the right side of the 3D imaging system, also referred to as the right imaging unit. Alternatively, the first camera can be the right-side camera, and the second camera can be the left-side camera; there are no restrictions on this.
[0046] A speckle projector, also known as a multi-speckle projection device, is used to project speckles (multiple speckles) onto the object being measured. Each time a speckle projector projects a speckle onto the object, a first image of the object is captured by a first camera, and a second image of the object is captured by a second camera.
[0047] The processor can be such as a CPU. The processor is used to acquire a first image from a first camera, acquire a second image from a second camera, and complete the image reconstruction of the object under test based on the first and second images. In other words, it obtains a three-dimensional reconstructed image of the object under test based on the first and second images.
[0048] In one possible implementation, see Figure 3 The diagram shows a schematic of the image acquisition process, which may include: turning on the speckle projector; after turning on the speckle projector, projecting speckles onto the object under test through the speckle projector (e.g., the speckle projector projects speckles onto the object under test based on the coded pattern K1); acquiring a first image L1 of the object under test through a first camera; acquiring a second image R1 of the object under test through a second camera; the acquisition time of the first image L1 and the second image R1 is the same.
[0049] After the first image L1 and the second image R1 are acquired, the speckle projector switches the coded pattern and projects speckles onto the object under test based on the switched coded pattern K2. The first image L2 of the object under test is acquired by the first camera, and the second image R2 of the object under test is acquired by the second camera.
[0050] After the first image L2 and the second image R2 are acquired, the speckle projector switches the coded pattern and projects speckles onto the object under test based on the switched coded pattern K3. The first image L3 of the object under test is acquired by the first camera, and the second image R3 of the object under test is acquired by the second camera.
[0051] This process continues until the speckle projector projects speckles onto the object under test based on the switched coded pattern Kn. The first image Ln of the object under test is acquired by the first camera, and the second image Rn of the object under test is acquired by the second camera. Then, the speckle projector can be turned off to complete the image acquisition process.
[0052] In summary, we can obtain n frames of the first image and n frames of the second image, namely, first image L1, first image L2, ..., first image Ln, second image R1, second image R2, ..., second image Rn.
[0053] For example, the speckle projector can be a single-light source speckle projector, meaning it projects specks onto the object being measured using a single light source. For instance, the single light source can project multiple specks onto the object, and the pattern formed by these multiple specks is the coded pattern. By changing the projection position of the single light source, different coded patterns can be generated. For example, when the projection position of the single light source is position A1, the speckle projector projects specks onto the object based on the coded pattern K1. Then, by moving the projection position of the single light source to position A2, the speckle projector projects specks onto the object based on the coded pattern K2. Similarly, by moving the projection position of the single light source to position An, the speckle projector projects specks onto the object based on the coded pattern Kn.
[0054] A speckle projector can also be a multi-source speckle projector, meaning it projects specks onto the object being measured using multiple light sources. For example, each light source can project multiple specks onto the object. The combination of these projected specks is the pattern projected onto the object by the speckle projector, and the resulting pattern is the coded pattern. By changing the projection positions of some or all light sources, different coded patterns can be generated. For instance, changing the projection positions of some or all light sources can produce coded patterns K1, K2, ..., Kn, etc., without limitation.
[0055] In summary, n coded patterns can be projected onto the object being measured. These n coded patterns can be obtained by projecting from a single light source and changing the position of the light source, or by projecting from multiple light sources individually or by combining them.
[0056] For example, when the speckle projector projects speckles onto the object under test based on the coded pattern K1, a first image L1 of the object under test is acquired by a first camera, and a second image R1 of the object under test is acquired by a second camera. In order to ensure that the acquisition time of the first image L1 and the second image R1 is the same, the processor can send a projection command to the speckle projector so that the speckle projector projects speckles onto the object under test based on the coded pattern K1. The processor can also send acquisition commands to the first camera and the second camera so that the first camera acquires the first image L1 of the object under test, and the second camera acquires the second image R1 of the object under test.
[0057] The first image L1 includes multiple speckles projected onto the object under test by the speckle projector based on the coded pattern K1, and the second image R1 also includes multiple speckles projected onto the object under test by the speckle projector based on the coded pattern K1. That is, the first image L1 and the second image R1 include the same multiple speckles.
[0058] Among them, the speckle projector projects multiple spots onto the object being measured. These multiple spots can be random speckle, i.e., random discrete spots; they can also be pseudo-random speckle, i.e., pseudo-random discrete spots; or they can be regular speckle, i.e., discrete spots in regular positions. There are no restrictions on which type they are.
[0059] For example, after the first image L1 and the second image R1 are acquired, the processor sends a projection command to the speckle projector so that the speckle projector projects speckles onto the object under test based on the coded pattern K2. At the same time, the processor sends an acquisition command to the first camera and the second camera so that the first camera acquires the first image L2 of the object under test and the second camera acquires the second image R2 of the object under test. In this way, the acquisition time of the first image L2 and the second image R2 can be the same, and so on.
[0060] In the above embodiments, n can be a positive integer greater than 1. The value of n is configured based on experience, and there is no restriction on the value of n, as long as n≥2. For example, n can be 3, 4, 5, 6, 7, 8, 9, etc.
[0061] Based on the above application scenarios, this application proposes an image reconstruction method that can be applied to three-dimensional imaging systems. See [link to relevant documentation]. Figure 4 The diagram shown is a flowchart of the method, which may include:
[0062] Step 401: Obtain n frames of the first image and n frames of the second image, where n can be a positive integer greater than 1.
[0063] For example, the first images of n frames are first image L1, first image L2, ..., first image Ln, and the second images of n frames are second image R1, second image R2, ..., second image Rn.
[0064] Step 402: Generate the first coded image based on the first image of n frames.
[0065] For example, when generating the first coded image, for each pixel in the first coded image, the coded value corresponding to that pixel in the first coded image is determined based on the pixel value corresponding to that pixel in each frame of the first image. For instance, for pixel (x, y) in the first coded image, the coded value corresponding to pixel (x, y) in the first coded image can be determined based on the pixel values corresponding to pixel (x, y) in first images L1, L2, ..., Ln. After obtaining the coded value corresponding to each pixel in the first coded image, the coded values corresponding to all pixels can be combined to obtain the first coded image; that is, the first coded image can include the coded value corresponding to each pixel.
[0066] In one possible implementation, the first coded image can be generated based on n frames of the first image in the following manner. Of course, the following manner is only a few examples and there is no limitation on this generation method.
[0067] Method 1: For each pixel in the first encoding map, the encoding value corresponding to the pixel in the first image in each frame is determined, and the first encoding map is generated based on the encoding value corresponding to each pixel. In other words, after obtaining the encoding value corresponding to each pixel in the first encoding map, the encoding values corresponding to all pixels are combined to obtain the first encoding map.
[0068] The method for obtaining the encoded value corresponding to each pixel is the same. Taking obtaining the encoded value corresponding to pixel (x, y) as an example, the following steps can be used to obtain the encoded value corresponding to pixel (x, y):
[0069] Step S11: Determine an intermediate value based on the n pixel values corresponding to pixel point (x, y) in the n frames of the first image. For example, the intermediate value can be the average of the n pixel values, or the minimum of the n pixel values, or the maximum of the n pixel values, or the median of the n pixel values; there is no limitation in this regard. For ease of description, in subsequent embodiments, the average of the n pixel values will be used as an example. The pixel values in this embodiment can all be grayscale values; of course, they can also be other types of numerical values, and there is no limitation in this regard.
[0070] For example, the pixel value corresponding to pixel point (x, y) in the first image L1 can be denoted as pixel value B1, the pixel value corresponding to pixel point (x, y) in the first image L2 can be denoted as pixel value B2, ..., and the pixel value corresponding to pixel point (x, y) in the first image Ln can be denoted as pixel value Bn. Based on this, the average value of pixel value B1, pixel value B2, ..., pixel value Bn can be used as the intermediate value.
[0071] Step S12: Determine n reference values corresponding to the n pixel values based on the n pixel values and the median value.
[0072] For example, based on pixel value B1 and the median value, a reference value C1 corresponding to pixel value B1 can be determined. For instance, the comparison result between pixel value B1 and the median value can be used as the reference value C1. Specifically, if pixel value B1 is greater than the median value, the comparison result can be the first value (e.g., 1), meaning the reference value C1 can be the first value; if pixel value B1 is not greater than the median value, the comparison result can be the second value (e.g., 0), meaning the reference value C1 can be the second value. Another example is the difference between pixel value B1 and the median value. Yet another example is the average of pixel value B1 and the median value. Yet another example is the sum of pixel value B1 and the median value. Of course, these are just a few examples and are not limiting, as long as the reference value C1 corresponding to pixel value B1 can be obtained.
[0073] Similarly, we can obtain the reference value C2 corresponding to pixel value B2, ..., and the reference value Cn corresponding to pixel value Bn.
[0074] Step S13: Determine the encoded value of pixel (x, y) in the first encoding image based on n reference values corresponding to pixel (x, y). For example, the encoded value of pixel (x, y) in the first encoding image is determined based on reference values C1, C2, ..., Cn corresponding to pixel (x, y). For example, the encoded value may include reference values C1, C2, ..., Cn.
[0075] For example, assuming the reference value is the result of comparing the pixel value with the median value, and n is 8, then the encoding value corresponding to the pixel (x, y) can be an 8-bit binary encoding value, such as 01110001. The last bit represents the reference value C1, the second to last bit represents the reference value C2, and so on, with the first bit representing the reference value C8. Alternatively, the last bit represents the reference value C8, and so on, with the first bit representing the reference value C1.
[0076] In summary, the encoded value corresponding to pixel (x, y) can be determined by the n pixel values corresponding to pixel (x, y) in the first image of n frames. For example, assuming the reference value is the comparison result between the pixel value and the intermediate value, and n is 8, then an 8-bit binary encoded value c(u, v, m) at pixel (u, v) in the first encoded image can be expressed as follows:
[0077]
[0078] In the above formula, m represents the number of bits of the encoding, that is, the number of bits of the encoding value corresponding to the pixel (u, v), such as an 8-bit encoding value, and c(u, v, m) represents the 8-bit binary encoding value corresponding to the pixel (u, v).
[0079] This represents the average of the n pixel values corresponding to pixel point (u, v) in the first n-frame image, i.e., the aforementioned intermediate value. p (u,v) represents the pixel value of pixel (u,v) in the first image. The value of p ranges from 1 to n. I1(u,v) represents the pixel value of pixel (u,v) in the first image L1, I2(u,v) represents the pixel value of pixel (u,v) in the first image L2, and so on.
[0080] like This indicates that the pixel value is greater than the median value, meaning the comparison result can be the first value 1, and the reference value is the first value 1. Conversely, if... This means that the pixel value is not greater than the median value, that is, the comparison result can be the second value 0, and the reference value is the second value 0.
[0081] For example, there are various ways to calculate the encoded value, and the number of bits in the encoded result also varies with the method. In addition to the comparison of pixel values mentioned above, there are also methods including but not limited to gray value difference, gray value mapping, gray value sorting order, and combinations of encoding results of various methods. This embodiment does not limit these methods.
[0082] Method 2: For each pixel in the first encoded image, determine the neighboring pixels corresponding to that pixel from its neighborhood window, and determine the encoded value corresponding to that pixel in the first encoded image based on the pixel value corresponding to that pixel in each frame of the first image and the pixel values corresponding to the neighboring pixels in each frame of the first image. Then, the first encoded image can be generated based on the encoded value corresponding to each pixel; that is, the encoded values corresponding to all pixels can be combined to obtain the first encoded image.
[0083] For example, for a pixel (x, y) in the first encoded image, based on the pixel value corresponding to pixel (x, y) in the first image L1, first image L2, ..., first image Ln, and the pixel values corresponding to the neighboring pixels of pixel (x, y) in the first image L1, first image L2, ..., first image Ln, the encoded value corresponding to pixel (x, y) in the first encoded image can be determined. After obtaining the encoded value corresponding to each pixel in the first encoded image, the encoded values corresponding to all pixels can be combined to obtain the first encoded image; that is, the first encoded image can include the encoded value corresponding to each pixel.
[0084] In summary, it can be seen that... Figure 5As shown, for the method of determining the encoding value of each pixel in the first encoding image, the encoding value corresponding to the pixel can be determined based on the pixel value of the pixel in different first images and the neighboring pixel value (i.e., the pixel value of adjacent pixels in different first images).
[0085] The method for obtaining the encoded value corresponding to each pixel is the same. Taking obtaining the encoded value corresponding to pixel (x, y) as an example, the following steps can be used to obtain the encoded value corresponding to pixel (x, y):
[0086] Step S21: Determine the neighboring pixels corresponding to pixel (x, y) from the neighborhood window of pixel (x, y). The neighborhood window of pixel (x, y) is a window centered on pixel (x, y).
[0087] For example, the neighborhood window of a pixel (x, y) can be an m*m cross-shaped window or an m*m rectangular window; there are no restrictions on the type of neighborhood window. For an m*m cross-shaped window, if m*m is 3*3, then the pixel (x, y) corresponds to four neighboring pixels: (x-1, y), (x+1, y), (x, y-1), and (x, y+1). For an m*m rectangular window, if m*m is 3*3, then the pixel (x, y) corresponds to 8 neighboring pixels, namely neighboring pixel (x-1, y-1), neighboring pixel (x-1, y), neighboring pixel (x-1, y+1), neighboring pixel (x, y-1), neighboring pixel (x, y+1), neighboring pixel (x+1, y-1), neighboring pixel (x+1, y), and neighboring pixel (x+1, y+1).
[0088] Of course, the above is just an example and is not a limitation. After the neighborhood window of pixel (x, y) is determined, the pixels within the neighborhood window can be regarded as the adjacent pixels of pixel (x, y).
[0089] Step S22: For each frame of the first image, based on the pixel value corresponding to the pixel point (x, y) in the first image and the pixel value corresponding to the adjacent pixel points in the first image, determine the reference value corresponding to the pixel point (x, y) in the first image, that is, obtain the reference value corresponding to the pixel point (x, y) in the n frames of the first image.
[0090] For example, for a first image L1, a reference value for pixel (x, y) in the first image L1 can be determined based on the pixel value corresponding to pixel (x, y) in the first image L1 and the pixel values corresponding to each of its neighboring pixels in the first image L1. For instance, the comparison result between the pixel value corresponding to pixel (x, y) and the pixel values corresponding to its neighboring pixels can be used as the reference value. Specifically, if the pixel value corresponding to a neighboring pixel is greater than the pixel value corresponding to pixel (x, y), the comparison result can be a first value (e.g., 1), meaning the reference value can be the first value; if the pixel value corresponding to a neighboring pixel is not greater than the pixel value corresponding to pixel (x, y), the comparison result can be a second value (e.g., 0), meaning the reference value can be the second value. Taking pixel (x, y) corresponding to four neighboring pixels as an example, the comparison result values corresponding to pixel (x, y) and each of the four neighboring pixels can be obtained. Therefore, the reference value corresponding to pixel (x, y) in the first image L1 can include four comparison result values.
[0091] For example, the difference between the pixel value corresponding to pixel point (x, y) and the pixel value corresponding to the neighboring pixel point can be used as a reference value. That is, the difference between pixel point (x, y) and the four neighboring pixels can be obtained. The reference value corresponding to pixel point (x, y) in the first image L1 can include four differences.
[0092] For example, the average value of the pixel value corresponding to pixel (x, y) and the pixel values corresponding to its four neighboring pixels can be used as a reference value. In other words, the average value of pixel (x, y) and its four neighboring pixels can be obtained. The reference value of pixel (x, y) in the first image L1 can include the four average values.
[0093] Of course, the above are just a few examples and are not limited to this. As long as we can obtain the reference value of pixel (x, y) in the first image L1, we can obtain the reference value of pixel (x, y) in the first image L2, ..., the reference value of pixel (x, y) in the first image Ln. That is, we can obtain the reference value of pixel (x, y) in the first image of n frames, i.e., n reference values.
[0094] Step S23: Based on the reference values corresponding to pixel (x, y) in the n frames of the first image, determine the encoded value corresponding to pixel (x, y) in the first encoded image. For example, based on the reference values corresponding to pixel (x, y) in the first image L1, the reference values corresponding to pixel (x, y) in the first image L2, ..., the reference values corresponding to pixel (x, y) in the first image Ln, determine the encoded value corresponding to pixel (x, y) in the first encoded image. For example, the encoded value may include the above multiple reference values.
[0095] For example, assuming the reference value is the comparison result value and n is 8, then the encoded value corresponding to pixel (x, y) can be a 32-bit binary code value. Each 4 bits represent a reference value, and there are a total of 8 reference values. Each reference value is a 4-bit binary value, and each bit of the binary value represents the comparison result value between pixel (x, y) and one adjacent pixel. Four adjacent pixels correspond to 4 bits.
[0096] In summary, the encoded value corresponding to pixel (x, y) can be determined by the n pixel values corresponding to pixel (x, y) in the first image of n frames, and there are no restrictions on the determination process.
[0097] For example, there are various ways to calculate the encoded value, and the number of bits in the encoded result also varies with the method. In addition to the comparison of pixel values mentioned above, there are also methods including but not limited to gray value difference, gray value mapping, gray value sorting order, and combinations of encoding results of various methods. This embodiment does not limit these methods.
[0098] Step 403: Generate a second coded image based on the second image of n frames.
[0099] For example, when generating the second coded image, for each pixel in the second coded image, the corresponding coded value in the second coded image is determined based on the pixel value corresponding to that pixel in each frame of the second image. For instance, for pixel (x, y) in the second coded image, the coded value corresponding to pixel (x, y) in the second coded image can be determined based on the pixel values corresponding to pixel (x, y) in second images R1, R2, ..., Rn. After obtaining the coded value corresponding to each pixel in the second coded image, the coded values corresponding to all pixels can be combined to obtain the second coded image; that is, the second coded image can include the coded value corresponding to each pixel.
[0100] In one possible implementation, the second coded image can be generated based on n frames of the second image in the following manner. Of course, the following manner is only a few examples and there is no limitation on this generation method.
[0101] Method 1: For each pixel in the second encoded image, determine the encoded value of that pixel in the second encoded image based on the pixel value corresponding to that pixel in each frame of the second image. Generate the second encoded image based on the encoded values of each pixel. For example, combine the encoded values of all pixels to obtain the second encoded image. The method for obtaining the encoded value of each pixel is the same. Taking obtaining the encoded value of pixel (x, y) as an example, the encoded value of pixel (x, y) can be obtained as follows: Determine intermediate values based on the n pixel values corresponding to pixel (x, y) in n frames of the second image; determine n reference values corresponding to the n pixel values based on the n pixel values and the intermediate values; determine the encoded value of pixel (x, y) in the second encoded image based on the n reference values corresponding to pixel (x, y).
[0102] For example, the median value can be the average of n pixel values; the reference value can be the comparison result of a pixel value and the median value, or the reference value can be the difference between a pixel value and the median value.
[0103] If the pixel value is greater than the median value, the comparison result between the pixel value and the median value is the first value; if the pixel value is not greater than the median value, the comparison result between the pixel value and the median value is the second value.
[0104] Method 2: For each pixel in the second encoded image, determine the neighboring pixels corresponding to that pixel from its neighborhood window, and determine the encoded value of that pixel in the second encoded image based on the pixel value corresponding to that pixel in each frame of the second image and the pixel values corresponding to the neighboring pixels in each frame of the second image. Then, the second encoded image can be generated based on the encoded value corresponding to each pixel; that is, the encoded values corresponding to all pixels can be combined to obtain the second encoded image.
[0105] The method for obtaining the encoded value corresponding to each pixel is the same. Taking the encoded value corresponding to pixel (x, y) as an example, it can be obtained as follows: Determine the neighboring pixels corresponding to pixel (x, y) from the neighborhood window of pixel (x, y), where the neighborhood window is a window centered on pixel (x, y). For each frame of the second image, based on the pixel value corresponding to pixel (x, y) in that second image and the pixel values corresponding to its neighboring pixels in that second image, determine the reference value corresponding to pixel (x, y) in that second image, thus obtaining the reference value corresponding to pixel (x, y) in n frames of the second image. Based on the reference values corresponding to pixel (x, y) in each of the n frames of the second image, determine the encoded value corresponding to pixel (x, y) in the second encoded image.
[0106] For example, the reference value can be the comparison result between the pixel value corresponding to pixel point (x, y) and the pixel value corresponding to the neighboring pixel point, or the difference between the pixel value corresponding to pixel point (x, y) and the pixel value corresponding to the neighboring pixel point. Wherein, if the pixel value corresponding to the neighboring pixel point is greater than the pixel value corresponding to pixel point (x, y), the comparison result can be the first value; if the pixel value corresponding to the neighboring pixel point is not greater than the pixel value corresponding to pixel point (x, y), the comparison result can be the second value.
[0107] In one possible implementation, based on n frames of the first image, a first coded image can be generated using a first coding strategy, which indicates how the coded value of each pixel in the first coded image is generated; based on n frames of the second image, a second coded image can be generated using a second coding strategy, which indicates how the coded value of each pixel in the second coded image is generated; wherein, the coded value generation method of the pixels indicated by the first coding strategy and the coded value generation method of the pixels indicated by the second coding strategy are the same coded value generation method, that is, the first coding strategy and the second coding strategy can be the same, that is, the first coded image and the second coded image use the same set of coding strategies.
[0108] Step 404: Determine the key point pair corresponding to each pixel in the first encoding map based on the first encoding map and the second encoding map (assuming there are M pixels in the first encoding map, then the M key point pairs corresponding to the M pixels can be determined); wherein, for each key point pair, the key point pair may include the first pixel in the first encoding map and the second pixel in the second encoding map, and the first pixel and the second pixel may be pixels corresponding to the same position on the object being measured.
[0109] For example, before steps 402 and 403, binocular correction can be performed on the n-frame first image and the n-frame second image to obtain the corrected n-frame first image and the corrected n-frame second image. In this way, in step 402, the first coded image is generated based on the corrected n-frame first image, and in step 403, the second coded image is generated based on the corrected n-frame second image.
[0110] For example, binocular correction can be performed on the first image L1 and the second image R1 (i.e., the first image and the second image at the same acquisition time) to obtain the corrected first image L1 and the corrected second image R1. Binocular correction can be performed on the first image L2 and the second image R2 to obtain the corrected first image L2 and the corrected second image R2. Similarly, binocular correction can be performed on the first image Ln and the second image Rn to obtain the corrected first image Ln and the corrected second image Rn. In summary, it can be seen that n corrected first images and n corrected second images can be obtained.
[0111] For example, binocular correction is used to ensure that the same point on the measured object has the same pixel height in both the corrected first image and the corrected second image. For instance, for the same point on the measured object, binocular correction corrects the first and second images to have the same pixel height, making matching easier as it can be performed directly within a single row. For example, matching corresponding points in two-dimensional space is very time-consuming. To reduce the matching search range, epipolar constraints can be used to reduce the matching of corresponding points from a two-dimensional search to a one-dimensional search. The role of binocular correction is to perform row correspondence between the first and second images, resulting in corrected first and second images where the epipolar lines of the corrected first and second images are exactly on the same horizontal line. Any point in the corrected first image will necessarily have the same row number as its corresponding point in the corrected second image, requiring only a one-dimensional search within that row. The specific search process affects the search process of the first and second coded images.
[0112] In steps 402 and 403, since the first coded image is generated based on the corrected n-frame first image and the second coded image is generated based on the corrected n-frame second image, the first coded image and the second coded image are also the first coded image and the second coded image after binocular correction. Based on the first coded image and the second coded image after binocular correction, the key point pair corresponding to each pixel in the first coded image is determined.
[0113] In one possible implementation, based on the first and second encoded maps, the keypoint pairs corresponding to each pixel in the first encoded map can be determined using the following steps:
[0114] Step S31: Select each pixel point from the first encoding image as the first pixel point.
[0115] For example, for each pixel in the first encoded image, that pixel can be taken as the first pixel, and subsequent steps can be performed on the first pixel to obtain the key point pair corresponding to the first pixel.
[0116] Step S32: For each first pixel in the first encoding image (the following explanation will use a single first pixel as an example), determine multiple candidate pixels corresponding to that first pixel in the second encoding image.
[0117] For example, based on the pixel height of the first pixel in the first encoding image, multiple candidate pixels corresponding to the first pixel are determined from the second encoding image; wherein the pixel height of each candidate pixel in the second encoding image is the same as the pixel height of the first pixel in the first encoding image. For example, assuming the pixel height of the first pixel in the first encoding image is h1, then all pixels or some pixels in the second encoding image with pixel height h1 are taken as multiple candidate pixels corresponding to the first pixel.
[0118] Step S33: For each candidate pixel, based on the encoding value of the first pixel in the first encoding image and the encoding value of the candidate pixel in the second encoding image, determine the similarity between the first pixel and the candidate pixel, that is, obtain the similarity between the first pixel and each candidate pixel.
[0119] For example, the encoding value corresponding to the first pixel in the first encoding image can be an 8-bit binary encoding value (or a 32-bit binary encoding value), and the encoding value corresponding to the candidate pixel in the second encoding image can be an 8-bit binary encoding value. Based on the encoding values corresponding to the first pixel in the first encoding image and the candidate pixel in the second encoding image, the similarity between the first pixel and the candidate pixel can be calculated, that is, the similarity between these two encoding values can be calculated. There are no restrictions on the method of calculating this similarity. For example, the distance similarity between the two encoding values can be calculated by the distance between them, or the cosine similarity between the two encoding values can be calculated using a cosine algorithm, or the Pearson correlation coefficient algorithm can be used to calculate the similarity between the two encoding values.
[0120] In summary, for each candidate pixel, the similarity between the first pixel and the candidate pixel can be determined, that is, the similarity between the first pixel and each candidate pixel can be obtained.
[0121] In summary, the first and second encoded images can be similarized using a similarity metric for encoded values to determine the correspondence between pixels. This metric can also depend on how the encoded values are calculated. For example, when using binary encoding (i.e., using the comparison result of pixel values as a reference value), the Hamming distance can be used to measure the similarity between the two encoded values. In step S33, based on the encoded value corresponding to the first pixel in the first encoded image and the encoded value corresponding to the candidate pixel in the second encoded image, the Hamming distance between these two encoded values can be calculated. This Hamming distance is used as the similarity between the first pixel and the candidate pixel.
[0122] For example, when using pixel value difference (i.e., the difference in pixel values as a reference value), the difference in encoded values can be used to measure similarity. In step S33, based on the encoded value corresponding to the first pixel in the first encoded image and the encoded value corresponding to the candidate pixel in the second encoded image, the difference between the two encoded values can be calculated, and this difference is used as the similarity between the first pixel and the candidate pixel.
[0123] Step S34: Based on the similarity between the first pixel and each candidate pixel, select a candidate pixel that matches the first pixel from multiple candidate pixels as the second pixel.
[0124] For example, based on the similarity between the first pixel and each candidate pixel, the candidate pixel corresponding to the maximum similarity can be used as the second pixel. That is, the candidate pixel corresponding to the maximum similarity is a candidate pixel that matches the first pixel, and this candidate pixel can be used as the second pixel.
[0125] Step S35: Generate a key point pair based on the first pixel and the second pixel. The key point pair may include the first pixel in the first encoding image and the second pixel in the second encoding image, and the first pixel and the second pixel may be pixels corresponding to the same position on the object being measured.
[0126] Clearly, for each pixel in the first coded image, a keypoint pair corresponding to that pixel can be obtained. This keypoint pair includes the first pixel in the first coded image and the second pixel in the second coded image. Thus, step 405 is completed, and the keypoint pair corresponding to each pixel in the first coded image can be determined based on the first and second coded images; that is, each pixel corresponds to one keypoint pair.
[0127] In another possible implementation, based on the first and second encoded maps, the keypoint pairs corresponding to each pixel in the second encoded map can be determined using the following steps:
[0128] Step S41: Select each pixel from the second encoding image as the second pixel.
[0129] Step S42: For each second pixel in the second encoding image (the following explanation will use a single second pixel as an example), determine multiple candidate pixels corresponding to that second pixel from the first encoding image.
[0130] For example, based on the pixel height of the second pixel in the second encoding image, multiple candidate pixels corresponding to the second pixel are determined from the first encoding image; wherein the pixel height of each candidate pixel in the first encoding image is the same as the pixel height of the second pixel in the second encoding image.
[0131] Step S43: For each candidate pixel, based on the encoding value of the second pixel in the second encoding image and the encoding value of the candidate pixel in the first encoding image, determine the similarity between the second pixel and the candidate pixel, that is, obtain the similarity between the second pixel and each candidate pixel.
[0132] Step S44: Based on the similarity between the second pixel and each candidate pixel, select a candidate pixel that matches the second pixel from multiple candidate pixels as the first pixel.
[0133] Step S45: Generate a key point pair based on the second pixel and the first pixel. The key point pair may include the second pixel in the second encoding image and the first pixel in the first encoding image, and the second pixel and the first pixel may be pixels corresponding to the same position on the object being measured.
[0134] Clearly, for each pixel in the second encoded image, a keypoint pair corresponding to that pixel can be obtained. This keypoint pair includes the first pixel in the first encoded image and the second pixel in the second encoded image. Thus, step 405 is completed, and the keypoint pair corresponding to each pixel in the second encoded image can be determined based on the first and second encoded images; that is, each pixel corresponds to one keypoint pair.
[0135] For example, steps S41-S45 can be referred to steps S31-S35, and will not be repeated here.
[0136] Step 405: Generate a 3D reconstructed image of the object under test based on key point pairs.
[0137] For example, a 3D reconstructed image of the object under test can be generated based on all key point pairs, or a 3D reconstructed image of the object under test can be generated based on some key point pairs; there are no restrictions on this.
[0138] For example, for each keypoint pair, which includes a first pixel in the first encoded image and a second pixel in the second encoded image, and where the first and second pixels correspond to the same location on the object being measured, triangulation can be used to determine the corresponding 3D point. Alternatively, other methods can be used to determine the corresponding 3D point; there are no restrictions on the method used. After obtaining the 3D points corresponding to all keypoint pairs, a 3D reconstructed image of the object being measured can be generated based on these 3D points.
[0139] As can be seen from the above technical solutions, in this embodiment, when the speckle projector projects speckles onto the object under test n times, n frames of first images and n frames of second images can be obtained. A first coded image is generated based on the n frames of first images, and a second coded image is generated based on the n frames of second images. Based on the first and second coded images, a three-dimensional reconstructed image corresponding to the object under test can be generated, thus achieving three-dimensional reconstruction. The three-dimensional reconstruction effect is very good, and an accurate and reliable three-dimensional reconstructed image can be obtained. Multiple frames of images can be used to synchronously construct the same coded value for matching. For example, since the coded values corresponding to the same location points on the object under test are highly similar in the first and second coded images, accurate and reliable key point pairs can be found from the first and second coded images. When generating a three-dimensional reconstructed image based on these key point pairs, an accurate and reliable three-dimensional reconstructed image can be obtained. The reconstruction accuracy of the three-dimensional reconstructed image is higher, and its robustness is better. Under the condition of achieving the same reconstruction effect, the above method has higher computational efficiency and is also flexible.
[0140] Based on the same concept as the above method, this application proposes an image reconstruction device applied to a three-dimensional imaging system. The three-dimensional imaging system includes a first camera, a second camera, and a speckle projector. (See also...) Figure 6 The diagram shown is a structural schematic of the device, which may include:
[0141] The acquisition module 61 is used to acquire n frames of first images and n frames of second images, where n is a positive integer greater than 1; wherein, each time the speckle projector projects speckles onto the object under test, the first image of the object under test is acquired by the first camera and the second image of the object under test is acquired by the second camera.
[0142] The generation module 62 is configured to generate a first coded image based on the n frames of first images and a second coded image based on the n frames of second images; wherein, for each pixel in the first coded image, the coded value corresponding to the pixel in the first coded image is determined based on the pixel value corresponding to the pixel in each frame of the first image; and for each pixel in the second coded image, the coded value corresponding to the pixel in the second coded image is determined based on the pixel value corresponding to the pixel in each frame of the second image.
[0143] The matching module 63 is used to determine key point pairs corresponding to each pixel in the first encoding map based on the first encoding map and the second encoding map; wherein, for each key point pair, the key point pair includes a first pixel in the first encoding map and a second pixel in the second encoding map, the first pixel and the second pixel are pixels corresponding to the same position point on the object under test, and a three-dimensional reconstructed image corresponding to the object under test is generated based on the key point pair.
[0144] For example, when generating a first coded image based on the n-frame first image, the generation module 62 generates a second coded image based on the n-frame second image, specifically: generating the first coded image based on the n-frame first image using a first encoding strategy, wherein the first encoding strategy is used to indicate the generation method of the encoding value of each pixel in the first coded image; generating the second coded image based on the n-frame second image using a second encoding strategy, wherein the second encoding strategy is used to indicate the generation method of the encoding value of each pixel in the second coded image; wherein the generation method of the encoding value of the pixel indicated by the first encoding strategy and the generation method of the encoding value of the pixel indicated by the second encoding strategy are the same encoding value generation method.
[0145] For example, when the generation module 62 determines the encoded value of the pixel in the first encoded image based on the pixel value corresponding to the pixel in each frame of the first image, it is specifically used to: determine an intermediate value based on n pixel values corresponding to the pixel in the n frames of the first image; determine n reference values corresponding to the n pixel values based on the n pixel values and the intermediate value; determine the encoded value of the pixel in the first encoded image based on the n reference values; wherein, the intermediate value is the average value of the n pixel values; the reference value is the comparison result of the pixel value and the intermediate value, or the reference value is the difference between the pixel value and the intermediate value; if the pixel value is greater than the intermediate value, the comparison result is a first value, and if the pixel value is not greater than the intermediate value, the comparison result is a second value.
[0146] For example, when the generation module 62 determines the encoding value of the pixel in the first encoding map based on the pixel value corresponding to the pixel in each frame of the first image, it is specifically used to: determine the neighboring pixels corresponding to the pixel from the neighborhood window corresponding to the pixel; and determine the encoding value of the pixel in the first encoding map based on the pixel value corresponding to the pixel in each frame of the first image and the pixel values corresponding to the neighboring pixels in each frame of the first image.
[0147] For example, when the generation module 62 determines the encoding value of the pixel in the first encoding map based on the pixel value corresponding to the pixel in each frame of the first image and the pixel values corresponding to the adjacent pixels in each frame of the first image, it specifically performs the following steps: For each frame of the first image, based on the pixel value corresponding to the pixel in the frame of the first image and the pixel values corresponding to the adjacent pixels in the frame of the first image, it determines a reference value corresponding to the pixel in the frame of the first image; based on the reference values corresponding to the pixel in each of the n frames of the first image, it determines the encoding value corresponding to the pixel in the first encoding map; wherein, the reference value is a comparison result value between the pixel value corresponding to the pixel and the pixel value corresponding to the adjacent pixels, or, the reference value is the difference between the pixel value corresponding to the pixel and the pixel value corresponding to the adjacent pixels; if the pixel value corresponding to the adjacent pixels is greater than the pixel value corresponding to the pixel, then the comparison result value is a first value; if the pixel value corresponding to the adjacent pixels is not greater than the pixel value corresponding to the pixel, then the comparison result value is a second value.
[0148] For example, when the matching module 63 determines the keypoint pair corresponding to each pixel in the first encoded image based on the first encoded image and the second encoded image, it specifically performs the following steps: for a first pixel in the first encoded image (where the first pixel is every pixel in the first encoded image), determine multiple candidate pixels corresponding to the first pixel from the second encoded image; for each candidate pixel, determine the similarity between the first pixel and the candidate pixel based on the encoding value corresponding to the first pixel in the first encoded image and the encoding value corresponding to the candidate pixel in the second encoded image; based on the similarity between the first pixel and each candidate pixel, select a candidate pixel that matches the first pixel from the multiple candidate pixels as the second pixel; and generate a keypoint pair based on the first pixel and the second pixel.
[0149] For example, when the generation module 62 generates a first coded image based on the n-frame first image and generates a second coded image based on the n-frame second image, it is specifically used to: perform binocular correction on the n-frame first image and the n-frame second image to obtain a corrected n-frame first image and a corrected n-frame second image; wherein, the binocular correction is used to ensure that the same position point on the measured object has the same pixel height in the corrected first image and the corrected second image; generate the first coded image based on the corrected n-frame first image and generate the second coded image based on the corrected n-frame second image; when the matching module 63 determines multiple candidate pixels corresponding to the first pixel in the second coded image for the first pixel in the first coded image, it is specifically used to: determine multiple candidate pixels corresponding to the first pixel in the second coded image based on the pixel height of the first pixel in the first coded image for the first pixel in the first coded image; wherein, the pixel height of each candidate pixel in the second coded image is the same as the pixel height of the first pixel in the first coded image.
[0150] Based on the same concept as the methods described above, this application proposes an electronic device that can be applied to a three-dimensional imaging system. In addition to this electronic device, the three-dimensional imaging system may also include a first camera, a second camera, and a speckle projector, etc. See also... Figure 7 As shown, the electronic device may include a processor 71 and a machine-readable storage medium 72, the machine-readable storage medium 72 storing machine-executable instructions that can be executed by the processor 71; wherein, the processor 71 is used to execute the machine-executable instructions to implement the image reconstruction method disclosed in the above example of this application.
[0151] Based on the same concept as the above method, this application also provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, can implement the image reconstruction method disclosed in the above examples of this application.
[0152] The aforementioned machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.
[0153] The systems, devices, modules, or units described in the above embodiments can be implemented by a computer entity or by a product with a certain function. A typical implementation device is a computer, which can be a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.
[0154] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0155] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0156] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0157] Furthermore, these computer program instructions can also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0158] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0159] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.
Claims
1. An image reconstruction method, characterized in that, Applied to a three-dimensional imaging system, the three-dimensional imaging system including a first camera, a second camera, and a speckle projector, the method includes: Acquire n frames of first image and n frames of second image, where n is a positive integer greater than 1; wherein, each time the speckle projector projects speckles onto the object under test, the first image of the object under test is acquired by the first camera and the second image of the object under test is acquired by the second camera; A first coded image is generated based on the n first images, and a second coded image is generated based on the n second images; wherein, for each pixel in the first coded image, the coded value corresponding to the pixel in the first coded image is determined based on the pixel value corresponding to the pixel in each first image; for each pixel in the second coded image, the coded value corresponding to the pixel in the second coded image is determined based on the pixel value corresponding to the pixel in each second image. Based on the first coded image and the second coded image, a key point pair corresponding to each pixel in the first coded image is determined; wherein, for each key point pair, the key point pair includes a first pixel in the first coded image and a second pixel in the second coded image, and the first pixel and the second pixel are pixels corresponding to the same position point on the object under test; A 3D reconstructed image of the object under test is generated based on key point pairs.
2. The method according to claim 1, characterized in that, The process of generating a first coded image based on the n first images and generating a second coded image based on the n second images includes: Based on the n frames of the first image, a first encoded image is generated using a first encoding strategy. The first encoding strategy is used to indicate the generation method of the encoded value of each pixel in the first encoded image. Based on the n frames of the second image, a second encoded image is generated using a second encoding strategy. The second encoding strategy is used to indicate the generation method of the encoded value of each pixel in the second encoded image. The encoding value generation method of the pixel indicated by the first encoding strategy is the same as the encoding value generation method of the pixel indicated by the second encoding strategy.
3. The method according to claim 1, characterized in that, Determining the encoded value of the pixel in the first encoded image based on the pixel value of the pixel in each frame of the first image includes: The intermediate value is determined based on the n pixel values corresponding to the pixel in the n-frame first image; Based on the n pixel values and the intermediate value, determine n reference values corresponding to the n pixel values; The encoding value of the pixel in the first encoding map is determined based on the n reference values; Wherein, the median value is the average of the n pixel values; the reference value is the comparison result between the pixel value and the median value, or the reference value is the difference between the pixel value and the median value; If the pixel value is greater than the median value, the comparison result between the pixel value and the median value is the first value; if the pixel value is not greater than the median value, the comparison result between the pixel value and the median value is the second value.
4. The method according to claim 1, characterized in that, Determining the encoded value of the pixel in the first encoded image based on the pixel value of the pixel in each frame of the first image includes: Determine the neighboring pixels corresponding to the pixel from the neighborhood window of the pixel; The encoding value of the pixel in the first encoding image is determined based on the pixel value corresponding to the pixel in each frame of the first image and the pixel value corresponding to the neighboring pixel in each frame of the first image.
5. The method according to claim 4, characterized in that, Determining the encoded value of the pixel in the first encoded image based on the pixel value corresponding to the pixel in each frame of the first image and the pixel values corresponding to the adjacent pixels in each frame of the first image includes: For each frame of the first image, a reference value corresponding to the pixel is determined based on the pixel value corresponding to the pixel in the first image of that frame and the pixel values corresponding to the adjacent pixels in the first image of that frame; and an encoded value corresponding to the pixel in the first encoded image is determined based on the reference values corresponding to the pixel in each of the n frames of the first image. The reference value is either a comparison result between the pixel value corresponding to the pixel point and the pixel value corresponding to the adjacent pixel point, or a difference between the pixel value corresponding to the pixel point and the pixel value corresponding to the adjacent pixel point. If the pixel value corresponding to the adjacent pixel point is greater than the pixel value corresponding to the pixel point, the comparison result value is a first value; if the pixel value corresponding to the adjacent pixel point is not greater than the pixel value corresponding to the pixel point, the comparison result value is a second value.
6. The method according to claim 1, characterized in that, The step of determining the keypoint pair corresponding to each pixel in the first encoded map based on the first encoded map and the second encoded map includes: For a first pixel in the first encoding image, where the first pixel is each pixel in the first encoding image, multiple candidate pixels corresponding to the first pixel are determined from the second encoding image; for each candidate pixel, the similarity between the first pixel and the candidate pixel is determined based on the encoding value corresponding to the first pixel in the first encoding image and the encoding value corresponding to the candidate pixel in the second encoding image. Based on the similarity between the first pixel and each candidate pixel, a candidate pixel that matches the first pixel is selected from the plurality of candidate pixels as the second pixel. Key point pairs are generated based on the first pixel and the second pixel.
7. The method according to claim 6, characterized in that, The step of generating a first coded image based on the n-frame first image and a second coded image based on the n-frame second image includes: performing binocular correction on the n-frame first image and the n-frame second image to obtain corrected n-frame first image and corrected n-frame second image; wherein, the binocular correction is used to ensure that the same position point on the measured object has the same pixel height in the corrected first image and the corrected second image; generating the first coded image based on the corrected n-frame first image and generating the second coded image based on the corrected n-frame second image; The step of determining multiple candidate pixels corresponding to a first pixel in the first encoded image from the second encoded image includes: determining multiple candidate pixels corresponding to a first pixel in the second encoded image based on the pixel height of the first pixel in the first encoded image; wherein the pixel height of each candidate pixel in the second encoded image is the same as the pixel height of the first pixel in the first encoded image.
8. An image reconstruction apparatus, characterized in that, Applied to a three-dimensional imaging system, the three-dimensional imaging system including a first camera, a second camera, and a speckle projector, the device includes: The acquisition module is used to acquire n frames of first images and n frames of second images, where n is a positive integer greater than 1; wherein, each time the speckle projector projects speckles onto the object under test, the first image of the object under test is acquired by the first camera and the second image of the object under test is acquired by the second camera; A generation module is configured to generate a first coded image based on the n frames of first images and a second coded image based on the n frames of second images; wherein, for each pixel in the first coded image, the coded value corresponding to the pixel in the first coded image is determined based on the pixel value corresponding to the pixel in each frame of the first image; and for each pixel in the second coded image, the coded value corresponding to the pixel in the second coded image is determined based on the pixel value corresponding to the pixel in each frame of the second image. The matching module is used to determine key point pairs corresponding to each pixel in the first coded image based on the first coded image and the second coded image; wherein, for each key point pair, the key point pair includes a first pixel in the first coded image and a second pixel in the second coded image, the first pixel and the second pixel are pixels corresponding to the same position point on the object under test, and a three-dimensional reconstructed image corresponding to the object under test is generated based on the key point pair.
9. The apparatus according to claim 8, Its features are, in, The generation module generates a first encoded image based on the n-frame first image, and when generating a second encoded image based on the n-frame second image, it is specifically used to: generate the first encoded image based on the n-frame first image using a first encoding strategy, wherein the first encoding strategy is used to indicate the generation method of the encoded value of each pixel in the first encoded image; Based on the n frames of the second image, a second encoding map is generated using a second encoding strategy. The second encoding strategy is used to indicate the encoding value generation method of each pixel in the second encoding map. The encoding value generation method of the pixel indicated by the first encoding strategy and the encoding value generation method of the pixel indicated by the second encoding strategy are the same encoding value generation method. Specifically, when the generation module determines the encoded value of the pixel in the first encoded image based on the pixel value corresponding to the pixel in each frame of the first image, it is used to: determine an intermediate value based on n pixel values corresponding to the pixel in the n frames of the first image; determine n reference values corresponding to the n pixel values based on the n pixel values and the intermediate value; and determine the encoded value of the pixel in the first encoded image based on the n reference values; wherein, the intermediate value is the average value of the n pixel values; the reference value is the comparison result of the pixel value and the intermediate value, or the reference value is the difference between the pixel value and the intermediate value; if the pixel value is greater than the intermediate value, the comparison result is a first value; if the pixel value is not greater than the intermediate value, the comparison result is a second value. Specifically, when the generation module determines the encoding value of the pixel in the first encoding map based on the pixel value corresponding to the pixel in each frame of the first image, it is used to: determine the neighboring pixels corresponding to the pixel from the neighborhood window corresponding to the pixel; and determine the encoding value of the pixel in the first encoding map based on the pixel value corresponding to the pixel in each frame of the first image and the pixel values corresponding to the neighboring pixels in each frame of the first image. Specifically, when the generation module determines the encoded value of the pixel in the first encoded image based on the pixel value corresponding to the pixel in each frame of the first image and the pixel values corresponding to the adjacent pixels in each frame of the first image, it is used to: for each frame of the first image, determine a reference value corresponding to the pixel in that frame of the first image based on the pixel value corresponding to the pixel in that frame of the first image and the pixel values corresponding to the adjacent pixels in that frame of the first image; determine the encoded value corresponding to the pixel in the first encoded image based on the reference values corresponding to the pixel in each of the n frames of the first image; wherein, the reference value is a comparison result between the pixel value corresponding to the pixel and the pixel value corresponding to the adjacent pixels, or, the reference value is the difference between the pixel value corresponding to the pixel and the pixel value corresponding to the adjacent pixels; if the pixel value corresponding to the adjacent pixels is greater than the pixel value corresponding to the pixel, then the comparison result is a first value; if the pixel value corresponding to the adjacent pixels is not greater than the pixel value corresponding to the pixel, then the comparison result is a second value. Specifically, when the matching module determines the keypoint pair corresponding to each pixel in the first encoded image based on the first encoded image and the second encoded image, it is used to: for a first pixel in the first encoded image (where the first pixel is every pixel in the first encoded image), determine multiple candidate pixels corresponding to the first pixel from the second encoded image; for each candidate pixel, determine the similarity between the first pixel and the candidate pixel based on the encoded value corresponding to the first pixel in the first encoded image and the encoded value corresponding to the candidate pixel in the second encoded image; based on the similarity between the first pixel and each candidate pixel, select a candidate pixel that matches the first pixel from the multiple candidate pixels as the second pixel; and generate a keypoint pair based on the first pixel and the second pixel. Specifically, when the generation module generates a first coded image based on the n-frame first image and a second coded image based on the n-frame second image, it performs the following: Binocular calibration is applied to the n-frame first image and the n-frame second image to obtain calibrated n-frame first images and calibrated n-frame second images; wherein the binocular calibration ensures that the same position point on the object being measured has the same pixel height in the calibrated first image and the calibrated second image; the first coded image is generated based on the calibrated n-frame first image, and the second coded image is generated based on the calibrated n-frame second image; when the matching module determines multiple candidate pixels corresponding to a first pixel in the second coded image, it performs the following: Based on the pixel height of the first pixel in the first coded image, multiple candidate pixels corresponding to the first pixel are determined in the second coded image; wherein the pixel height of each candidate pixel in the second coded image is the same as the pixel height of the first pixel in the first coded image.
10. An electronic device, characterized in that, include: A processor and a machine-readable storage medium storing machine-executable instructions that can be executed by the processor; wherein the processor is configured to execute the machine-executable instructions to implement the method of any one of claims 1-7.
Citation Information
Patent Citations
Image reconstruction method, device and equipment
CN114820939A
3D object model reconstruction from 2d images
US20210398351A1