Method, apparatus, terminal, imaging system and medium for obtaining depth image
By combining monocular structured light system and binocular stereo vision system, color cameras and IR cameras are used to obtain color and speckle images, the accuracy and cost problems of depth image acquisition in the prior art are solved, and high-precision three-dimensional reconstruction of weak texture, long distance and low reflectivity surfaces are achieved.
Patent Information
- Application Number
- CN202210092986.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-01-26
AI Technical Summary
Existing depth image acquisition techniques are difficult to obtain reliable depth estimates in weak or non-textured surfaces, low reflectivity surfaces, long-distance objects and outdoor environments, and are highly cost-effective, making them prone to hollows and occlusion areas.
The combination of monocular structured light system and binocular stereo vision system is adopted. Through the color camera and IR camera combined with the speckle projector, color images and speckle images are obtained, and image feature matching and neural network models are used to fuse depth images, reducing hardware costs and improving accuracy.
A good three-dimensional reconstruction of weak or textureless targets is achieved, with advantages over long-distance and low-reflectivity surfaces, reducing hardware costs, avoiding hollows and occlusion areas, and improving the accuracy of depth images.
Smart Images

Figure CN114511608B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of image processing, and particularly relates to a method, device, terminal, imaging system and medium for obtaining a depth image. Background Art
[0002] A depth image, also known as a range image, refers to an image that uses the distance (depth) from an image collector to each point in a scene as pixel values, which directly reflects the geometric shape of the visible surface of a scene. Currently, depth images are often obtained through passive binocular stereo vision systems, monocular speckle structured light systems, and binocular speckle structured light systems.
[0003] A passive binocular stereo vision system consists of two cameras and is easily affected by the texture condition of the surface of an object to be measured. When the surface of the object to be measured is a weak texture or textureless surface, it is difficult to obtain a reliable depth estimate.
[0004] A monocular speckle structured light system generally consists of an infrared (IR) camera and a near-infrared speckle projector. This active three-dimensional imaging method projects random speckle patterns invisible to the human eye onto the surface of the object to be measured. For an object to be measured with a weak texture or textureless surface, a reliable depth estimation result can still be obtained. To obtain the color information of the object to be measured, an RGB camera also needs to be configured. However, the monocular speckle structured light system also has its limitations. For an object to be measured with a low reflectivity surface, since the light reflected from the surface of the object to be measured to the IR camera is very weak, holes are likely to appear in its depth image. For a distant object, due to the attenuation of light, the IR camera can hardly capture the speckle pattern projected by the speckle projector onto the surface of the object to be measured, which will also cause holes to appear in the depth image. In addition, the monocular speckle structured light system is easily affected by sunlight outdoors.
[0005] A binocular speckle structured light system consists of two IR cameras and a speckle projector. On the one hand, the binocular speckle structured light can use the speckle pattern projected by the speckle projector for image matching, and at the same time, it can also use the texture information of the object surface itself for image matching, and can be used indoors and outdoors. Similarly, to obtain the color information of an object, an RGB camera also needs to be configured. Compared with the monocular speckle structured light system, it requires an additional IR camera, which increases the hardware cost. In addition, to obtain a depth image in the RGB camera coordinate system, the depth image needs to be reprojected onto the RGB camera image plane. However, since there is a certain distance between the optical centers of the RGB camera and the IR camera, when there are three-dimensional targets in the scene, the occluded area will cause new holes to appear in the depth reprojected onto the RGB camera coordinate system, which is not desirable for practical applications. Summary of the Invention
[0006] An embodiment of the present application provides a method, apparatus, terminal, imaging system, and medium for obtaining a depth image, which can improve the accuracy of the depth image while reducing the hardware cost.
[0007] In a first aspect of the embodiments of the present application, a method for obtaining a depth image is provided, including:
[0008] Obtaining a color image obtained by a first camera photographing a target object;
[0009] After controlling a speckle projector to project a preset speckle pattern onto the target object, obtaining a speckle image obtained by a second camera photographing the target object;
[0010] Based on the speckle pattern projected onto the target object in the speckle image, determining a first depth image of the target object;
[0011] Extracting and matching the image features of the color image and the speckle image, and obtaining a second depth image of the target object based on the matched image features and the first depth image.
[0012] In a second aspect of the embodiments of the present application, a device for obtaining a depth image is provided, including:
[0013] A color image acquisition unit for obtaining a color image obtained by a first camera photographing a target object;
[0014] A speckle image acquisition unit for, after controlling a speckle projector to project a preset speckle pattern onto the target object, obtaining a speckle image obtained by a second camera photographing the target object;
[0015] A monocular structured light unit for determining a first depth image of the target object based on the speckle pattern projected onto the target object in the speckle image;
[0016] A depth image acquisition unit for extracting and matching the image features of the color image and the speckle image, and obtaining a second depth image of the target object according to the matched image features and the first depth image.
[0017] In a third aspect of the embodiments of the present application, a terminal is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of the above method are implemented.
[0018] In a fourth aspect of the embodiments of the present application, a depth imaging system is provided, including a first camera, a second camera, a speckle projector, and a terminal provided in the third aspect of the embodiments of the present application. The first camera, the second camera, and the speckle projector are arranged on the same plane;
[0019] The speckle projector is configured to project a speckle pattern onto a target object;
[0020] The first camera is configured to capture a color image of the target object;
[0021] The second camera is configured to capture a speckle image of the target object;
[0022] The terminal is configured to obtain a depth image of the target object by using the color image and the speckle image.
[0023] In a fifth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0024] In a sixth aspect of the embodiments of the present application, a computer program product is provided. When the computer program product runs on a terminal, the terminal is caused to execute the steps of the method.
[0025] In the embodiments of the present application, by obtaining a color image captured by a first camera of a target object and a speckle image captured by a second camera of the target object, and based on the speckle pattern projected onto the target object in the speckle image, a first depth image of the target object is determined. At this time, the first depth image is a depth image obtained based on a monocular structured light subsystem in the imaging system, and has a good 3D reconstruction effect for weakly textured or textureless target objects; the color image and the speckle image can be used as the left image and the right image in a binocular stereo vision system respectively, which has advantages for measuring long-distance target objects, and also has a certain reconstruction ability for target objects with low-reflectivity surfaces; a second depth image of the target object is obtained based on the color image, the speckle image, and the first depth image, which improves the accuracy of the depth image while reducing the number of cameras. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1It is a schematic structural diagram of a depth imaging system provided by an embodiment of the present application;
[0028] Figure 2 It is a schematic implementation flowchart of a method for obtaining a depth image provided by an embodiment of the present application;
[0029] Figure 3 It is a schematic specific implementation flowchart of step S204 provided by an embodiment of the present application;
[0030] Figure 4 It is a schematic flowchart of a stereo matching network model provided by an embodiment of the present application;
[0031] Figure 5a It is a schematic diagram of a color image collected by a first camera provided by an embodiment of the present application;
[0032] Figure 5b It is a schematic diagram of a speckle image collected by a second camera provided by an embodiment of the present application;
[0033] Figure 5c It is a schematic diagram of a first depth image provided by an embodiment of the present application;
[0034] Figure 5d It is a schematic diagram of a second depth image provided by an embodiment of the present application;
[0035] Figure 6 It is a schematic structural diagram of a device for obtaining a depth image provided by an embodiment of the present application;
[0036] Figure 7 It is a schematic structural diagram of a terminal provided by an embodiment of the present application. Detailed implementation manners
[0037] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the protection scope of the present application.
[0038] The passive binocular stereo vision system consists of two cameras and is easily affected by the texture condition of the surface of the object to be measured. When the surface of the object to be measured is a weak texture or textureless surface, it is difficult to obtain a reliable depth estimate.
[0039] A monocular speckle structured light system generally consists of an infrared (IR) camera and a near-infrared speckle projector. This active three-dimensional imaging method projects random speckles that are invisible to the human eye on the surface of the object being measured. For objects with weak or no texture, reliable depth estimation results can still be obtained. In order to obtain the color information of the object being measured, an RGB camera is also required. However, the monocular speckle structured light system also has its limitations. For objects with low reflectivity surfaces, holes are likely to appear in the depth image because the light reflected from the surface of the object to be measured to the IR camera is very weak. For distant objects, due to light attenuation, the IR camera can hardly capture the speckle pattern projected by the speckle projector on the surface of the object to be measured, which will also cause holes in the depth image. In addition, the monocular speckle structured light system is easily affected by sunlight outdoors.
[0040] A binocular speckle structured light system consists of two IR cameras and a speckle projector. Binocular speckle structured light can perform image matching using both the speckle pattern projected by the speckle projector and the texture information of the object's surface itself, making it compatible with both indoor and outdoor use. Similarly, to obtain the object's color information, an RGB camera is also required. Compared to a monocular speckle structured light system, it requires the addition of an IR camera, which increases hardware costs. Furthermore, to obtain a depth image in the RGB camera coordinate system, the depth image must be reprojected onto the RGB camera image plane. However, due to the distance between the optical centers of the RGB and IR cameras, when a three-dimensional target is present in the scene, occluded areas will result in new depth holes in the reprojected RGB camera coordinate system, which is undesirable for practical applications.
[0041] Therefore, an imaging system and a corresponding depth image acquisition method are needed that can combine the advantages of a passive binocular stereo vision system, a monocular speckle structured light system, and a binocular speckle structured light system.
[0042] In order to illustrate the technical solution of the present application, specific embodiments are provided below.
[0043] Figure 1 A depth imaging system provided by the present application is shown, which is suitable for improving the accuracy of depth images while reducing hardware costs.
[0044] The imaging system may include a first camera 11, a second camera 12, a speckle projector 13 and a terminal (not shown in the figure), wherein the first camera 11, the second camera 12 and the speckle projector 13 are arranged on the same plane.
[0045] Among them, the speckle projector 13 is used to project a speckle pattern onto the target object, and the spectral range corresponding to the projected speckle pattern can be the near-infrared spectrum. The first camera 11 is used to collect the color image of the target object; the second camera 12 is used to collect the speckle image of the target object. The above terminal is used to obtain the depth image of the target object by using the color image and the speckle image.
[0046] It should be noted that there is a distance of the baseline length between the second camera 12 and the speckle projector 13, and there is another distance of the baseline length between the first camera 11 and the second camera 12. In some embodiments of the present application, the first camera 11 and the speckle projector 13 should be as close as possible. Ideally, the first camera 11 can be attached to the speckle projector 13.
[0047] In some embodiments of the present application, the first camera 11 can be an RGB camera, and the spectral range in which it operates can be the visible light spectrum.
[0048] In some embodiments of the present application, the second camera 12 can be an IR camera, and its spectral range of operation includes the infrared light spectrum and the visible light spectrum. In a conventional monocular speckle structured light system, the IR camera needs to be installed with a narrowband filter to filter out non-effective ambient light, that is, the IR camera can only sense near-infrared light. However, in the embodiments of the present application, the second camera 12 does not need to be installed with a narrowband filter.
[0049] In some embodiments of the present application, the terminal in the above imaging system can include a processor, a memory, and a computer program stored in the memory and executable on the processor, which can perform image processing on the color image collected by the first camera 11 and the speckle image collected by the second camera 12 to obtain the depth image of the target object.
[0050] In the embodiments of the present application, the second camera 12 and the speckle projector 13 can form a monocular structured light subsystem, and the first camera 11 and the second camera 12 form a pair of binocular stereo vision subsystems. The monocular structured light subsystem has a good three-dimensional reconstruction effect for weakly textured or textureless target objects, while the binocular stereo vision subsystem has an advantage in measuring long-distance target objects, and also has a certain reconstruction ability for target objects with low-reflectivity surfaces. At the same time, since the second camera 12 has the photosensitive capabilities of both an IR camera and an RGB camera, the number of IR cameras can be reduced, and the finally output depth image and color image can also be naturally aligned pixel by pixel.
[0051] Therefore, the embodiments of the present application improve the accuracy of the depth image while reducing the hardware cost, integrating the respective advantages of the monocular speckle structured light system and the passive binocular stereo vision system, and solving the disadvantages of the binocular speckle structured light system.
[0052] Specifically, Figure 2 FIG. shows a schematic implementation flow of a method for obtaining a depth image provided by an embodiment of the present application. This method can be applied to a terminal and is applicable to situations where it is necessary to improve the accuracy of the depth image while reducing the hardware cost.
[0053] Among them, the above terminal can be the terminal in the aforementioned imaging system or a terminal separated from the imaging system. Specifically, the above terminal can be a device such as a computer, a tablet computer, or a smart phone.
[0054] Specifically, the above method for obtaining a depth image may include the following steps S201 to S204.
[0055] Step S201: Obtain a color image obtained by the first camera photographing a target object.
[0056] Step S202: After controlling the speckle projector to project a preset speckle pattern onto the target object, obtain a speckle image obtained by the second camera photographing the target object.
[0057] In an embodiment of the present application, the terminal can generate a first control signal to control the speckle projector 13 to project a preset speckle pattern onto the target object.
[0058] Then, the terminal can generate a second control signal to control the first camera 11 and the second camera 12 to perform image acquisition respectively, and obtain the images acquired by the first camera 11 and the second camera 12. The terminal can also obtain the images actively acquired and uploaded by the first camera 11 and the second camera 12 after projecting the preset speckle pattern onto the target object.
[0059] In some specific embodiments of the present application, in a scene with a moving object, the terminal can control the first camera 11 and the second camera 12 to perform image acquisition at the same frequency, and analyze the depth image of the target object based on the color image and the speckle image acquired at the same sampling moment; while in a scene with a static object, the terminal can analyze the depth image of the target object by using the color image and the speckle image acquired at different acquisition moments.
[0060] Step S203: Determine a first depth image of the target object based on the speckle pattern projected onto the target object in the speckle image.
[0061] In some embodiments of the present application, after the speckle projector 13 projects a preset speckle pattern onto the target object, a speckle pattern will be formed on the surface of the target object, and the texture on the surface of the target object will affect parameters such as the shape and distance of each speckle in the speckle pattern on its surface. Therefore, based on the speckle pattern on the target object in the speckle image, the terminal can determine the first depth image of the target object. For a target object with weak texture or no texture, the first depth image can still better reflect the depth information of its surface.
[0062] Specifically, the terminal can identify each speckle in the speckle pattern in the speckle image, and calculate the first depth value of each pixel point corresponding to the surface of the target object in the speckle image based on the diffraction grating theorem and the triangulation method. The terminal can also use an existing neural network model, input the speckle image into the neural network model, and obtain the first depth image corresponding to the speckle image. The terminal can also match each speckle in the speckle pattern in the speckle image with the preset speckle pattern to obtain the first depth image corresponding to the speckle image.
[0063] It should be noted that other methods for obtaining the first depth image applied to the monocular speckle structured light system are also applicable to the present application. The specific manner adopted in step S203 can be selected according to the actual situation, and the present application does not limit this.
[0064] Step S204, extract and match the image features of the color image and the speckle image, and obtain the second depth image of the target object based on the matched image features and the first depth image.
[0065] Specifically, as Figure 3 shown, the above step S204 may include the following steps S301 to S303.
[0066] Step S301, extract the first feature image of the color image and the second feature image of the speckle image.
[0067] Specifically, in the embodiments of the present application, the terminal can extract the image features of the color image through the first feature extraction algorithm to obtain the first feature image. The pixel points of the first feature image can represent the image features of the pixel points at the same position in the color image.
[0068] Similarly, the image features of the speckle image can be extracted through the second feature extraction algorithm to obtain the second feature image. The pixel points of the second feature image can represent the image features of the pixel points at the same position in the speckle image.
[0069] Among them, the first feature extraction algorithm and the second feature extraction algorithm can be the same or different.
[0070] Since the spectral range in which the second camera 12 operates includes the infrared light spectrum and the visible light spectrum, the second feature image can characterize the image features of the visible light spectrum, and the first feature image and the second feature image can be used as the feature images of the left image and the right image in the passive binocular stereo vision system, respectively. At the same time, the second feature image can also characterize the image features of the near-infrared spectrum, and these image features include the features of the speckle pattern projected by the speckle projector 13 and formed on the target object.
[0071] It should be noted that, in order to ensure the accuracy of feature extraction, before performing feature extraction on the color image and the speckle image, the terminal can preprocess the color image and the speckle image respectively. The preprocessing can specifically include background removal, alignment, incremental processing, etc.
[0072] Step S302: Match the first feature image and the second feature image to obtain a first matching image, and obtain a first fused image by fusing the first matching image with the first depth image.
[0073] In some embodiments of the present application, the above terminal can calculate the matching features of the first pixel point in the first feature image and the second pixel point corresponding to the first pixel point in the second feature image at each parallax respectively, to obtain the first matching image, that is, the cost space c d . Wherein, the first pixel point is any pixel point in the first feature image. By traversing each pixel point of the first feature image, the first matching image can be finally obtained.
[0074] In some embodiments of the present application, the first matching image Wherein, <f RGB (x,y), f IR (x - d,y)> represents the inner product of f RGB (x,y) and f IR (x - d,y), N c is the number of channels of the feature, f RGB (x,y) represents the pixel value of the first pixel point in the first feature image and f IR (x - d,y) represents the second pixel point corresponding to the first pixel point in the second feature image under the condition that the parallax is d; an inner product operation needs to be performed for each parallax d.
[0075] In some other embodiments of the present application, the first matching image C concat (d,x,y) = Concat{f RGB (x,y), f IR (x - d,y)}, wherein, Concat{f RGB (x,y), f IR(x - d, y) represents the Concat operation on the first feature image and the second feature image at different disparities d.
[0076] Further, after obtaining the first matching image, the terminal can perform feature fusion on the first matching image and the first depth image, and adjust the first matching image according to the first depth image to obtain the first fused image.
[0077] It should be noted that the first matching image, that is, the cost space c d expresses the correlation degree of the first feature image and the second feature image during stereo matching at different disparities. Since the establishment of this cost space is based on a passive binocular stereo matching system, the accuracy is not as good as that of the structured light system when dealing with textureless areas. Therefore, the first depth image can be used to adjust the cost space to achieve the purpose of fusing the first matching image and the first depth image.
[0078] In one embodiment, there are multiple ways to perform feature fusion on the first matching image and the first depth image. Preferably, the distribution of the first matching image is adjusted using a Gaussian distribution to obtain the first fused image. The specific formula is:
[0079]
[0080] where h d is the adjusted cost space, that is, the first fused image, g ij is the corresponding disparity map obtained by reprojection of the first depth image, v ij is a binary mask. When g ij ≠0, v ij = 1, and k and σ are predetermined parameters that determine the shape of the Gaussian distribution.
[0081] Step S303, perform disparity estimation on the first fused image to obtain the disparity map of the target object, and determine the second depth image of the target object based on the disparity map.
[0082] Specifically, the above-mentioned disparity map can be obtained by performing a soft argmin operation on the first fused image. The disparity values in the disparity map where D max is the maximum disparity value, d ∈ [0, D max ). σ(c d ) represents performing a softmax operation on the cost space c d . After obtaining the disparity map, the second depth image can be obtained through a linear transformation. It should be understood that the disparity map can be obtained by performing a single softmax operation, or can be obtained through multiple iterations under the supervision of a loss function after performing the softmax operation. There is no limitation here.
[0083] In some embodiments of the present application, as Figure 4 shown, the above step S204 can be implemented by a trained stereo matching network model. Specifically, the terminal can input the color image, the speckle image, and the first depth image into the stereo matching network model, and obtain the second depth image output by the stereo matching network model.
[0084] Specifically, the stereo matching network may include a first feature extraction layer, a second feature extraction layer, a feature matching layer, a depth feature fusion layer, and a disparity estimation layer.
[0085] Among them, the first feature extraction layer can extract the first feature image of the color image based on the feature extraction algorithm. The second feature extraction layer can extract the second feature image of the IR speckle image based on the feature extraction algorithm.
[0086] The feature matching layer can calculate the matching features of the first pixel point in the first feature image and the second pixel point corresponding to the first pixel point in the second feature image at each disparity, and obtain the first matching image.
[0087] After obtaining the first matching image, the depth feature fusion layer can perform feature fusion on the first matching image and the first depth image to obtain the first fusion image.
[0088] The disparity estimation layer can obtain the disparity map of the target object by performing disparity estimation on the first fusion image, and determine the second depth image based on the disparity map.
[0089] In some embodiments of the present application, the training process of the above stereo matching network model may include: obtaining a sample set, where each sample in the sample set includes a sample color image, a sample speckle image, and a sample first depth image. Randomly extract samples from the sample set and input them into the neural network model to be trained, calculate the error between the output depth image and the reference depth image corresponding to the sample, and adjust the parameters and weights in the neural network model according to the error size, and re-extract samples from the sample set for training until the number of training iterations reaches a preset threshold to complete the training. And use the neural network model obtained after the training is completed as the above stereo matching network model.
[0090] In the embodiments of the present application, a color image obtained by the first camera 11 photographing a target object and a speckle image obtained by the second camera 12 photographing the target object are acquired. Based on the speckle pattern projected onto the target object in the speckle image, a first depth image of the target object is determined. At this time, the first depth image is a depth image obtained based on the monocular structured light photon system in the imaging system, and has a good 3D reconstruction effect for target objects with weak or no texture. The color image and the speckle image can be used as the left image and the right image in a binocular stereo vision system respectively, which has advantages in measuring long-distance target objects, and also has a certain reconstruction ability for target objects with low-reflectivity surfaces. A second depth image of the target object is obtained based on the color image, the speckle image and the first depth image, which improves the accuracy of the depth image while reducing the number of cameras.
[0091] That is, compared with the passive binocular stereo vision system, the solution provided by the present application can output reliable depth information for both target objects with weak or no texture to be measured and target objects with rich texture. Compared with the monocular speckle structured light system, the solution provided by the present application has a larger measurement distance range, better performance for low-albedo targets, and at the same time, can degrade into an ordinary passive binocular stereo vision system to work in an outdoor environment with strong sunlight interference. Compared with the binocular speckle structured light system, the solution provided by the present application can reduce one IR camera. And, since the second camera 12 has the ability to sense visible light and infrared light, which is equivalent to the optical centers of the RGB camera and the IR camera coinciding, there is no occlusion area in the solution provided by the present application. Therefore, there will be no holes in the depth image, and the depth image is naturally aligned with the RGB image pixel by pixel.
[0092] Figure 5a shows the color image collected by the first camera 11 and Figure 5b shows the speckle image collected by the second camera 12, Figure 5c shows the first depth image obtained by processing the monocular structured light photon system, Figure 5d shows the second depth image output by the stereo matching network model. It can be seen from the images that the imaging system provided by the present application and the acquisition of the corresponding depth image can obtain a depth image with higher accuracy.
[0093] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences.
[0094] Such as Figure 6The following is a schematic structural diagram of an apparatus 600 for acquiring a depth image provided by an embodiment of the present application. The apparatus 600 for acquiring a depth image is configured on a terminal.
[0095] Specifically, the apparatus 600 for acquiring a depth image may include:
[0096] A color image acquisition unit 601, configured to acquire a color image obtained by a first camera photographing a target object;
[0097] A speckle image acquisition unit 602, configured to, after controlling a speckle projector to project a preset speckle pattern onto the target object, acquire a speckle image obtained by a second camera photographing the target object;
[0098] A monocular structured light unit 603, configured to determine a first depth image of the target object based on the speckle pattern projected onto the target object in the speckle image;
[0099] A depth image acquisition unit 604, configured to extract and match image features of the color image and the speckle image, and obtain a second depth image of the target object based on the matched image features and the first depth image.
[0100] In some embodiments of the present application, the above-mentioned depth image acquisition unit 604 may specifically be configured to: extract a first feature image of the color image and a second feature image of the speckle image; perform feature matching on the first feature image and the second feature image, and perform feature fusion on the first matching image obtained after feature matching and the depth image to obtain a first fusion image; perform disparity estimation on the first fusion image to obtain a disparity map of the target object, and determine the second depth image based on the disparity map.
[0101] In some embodiments of the present application, the above-mentioned depth image acquisition unit 604 may specifically be configured to: extract and match image features of the color image and the speckle image through a preset stereo matching network model, and obtain a second depth image of the target object based on the matched image features and the first depth image, where the preset stereo matching network model includes a first feature extraction layer, a second feature extraction layer, a feature matching layer, a depth feature fusion layer, and a disparity estimation layer.
[0102] In some embodiments of the present application, the above-mentioned depth image acquisition unit 604 may specifically be configured to: calculate matching features of a first pixel point in the first feature image and a second pixel point corresponding to the first pixel point in the second feature image at each disparity, to obtain the first matching image.
[0103] In some embodiments of the present application, the above-mentioned depth image acquisition unit 604 may specifically be configured to: according to the first depth image, adjust the distribution of the first matching image through a Gaussian distribution to obtain the first fusion image, and the adjustment formula is: where h d is the adjusted cost space, that is, the first fusion image, g ij is the corresponding disparity map obtained by reprojection of the first depth image, v ij is a binary mask, when g ij ≠0, v ij =1, k and σ are predetermined parameters that determine the shape of the Gaussian distribution.
[0104] In some embodiments of the present application, the spectral range of the above-mentioned second camera includes an infrared light spectrum and a visible light spectrum.
[0105] It should be noted that for the convenience and brevity of description, the specific working process of the above-mentioned depth image acquisition device 600 can refer to Figures 1 to 4 and Figures 5a to 5d the corresponding process of the method, which will not be elaborated here.
[0106] As Figure 7 shown, it is a schematic diagram of a terminal provided by an embodiment of the present application. The terminal 7 may include: a processor 70, a memory 71, and a computer program 72 stored in the memory 71 and executable on the processor 70, such as a depth image acquisition program. When the processor 70 executes the computer program 72, the steps in the above-mentioned embodiments of the depth image acquisition method are implemented, such as Figure 1 the steps S201 to S204 shown. Alternatively, when the processor 70 executes the computer program 72, the functions of each module / unit in the above-mentioned device embodiments are implemented, such as Figure 6 the color image acquisition unit 601, the speckle image acquisition unit 602, the monocular structured light unit 603, and the depth image acquisition unit 604 shown.
[0107] The computer program may be divided into one or more modules / units. The one or more modules / units are stored in the memory 71 and executed by the processor 70 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal.
[0108] For example, the computer program may be divided into: a color image acquisition unit, a speckle image acquisition unit, a monocular structured light unit, and a depth image acquisition unit.
[0109] The specific functions of each unit are as follows: a color image acquisition unit for acquiring a color image obtained by photographing a target object with a first camera; a speckle image acquisition unit for controlling a speckle projector to project a preset speckle pattern onto the target object and then acquiring a speckle image obtained by photographing the target object with a second camera; a monocular structured light unit for determining a first depth image of the target object based on the speckle pattern projected onto the target object in the speckle image; a depth image acquisition unit for extracting and matching the image features of the color image and the speckle image and obtaining a second depth image of the target object according to the matched image features and the first depth image.
[0110] The terminal may include, but is not limited to, a processor 70 and a memory 71. Those skilled in the art can understand that Figure 7 merely examples of the terminal, which do not constitute a limitation on the terminal, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the terminal may also include input / output devices, network access devices, buses, etc.
[0111] The so-called processor 70 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0112] The memory 71 may be an internal storage unit of the terminal, such as the hard disk or memory of the terminal. The memory 71 may also be an external storage device of the terminal, such as a plug-in hard disk equipped on the terminal, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 71 may also include both the internal storage unit and the external storage device of the terminal. The memory 71 is used to store the computer program and other programs and data required by the terminal. The memory 71 may also be used to temporarily store data that has been output or will be output.
[0113] It should be noted that for the convenience and brevity of description, the structure of the above terminal may also refer to the specific description of the structure in the method embodiments, which will not be elaborated here.
[0114] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated here.
[0115] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0116] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0117] In the embodiments provided in the present application, it should be understood that the disclosed device / terminal and method can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For example, the division of the module or unit is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0118] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0119] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0120] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0121] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for obtaining a depth image, characterized in that, Including: Obtain a color image obtained by a first camera photographing a target object; After controlling a speckle projector to project a preset speckle pattern onto the target object, obtain a speckle image obtained by a second camera photographing the target object, where the spectral range in which the second camera operates includes the infrared light spectrum and the visible light spectrum; Based on the speckle pattern on the target object in the speckle image, determine a first depth image of the target object; Extract and match the image features of the color image and the speckle image, and obtain a second depth image of the target object according to the matched image features and the first depth image; Among them, the extracting and matching the image features of the color image and the speckle image, and obtaining the second depth image of the target object according to the matched image features and the first depth image includes: Extract a first feature image of the color image and a second feature image of the speckle image; Match the first feature image and the second feature image to obtain a first matching image, and obtain a first fused image by fusing the first matching image and the first depth image; Perform parallax estimation on the first fused image to obtain a parallax map of the target object, and determine the second depth image of the target object according to the parallax map.
2. The method for obtaining a depth image according to claim 1, wherein The extracting and matching the image features of the color image and the speckle image, and obtaining the second depth image of the target object according to the matched image features and the first depth image includes: Extract and match the image features of the color image and the speckle image through a preset stereo matching network model, and obtain a second depth image of the target object based on the matched image features and the first depth image, where the preset stereo matching network model includes a first feature extraction layer, a second feature extraction layer, a feature matching layer, a depth feature fusion layer, and a parallax estimation layer.
3. The method for obtaining a depth image according to claim 1, characterized in that, The matching the first feature image and the second feature image to obtain a first matching image includes: Calculate the matching features of the first pixel point in the first feature image and the second pixel point corresponding to the first pixel point in the second feature image at each parallax respectively, to obtain the first matching image.
4. The method for obtaining a depth image according to claim 1, wherein The obtaining the first fused image by fusing the first matching image and the first depth image includes: According to the first depth image, adjust the distribution of the first matching image through a Gaussian distribution to obtain the first fused image, and the adjustment formula is: ; Among them, is the adjusted cost space, that is, the first fused image, is the corresponding disparity map obtained by reprojection of the first depth image, is a binary mask. When is true, , and are predetermined parameters that determine the shape of the Gaussian distribution, is the cost space before adjustment, that is, the first matching image.
5. An apparatus for obtaining a depth image, characterized in that, Including: A color image acquisition unit, configured to obtain a color image obtained by a first camera photographing a target object; A speckle image acquisition unit, configured to, after controlling a speckle projector to project a preset speckle pattern onto the target object, obtain a speckle image obtained by a second camera photographing the target object, where the spectral range in which the second camera operates includes the infrared light spectrum and the visible light spectrum; A monocular structured light unit, configured to determine a first depth image of the target object based on the speckle pattern projected onto the target object in the speckle image; A depth image acquisition unit, configured to extract and match image features of the color image and the speckle image, and obtain a second depth image of the target object according to the matched image features and the first depth image; Wherein, the depth image acquisition unit is specifically configured to: extract a first feature image of the color image and a second feature image of the speckle image; match the first feature image and the second feature image to obtain a first matching image, and obtain a first fused image by fusing the first matching image and the first depth image; Perform disparity estimation on the first fused image to obtain a disparity map of the target object, and determine the second depth image of the target object according to the disparity map.
6. A terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.
7. A depth imaging system, characterized in that, Including a first camera, a second camera, a speckle projector and the terminal according to claim 6, wherein the first camera, the second camera and the speckle projector are arranged on the same plane; The speckle projector is configured to project a preset speckle pattern onto the target object; The first camera is configured to collect a color image of the target object; The second camera is configured to collect a speckle image of the target object; The terminal is configured to obtain a depth image of the target object by using the color image and the speckle image.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Binocular depth camera system and depth image generating method
CN108234984A
Depth image acquisition method, terminal and computer readable storage medium
CN112150528A
Depth map colorization method based on monocular depth camera
CN112700484A
Image generation method and electronic equipment
CN113592754A