Image generation method, electronic equipment and computer readable storage medium
By combining a binocular camera and a TOF camera, image interpolation and depth value correction are performed, solving the problem of inaccurate depth measurement in complex environments using traditional binocular depth measurement methods, and achieving high-precision depth estimation in complex environments.
Patent Information
- Application Number
- CN202511063484.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional binocular depth measurement methods are prone to mapping errors in scenarios with repetitive textures, weak textures, and low-light environments, resulting in inaccurate image depth.
By combining a stereo camera and a Time-of-Flight (TOF) camera, stereo images and TOF depth images are acquired, and interpolation processing and depth value correction are performed. A continuous depth surface is generated using TOF pixels to improve image resolution, and the TOF depth image is corrected using stereo images to improve depth accuracy.
It provides reliable depth values in scenarios with repetitive textures, weak textures, and low-light environments, significantly improving the depth accuracy of target depth images, especially the accuracy of object edges and boundary areas.
Smart Images

Figure CN120997271A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to an image generation method, an electronic device and a computer readable storage medium. BACKGROUND
[0002] With the development of computer vision technology, depth information estimation technology has gradually developed into one of the research hotspots in the field of computer vision, and is widely used in navigation system analysis, stereoscopic video generation, virtual reality and other systems.
[0003] At present, a binocular depth measurement method is usually used to obtain depth information, that is, a binocular camera is used to obtain an image, and a depth image is recovered from the image. However, the traditional binocular depth measurement method is prone to mapping errors in scenes such as repetitive texture, weak texture and weak light environment, resulting in inaccurate image depth. SUMMARY
[0004] To solve the above technical problems, the present application provides an image generation method, an electronic device and a computer readable storage medium.
[0005] The first aspect of the present application provides an image generation method, which comprises the steps of: acquiring a binocular image of a photographed object by a binocular camera, and acquiring a first TOF depth image of the photographed object by a TOF camera; determining a binocular depth image based on the binocular image; performing interpolation processing on the first TOF depth image to obtain a second TOF depth image; correcting the depth value of the second TOF depth image based on the binocular image to obtain a third TOF depth image; and correcting the depth value of the binocular depth image based on the third TOF depth image to obtain a target depth image.
[0006] The image generation method provided by the present application can generate a continuous depth surface by interpolation processing on the first TOF depth image, improve the image resolution, and facilitate subsequent correction of the binocular depth image. By using the binocular image to correct the depth value of the second TOF depth image, the depth accuracy of the object edge region and the object boundary region in the second TOF depth image can be improved. The TOF depth measurement method can provide reliable depth values in scenes such as repetitive texture, weak texture and weak light environment. By using the third TOF depth image to correct the binocular depth image, the accuracy of the depth value of the binocular depth image can be improved, thereby improving the depth accuracy of the target depth image obtained in scenes such as repetitive texture, weak texture, object edge and weak light environment.
[0007] The second aspect of the present application provides an electronic device, comprising a binocular camera, a TOF camera, a processor and a memory, the memory stores a computer program, and the processor executes the computer program to perform the image generation method provided in the first aspect.
[0008] The electronic device provided in the second aspect has the corresponding features of the image generation method provided in the first aspect, and thus can achieve the same or corresponding beneficial effects as the method provided in the first aspect, which will not be described here.
[0009] The third aspect of the present application provides a computer readable storage medium, comprising a processor and a memory, the memory stores a computer program, and the processor executes the computer program to perform the image generation method provided in the first aspect.
[0010] The computer readable storage medium provided in the third aspect has the corresponding features of the image generation method provided in the first aspect, and thus can achieve the same or corresponding beneficial effects as the method provided in the first aspect, which will not be described here. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0012] Figure 1 The flowchart of the image generation method in some embodiments of the present application.
[0013] Figure 2 The schematic diagram of the pixel points of the first TOF depth image and the second TOF depth image in some embodiments of the present application.
[0014] Figure 3 The schematic diagram of the pixel points of the first TOF depth image and the second TOF depth image in some embodiments of the present application.
[0015] Figure 4 The sub-flowchart of step S14 in some embodiments.
[0016] Figure 5 The sub-flowchart of step S21 in some embodiments.
[0017] Figure 6 The schematic diagram of the pixel points of the target image in some embodiments of the present application.
[0018] Figure 7 The schematic diagram of the target image in some embodiments of the present application.
[0019] Figure 8 A schematic diagram of a third TOF depth image in some embodiments of the present application.
[0020] Figure 9 A schematic diagram of a pixel of a binocular depth image and a third TOF depth image in some embodiments of the present application.
[0021] Figure 10 A subflowchart of step S15 in some embodiments.
[0022] Figure 11 A subflowchart of step S42 in some embodiments.
[0023] Figure 12 A subflowchart of step S12 in some embodiments.
[0024] Figure 13 A schematic diagram of a binocular depth image in some embodiments of the present application.
[0025] Figure 14 A schematic diagram of a third TOF depth image in some embodiments of the present application.
[0026] Figure 15 A schematic diagram of an RGBD image in some embodiments of the present application.
[0027] Figure 16 A schematic diagram of a first trigger signal of a binocular camera and a second trigger signal of a TOF camera in some embodiments of the present application.
[0028] Figure 17 A block diagram of an electronic device in some embodiments of the present application.
[0029] Figure 18 A schematic diagram of a TOF camera and a binocular camera in some embodiments of the present application.
[0030] Figure 19 A schematic diagram of a dot array light source of a TOF camera in some embodiments of the present application.
[0031] Figure 20 A schematic diagram of an electronic device in some embodiments of the present application.
[0032] Figure 21 A schematic diagram of an electronic device in some other embodiments of the present application. DETAILED DESCRIPTION
[0033] With reference to the drawings and embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort belong to the scope of the present application.
[0034] In the description of the present application, the terms "first", "second", "third", "fourth" and the like are used to distinguish different objects, and are not used to describe a specific sequence, and therefore cannot be understood as a limitation on the present application.
[0035] In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, a / b can represent a or b; "and / or" in the text only describes the association relationship of the associated objects, which means that there can be three relationships, for example, a and / or b can represent: a exists alone, a and b exist together, and b exists alone, in addition, in the description of the embodiments of the present application, "multiple" means two or more than two, and "multiple" means two or more than two.
[0036] In the present application, "embodiment" means that the specific features, structures or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase is shown at various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment to other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in the present application can be combined with other embodiments.
[0037] The embodiments of the present application provide an image generation method, which is applied to an electronic device, and the electronic device includes a binocular camera and a TOF camera.
[0038] The electronic device can be various types of devices, for example, can be a humanoid robot, an industrial robot, a sweeping machine, a drone, an automatic driving system, a smart phone, a tablet, a computer, a virtual reality (VR) device, an augmented reality (AR) device, etc.
[0039] The binocular camera includes two cameras spaced apart by a certain distance. The TOF camera can be an iTOF camera including a sensor based on indirect time-of-flight depth sensing technology or a dTOF camera including a sensor based on direct time-of-flight depth sensing technology.
[0040] Please refer to Figure 1 , the flowchart of the image generation method in some embodiments of the present application. As Figure 1 shown, the image generation method includes the following steps: S11: acquiring a binocular image of a photographed object by the binocular camera and acquiring a first TOF depth image of the photographed object by the TOF camera.
[0041] S12: determining a binocular depth image based on the binocular image.
[0042] S13: performing interpolation processing on the first TOF depth image to obtain a second TOF depth image.
[0043] S14: correcting a depth value of the second TOF depth image based on the binocular image to obtain a third TOF depth image.
[0044] S15: correcting a depth value of the binocular depth image based on the third TOF depth image to obtain a target depth image.
[0045] The image generation method provided by the embodiment of the present application can generate a continuous depth surface by performing interpolation processing on the first TOF depth image, improve the image resolution, and facilitate subsequent correction of the binocular depth image. The depth precision of an object edge region and an object junction region in the second TOF depth image can be improved by correcting the depth value of the second TOF depth image using the binocular image. The TOF depth measurement method can provide reliable depth values in scenes such as repetitive textures, weak textures, and weak light environments. The depth value precision of the binocular depth image can be improved by correcting the binocular depth image using the third TOF depth image, thereby improving the depth precision of the target depth image obtained in scenes such as repetitive textures, weak textures, object edges, and weak light environments.
[0046] Since the resolution of the TOF camera is lower than that of the binocular camera, the number of pixels of the first TOF depth image is less than that of the binocular image and the binocular depth image. The number of pixels can be increased by performing interpolation processing on the first TOF depth image, the difference in the number of pixels between the second TOF depth image and the binocular image and the binocular depth image can be reduced, that is, the resolution difference between the second TOF depth image and the binocular depth image can be reduced, thereby facilitating subsequent pixel-level correction of the depth value of the second TOF depth image based on the binocular image and pixel-level correction of the depth value of the binocular depth image based on the third TOF depth image.
[0047] In step S11, the binocular camera and the TOF camera synchronously acquire images, that is, the binocular image and the first TOF depth image are acquired at the same time.
[0048] The first TOF depth image includes a depth value of each pixel point of the first TOF depth image, and the binocular depth image includes a depth value of each pixel point of the binocular depth image.
[0049] In some embodiments, the first TOF depth image includes a plurality of TOF pixel points. The step S13 of performing interpolation processing on the first TOF depth image to obtain a second TOF depth image includes: performing interpolation processing on the first TOF depth image to obtain a plurality of interpolation pixel points, so as to obtain the second TOF depth image, wherein the second TOF depth image includes the plurality of TOF pixel points and the plurality of interpolation pixel points.
[0050] By inserting the plurality of interpolation pixel points in the first TOF depth image, the second TOF depth image is obtained, and the resolution of the second TOF depth image is improved, which is beneficial to subsequent pixel mapping operations of the second TOF depth image and a target image.
[0051] For example, refer to Figure 2 , Figure 2 The figure is a schematic diagram of the first TOF depth image Image_TOF1 and the second TOF depth image Image_TOF2 in some embodiments of the present application. Figure 2 The (a) and (b) in FIG. 1 respectively show the first TOF depth image Image_TOF1 and the second TOF depth image Image_TOF2. The number and resolution of the pixel points of the second TOF depth image Image_TOF2 are greater than those of the first TOF depth image Image_TOF1.
[0052] For example, refer to Figure 3 , Figure 3 The figure is a schematic diagram of the pixel points of the first TOF depth image Image_TOF1 and the second TOF depth image Image_TOF2 in some embodiments of the present application. Figure 3 The (a) and (b) in FIG. 2 respectively show the pixel points of the first TOF depth image Image_TOF1 and the second TOF depth image Image_TOF2. The first TOF depth image Image_TOF1 includes a plurality of TOF pixel points P_TOF, and the second TOF depth image Image_TOF2 includes the plurality of TOF pixel points P_TOF and a plurality of interpolation pixel points P_in.
[0053] In some embodiments, the aforementioned interpolating the first TOF depth image Image_TOF1 to obtain the plurality of interpolated pixel points P_in includes: taking a plurality of TOF pixel points P_TOF adjacent to a pixel point to be interpolated as a reference point set; determining an interpolation weight corresponding to a depth value of each TOF pixel point P_TOF in the reference point set according to a distance between the TOF pixel point P_TOF and the pixel point to be interpolated; and performing weighted fusion on the depth values of the plurality of TOF pixel points P_TOF in the reference point set according to the corresponding interpolation weights to obtain the depth value of the pixel point to be interpolated, and generating the interpolated pixel point P_in based on the depth value of the pixel point to be interpolated. The depth value of the pixel point to be interpolated obtained by the weighted fusion is the depth value of the interpolated pixel point P_in inserted at the pixel point to be interpolated.
[0054] By assigning the weight of the depth value of the TOF pixel point P_TOF according to the distance between the TOF pixel point P_TOF and the pixel point to be interpolated, more accurate interpolation processing can be achieved.
[0055] In some embodiments, the interpolation weight corresponding to the depth value of the TOF pixel point P_TOF is negatively correlated with the distance between the TOF pixel point P_TOF and the pixel point to be interpolated. That is, the smaller the distance, the greater the interpolation weight; otherwise, the smaller the interpolation weight.
[0056] In some embodiments, a radial basis function can be used to interpolate the first TOF depth image Image_TOF1.
[0057] In some embodiments, the method further includes determining a confidence degree of the depth value of each TOF pixel point P_TOF in the reference point set.
[0058] In some embodiments, the determination of the confidence degree of the depth value of each TOF pixel point P_TOF in the reference point set includes determining the confidence degree according to a reflection signal intensity value and / or a depth value stability of each TOF pixel point P_TOF.
[0059] In some embodiments, the depth value of the TOF pixel point P_TOF in the first TOF depth image of a plurality of continuous frames can be averaged, and a standard deviation can be calculated, and then the depth value stability can be determined based on the standard deviation. The smaller the standard deviation, the higher the depth value stability; otherwise, the lower the depth value stability.
[0060] In some embodiments, the higher the reflection signal intensity value, the higher the confidence degree; otherwise, the lower the confidence degree. The smaller the standard deviation, the higher the confidence degree; otherwise, the lower the confidence degree.
[0061] The foregoing determining the interpolation weight corresponding to the depth value of each TOF pixel point P_TOF according to the distance between each TOF pixel point P_TOF in the reference point set and the pixel point to be interpolated includes: determining the interpolation weight corresponding to the depth value of each TOF pixel point P_TOF according to the distance between each TOF pixel point P_TOF in the reference point set and the pixel point to be interpolated and the confidence degree of the depth value of each TOF pixel point P_TOF.
[0062] In some embodiments, the weight corresponding to the depth value of the TOF pixel point P_TOF is positively correlated with the confidence degree, which is beneficial to improving the accuracy of interpolation.
[0063] Please refer to Figure 4 , Figure 4 The flowchart in FIG. 14 is a sub-flowchart of step S14 in some embodiments. In some embodiments, as shown in FIG. 15, the step S14 of “correcting the depth value of the second TOF depth image based on the binocular image” includes: Figure 4 S21: determining a target TOF pixel point adjacent to the interpolation pixel point and located in the same spatial structure in the second TOF depth image Image_TOF2 based on the binocular image.
[0064] S22: when it is determined that the interpolation pixel point P_in is a to-be-corrected interpolation point based on the interpolation pixel point P_in and the target TOF pixel point, determining a first target depth value based on the depth value of the target TOF pixel point adjacent to the to-be-corrected interpolation pixel point and located in the same spatial structure.
[0065] S23: correcting the depth value of the to-be-corrected interpolation pixel point to the first target depth value.
[0066] The spatial structure can include one or more of an object surface, an object edge region, and an object intersection region.
[0067] For example, the shooting object is a first wall, a second wall, and a wall corner formed by the intersection of the first wall and the second wall, the surface of the first wall is an object surface, the surface of the second wall is an object surface, and the wall corner is an object intersection region. When the interpolation pixel point is located on the first wall, the target TOF pixel point is also located on the first wall. When the interpolation pixel point is located on the wall corner, the target TOF pixel point is also located on the wall corner.
[0068] As another example, the subjects being photographed are a wall and a decorative painting hanging on the wall. The surface of the wall is an object surface, the surface of the decorative painting facing away from the wall is an object surface, and the edge of the frame of the decorative painting is an object edge region. When the interpolated pixel is located on the wall, the target TOF pixel is also located on the wall; when the interpolated pixel is located at the edge of the frame, the target TOF pixel is also located at the edge of the frame.
[0069] The target TOF pixel can be one or more TOF pixels, that is, the number of target TOF pixels can be one or more.
[0070] In some embodiments, in step S22, the first target depth value can be calculated by bilinear interpolation, bicubic interpolation, natural neighborhood interpolation, mean interpolation, distance-weighted average interpolation, etc.
[0071] This application embodiment utilizes the physical constraints of three-dimensional space, using the depth values of target TOF pixels located in the same spatial structure to correct the depth values of the interpolated pixels to be corrected. This can avoid the problem of depth distortion, especially for object edge regions and object boundary regions, where depth abrupt changes may cause depth distortion after interpolation. Using the above correction method, the depth values of the interpolated pixels to be corrected can be corrected to be close to the true depth values, thereby significantly improving the depth accuracy of the second TOF depth image Image_TOF2.
[0072] For example, such as Figure 3 As shown in (b), the second TOF depth image Image_TOF2 includes the interpolation pixel P_in_cor to be corrected, and the target TOF pixel P_TOF1 that is adjacent to the interpolation pixel P_in_cor and located in the same spatial structure.
[0073] In some embodiments, the binocular image includes a first image and a second image acquired simultaneously by two cameras, that is, the first image and the second image acquired by the two cameras at the same time.
[0074] Please see Figure 5 , Figure 5 This is a sub-flowchart of step S21 in some embodiments. In some embodiments, such as... Figure 5 As shown, step S21, "based on the binocular image and the second TOF depth image Image_TOF2, determining the interpolation pixel to be corrected in the second TOF depth image Image_TOF2 and the target TOF pixel located in the same spatial structure as the interpolation pixel to be corrected," includes: S31: determining a target image based on the binocular image, wherein the target image is obtained based on one of the first image, the second image and a third image, and the third image is obtained by fusing the first image and the second image.
[0075] S32: performing pixel point mapping on the target image and the second TOF depth image Image_TOF2 to determine a first mapping pixel point of the target image corresponding to each interpolation pixel point P_in and a second mapping pixel point of the target image corresponding to each TOF pixel point P_TOF.
[0076] S33: determining a second mapping pixel point adjacent to the first mapping pixel point and located in a same spatial structure based on the target image.
[0077] S34: taking the TOF pixel point P_TOF corresponding to the determined second mapping pixel point as a target TOF pixel point adjacent to the interpolation pixel point P_in and located in the same spatial structure.
[0078] The foregoing determination of the interpolation pixel point P_in as the to-be-corrected interpolation pixel point P_in_cor based on the interpolation pixel point P_in and the target TOF pixel point includes: when an absolute value of a difference between a depth value of the interpolation pixel point P_in and a depth value of the target TOF pixel point is greater than a first preset depth difference, determining the interpolation pixel point P_in as the to-be-corrected interpolation pixel point P_in_cor, and determining the target TOF pixel point adjacent to the interpolation pixel point P_in and located in the same spatial structure as the target TOF pixel point P_TOF1 adjacent to the to-be-corrected interpolation pixel point P_in_cor and located in the same spatial structure.
[0079] When an absolute value of a difference between the depth value of the interpolation pixel point P_in and a depth value of the target TOF pixel point adjacent to the interpolation pixel point P_in and located in the same spatial structure is less than or equal to the first preset depth difference, it is determined that the depth value of the interpolation pixel point P_in does not need to be corrected and remains an original value.
[0080] The number of the second mapping pixel points adjacent to the first mapping pixel point and located in the same spatial structure with the first mapping pixel point can be one or more, that is, the number of the target TOF pixel points can be one or more. When the number of the target TOF pixel points is multiple, the depth value of the interpolation pixel point P_in can be compared with an arithmetic mean value or a weighted mean value of the depth values of the multiple target TOF pixel points.
[0081] The first preset depth difference can be set according to actual requirements.
[0082] In some embodiments, the first image and the second image can be fused to obtain the third image by a weighted average fusion method, a pyramid-based multi-scale fusion method, a deep learning-based fusion method, or the like. For example, the deep learning-based fusion method uses a CNN (Convolutional Neural Network) or a Transformer architecture, inputs the first image and the second image, and directly outputs the fused third image.
[0083] In some embodiments, the resolution of the second TOF depth image Image_TOF2 is the same as that of the target image. Each interpolation pixel point P_in corresponds to a first mapping pixel, and each TOF pixel point P_TOF corresponds to a second mapping pixel.
[0084] In some embodiments, the resolution of the second TOF depth image Image_TOF2 is the same as that of the first image, the second image, or the third image, and the target image is one of the first image, the second image, or the third image. In the interpolation process, the number of interpolation pixel points P_in can be controlled to make the resolutions the same. In other embodiments, the resolution of the second TOF depth image Image_TOF2 is different from that of the first image, the second image, or the third image. In this case, one of the first image, the second image, or the third image can be up-sampled or down-sampled to obtain the target image, which has the same resolution as the second TOF depth image Image_TOF2.
[0085] In other embodiments, the resolution of the second TOF depth image Image_TOF2 is different from that of the target image. Each interpolation pixel point P_in corresponds to at least one first mapping pixel adjacent to the interpolation pixel point P_in, and each TOF pixel point corresponds to at least one second mapping pixel adjacent to the interpolation pixel point P_in.
[0086] In some embodiments, the method further includes calibrating the intrinsic parameters of the binocular camera and the intrinsic parameters of the TOF camera, respectively, to correct the lens distortion of the binocular camera and the TOF camera, and to establish the mapping relationship between the pixel coordinate system and the camera coordinate system of each camera.
[0087] In some embodiments, the method further includes calibrating the intrinsic parameters of the binocular camera and the intrinsic parameters of the TOF camera, respectively, to correct the lens distortion of the binocular camera and the TOF camera, and to establish the mapping relationship between the pixel coordinate system and the camera coordinate system of each camera.
[0088] Mapping the pixel points of the second TOF depth image Image_TOF2 to the target image can include: converting the pixel points of the second TOF depth image Image_TOF2 from the camera coordinate system of the TOF camera to the camera coordinate system of the binocular camera based on the spatial position relationship, and mapping the pixel points of the second TOF depth image Image_TOF2 in the camera coordinate system of the binocular camera to the pixel coordinate system of the binocular camera by using the mapping relationship, so as to realize pixel point mapping. The process of mapping the pixel points of the target image to the second TOF depth image Image_TOF2 is opposite.
[0089] In some embodiments, the determining, based on the target image, the second mapped pixel points adjacent to the first mapped pixel point and located in the same spatial structure includes: extracting image features based on the target image, the image features including one or more of color features, texture features, and semantic features; dividing the target image into a plurality of spatial structures based on the image features; and determining, based on the spatial structure in which the first mapped pixel point is located, the second mapped pixel points adjacent to the first mapped pixel point and located in the same spatial structure.
[0090] For example, the photographed object is a first wall, a second wall, and a wall corner formed by the intersection of the first wall and the second wall. By performing image feature extraction on the target image, the first wall, the second wall, and the wall corner can be recognized, and the target image can be divided into the three spatial structures.
[0091] Using the color, texture, semantic, and other information provided by the high-resolution target image as guidance, the object surface, object edge region, object boundary region, and other objects in the photographed scene can be accurately recognized, i.e., each spatial structure in the target image can be accurately recognized. By recognizing the spatial structure of the target image, the target TOF pixel point located in the same spatial structure as the interpolation pixel point P_in and adjacent to the interpolation pixel point P_in can be accurately found, and then the depth value of the inserted interpolation pixel point P_in can be accurately judged by the TOF true value point as to whether the depth value is distorted, i.e., whether the error is too large, so that the depth value of the interpolation pixel point P_in can be corrected by using the depth value of the target TOF pixel point when the depth value is distorted, thereby improving the accuracy of the depth value of the interpolation pixel point P_in.
[0092] For example, please refer to Figure 6 , Figure 6Fig. 1 is a schematic diagram of a target image Image_t in some embodiments of the present application. After pixel mapping of the target image Image_t and the second TOF depth image Image_TOF2, the pixels of the target image Image_t are divided into a plurality of first mapped pixels P_m1 and a plurality of second mapped pixels P_m2. Image features of the target image Image_t are extracted, and the target image Image_t is divided into an image region A corresponding to a first wall, an image region B corresponding to a second wall, and an image region C corresponding to a corner based on the image features. Please refer to Figure 3 Fig. 1(b) and Figure 6 Fig. 1(c), the corner interpolation pixel P_in' corresponds to the corner first mapped pixel P_m1', a plurality of corner second mapped pixels P_m2' adjacent to the corner first mapped pixel P_m1' and in the same spatial structure, and the plurality of corner second mapped pixels P_m2' correspond to a plurality of adjacent corner interpolation pixels P_in' and in the same spatial structure.
[0093] When the absolute value of the difference between the depth value of the corner interpolation pixel P_in' and the depth values of the plurality of adjacent corner interpolation pixels P_in' and in the same spatial structure is greater than a first preset depth difference, the corner interpolation pixel P_in' is a to-be-corrected interpolation pixel P_in_cor, and the plurality of adjacent corner interpolation pixels P_in' and in the same spatial structure are a plurality of adjacent to-be-corrected interpolation pixels P_in_cor and in the same spatial structure. The first target depth value is determined based on the depth values of the plurality of adjacent to-be-corrected interpolation pixels P_in_cor and in the same spatial structure, and the depth value of the to-be-corrected interpolation pixel P_in_cor is corrected to the first target depth value.
[0094] Fig. 2 is a schematic diagram of a target image Image_t in some embodiments of the present application, Figure 7 Fig. 2(b) and Figure 8 Fig. 2(c), Figure 7 Fig. 2 is a schematic diagram of a target image Image_t in some embodiments of the present application, Figure 8 Fig. 3 is a schematic diagram of a third TOF depth image Image_TOF3 in some embodiments of the present application. Figure 7 Fig. 3(a) shows an image region A corresponding to a first wall, an image region B corresponding to a second wall, and an image region C corresponding to a corner. As Figure 2 Fig. 3(b), Figure 8As shown, compared with the first TOF depth image Image_TOF1 and the second TOF depth image Image_TOF2, the area corresponding to the corner of the wall in the third TOF depth image Image_TOF3 is clearer, thereby improving the depth accuracy of the corner of the wall.
[0095] Referring to Figure 9 , Figure 9 FIG. 2 is a schematic diagram of pixels of the binocular depth image Image_dep and the third TOF depth image Image_TOF3 in some embodiments of the present application. Figure 9 (a) and (b) of FIG. 2 respectively show the pixels of the binocular depth image Image_dep and the pixels of the third TOF depth image Image_TOF3. The binocular depth image Image_dep includes a plurality of first pixels P_1, and the third TOF depth image includes a plurality of second pixels P_2, wherein the plurality of second pixels P_2 includes the plurality of TOF pixels P_TOF and the plurality of interpolation pixels P_in. The depth value of the first pixel P_1 is a first depth value, and the depth value of the second pixel P_2 is a second depth value.
[0096] Referring to Figure 10 , Figure 10 FIG. 3 is a sub-flowchart of step S15 in some embodiments. As shown in Figure 10 some embodiments, the step S15 of “correcting the depth value of the binocular depth image Image_dep based on the third TOF depth image” includes: S41: performing pixel mapping on the third TOF depth image Image_TOF3 and the binocular depth image Image_dep to determine a target second pixel P_2_t corresponding to the first pixel P_1.
[0097] S42: determining a second target depth value according to the difference between the first depth value of the first pixel P_1 and the second depth value of the target second pixel P_2_t corresponding to the first pixel P_1 and / or the shooting distance between the shooting object and the camera, wherein the camera is the binocular camera and / or the TOF camera; S43: correcting the first depth value of the first pixel P_1 to the second target depth value.
[0098] The photographing distance can be determined based on one of the first image, the second image, the third image, and the first TOF depth image Image_TOF1. In some embodiments, when the photographing distance is determined to be close to a first preset distance range based on the images, the photographing distance is determined based on the first TOF depth image Image_TOF1; when the photographing distance is determined to be close to a second preset distance range based on the images, the photographing distance is determined based on the first TOF depth image Image_TOF1; and when the photographing distance is determined to be close to a third preset distance range based on the images, the photographing distance is determined based on one of the first image, the second image, and the third image. The upper limit of the first preset distance range is less than the lower limit of the second preset distance range, and the upper limit of the second preset distance range is less than the lower limit of the third preset distance range. The three distance ranges can be set according to actual needs.
[0099] By correcting the binocular depth value based on the difference between the binocular depth value and the TOF depth value at the same position and / or the photographing distance, the accuracy of the binocular depth value can be improved.
[0100] In some embodiments, the resolution of the third TOF depth image Image_TOF3 is the same as that of the binocular depth image Image_dep. Each first pixel point P_1 corresponds to a target second pixel point P_2_t.
[0101] In other embodiments, the resolution of the third TOF depth image Image_TOF3 is different from that of the binocular depth image Image_dep. If the first pixel point P_1 maps to a position on the third TOF depth image Image_TOF3 where the second pixel point P_2 does not exist, the first depth value of the first pixel point P_1 is retained and not corrected. If the first pixel point P_1 maps to a position on the third TOF depth image Image_TOF3 where the second pixel point P_2 (i.e., the target second pixel point P_2_t) exists, steps S42-S43 are performed.
[0102] The pixel mapping process of the third TOF depth image Image_TOF3 and the binocular depth image Image_dep can refer to the pixel mapping process of the second TOF depth image Image_TOF2 and the target image described above, which will not be described in detail here.
[0103] Please refer to Figure 11 , Figure 11 The flowchart in FIG. 6 is a sub-flowchart of step S42 in some embodiments. In some embodiments, as shown in FIG. 6, step S42 includes steps S601-S604. Figure 11As shown, step S42 "determining a second target depth value according to a difference between the first depth value of the first pixel point P_1 and the second depth value of the target second pixel point P_2_t corresponding to the first pixel point P_1 and / or the shooting distance of the shooting object from the camera", includes: S51: determining a first weight corresponding to the first depth value and a second weight corresponding to the second depth value according to a difference between the first depth value of the first pixel point P_1 and the second depth value of the target second pixel point P_2_t corresponding to the first pixel point P_1 and / or the shooting distance.
[0104] S52: obtaining the second target depth value by weighted fusion based on the first depth value of the first pixel point P_1, the first weight, the second depth value of the target second pixel point P_2_t corresponding to the first pixel point P_1, and the second weight.
[0105] By assigning respective weights to the first depth value and the second depth value according to a difference between the first depth value of the first pixel point P_1 and the second depth value of the target second pixel point P_2_t corresponding to the first pixel point P_1 and / or the shooting distance, and performing weighted fusion according to the respective weights, a second target depth value that combines the precision advantages of binocular ranging method and TOF ranging method can be obtained.
[0106] For example, the first depth value of the first pixel point P_1 is D1, the first weight is W1, the second depth value of the target second pixel point P_2_t corresponding to the first pixel point P_1 is D2, and the second weight is W2. The second target depth value is calculated according to the formula D_t=(D1*W1+D2*W2) / (W1+W2), where D_t is the second target depth value.
[0107] In some embodiments, the sum of the first weight and the second weight is equal to 1.
[0108] In some embodiments, the image generation method further includes: determining a pattern region and a non-pattern region based on the binocular image.
[0109] In some embodiments, the determining a pattern region and a non-pattern region based on the binocular image includes: performing image feature extraction based on the third image, the image features including one or more of color features, texture features, and semantic features; and determining the pattern region and the non-pattern region based on the obtained image features.
[0110] For example, the pattern region is an "O" shaped region, and the non-pattern region is a region outside the "O" shaped region.
[0111] When it is determined that a first pixel point P_1 is located in the pattern region, it is determined that a target second pixel point P_2_t corresponding to the first pixel point P_1 is located in the pattern region. Similarly, when it is determined that a first pixel point P_1 is located in the non-pattern region, it is determined that a target second pixel point P_2_t corresponding to the first pixel point P_1 is located in the non-pattern region.
[0112] In some embodiments, the image generation method further includes determining a high reflection region based on the binocular image. The high reflection region can include one or more of a mirror region and a metal region.
[0113] In some embodiments, the determination of the high reflection region based on the binocular image includes obtaining the luminance value of the third image and determining a region with a luminance value greater than a preset luminance value as the high reflection region, or determining a region with a high matching cost based on the first image and the second image and taking the region as the high reflection region. In other embodiments, the high reflection region can be determined in other ways, for example, by identifying the high reflection region based on a reflection component separation method or by identifying the high reflection region based on a deep learning method.
[0114] When it is determined that a first pixel point P_1 is located in the high reflection region, it is determined that a target second pixel point P_2_t corresponding to the first pixel point P_1 is located in the high reflection region.
[0115] In some embodiments, step S51 includes: if the absolute value of the difference between the first depth value of the first pixel point P_1 and the second depth value of the target second pixel point P_2_t corresponding to the first pixel point P_1 is greater than a second preset depth difference value and / or the shooting distance is within a first preset distance range, determining that the first weight is 0 and the second weight is 1; and / or, if the absolute value of the difference between the first depth value of the first pixel point P_1 and the second depth value of the target second pixel point P_2_t corresponding to the first pixel point P_1 is less than or equal to the second preset depth difference value, determining the first weight corresponding to the first depth value and the second weight corresponding to the second depth value according to one or more of the shooting distance, the first pixel point P_1, and a target region in which the target second pixel point P_2_t is located, wherein the target region includes one or more of the pattern region, the non-pattern region, and the high reflection region.
[0116] The second preset depth difference value can be set according to actual needs.
[0117] In some embodiments, the first preset distance range is (0, 0.2m]. In other embodiments, the first preset distance range can be other distance ranges.
[0118] When the difference between the first depth value of the first pixel point P_1 and the second depth value of the target second pixel point P_2_t corresponding to the first pixel point P_1 is large, it indicates that the binocular depth error is large, and by replacing the first depth value of the first pixel point P_1 with the second depth value of the target second pixel point P_2_t, the accuracy of the depth value of the target depth image can be improved.
[0119] Exemplarily, as shown in Figure 9 When the absolute value of the difference between the first depth value of a certain first pixel point P_1’ and the second depth value of the target second pixel point P_2_t’ corresponding to the certain first pixel point P_1’ is less than or equal to the second preset depth difference, the first depth value of the first pixel point P_t’ is corrected according to the shooting distance.
[0120] When the shooting distance is within the first preset distance range, since the shooting distance is very close, it may be impossible to perform stereo matching, which is not conducive to binocular depth calculation. It can be understood that the binocular camera is in a binocular blind area. By using the second depth value to replace the first depth value, the accuracy of the depth value of the target depth image can be improved.
[0121] When the difference between the first depth value of the first pixel point P_1 and the second depth value of the target second pixel point P_2_t is small, the second target depth value can be obtained by combining the first depth value and the second depth value, which can reduce the error and improve the measurement stability.
[0122] Since the binocular ranging method and the TOF ranging method have different accuracy performances at different shooting distances, by performing depth correction according to the shooting distance, binocular depth and TOF depth, the advantages of the two ranging methods are dynamically fused, so that the accuracy of the second target depth value can be further improved.
[0123] In some embodiments, the foregoing determination of the first weight corresponding to the first depth value and the second weight corresponding to the second depth value according to one or more of the shooting distance, the target region in which the first pixel point P_1 and the target second pixel point P_2_t are located, includes: when the shooting distance is within the second preset distance range, the first weight is greater than the second weight; when the shooting distance is within the third preset distance range, the first weight is less than the second weight; when the shooting distance is within the fourth preset distance range, the first weight is greater than the second weight.
[0124] The lower limit value of the second preset distance range is greater than the upper limit value of the first preset distance range, the lower limit value of the third preset distance range is greater than the upper limit value of the second preset distance range, and the lower limit value of the fourth preset distance range is greater than the upper limit value of the third preset distance range.
[0125] When the shooting distance is small, i.e., close-range shooting, the stereo matching of the binocular camera can reach a sub-pixel level, and the accuracy of the binocular depth is higher than that of the TOF depth. By setting the first weight corresponding to the binocular depth to be greater than the second weight corresponding to the TOF depth, the accuracy of the determined second target depth value can be improved. When the shooting distance is medium, the accuracy of the TOF depth is higher than that of the binocular depth. By setting the first weight corresponding to the binocular depth to be less than the second weight corresponding to the TOF depth, the accuracy of the determined second target depth value can be improved. When the shooting distance is far, the emission light energy of the TOF camera decays greatly, resulting in a decrease in the intensity of the reflected light, and thus the accuracy of the TOF depth is reduced. By setting the second weight corresponding to the TOF depth to be less than the first weight corresponding to the binocular depth, the accuracy of the determined second target depth value can be improved.
[0126] Exemplarily, the second preset distance range is (0.2m, 1m], the third preset distance range is (1m, 10m], and the fourth preset distance range is greater than 10m. When the shooting distance is a value within (0.2m, 1m], the first weight is 0.7 and the second weight is 0.3; when the shooting distance is a value within (1m, 10m], the first weight is 0.4 and the second weight is 0.6; and when the shooting distance is greater than 10m, the first weight is 0.8 and the second weight is 0.2.
[0127] Obviously, in other embodiments, the second preset distance range, the third preset distance range, and the fourth preset distance range can be other distance ranges, and the first weight and the second weight can be other values.
[0128] In some embodiments, the determining the first weight corresponding to the first depth value and the second weight corresponding to the second depth value according to one or more of the shooting distance, the first pixel point P_1 and the target second pixel point P_2_t located in a target region includes: if the first pixel point P_1 and the target second pixel point P_2_t corresponding to the first pixel point P_1 are located in the pattern region, determining that the first weight corresponding to the first depth value is greater than the second weight corresponding to the second depth value; if the first pixel point P_1 and the target second pixel point P_2_t corresponding to the first pixel point P_1 are located in the non-pattern region, determining that the first weight corresponding to the first depth value is less than the second weight corresponding to the second depth value; and / or, determining a first sub-weight corresponding to the first depth value and a second sub-weight corresponding to the second depth value according to the target region where the first pixel point P_1 and the target second pixel point P_2_t are located, and determining a third sub-weight corresponding to the first depth value and a fourth sub-weight corresponding to the second depth value according to the shooting distance, and determining the second weight according to the first sub-weight and the third sub-weight, and determining the second weight according to the second sub-weight and the fourth sub-weight.
[0129] wherein the sum of the first sub-weight and the second sub-weight is 1, and the sum of the third sub-weight and the fourth sub-weight is 1.
[0130] In the pattern region and the non-pattern region, the stereo matching effect of the binocular distance measurement method will be different, wherein in the pattern region, i.e. the region with complex texture, the stereo matching effect is better, so that the accuracy of the binocular depth is higher, while in the non-pattern region, i.e. the weak texture region, the accuracy of the TOF depth is higher. By identifying the pattern region and the non-pattern region and assigning different weights based on the region where the pixel point is located, the accuracy of the second target depth value can be improved.
[0131] For example, when the first pixel point P_1 and the target second pixel point P_2_t corresponding to the first pixel point P_1 are located in the pattern region, the first weight is 0.7 and the second weight is 0.3; when the first pixel point P_1 and the target second pixel point P_2_t corresponding to the first pixel point P_1 are located in the non-pattern region, the first weight is 0.4 and the second weight is 0.6.
[0132] In some embodiments, by combining the region where the pixel point is located and the shooting distance, the accuracy of the second target depth value can be further improved.
[0133] The determining the first weight and the second weight according to one or more of the shooting distance, the target region where the first pixel point and the target second pixel point are located, in the foregoing embodiment, comprises: multiplying the first sub-weight and the third sub-weight to obtain a first product, multiplying the second sub-weight and the fourth sub-weight to obtain a second product; normalizing the first product and the second product to obtain the first weight and the second weight.
[0134] Exemplarily, when the target region where the first pixel point P_1 and the target second pixel point P_2_t are located is the pattern region, it is determined that the first sub-weight corresponding to the first depth value is 0.7, and the second sub-weight corresponding to the second depth value is 0.3, when the shooting distance is a value in (0.2m, 1m], the third sub-weight corresponding to the first depth value is 0.7, and the fourth sub-weight corresponding to the second depth value is 0.3, so that the first product is 0.49, the second product is 0.09, and the normalization obtains the first weight as 0.84 and the second weight as 0.16.
[0135] Exemplarily, when the target region where the first pixel point P_1 and the target second pixel point P_2_t are located is the non-pattern region, it is determined that the first sub-weight corresponding to the first depth value is 0.4, and the second sub-weight corresponding to the second depth value is 0.6, when the shooting distance is a value in (0.2m, 1m], the third sub-weight corresponding to the first depth value is 0.7, and the fourth sub-weight corresponding to the second depth value is 0.3, so that the first product is 0.28, the second product is 0.18, and the normalization obtains the first weight as 0.6 and the second weight as 0.4.
[0136] Obviously, in other embodiments, the first sub-weight, the third sub-weight, the second sub-weight, and the fourth sub-weight can be other values.
[0137] In some embodiments, the determining the first weight corresponding to the first depth value and the second weight corresponding to the second depth value according to one or more of the shooting distance, the target region where the first pixel point and the target second pixel point are located, in the foregoing embodiment, comprises: if the first pixel point P_1 and the target second pixel point P_2_t corresponding to the first pixel point P_1 are located in the high-reflection region, determining that the first weight is less than the second weight, and the first weight is greater than or equal to 0.
[0138] In the high-reflection region, the stereo matching of the binocular distance measurement method can fail, and the accuracy of the binocular depth is reduced. By setting the weight corresponding to the first pixel point P_1 in the high-reflection region to be lower than the target second pixel point P_2_t, the accuracy of the second target depth value is improved.
[0139] For example, when a first pixel point P_1 and the target second pixel point P_2_t corresponding to the first pixel point P_1 are located in the high reflection region, the first weight corresponding to the first depth value of the first pixel point P_1 is 0.1, and the second weight corresponding to the second depth value of the target second pixel point P_2_t is 0.9. Obviously, the first weight and the second weight can be other values.
[0140] Referring to Figure 12 , Figure 12 is a subflowchart of step S12 in some embodiments. In some embodiments, as shown in Figure 12 , step S12 "determining a binocular depth image based on the binocular image" includes: S81: performing feature extraction on the first image to obtain a plurality of layers of features corresponding to the first image.
[0141] S82: performing feature extraction on the second image to obtain a plurality of layers of features corresponding to the second image.
[0142] S83: fusing the plurality of layers of features corresponding to the first image based on a target dilated convolution to obtain a first feature fusion image.
[0143] S84: fusing the plurality of layers of features corresponding to the second image based on the target dilated convolution to obtain a second feature fusion image.
[0144] S85: performing pixel point matching on the first feature fusion image and the second feature fusion image to determine a disparity map.
[0145] S86: determining the binocular depth image Image_dep based on the disparity map, the focal length of the binocular camera, and the baseline distance.
[0146] The baseline distance is the distance between the optical centers of the two cameras of the binocular camera. The disparity map includes a disparity value of each pixel point of the disparity map, and the disparity value is the offset amount of the positions of the same object in the first image and the second image.
[0147] Using dilated convolution can expand the receptive field of the convolution kernel without increasing the number of parameters and reducing the resolution, so that the neural network can capture more semantic information in a larger range, thereby accurately matching feature points and further obtaining a more accurate and clearer binocular depth image Image_dep.
[0148] In some embodiments, the multi-layer feature includes a shallow layer feature, a middle layer feature, and a deep layer feature. Step S81 includes inputting the first image into a convolutional neural network with shared weights, extracting features through three levels of convolutional layers in a shallow layer, a middle layer, and a deep layer in sequence, and retaining the features of each level of convolutional layer by using a skip connection. Step S82 includes inputting the second image into a convolutional neural network with shared weights, extracting features through three levels of convolutional layers in a shallow layer, a middle layer, and a deep layer in sequence, and retaining the features of each level of convolutional layer by using a skip connection.
[0149] In some embodiments, step S83 and step S84 each include upsampling the deep layer feature, fusing the upsampled deep layer feature with the middle layer feature to obtain fused features, upsampling the fused features, and fusing the upsampled fused features with the shallow layer feature to construct a feature pyramid, thereby forming a sequence of feature maps with increasing resolutions, and performing context enhancement processing on the sequence of feature maps based on the target dilated convolution to capture multi-scale context information, thereby obtaining a feature fusion map. That is, the first feature fusion map and the second feature fusion map can be obtained respectively.
[0150] In some embodiments, step S83 and step S84 each include upsampling the deep layer feature based on the target dilated convolution, fusing the upsampled deep layer feature with the middle layer feature to obtain fused features, upsampling the fused features based on the target dilated convolution, and fusing the upsampled fused features with the shallow layer feature after fusion to construct a feature pyramid, thereby forming a sequence of feature maps with increasing resolutions, and obtaining a feature fusion map. That is, the first feature fusion map and the second feature fusion map can be obtained respectively.
[0151] The resolutions of the sequence of feature maps can be 1 / 8, 1 / 4, and 1 / 2 respectively.
[0152] In some embodiments, the neural network architecture can be used to determine the binocular depth image Image_dep, for example, a ResNet-50 architecture.
[0153] In some embodiments, the first feature fusion map and / or the second feature fusion map can be used to determine the pattern region, the non-pattern region, and the high reflection region.
[0154] In some embodiments, the method further includes determining the target dilated convolution based on the shooting distance of the shooting object from the binocular camera and / or the TOF camera.
[0155] In some embodiments, the target dilated convolution corresponding to the first preset distance range, the second preset distance range, and the third preset distance range decreases in sequence.
[0156] The parallax is negatively correlated with the binocular depth. When a photograph is taken at a close distance, the parallax is large, a large target hole convolution is used to obtain a large receptive field, so as to cover a large parallax range, and more rich context information can be obtained to facilitate feature point matching of large parallax. When a photograph is taken at a long distance, the parallax is small, and when feature point matching is performed, local fine features and slight position differences need to be relied on. A small target hole convolution is used to not only cover a small parallax range, but also greatly retain local details and edge sharpness, so as to facilitate feature point matching of small parallax.
[0157] Exemplarily, refer to Figure 13 and Figure 14 , Figure 13 FIG. 1 is a schematic diagram of a binocular depth image Image_dep in some embodiments of the present application, Figure 14 FIG. 2 is a schematic diagram of a third TOF depth image Image_TOF3 in some embodiments of the present application.
[0158] In some embodiments, the image generation method further comprises: acquiring a motion parameter of the binocular camera, and / or determining parallax change information based on the binocular image; when the motion parameter and / or the parallax change information meet a preset condition, determining that the current scene is a dynamic scene; when the motion parameter and / or the parallax change information do not meet the preset condition, determining that the current scene is a static scene.
[0159] In some embodiments, the parallax change information comprises a proportion of pixel points whose change amount of parallax value is greater than a preset change amount. The preset condition comprises that the motion parameter is greater than a preset motion parameter and / or the proportion is greater than a preset proportion.
[0160] The step S11 of “acquiring, by the TOF camera, a first TOF depth image of the photographed object” comprises: when the current scene is a dynamic scene, starting the TOF camera and acquiring the first TOF depth image by the TOF camera. The dynamic scene comprises a scene in which the binocular camera is dynamic and / or the photographed object is dynamic.
[0161] When the current scene is a dynamic scene, only distance measurement by the binocular camera can cause a large error of the depth value of the image. By combining binocular distance measurement and TOF distance measurement, the accuracy of the depth value of the target depth image can be improved.
[0162] The judgment of whether the current scene is a dynamic scene according to the motion parameter and the parallax change information can improve the judgment accuracy. The applicant finds in the research that in some cases, when the binocular camera is translated (for example, forward and backward, up and down, left and right translation), the parallax change may not be obvious, and if the current scene is judged only by the parallax change information, the accuracy of the judgment may be affected. By combining the motion parameter of the binocular camera, the judgment accuracy can be improved. When the binocular camera is rotated (for example, left and right, up and down), the parallax change may be more obvious when the photographed object is dynamic.
[0163] In some embodiments, when the current scene is a static scene, the TOF camera is turned off. Thus, distance measurement is performed only by the binocular camera, which can reduce power consumption.
[0164] In some embodiments, when the current scene is a static scene or a dynamic scene, the TOF camera is turned on.
[0165] In some embodiments, the foregoing obtaining the motion parameter of the binocular camera comprises: obtaining the speed and the angular acceleration of the binocular camera at the current time; and determining the motion parameter according to the speed and the angular acceleration of the binocular camera, a first weight coefficient corresponding to the speed, and a second weight coefficient corresponding to the angular acceleration, wherein the second weight coefficient is greater than the first weight coefficient.
[0166] In some embodiments, the applicant finds in the research that compared with the translation (for example, forward and backward, up and down, left and right translation) of the binocular camera, the rotation (for example, left and right, up and down) of the binocular camera has a greater influence on the motion state of the binocular camera on the calculation of binocular depth. By setting the first weight coefficient corresponding to the speed to be less than the second weight coefficient corresponding to the angular acceleration, the influence of the current motion state of the binocular camera on the binocular depth error can be more accurately quantified by the motion parameter, thereby facilitating more accurate judgment of whether the TOF camera needs to be turned on.
[0167] In some embodiments, the speed comprises x-axis speed, y-axis speed, and z-axis speed, and the angular acceleration comprises x-axis angular acceleration, y-axis angular acceleration, and z-axis angular acceleration. The determination of the motion parameter according to the speed and the angular acceleration of the binocular camera, the first weight coefficient corresponding to the speed, and the second weight coefficient corresponding to the angular acceleration comprises: calculating the motion parameter according to the formula M wherein M is the motion parameter, is the first weight coefficient, is the second weight coefficient, is the x-axis speed, is the y-axis speed, is the z-axis speed, is the x-axis angular acceleration, is a z-axis angular acceleration. is a z-axis angular acceleration.
[0168] In some embodiments, the first weight coefficient is 1, and the second weight coefficient is greater than 1.
[0169] In some embodiments, the method further comprises: when the binocular camera is in translation, capturing a preset shooting object to obtain a first test binocular image; determining a first test binocular depth image based on the first test binocular image; determining a first depth error based on a depth value of the first test binocular depth image and depth information of the preset shooting object; when the binocular camera is in rotation, capturing the preset shooting object to obtain a second test binocular image; determining a second test binocular depth image based on the second test binocular image; determining a second depth error based on a depth value of the second test binocular depth image and the depth information of the preset shooting object; and determining the first weight coefficient and the second weight coefficient based on the first depth error and the second depth error.
[0170] wherein the depth information of the preset shooting object is known, and the preset shooting object can be a checkerboard calibration board, a circular array calibration board, or other shooting objects with known depth information.
[0171] By determining the corresponding depth errors when the binocular camera is in translation and rotation respectively, the influence of the translation and rotation motion states on the accuracy of binocular depth can be obtained, and then the first weight coefficient and the second weight coefficient can be obtained.
[0172] In some embodiments, the method further comprises: obtaining a plurality of test motion parameters of the binocular camera in different first preset scenes, wherein one test motion parameter is obtained in each first preset scene; and determining the preset motion parameter according to the plurality of test motion parameters.
[0173] wherein the first preset scene can include a scene in which the binocular camera is in dynamic state, such as rotating the binocular camera up and down, left and right, translating the binocular camera forward and backward, and translating the binocular camera left and right. Alternatively, the first preset scene can further include a scene in which the binocular camera is in static state.
[0174] In some embodiments, determining the preset motion parameter according to the plurality of test motion parameters comprises: fitting the plurality of test motion parameters to obtain a motion parameter curve; and determining the preset motion parameter based on the motion parameter curve. For example, the motion parameter curve is differentiated, and the motion parameter corresponding to the maximum value is taken as the preset motion parameter. Other ways of determining the preset motion parameter based on the motion parameter curve can also be used.
[0175] In some embodiments, the determining the preset motion parameter according to the plurality of test motion parameters comprises: taking an arithmetic mean or a weighted mean of the plurality of test motion parameters to obtain the preset motion parameter.
[0176] In some embodiments, the electronic device comprises an IMU (Inertial Measurement Unit) configured to obtain a velocity and an angular acceleration of the binocular camera.
[0177] In some embodiments, the IMU (Inertial Measurement Unit) can comprise a three-axis accelerometer and / or a three-axis gyroscope. In some embodiments, the IMU (Inertial Measurement Unit) can comprise a nine-axis sensor.
[0178] In some embodiments, the determining the disparity change information based on the binocular images comprises: determining a disparity map at a previous time based on a binocular image at the previous time; determining a disparity map at a current time based on a binocular image at the current time; and comparing disparity values of pixel points in the disparity map at the current time and the disparity map at the previous time to determine a proportion of pixel points with a disparity value change greater than a preset change amount.
[0179] In some embodiments, the preset change amount can be set according to actual requirements.
[0180] In some embodiments, the determining the disparity map at the previous time and the disparity map at the current time can refer to the steps S81-S85 described above, which will not be described in detail here.
[0181] The applicant has found in research that when the binocular camera and / or the photographed object is in a dynamic state, the disparity value of the disparity map will change. The embodiments of the present application can determine whether the binocular camera and / or the photographed object is in a dynamic state by determining the proportion of pixel points with a disparity value change greater than a preset change amount, which is conducive to determining whether to start the TOF camera.
[0182] The preset motion parameter is M0, the preset proportion is R0, and the proportion of pixel points with a disparity value change greater than a preset change amount is R. In some embodiments, when M>M0 and / or R>R0, it is determined that the current scene is a dynamic scene, and the TOF camera is started; when M
[0183] In some embodiments, the method further comprises: in different second preset scenes, obtaining a parallax map of a previous time and a current time in each second preset scene; comparing the parallax values of the pixel points in the parallax map of the current time and the parallax map of the previous time in each second preset scene to determine a test proportion of the pixel points with a parallax value change greater than a preset change, and then obtaining a plurality of test proportions; and determining the preset proportion based on the plurality of test proportions.
[0184] The second preset scene can include a scene in which the binocular camera and / or the photographed object is dynamic, such as rotating the binocular camera up and down or left and right, translating the binocular camera forward and backward or left and right, or the photographed object moving, and the like. Alternatively, the second preset scene can further include a scene in which the binocular camera and / or the photographed object is static.
[0185] In some embodiments, determining the preset proportion based on the plurality of test proportions comprises: fitting the plurality of test proportions to obtain a proportion curve; and determining the preset proportion based on the proportion curve. For example, the proportion curve is differentiated, and the proportion corresponding to the maximum value is taken as the preset proportion. Other ways of determining the preset proportion based on the proportion curve can also be used.
[0186] In other embodiments, determining the preset proportion based on the plurality of test proportions comprises: taking an arithmetic mean or a weighted mean of the plurality of proportions to obtain the proportion.
[0187] Determining the preset motion parameter and the preset proportion by testing in a plurality of first preset scenes and a plurality of second preset scenes can reduce the probability of mistakenly starting or failing to start the TOF camera.
[0188] In some embodiments, the method comprises: obtaining the motion parameter of the binocular camera; determining whether the motion parameter is greater than the preset motion parameter; starting the TOF camera when the motion parameter is greater than the preset motion parameter; and obtaining the parallax change information when the motion parameter is less than or equal to the preset motion parameter.
[0189] In some embodiments, the motion parameter comprises the speed and / or the angular acceleration.
[0190] In some embodiments, the IMU obtains the angle and the angular acceleration at the current time, and the binocular camera synchronously obtains the binocular image at the current time.
[0191] In some embodiments, the two cameras are both color cameras, and the first image, the second image, and the third image are all color images.
[0192] In some embodiments, the image generation method further comprises: fusing the target depth image and the target image to generate an RGBD image. The RGBD image comprises depth information, color information, texture information, semantic information, etc.
[0193] For example, refer to Figure 15 , Figure 15 For example, refer to
[0194] In some embodiments, before step S11, the image generation method further comprises: aligning timestamps of the binocular camera and the TOF camera. In this way, the binocular camera and the TOF camera can synchronously capture images.
[0195] In some embodiments, the binocular camera is triggered to capture images by a first trigger signal output by a processor of the electronic device, and the TOF camera is triggered to capture images by a second trigger signal output by the processor.
[0196] For example, refer to Figure 16 , Figure 16 For example, refer to Figure 16 In some embodiments, as shown in FIG. 6, the voltage of the first trigger signal and the second trigger signal is greater than or equal to 1.8V, the width of the first trigger signal and the second trigger signal is greater than or equal to 1ms, and the time interval of the first trigger signal and the second trigger signal is less than or equal to 0.1ms.
[0197] In some embodiments, the method further comprises: aligning timestamps of the IMU and the binocular camera. In this way, when the IMU obtains an angle and an angular acceleration, the binocular camera synchronously obtains the binocular image.
[0198] In some embodiments, the image generation method further comprises: mutually calibrating intrinsic parameters and extrinsic parameters of the two cameras of the binocular camera and the TOF camera, so that the mutual calibration error of the two cameras is less than or equal to 0.5 pixel, and the mutual calibration error of the camera and the TOF camera is less than or equal to 1 pixel.
[0199] For example, refer to Figure 17 , Figure 17 For example, refer to Figure 17 As shown in FIG. 8, the electronic device 100 comprises a binocular camera 10, a TOF camera 20, a processor 30 and a memory 40. The memory 40 stores a computer program, and the processor 30 runs the computer program to execute the image generation method of any of the preceding embodiments.
[0200] The memory 40 can be a ROM or random access memory (RAM), a magnetic disk or an optical disk, or any other medium capable of storing program code. The processor 30 can be a microcontroller, CPU (central processing unit), DSP (digital signal processing), NPU (Neural Processing Unit), or a hardware unit within such a microcontroller, CPU, or DSP. Alternatively, it can be a software program module burned into such a microcontroller, CPU, or DSP.
[0201] In some embodiments, the resolution of the TOF camera 20 is lower than the resolution of the stereo camera 10, that is, the number of pixels in the first TOF depth image Image_TOF1 is less than the number of pixels in the stereo depth image Image_dep. In other embodiments, the resolution of the TOF camera 20 may be greater than or equal to the resolution of the stereo camera 10.
[0202] In some embodiments, the number of pixels in the second TOF depth image Image_TOF2 and the third TOF depth image Image_TOF3 may be equal to or different from the number of pixels in the binocular depth image Image_dep.
[0203] Please see Figure 18 , Figure 18 This is a schematic diagram illustrating the structure of the TOF camera 20 and the binocular camera 10 in some embodiments of this application. In some embodiments, such as... Figure 18 As shown, the TOF camera 20 is located between the two cameras 11 of the binocular camera 10. The distance between the two cameras 11 is 50mm-60mm, that is, the baseline distance BL is 50mm-60mm, which facilitates the processor 30 to calculate the binocular depth.
[0204] In some other embodiments, the positional relationship between the TOF camera 20 and the two cameras 11 can be adjusted, and the distance between the two cameras 11 can also be adjusted.
[0205] Please see Figure 19 , Figure 19 This is a schematic diagram of the dot matrix light source 21 of the TOF camera 20 in some embodiments of this application. In some embodiments, such as Figure 19 As shown, the TOF camera 20 includes a dot matrix light source 21, the dot matrix light source 21 including light sources along a first direction (e.g., ...). Figure 19a plurality of light source units 211 arranged along a second direction (e.g. Y direction) shown, each of the light source units 211 comprises a plurality of light source groups 211a arranged along a first direction (e.g. X direction) shown, each of the light source groups 211a comprises at least two light sources 2111 arranged along a second direction (e.g. Y direction) shown, the projections of the light sources 2111 of two adjacent light source groups 211a along the first direction are spaced from each other, i.e. the two adjacent light source groups 211a are staggered, wherein the second direction is perpendicular to the first direction. Figure 19 a plurality of light source units 211 arranged along a second direction (e.g. Y direction) shown, each of the light source units 211 comprises a plurality of light source groups 211a arranged along a first direction (e.g. X direction) shown, each of the light source groups 211a comprises at least two light sources 2111 arranged along a second direction (e.g. Y direction) shown, the projections of the light sources 2111 of two adjacent light source groups 211a along the first direction are spaced from each other, i.e. the two adjacent light source groups 211a are staggered, wherein the second direction is perpendicular to the first direction.
[0206] exemplarily, as shown in FIG. 2, each of the light source units 211 comprises two light source groups 211a arranged along the first direction, respectively a first light source group 211a’ and a second light source group 211a’’, Figure 19 exemplarily, as shown in FIG. 2, each of the light source units 211 comprises two light source groups 211a arranged along the first direction, respectively a first light source group 211a’ and a second light source group 211a’’,
[0207] exemplarily, as shown in FIG. 2, each of the light source units 211 comprises two light source groups 211a arranged along the first direction, respectively a first light source group 211a’ and a second light source group 211a’’,
[0208] In other embodiments, the number of the light source groups 211a comprised by the light source unit 211 can be other values.
[0209] By setting the two adjacent light source groups 211a staggered, the TOF pixel points corresponding to the two adjacent light source groups 211a in the first TOF depth image Image_TOF1 can also be staggered, this kind of pixel arrangement can make the to-be-interpolated pixel points have more TOF reference points in the same distance range, thus being beneficial to more accurate interpolation processing.
[0210] In some embodiments, the included angle between the line connecting the two adjacent light sources 2111 of each light source group 211a and the center of the TOF camera 20 is 1.75°, i.e. the included angle between the line connecting one of the two adjacent light sources 2111 and the center and the line connecting the other of the two adjacent light sources 2111 and the center is 1.75°, which can make the TOF camera 20 have higher resolution and the number of light sources 2111 will not be too much to reduce the cost. In other embodiments, the included angle can be other values.
[0211] In some embodiments, the emission angle of each light source 2111 is 0.2°-0.4°, which can make the diffusion range of the light pulse in the propagation process smaller, so that the light pulse can irradiate more accurately to the specific point on the shooting object, thus reducing the measurement error caused by the diffusion of the light pulse and improving the ranging accuracy. In other embodiments, the emission angle can be other values.
[0212] Please refer toFigure 20 , Figure 20 Figure 1 is a schematic diagram of an electronic device 100 according to some embodiments of the present application. As shown in Figure 1, the electronic device 100 comprises a processor 30, and the binocular camera 10, the TOF camera 20, the processor 30, the DDR 41 (dynamic random access memory), the Flash 42 (non-volatile flash memory), the IMU 50, the PHY 61 (Ethernet physical layer chip), the Debug 62 (development debugging interface), the serializer 63, the USB 64 electrically connected to the processor 30. The binocular camera 10 comprises two cameras 11. The processor 30 can comprise an NPU unit and a CPU unit. Figure 20
[0213] The image data collected by the binocular camera 10 is transmitted to the processor 30 at high speed through the MIPI CSI interface, and the data collected by the IMU 50 is transmitted to the processor 30 through the SPI bus. The CPU unit of the processor 30 is used for TOF depth calculation and coordination task scheduling, and the NPU unit is used for AI base binocular depth model calculation, including image fusion and training processing. The serializer 63 is used to convert the MIPI CSI-2 signal into a single-link GMSL2 signal, allowing data to be transmitted over a long distance on a coaxial cable or a shielded twisted pair (STP), and the transmission data rate can reach 3G / 6Gbps.
[0214] Figure 2 is a schematic diagram of an electronic device 100 according to some other embodiments of the present application. As shown in Figure 2, the electronic device 100 comprises an adapter power board 70, and the binocular camera 10, the TOF camera 20, the IMU 50, the serializer 63, the connector 65 electrically connected to the adapter power board 70. Figure 21 , Figure 21 Figure 21 The adapter power board 70 is used to provide power input for the binocular camera 10, the TOF camera 20, the serializer 63, the IMU 50, and the connector 65. The serializer 63 is used to convert the data (MIPI CSI-2 signal) collected by the binocular camera 10, the TOF camera 20, and the IMU 50 into a high-speed serial signal (GMSL2), achieving long-distance transmission (>1m). The connector 65 is used to directly output the data collected by the binocular camera 10, the TOF camera 20, and the IMU 50 through the MIPI CSI interface.
[0215]
[0216] It should be noted that the function operations performed by the electronic device 100 correspond to the aforementioned image generation method, for example, the aforementioned step of acquiring the binocular image can be performed by the binocular camera 10, the step of acquiring the first TOF depth image Image_TOF1 can be performed by the TOF camera 20, and other steps can be performed by the processor 30. For more detailed description, please refer to the content of each embodiment of the aforementioned image generation method. The electronic device 100 and the content of the aforementioned image generation method can also be mutually referred to.
[0217] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than illustrated, or combine certain components, or split certain components, or different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware. In the above method embodiments, the execution order of the steps can be adjusted according to the situation.
[0218] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program for execution by a processor to implement the image generation method of any of the aforementioned embodiments.
[0219] The computer readable storage medium includes ROM or random access memory (RAM), magnetic or optical disks, and various media that can store program codes.
[0220] The above is the implementation of the embodiments of the present application. It should be noted that for those skilled in the art, without departing from the principles of the embodiments of the present application, a number of improvements and refinements can be made, which are also considered within the scope of protection of the present application.
Claims
1. An image generation method, applied to electronic devices, characterized in that, The electronic device includes a binocular camera and a TOF camera; the method includes: The binocular camera acquires binocular images of the subject, and the TOF camera acquires a first TOF depth image of the subject. Determine a binocular depth image based on the binocular image; The second TOF depth image is obtained by interpolating the first TOF depth image. The depth value of the second TOF depth image is corrected based on the binocular image to obtain the third TOF depth image; The depth value of the binocular depth image is corrected based on the third TOF depth image to obtain the target depth image.
2. The image generation method according to claim 1, characterized in that, The first TOF depth image includes multiple TOF pixels; the step of interpolating the first TOF depth image to obtain the second TOF depth image includes: The first TOF depth image is interpolated to obtain multiple interpolated pixels, thereby obtaining the second TOF depth image, wherein the second TOF depth image includes the multiple TOF pixels and the multiple interpolated pixels; The step of correcting the depth value of the second TOF depth image based on the stereo image includes: Based on the binocular image, determine the target TOF pixel in the second TOF depth image that is adjacent to the interpolated pixel and located in the same spatial structure; When the interpolated pixel is determined to be an interpolated pixel to be corrected based on the interpolated pixel and the target TOF pixel, a first target depth value is determined based on the depth value of the target TOF pixel that is adjacent to the interpolated pixel to be corrected and located in the same spatial structure. The depth value of the interpolated pixel to be corrected is corrected to the first target depth value.
3. The image generation method according to claim 2, characterized in that, The binocular image includes a first image and a second image acquired simultaneously by two cameras; determining the target TOF pixel in the second TOF depth image that is adjacent to the interpolated pixel and located in the same spatial structure based on the binocular image includes: A target image is determined based on the binocular images, wherein the target image is obtained based on one of the first image, the second image, and the third image, and the third image is obtained by fusing the first image and the second image; Pixel mapping is performed on the target image and the second TOF depth image to determine the first mapped pixel of the target image corresponding to each interpolated pixel, and the second mapped pixel of the target image corresponding to each TOF pixel; Based on the target image, determine a second mapped pixel that is adjacent to the first mapped pixel and located in the same spatial structure; The TOF pixel corresponding to the determined second mapped pixel is taken as the target TOF pixel that is adjacent to the interpolated pixel and located in the same spatial structure; The step of determining the interpolated pixel as the interpolated pixel to be corrected based on the interpolated pixel and the target TOF pixel includes: When the absolute value of the difference between the depth value of the interpolated pixel and the depth value of the target TOF pixel is greater than a first preset depth difference, the interpolated pixel is determined to be the interpolated pixel to be corrected.
4. The image generation method according to claim 2, characterized in that, The process of interpolating the first TOF depth image to obtain multiple interpolated pixels includes: Use multiple TOF pixels adjacent to the pixel to be interpolated as a reference point set; Based on the distance between each TOF pixel in the reference point set and the pixel to be interpolated, or further based on the confidence level of the depth value of each TOF pixel, the interpolation weight corresponding to the depth value of each TOF pixel is determined. The depth values of multiple TOF pixels in the reference point set are weighted and fused according to the corresponding interpolation weights to obtain the depth value of the pixel to be interpolated, and the interpolated pixel is generated based on the depth value of the pixel to be interpolated.
5. The image generation method according to claim 1, characterized in that, The binocular depth image includes multiple first pixels, and the third TOF depth image includes multiple second pixels; the correction of the depth value of the binocular depth image based on the third TOF depth image includes: Pixel mapping is performed on the third TOF depth image and the binocular depth image to determine the target second pixel corresponding to the first pixel. The second target depth value is determined based on the difference between the first depth value of the first pixel and the second depth value of the target second pixel corresponding to the first pixel and / or the shooting distance between the shooting object and the camera, wherein the camera is the binocular camera and / or the TOF camera; The first depth value of the first pixel is corrected to the second target depth value.
6. The image generation method according to claim 5, characterized in that, The step of determining the second target depth value based on the difference between the first depth value of the first pixel and the second depth value of the target second pixel corresponding to the first pixel and / or the shooting distance between the object being photographed and the camera includes: Based on the difference between the first depth value of the first pixel and the second depth value of the target second pixel corresponding to the first pixel and / or the shooting distance, determine the first weight corresponding to the first depth value and the second weight corresponding to the second depth value. The second target depth value is obtained by weighted fusion based on the first depth value of the first pixel, the first weight, the second depth value of the target second pixel corresponding to the first pixel, and the second weight.
7. The image generation method according to claim 6, characterized in that, The step of determining the first weight corresponding to the first depth value and the second weight corresponding to the second depth value based on the difference between the first depth value of the first pixel and the second depth value of the target second pixel corresponding to the first pixel and / or the shooting distance includes: If the absolute value of the difference between the first depth value of the first pixel and the second depth value of the target second pixel corresponding to the first pixel is greater than the second preset depth difference, and / or the shooting distance is within the first preset distance range, then the first weight is determined to be 0 and the second weight is determined to be 1. And / or, if the absolute value of the difference between the first depth value of the first pixel and the second depth value of the target second pixel corresponding to the first pixel is less than or equal to the second preset depth difference, then a first weight corresponding to the first depth value and a second weight corresponding to the second depth value are determined according to one or more of the shooting distance, the target region where the first pixel and the target second pixel are located, wherein the target region includes one or more of the patterned region, the non-patterned region, and the highly reflective region.
8. The image generation method according to claim 7, characterized in that, The step of determining a first weight corresponding to the first depth value and a second weight corresponding to the second depth value based on one or more of the shooting distance, the first pixel, and the target region where the second pixel is located includes: If the first pixel and the target second pixel corresponding to the first pixel are located in the pattern area, then it is determined that the first weight corresponding to the first depth value is greater than the second weight corresponding to the second depth value. If the first pixel and the target second pixel corresponding to the first pixel are located in the non-pattern area, then it is determined that the first weight corresponding to the first depth value is less than the second weight corresponding to the second depth value. And / or, based on the target region where the first pixel and the target second pixel are located, determine the first sub-weight corresponding to the first depth value and the second sub-weight corresponding to the second depth value, and based on the shooting distance, determine the third sub-weight corresponding to the first depth value and the fourth sub-weight corresponding to the second depth value, and determine the second weight based on the first sub-weight and the third sub-weight, and determine the second weight based on the second sub-weight and the fourth sub-weight.
9. The image generation method according to claim 1, characterized in that, The binocular image includes a first image and a second image acquired simultaneously by two cameras; the step of determining the binocular depth image based on the binocular image includes: Feature extraction is performed on the first image to obtain multi-layer features corresponding to the first image; Feature extraction is performed on the second image to obtain the multi-layer features corresponding to the second image; The first feature fusion map is obtained by fusing the multi-layer features corresponding to the first image based on the target dilated convolution. Based on the target dilated convolution, the multi-layer features corresponding to the second image are fused to obtain the second feature fusion map; Pixel matching is performed on the first feature fusion map and the second feature fusion map to determine the disparity map; The binocular depth image is determined based on the disparity map, the focal length of the binocular camera, and the baseline distance.
10. The image generation method according to claim 9, characterized in that, The method further includes: The target dilated convolution is determined based on the shooting distance between the object being photographed and the binocular camera and / or the TOF camera.
11. The image generation method according to claim 1, characterized in that, The method further includes: Acquire the motion parameters of the binocular camera, and / or determine disparity change information based on the binocular images; When the motion parameters and / or the disparity change information meet preset conditions, the current scene is determined to be a dynamic scene; The acquisition of the first TOF depth image of the subject by the TOF camera includes: When the current scene is a dynamic scene, the TOF camera is activated, and the first TOF depth image is acquired through the TOF camera.
12. The image generation method according to claim 11, characterized in that, The acquisition of motion parameters of the binocular camera includes: Obtain the velocity and angular acceleration of the stereo camera at the current moment; The motion parameters are determined based on the velocity, angular acceleration, first weighting coefficient corresponding to the velocity, and second weighting coefficient corresponding to the angular acceleration of the binocular camera, wherein the second weighting coefficient is greater than the first weighting coefficient, and the preset condition includes that the motion parameters are greater than preset motion parameters. And / or, the disparity change information includes the proportion of pixels whose disparity value changes by a greater than a preset amount; the step of determining the disparity change information based on the binocular image includes: Determine the disparity map of the previous time step based on the binocular image of the previous time step; Determine the disparity map at the current moment based on the binocular images at the current moment; By comparing the disparity values of pixels in the current disparity map with those in the previous disparity map, the proportion of pixels whose disparity value changes by a greater than a preset change is determined, wherein the preset condition includes the proportion being greater than a preset proportion.
13. An electronic device, characterized in that, The electronic device includes a binocular camera, a TOF camera, a processor, and a memory, wherein the memory stores a computer program, and the processor runs the computer program to perform the image generation method according to any one of claims 1-12.
14. The electronic device according to claim 13, characterized in that, The TOF camera includes a dot matrix light source, which includes multiple light source units arranged along a first direction. Each light source unit includes multiple light source groups arranged along the first direction. Each light source group includes at least two light sources arranged along a second direction. The projections of two adjacent light source groups along the first direction are spaced apart from each other. The second direction is perpendicular to the first direction.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a processor and a memory, the memory storing a computer program, and the processor running the computer program to perform the image generation method according to any one of claims 1-12.
Citation Information
Patent Citations
Multi-vision sensor fusion device and method and electronic equipment
CN112802114A
Image demosaicing for hybrid optical sensor arrays
US20180197275A1
Systems and methods for molecular imaging
US20240037834A1
Image processing method and apparatus, electronic device, and computer readable storage medium
WO2021017811A1