Image sensor, image generation method and imaging apparatus

By introducing a pixel array of ToF pixels and image sensing pixels into the image sensor, and through the cooperation of the controller and signal processor, it is possible to simultaneously acquire RGB images and depth maps on a single sensor. This solves the problems of high integration difficulty and high cost in existing technologies and improves the accuracy of depth maps.

WO2025223185A1PCT designated stage Publication Date: 2025-10-30ZHUHAI MOJIE TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/087546
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-25
Filing Date
2025-04-07
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing image sensors cannot simultaneously acquire RGB images and depth information, resulting in high integration difficulty and manufacturing costs in applications such as virtual reality, augmented reality, and 3D reconstruction.

Method used

An image sensor is used, comprising a pixel array of image sensing pixels and ToF pixels. The ToF pixels are controlled by a controller to measure distance and acquire depth information. The voltage signal is then combined with the exposure output voltage signal of the image sensing pixels. The signal processor processes the voltage signal to generate an RGB image and a depth map.

Benefits of technology

This technology enables the simultaneous acquisition of RGB images and depth maps on a single image sensor, reducing integration difficulty and manufacturing costs while improving the accuracy of the depth maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025087546_30102025_PF_FP_ABST
    Figure CN2025087546_30102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are an image sensor, an image generation method and an imaging apparatus. The image sensor comprises: a pixel array, which comprises a plurality of pixels arranged in multiple rows and columns, wherein the plurality of pixels comprise a first number of image sensing pixels and a second number of ToF pixels, the image sensing pixels being used for exposure output of a voltage signal, and the ToF pixels being used for determining depth information; a controller, which is used for controlling the second number of ToF pixels to perform ranging, so as to obtain the depth information, and controlling the first number of image sensing pixels to perform exposure, so as to output the voltage signal; and a signal processor, which is used for processing the voltage signal to obtain a target RGB image, wherein the controller is also used for generating a target depth map on the basis of the target RGB image and the depth information. The present application realizes the simultaneous acquisition of an RGB image and a depth map on one image sensor, thereby effectively reducing the integration difficulty and manufacturing cost of the image sensor, and also ensuring the accuracy of the depth map.
Need to check novelty before this filing date? Find Prior Art

Description

Image sensor, image generation method and imaging device

[0001] This application claims priority to Chinese Patent Application No. 2024105097986, filed on April 25, 2024, entitled “Image Sensor, Image Generation Method and Imaging Device”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of image sensor technology, and in particular to an image sensor, an image generation method, and an imaging device. Background Technology

[0003] Currently, image sensors are widely used in fields such as photography, surveillance, and machine vision. However, for applications such as virtual reality, augmented reality, and 3D reconstruction, it is necessary to obtain depth information of objects in the scene. However, current image sensors cannot directly obtain both image and depth information simultaneously. Therefore, depth estimation is usually performed using images from multiple image sensors, or by fusing image sensors and depth measurement sensors to obtain both images and depth information. However, this increases the number and size of sensors, leading to high integration difficulty and high manufacturing costs. Summary of the Invention

[0004] This application provides an image sensor, an image generation method, and an imaging device, which aim to reduce the integration difficulty and manufacturing cost of the image sensor while simultaneously acquiring RGB images and depth maps.

[0005] In a first aspect, embodiments of this application provide an image sensor, including:

[0006] A pixel array includes multiple pixels arranged in multiple rows and columns, the multiple pixels including a first number of image sensing pixels and a second number of ToF pixels, the image sensing pixels being used to expose output voltage signals, and the ToF pixels being used to determine depth information;

[0007] The controller is used to control the second number of ToF pixels to perform ranging to obtain depth information; and to control the first number of image sensing pixels to perform exposure in order to output a voltage signal.

[0008] A signal processor is used to process the voltage signal to obtain a target RGB image;

[0009] The controller is also configured to generate a target depth map based on the target RGB image and the depth information.

[0010] Secondly, embodiments of this application also provide an image generation method, applied to the image sensor as described in the first aspect, the method comprising:

[0011] The voltage signal and depth information are acquired. The voltage signal is obtained by exposure of the image sensing pixels contained in the image sensor, and the depth information is obtained by ranging of the ToF pixels contained in the image sensor.

[0012] The voltage signal is processed to obtain a target RGB image, and a target depth map is generated based on the target RGB image and the depth information.

[0013] Thirdly, embodiments of this application also provide an imaging device, the imaging device including a light emitter and an image sensor as described in the first aspect.

[0014] This application provides an image sensor, an image generation method, and an imaging device. The image sensor includes a pixel array containing multiple image sensing pixels and multiple time-of-flight (ToF) pixels, enabling the image sensor to acquire RGB images and depth information. Based on the RGB images and depth information, a depth map can be generated, thereby achieving the simultaneous acquisition of RGB images and depth maps on a single image sensor. This effectively reduces the integration difficulty and manufacturing cost of the image sensor. Furthermore, since the depth map is generated by fusing actual depth information and RGB images, the accuracy of the depth map can also be guaranteed. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 is a schematic block diagram of an image sensor provided in an embodiment of this application;

[0017] Figure 2 is a schematic diagram of a pixel array in an embodiment of this application;

[0018] Figure 3 is another schematic diagram of the pixel array structure in an embodiment of this application;

[0019] Figure 4 is another structural schematic diagram of the pixel array in an embodiment of this application;

[0020] Figure 5 is a flowchart illustrating an image generation method provided in an embodiment of this application;

[0021] Figure 6 is a schematic block diagram of an imaging device provided in an embodiment of this application. Detailed Implementation

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0024] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0025] Currently, image sensors are widely used in fields such as photography, surveillance, and machine vision. However, for applications such as virtual reality, augmented reality, and 3D reconstruction, it is necessary to obtain depth information of objects in the scene. However, current image sensors cannot directly obtain both image and depth information simultaneously. Therefore, depth estimation is usually performed using images from multiple image sensors, or by fusing image sensors and depth measurement sensors to obtain both images and depth information. However, this increases the number and size of sensors, leading to high integration difficulty and high manufacturing costs.

[0026] To address the aforementioned issues, this application provides an image sensor, an image generation method, and an imaging device. The image sensor includes a pixel array containing multiple image sensing pixels and multiple ToF pixels, enabling the image sensor to acquire RGB images and depth information. Based on the RGB images and depth information, a depth map can be generated, thereby achieving the simultaneous acquisition of RGB images and depth maps on a single image sensor. This effectively reduces the integration difficulty and manufacturing cost of the image sensor. Furthermore, since the depth map is generated by fusing actual depth information and RGB images, the accuracy of the depth map can also be guaranteed.

[0027] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0028] Please refer to Figure 1, which is a schematic block diagram of the structure of an image sensor provided in an embodiment of this application.

[0029] As shown in Figure 1, the image sensor 100 includes a pixel array 110, a controller 120, and a signal processor 130. The pixel array 110 includes multiple pixels arranged in multiple rows and columns. These pixels include a first number of image sensing pixels and a second number of ToF pixels. The image sensing pixels are used to expose and output a voltage signal, while the ToF pixels are used to determine depth information. The controller 120 controls the second number of ToF pixels to perform ranging to obtain depth information and controls the first number of image sensing pixels to perform exposure to output a voltage signal. The signal processor 130 processes the voltage signal to obtain a target RGB image. The controller 120 also generates a target depth map based on the target RGB image and the depth information.

[0030] In some embodiments, the first number of image sensing pixels in the pixel array 110 is greater than the second number of ToF pixels, and the ratio of the first number to the second number is a preset ratio, and the ToF pixels are uniformly distributed within the pixel array. The preset ratio can be set based on actual conditions, and this embodiment does not specifically limit it. For example, a preset ratio of 63:1 means that there is one ToF pixel in every 64 pixels in the pixel array 110. Another example is a preset ratio of 127:1, meaning that there is one ToF pixel in every 128 pixels in the pixel array 110.

[0031] In some embodiments, the area of ​​a ToF pixel is the same as the area of ​​an image sensing pixel, or the area of ​​a ToF pixel is the same as the sum of the areas of multiple image sensing pixels. The image sensing pixels include any one of R pixels, G pixels, and B pixels. For example, as shown in FIG2, the pixel array 110 includes 12 ToF pixels, and the area of ​​each of these 12 ToF pixels is the same as the area of ​​an image sensing pixel.

[0032] In some embodiments, the areas of ToF pixels at different locations may be the same or different. For example, as shown in FIG3, the pixel array 110 includes 5 ToF pixels, and the area of ​​each of these 5 ToF pixels is the same as the sum of the areas of 4 image sensing pixels. As another example, as shown in FIG4, the pixel array 110 includes 9 ToF pixels, and among these 9 ToF pixels, the area of ​​the ToF pixel located in the center of the pixel array 110 is the same as the sum of the areas of 4 image sensing pixels, while the area of ​​each of the remaining ToF pixels is the same as the sum of the areas of 2 image sensing pixels. Therefore, the area of ​​the ToF pixel located in the center of the pixel array 110 is different from the area of ​​each of the remaining ToF pixels.

[0033] In some embodiments, the ToF pixel includes a first microlens, a first color filter, a ToF photosensitive device, and a time digital conversion circuit (TDC), and the image sensing pixel includes a second microlens, a second color filter, and a photosensitive unit; the first color filter is configured to allow light in a first wavelength range to pass through, the second color filter is configured to allow light in a second wavelength range to pass through, and the first wavelength range is different from the second wavelength range. Among them, the second wavelength range includes any one of the wavelength ranges of red light, green light, and blue light, and the first wavelength range includes any one of the wavelength ranges of infrared light and laser light. The ToF photosensitive device includes at least one of a silicon-based single photon avalanche diode (SPAD), an avalanche photodiode (APD), and a silicon photomultiplier (SiPM). Since the wavelength range of light allowed to pass through by the color filter in the ToF pixel is different from the wavelength range of light allowed to pass through by the color filter in the image sensing pixel, the light of the remaining wavelengths in the outside world will not affect the operation of the ToF pixel and the image sensing pixel. Therefore, the ToF pixel and the image sensing pixel can work simultaneously, ensuring the consistency of the generation time of the depth map and the RGB image.

[0034] For example, while an infrared light emitter located outside the image sensor 100 emits infrared light, the controller 120 in the image sensor 100 controls a second number of ToF pixels to perform ranging to obtain depth information, and controls a first number of image sensing pixels to perform exposure to obtain voltage signals; the signal processor 130 processes the voltage signals to obtain a target RGB image; the controller 120 generates a target depth map based on the target RGB image and the depth information. Among them, the infrared light emitted by the light emitter is emitted by the measured object, and the reflected infrared light passes through the first microlens and the first water filter and reaches the ToF photosensitive device. The infrared light is sensed by the ToF photosensitive device and outputs a current signal to the TDC. The TDC records the received time. According to the time when the light emitter emits infrared light and the received time, the light flight time is determined. After N times of transmission and reception, the TDC can record n times (n < N) of light flight times. Based on the n times of light flight times, a histogram of the light flight time distribution is generated, and the light flight time with the highest occurrence frequency is obtained to get the target light flight time. According to the target light flight time, the distance between the ToF pixel and the measured object (the depth value of the pixel point corresponding to the ToF pixel) is calculated, that is, multiplying the target light flight time by the speed of light and then dividing by 2, so as to obtain the distance between the ToF pixel and the measured object.

[0035] In some embodiments, processing the voltage signal to obtain the target RGB image may include: converting the voltage signal into a digital signal, generating an initial RGB image based on the digital signal, and determining the initial RGB image as the target RGB image. Alternatively, the voltage signal may be converted into a digital signal, an initial RGB image generated based on the digital signal, the target RGB value corresponding to the ToF pixel determined based on the initial RGB image and the position of the ToF pixel in the pixel array, and the initial RGB image compensated based on the target RGB value corresponding to each ToF pixel to obtain the target RGB image. Since the pixel array 110 contains ToF pixels, the positions corresponding to the ToF pixels in the initial RGB image are missing colors. This embodiment can determine the target RGB image at the position of the ToF pixel in the initial RGB image by the position of the ToF pixel in the pixel array, and compensate the initial RGB image based on the target RGB image, thereby filling in the missing colors at the positions corresponding to the ToF pixels and improving the quality of the RGB image.

[0036] In some embodiments, compensating the initial RGB image based on the target RGB value corresponding to each ToF pixel to obtain the target RGB image may include: determining multiple compensation positions in the initial RGB image based on the position of each ToF pixel in the pixel array and a preset position transformation relationship, with one ToF pixel corresponding to one compensation position; filling each compensation position in the initial RGB image with the corresponding target RGB value to obtain the target RGB image. The preset position transformation relationship can be set based on actual conditions, and this embodiment does not specifically limit it.

[0037] In some embodiments, determining the target RGB value corresponding to a ToF pixel based on an initial RGB image and the position of the ToF pixel in the pixel array may include: obtaining the neighborhood corresponding to the ToF pixel from the initial RGB image based on the position of the ToF pixel in the pixel array; and determining the target RGB value corresponding to the ToF pixel based on the RGB values ​​of the pixels in the neighborhood. The size of the neighborhood can be set based on actual conditions; for example, the size of the neighborhood can be 3*3 or 2*2. The RGB values ​​of the pixels in the neighborhood refer to the RGB values ​​of the pixels located near the corresponding position of the ToF pixel in the initial RGB image. By using the RGB values ​​of the pixels near the corresponding position of the ToF pixel in the initial RGB image, the missing color at the corresponding position of the ToF pixel in the initial RGB image can be accurately determined.

[0038] In some embodiments, determining the target RGB value corresponding to the ToF pixel based on the RGB values ​​of the pixels in the neighborhood may include: calculating the average of the RGB values ​​of the pixels in the neighborhood and determining the average as the target RGB value corresponding to the ToF pixel. Alternatively, it may involve determining the weighting coefficient of each pixel in the neighborhood based on the distance between the location of each pixel in the neighborhood and the location corresponding to the ToF pixel; calculating the product of the RGB value of each pixel in the neighborhood and its corresponding weighting coefficient to obtain the weighted RGB value of each pixel in the neighborhood; and summing the weighted RGB values ​​of each pixel in the neighborhood to obtain the target RGB value corresponding to the ToF pixel. The weighting coefficient of a pixel is inversely proportional to the distance between the location of the pixel and the location corresponding to the ToF pixel; that is, the closer the location of a pixel in the neighborhood is to the location corresponding to the ToF pixel, the larger the weighting coefficient of the pixel; and the farther the location of a pixel in the neighborhood is to the location corresponding to the ToF pixel, the smaller the weighting coefficient of the pixel.

[0039] In some embodiments, generating a target depth map based on a target RGB image and depth information may include: inputting the target RGB image into a preset first depth map generation model for processing to obtain a first depth map; and optimizing the first depth map based on the depth information to obtain the target depth map. The first depth map generation model is obtained by iteratively training a neural network model based on a first training sample set. The first training sample set includes multiple first training samples, each including a first sample RGB image and annotated real depth maps. This embodiment obtains an initial first depth map by performing depth estimation on the target RGB image, and then optimizes the first depth map based on depth information obtained from actual ranging, resulting in a target depth map with higher accuracy and precision, thus improving the accuracy of the depth map.

[0040] In some embodiments, the training process of the first depth map generation model is as follows: obtaining a first training sample from the first training sample set as the current training sample; inputting the first sample RGB image in the current training sample into a preset neural network model for depth estimation processing to obtain a predicted depth map; determining the model loss value based on the predicted depth map and the real depth map in the current training sample; if the model loss value is greater than the preset loss value, updating the parameters of the neural network model and returning to the step of obtaining a first training sample from the first training sample set as the current training sample; if the model loss value is less than or equal to the preset loss value, ending the training and determining the current neural network model as the first depth map generation model.

[0041] In some embodiments, the depth information includes measured depth values ​​of one or more target pixels corresponding to each ToF pixel. A target pixel is any pixel in the target RGB image. Optimizing the first depth map based on the depth information to obtain a target depth map may include: determining a target depth value for each pixel in the first depth map based on the measured depth values ​​of one or more target pixels corresponding to each ToF pixel and the current depth value of each pixel in the first depth map; and optimizing the first depth map based on the target depth value of each pixel in the first depth map to obtain the target depth map. Specifically, this involves replacing the current depth value of each pixel in the first depth map with the corresponding target depth value to obtain the target depth map.

[0042] In some embodiments, determining the target depth value of each pixel in the first depth map based on the measured depth value of one or more target pixels corresponding to each ToF pixel and the current depth value of each pixel in the first depth map may include: for each pixel in the first depth map, if the pixel is a target pixel, performing a weighted summation of the current depth value of the pixel and the corresponding measured depth value to obtain the target depth value of the pixel; if the pixel is not a target pixel, determining the current depth value of the pixel as the target depth value.

[0043] Furthermore, the weighted summation of the current depth value and the corresponding measured depth value of the pixel may include: calculating the product of the current depth value of the pixel and a first preset coefficient to obtain a first weighted depth value; calculating the product of the corresponding measured depth value of the pixel and a second preset coefficient to obtain a second weighted depth value; and summing the first weighted depth value and the second weighted depth value to obtain the target depth value of the pixel. Wherein, the first preset coefficient is less than or equal to the second preset coefficient, and the sum of the first preset coefficient and the second preset coefficient is 1. For example, the first preset coefficient is 0.5, and the second preset coefficient is 0.5. Or the first preset coefficient is 0.4, and the second preset coefficient is 0.6.

[0044] In some embodiments, generating a target depth map based on a target RGB image and depth information may include: inputting the depth information and the target RGB image into a preset second depth map generation model for processing to obtain the target depth map. The preset second depth map generation model is obtained by iteratively training a neural network model based on a second training sample set. The second training sample set includes multiple second training samples, which include second sample RGB images, labeled real depth maps, and actual depth values ​​of some pixels in the second sample RGB images. This embodiment, based on depth information obtained from actual ranging and the target RGB image, can generate target depth maps with higher precision and accuracy, improving the accuracy of the depth map.

[0045] In some embodiments, the training process of the second depth map generation model is as follows: A second training sample is obtained from the second training sample set as the current training sample; the actual depth values ​​of some pixels in the second sample RGB image and the second sample RGB image in the current training sample are input into a preset neural network model for depth estimation processing to obtain a predicted depth map; based on the predicted depth map and the real depth map in the current training sample, a model loss value is determined; if the model loss value is greater than a preset loss value, the parameters of the neural network model are updated, and the process returns to the step of obtaining a second training sample from the second training sample set as the current training sample; if the model loss value is less than or equal to the preset loss value, training ends, and the current neural network model is determined as the second depth map generation model.

[0046] In some embodiments, each ToF pixel in the second number of ToF pixels can independently turn its ranging function on and off. The ranging function of the ToF pixels can be manually turned on or off by the user, or it can be automatically turned on or off by the image sensor 100; this application embodiment does not specifically limit this.

[0047] Please refer to Figure 5, which is a flowchart illustrating an image generation method provided in an embodiment of this application.

[0048] As shown in Figure 5, the image generation method includes steps S101 to S102.

[0049] Step S101: Acquire voltage signal and depth information.

[0050] In this embodiment, the voltage signal is obtained by exposure of the image sensing pixels included in the image sensor, and the depth information is obtained by ranging of the ToF pixels included in the image sensor. For example, while the light emitter located outside the image sensor 100 emits infrared light, the controller 120 in the image sensor 100 controls a second number of ToF pixels to perform ranging to obtain depth information, and controls a first number of image sensing pixels to perform exposure to obtain a voltage signal.

[0051] Step S102: Process the voltage signal to obtain the target RGB image, and generate the target depth map based on the target RGB image and depth information.

[0052] In this embodiment, the voltage signal can be processed by the signal processor 130 in the image sensor 100 to obtain the target RGB image. Alternatively, the voltage signal can be processed by the controller 120 in the image sensor 100 to obtain the target RGB image; this embodiment does not specifically limit the method used.

[0053] In some embodiments, processing the voltage signal to obtain the target RGB image may include: converting the voltage signal into a digital signal, generating an initial RGB image based on the digital signal, and determining the initial RGB image as the target RGB image; or, converting the voltage signal into a digital signal, generating an initial RGB image based on the digital signal; determining the target RGB value corresponding to the ToF pixel based on the initial RGB image and the position of the ToF pixel in the pixel array; and performing compensation processing on the initial RGB image based on the target RGB value corresponding to each ToF pixel to obtain the target RGB image. Since the pixel array 110 contains ToF pixels, the positions corresponding to the ToF pixels in the initial RGB image are missing colors. In this embodiment, the position of the ToF pixel in the pixel array can determine the target RGB image at the position corresponding to the ToF pixel in the initial RGB image, and the initial RGB image is compensated based on the target RGB image, thereby filling in the missing colors at the positions corresponding to the ToF pixels and improving the quality of the RGB image.

[0054] In some embodiments, compensating the initial RGB image based on the target RGB value corresponding to each ToF pixel to obtain the target RGB image may include: determining multiple compensation positions in the initial RGB image based on the position of each ToF pixel in the pixel array and a preset position transformation relationship, with one ToF pixel corresponding to one compensation position; filling each compensation position in the initial RGB image with the corresponding target RGB value to obtain the target RGB image. The preset position transformation relationship can be set based on actual conditions, and this embodiment does not specifically limit it.

[0055] In some embodiments, determining the target RGB value corresponding to a ToF pixel based on an initial RGB image and the position of the ToF pixel in the pixel array may include: obtaining the neighborhood corresponding to the ToF pixel from the initial RGB image based on the position of the ToF pixel in the pixel array; and determining the target RGB value corresponding to the ToF pixel based on the RGB values ​​of the pixels in the neighborhood. The size of the neighborhood can be set based on actual conditions; for example, the size of the neighborhood can be 3*3 or 2*2. The RGB values ​​of the pixels in the neighborhood refer to the RGB values ​​of the pixels located near the corresponding position of the ToF pixel in the initial RGB image. By using the RGB values ​​of the pixels near the corresponding position of the ToF pixel in the initial RGB image, the missing color at the corresponding position of the ToF pixel in the initial RGB image can be accurately determined.

[0056] In some embodiments, determining the target RGB value corresponding to the ToF pixel based on the RGB values ​​of the pixels in the neighborhood may include: calculating the average of the RGB values ​​of the pixels in the neighborhood and determining the average as the target RGB value corresponding to the ToF pixel. Alternatively, it may involve determining the weighting coefficient of each pixel in the neighborhood based on the distance between the location of each pixel in the neighborhood and the location corresponding to the ToF pixel; calculating the product of the RGB value of each pixel in the neighborhood and its corresponding weighting coefficient to obtain the weighted RGB value of each pixel in the neighborhood; and summing the weighted RGB values ​​of each pixel in the neighborhood to obtain the target RGB value corresponding to the ToF pixel. The weighting coefficient of a pixel is inversely proportional to the distance between the location of the pixel and the location corresponding to the ToF pixel; that is, the closer the location of a pixel in the neighborhood is to the location corresponding to the ToF pixel, the larger the weighting coefficient of the pixel; and the farther the location of a pixel in the neighborhood is to the location corresponding to the ToF pixel, the smaller the weighting coefficient of the pixel.

[0057] In some embodiments, generating a target depth map based on a target RGB image and depth information may include: inputting the target RGB image into a preset first depth map generation model for processing to obtain a first depth map; and optimizing the first depth map based on the depth information to obtain the target depth map. The first depth map generation model is obtained by iteratively training a neural network model based on a first training sample set. The first training sample set includes multiple first training samples, each including a first sample RGB image and annotated real depth maps. This embodiment obtains an initial first depth map by performing depth estimation on the target RGB image, and then optimizes the first depth map based on depth information obtained from actual ranging, resulting in a target depth map with higher accuracy and precision, thus improving the accuracy of the depth map.

[0058] In some embodiments, the training process of the first depth map generation model is as follows: obtaining a first training sample from the first training sample set as the current training sample; inputting the first sample RGB image in the current training sample into a preset neural network model for depth estimation processing to obtain a predicted depth map; determining the model loss value based on the predicted depth map and the real depth map in the current training sample; if the model loss value is greater than the preset loss value, updating the parameters of the neural network model and returning to the step of obtaining a first training sample from the first training sample set as the current training sample; if the model loss value is less than or equal to the preset loss value, ending the training and determining the current neural network model as the first depth map generation model.

[0059] In some embodiments, the depth information includes measured depth values ​​of one or more target pixels corresponding to each ToF pixel. A target pixel is any pixel in the target RGB image. Optimizing the first depth map based on the depth information to obtain a target depth map may include: determining a target depth value for each pixel in the first depth map based on the measured depth values ​​of one or more target pixels corresponding to each ToF pixel and the current depth value of each pixel in the first depth map; and optimizing the first depth map based on the target depth value of each pixel in the first depth map to obtain the target depth map. Specifically, this involves replacing the current depth value of each pixel in the first depth map with the corresponding target depth value to obtain the target depth map.

[0060] In some embodiments, determining the target depth value of each pixel in the first depth map based on the measured depth value of one or more target pixels corresponding to each ToF pixel and the current depth value of each pixel in the first depth map may include: for each pixel in the first depth map, if the pixel is a target pixel, performing a weighted summation of the current depth value of the pixel and the corresponding measured depth value to obtain the target depth value of the pixel; if the pixel is not a target pixel, determining the current depth value of the pixel as the target depth value.

[0061] Furthermore, the weighted summation of the current depth value and the corresponding measured depth value of the pixel may include: calculating the product of the current depth value of the pixel and a first preset coefficient to obtain a first weighted depth value; calculating the product of the corresponding measured depth value of the pixel and a second preset coefficient to obtain a second weighted depth value; and summing the first weighted depth value and the second weighted depth value to obtain the target depth value of the pixel. Wherein, the first preset coefficient is less than or equal to the second preset coefficient, and the sum of the first preset coefficient and the second preset coefficient is 1. For example, the first preset coefficient is 0.5, and the second preset coefficient is 0.5. Or the first preset coefficient is 0.4, and the second preset coefficient is 0.6.

[0062] In some embodiments, generating a target depth map based on a target RGB image and depth information may include: inputting the depth information and the target RGB image into a preset second depth map generation model for processing to obtain the target depth map. The preset second depth map generation model is obtained by iteratively training a neural network model based on a second training sample set. The second training sample set includes multiple second training samples, which include second sample RGB images, labeled real depth maps, and actual depth values ​​of some pixels in the second sample RGB images. This embodiment, based on depth information obtained from actual ranging and the target RGB image, can generate target depth maps with higher precision and accuracy, improving the accuracy of the depth map.

[0063] In some embodiments, the training process of the second depth map generation model is as follows: A second training sample is obtained from the second training sample set as the current training sample; the actual depth values ​​of some pixels in the second sample RGB image and the second sample RGB image in the current training sample are input into a preset neural network model for depth estimation processing to obtain a predicted depth map; based on the predicted depth map and the real depth map in the current training sample, a model loss value is determined; if the model loss value is greater than a preset loss value, the parameters of the neural network model are updated, and the process returns to the step of obtaining a second training sample from the second training sample set as the current training sample; if the model loss value is less than or equal to the preset loss value, training ends, and the current neural network model is determined as the second depth map generation model.

[0064] The image generation method provided in this application embodiment can use an image sensor to simultaneously acquire RGB images and depth maps. The depth map is obtained based on the depth information obtained from actual ranging and the target RGB image, which effectively improves the accuracy of the depth map.

[0065] Please refer to Figure 6, which is a schematic block diagram of an imaging device provided in an embodiment of this application.

[0066] As shown in Figure 6, the imaging device 200 includes a light emitter 210 and an image sensor 100. The imaging device 200 can be applied to electronic devices, including mobile phones, tablets, laptops, desktop computers, personal digital assistants, head-mounted displays, aircraft, or robot vacuums. Head-mounted displays can include augmented reality (AR) glasses, AR helmets, mixed reality (MR) glasses, and MR helmets.

[0067] In some embodiments, the light emitter 210 is used to emit infrared light or laser light, and the light emitter 210 may include a vertical-cavity surface-emitting laser (VCSEL). As shown in FIG1, the image sensor 100 includes a pixel array 110, a controller 120, and a signal processor 130. The pixel array 110 includes a plurality of pixels arranged in multiple rows and columns, the plurality of pixels including a first number of image sensing pixels and a second number of ToF pixels. The image sensing pixels are used to expose and output voltage signals, and the ToF pixels are used to determine depth information. The controller 120 is used to control the second number of ToF pixels to perform ranging to obtain depth information; and to control the first number of image sensing pixels to perform exposure to output voltage signals. The signal processor 130 is used to process the voltage signals to obtain a target RGB image; the controller 120 is also used to generate a target depth map based on the target RGB image and the depth information.

[0068] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the imaging device described above can be referred to the corresponding process in the aforementioned image generation method embodiments, and will not be repeated here.

[0069] This application also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs that can be executed by one or more processors to implement any of the image generation methods provided in the specification of this application.

[0070] The storage medium can be an internal storage unit of the imaging device described in the foregoing embodiments, such as the hard disk or memory of the imaging device. Alternatively, the storage medium can be an external storage device of the imaging device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card.

[0071] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware embodiments, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0072] It should be understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0073] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. The above descriptions are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image sensor, wherein, include: A pixel array includes a plurality of pixels arranged in multiple rows and columns, the plurality of pixels including a first number of image sensing pixels and a second number of ToF pixels, the image sensing pixels being used to expose an output voltage signal and the ToF pixels being used to determine depth information; The controller is used to control the second number of ToF pixels to perform ranging and obtain depth information; And control the first number of image sensing pixels to expose so as to output a voltage signal; A signal processor is used to process the voltage signal to obtain a target RGB image; The controller is also configured to generate a target depth map based on the target RGB image and the depth information.

2. The image sensor according to claim 1, wherein, The step of generating a target depth map based on the target RGB image and the depth information includes: The target RGB image is input into a preset first depth map generation model for processing to obtain a first depth map; Based on the depth information, the first depth map is optimized to obtain the target depth map.

3. The image sensor according to claim 2, wherein, The depth information includes measured depth values ​​of one or more target pixels corresponding to each ToF pixel, where the target pixel is any pixel in the target RGB image. The step of optimizing the first depth map based on the depth information to obtain a target depth map includes: The target depth value of each pixel in the first depth map is determined based on the measured depth value of one or more target pixels corresponding to each ToF pixel and the current depth value of each pixel in the first depth map; Based on the target depth value of each pixel in the first depth map, the first depth map is optimized to obtain the target depth map.

4. The image sensor according to claim 3, wherein, The step of determining the target depth value of each pixel in the first depth map based on the measured depth value of one or more target pixels corresponding to each ToF pixel and the current depth value of each pixel in the first depth map includes: For each pixel in the first depth map, when the pixel is the target pixel, the current depth value of the pixel and the corresponding measured depth value are weighted and summed to obtain the target depth value of the pixel. If the pixel is not the target pixel, the current depth value of the pixel is determined as the target depth value.

5. The image sensor according to claim 3, wherein, The step of optimizing the first depth map based on the target depth value of each pixel in the first depth map to obtain the target depth map includes: Replace the current depth value of each pixel in the first depth map with the corresponding target depth value to obtain the target depth map.

6. The image sensor according to claim 1, wherein, The step of generating a target depth map based on the target RGB image and the depth information includes: The depth information and the target RGB image are input into a preset second depth map generation model for processing to obtain the target depth map.

7. The image sensor according to any one of claims 1-6, wherein, The ToF pixel includes a first microlens, a first color filter, a ToF photosensitive device, and a time-to-digital conversion circuit; the image sensing pixel includes a second microlens, a second color filter, and a photosensitive unit. The first color filter is used to allow light of a first wavelength range to pass through, and the second color filter is used to allow light of a second wavelength range to pass through, wherein the first wavelength range is different from the second wavelength range.

8. The image sensor according to claim 7, wherein, The ToF photosensitive device includes at least one of silicon-based single-photon avalanche diode, avalanche photodiode, and silicon photomultiplier tube.

9. The image sensor according to any one of claims 1-6, wherein, The process of processing the voltage signal to obtain the target RGB image includes: The voltage signal is converted into a digital signal, and an initial RGB image is generated based on the digital signal. Based on the initial RGB image and the position of the ToF pixel in the pixel array, determine the target RGB value corresponding to the ToF pixel; and The initial RGB image is compensated based on the target RGB value corresponding to each ToF pixel to obtain the target RGB image.

10. The image sensor according to claim 9, wherein, The step of determining the target RGB value corresponding to the ToF pixel based on the initial RGB image and the position of the ToF pixel in the pixel array includes: Based on the position of the ToF pixel in the pixel array, the neighborhood corresponding to the ToF pixel is obtained from the initial RGB image; The target RGB value corresponding to the ToF pixel is determined based on the RGB values ​​of the pixels in the neighborhood.

11. The image sensor according to claim 10, wherein, Determining the target RGB value corresponding to the ToF pixel based on the RGB values ​​of the pixels in the neighborhood includes: Calculate the average RGB value of the pixels in the neighborhood, and determine the average value as the target RGB value corresponding to the ToF pixel.

12. The image sensor according to claim 10, wherein, Determining the target RGB value corresponding to the ToF pixel based on the RGB values ​​of the pixels in the neighborhood includes: The weighting coefficient of each pixel in the neighborhood is determined based on the distance between the position of each pixel in the neighborhood corresponding to the ToF pixel and the position corresponding to the ToF pixel. The weighted RGB value of each pixel in the neighborhood is obtained by multiplying the RGB value of each pixel in the neighborhood by the corresponding weighting coefficient. The weighted RGB values ​​of each pixel in the neighborhood are summed to obtain the target RGB value corresponding to the ToF pixel.

13. The image sensor according to claim 9, wherein, The step of compensating the initial RGB image based on the target RGB value corresponding to each ToF pixel to obtain the target RGB image includes: Based on the position of each ToF pixel in the pixel array and a preset position transformation relationship, multiple positions to be compensated in the initial RGB image are determined, with one ToF pixel corresponding to one position to be compensated; The target RGB image is obtained by filling each of the locations to be compensated in the initial RGB image with the corresponding target RGB value.

14. An image generation method, wherein, Applied to an image sensor as described in any one of claims 1-13, the method comprises: The voltage signal and depth information are acquired. The voltage signal is obtained by exposure of the image sensing pixels contained in the image sensor, and the depth information is obtained by ranging of the ToF pixels contained in the image sensor. The voltage signal is processed to obtain a target RGB image, and a target depth map is generated based on the target RGB image and the depth information.

15. The image generation method according to claim 14, wherein, The step of generating a target depth map based on the target RGB image and the depth information includes: The target RGB image is input into a preset first depth map generation model for processing to obtain a first depth map; Based on the depth information, the first depth map is optimized to obtain the target depth map.

16. The image generation method according to claim 15, wherein, The depth information includes measured depth values ​​of one or more target pixels corresponding to each ToF pixel, where the target pixel is any pixel in the target RGB image. The step of optimizing the first depth map based on the depth information to obtain a target depth map includes: The target depth value of each pixel in the first depth map is determined based on the measured depth value of one or more target pixels corresponding to each ToF pixel and the current depth value of each pixel in the first depth map; Based on the target depth value of each pixel in the first depth map, the first depth map is optimized to obtain the target depth map.

17. The image generation method according to claim 14, wherein, The step of generating a target depth map based on the target RGB image and the depth information includes: The depth information and the target RGB image are input into a preset second depth map generation model for processing to obtain the target depth map.

18. The image generation method according to claim 14, wherein, The process of processing the voltage signal to obtain the target RGB image includes: The voltage signal is converted into a digital signal, and an initial RGB image is generated based on the digital signal. Based on the initial RGB image and the position of the ToF pixel in the pixel array, determine the target RGB value corresponding to the ToF pixel; and The initial RGB image is compensated based on the target RGB value corresponding to each ToF pixel to obtain the target RGB image.

19. The image generation method according to claim 18, wherein, The step of determining the target RGB value corresponding to the ToF pixel based on the initial RGB image and the position of the ToF pixel in the pixel array includes: Based on the position of the ToF pixel in the pixel array, the neighborhood corresponding to the ToF pixel is obtained from the initial RGB image; The target RGB value corresponding to the ToF pixel is determined based on the RGB values ​​of the pixels in the neighborhood.

20. An imaging device, wherein, It includes a light emitter and an image sensor as described in any one of claims 1-13.

Citation Information

Patent Citations

  • Depth estimation method, storage medium and computer equipment

    CN113343973A

  • 3D sensing system and method for providing image based on hybrid sensing array

    CN113542715A

  • Depth and image sensor device, manufacturing method and depth and image sensor chip

    CN114284306A

  • Depth compensation method and device

    CN117294829A

  • Image sensor, image generation method, and imaging apparatus

    CN118338151A