Image acquisition method and device, equipment and storage medium

By determining different exposure times according to the current light source environment in the original image sequence output by the image sensor and exposing them with pixels as particle size, the problem in the prior art that the Tone Mapping algorithm is difficult to calculate excellent tone images when dynamics are insufficient or overexposed is achieved, and a more natural tone relationship and rich image levels are achieved.

CN120186476APending Publication Date: 2025-06-20GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510402505.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When the prior art uses the Tone Mapping algorithm to improve the tone effect, due to the original image exposed by the image sensor, it is difficult to calculate an excellent tone image when the dynamics are insufficient or overexposed.

Method used

By obtaining the original image sequence output by the image sensor, a first exposure time related to the current light source environment is determined, the second exposure information is determined based on the first exposure information, and the original image sequence is exposed at a pixel size to obtain an image with a more natural tone relationship.

Benefits of technology

It realizes a richer and more natural presentation of image levels and details, presents a more realistic tone relationship, and eliminates the complex Tone Mapping algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186476A_ABST
    Figure CN120186476A_ABST
Patent Text Reader

Abstract

The invention discloses an image acquisition method and device, equipment and a storage medium, an original image sequence output by an image sensor is obtained, first exposure time corresponding to the original image sequence is determined, and the first exposure time is related to a current light source environment; second exposure information is determined based on the first exposure time and first exposure information of the image sensor, the first exposure information represents the exposure degree corresponding to each pixel of the image sensor, and the second exposure information comprises second exposure time corresponding to each pixel of the image sensor; based on the second exposure information, exposing the original image sequence by taking pixels as granularity to obtain a first image; and performing format conversion on the first image to obtain a second image, the second image being an image with a browsable format.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to multimedia processing technologies, and in particular, to an image acquisition method, apparatus, device, and storage medium. Background Art

[0002] Tone is one of the important expressive forces of a photo or video, referring to the relationships such as the light and dark levels and the virtual and real contrasts in the picture. Through these relationships, the viewer can feel the flow and change of light. What it emphasizes is the light and dark of the picture rather than the color.

[0003] In the related art, various tone mapping algorithms are used in mobile phones or cameras to improve the tone effect, including local and global ones. However, using the Tone Mapping algorithm is limited by the raw image exposed by the image sensor. If the dynamic range of the raw image is insufficient or overexposed, it is difficult for the Tone Mapping algorithm to calculate an image with excellent tone. Summary of the Invention

[0004] Embodiments of the present application provide an image acquisition method, apparatus, device, and storage medium, which can make the levels and details of the image richer and more natural, presenting a more real tone relationship.

[0005] The technical solution of the embodiments of the present application is implemented as follows:

[0006] Embodiments of the present application provide an image acquisition method, the method including:

[0007] Obtain a sequence of raw images output by an image sensor, and determine a first exposure time corresponding to the sequence of raw images, where the first exposure time is related to the current light source environment;

[0008] Based on the first exposure time and first exposure information of the image sensor, determine second exposure information, where the first exposure information represents the exposure degree corresponding to each pixel of the image sensor, and the second exposure information includes a second exposure time corresponding to each pixel of the image sensor;

[0009] Based on the second exposure information, expose the sequence of raw images in pixel granularity to obtain a first image;

[0010] Perform format conversion on the first image to obtain a second image, where the second image is an image with a browsable format.

[0011] Embodiments of the present application provide an image acquisition apparatus, the image acquisition apparatus including: a processor and an image sensor, where the processor is configured to:

[0012] Obtain the original image sequence output by the image sensor, and determine the first exposure time corresponding to the original image sequence, where the first exposure time is related to the current light source environment;

[0013] Based on the first exposure time and the first exposure information of the image sensor, determine the second exposure information, where the first exposure information represents the exposure degree corresponding to each pixel of the image sensor, and the second exposure information includes the second exposure time corresponding to each pixel of the image sensor;

[0014] Based on the second exposure information, perform exposure on the original image sequence pixel by pixel to obtain a first image;

[0015] Perform format conversion on the first image to obtain a second image, where the second image is an image with a browsable format.

[0016] An embodiment of the present application provides an electronic device, including the above image acquisition device.

[0017] An embodiment of the present application provides a computer-readable storage medium, that is, a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above image acquisition method is implemented.

[0018] The chip provided by an embodiment of the present application is used to implement the above image acquisition method. The chip includes: a processor, configured to call and run a computer program from a memory, so that a device installed with the chip executes the above image acquisition method.

[0019] The image acquisition method, device, equipment, and storage medium provided by an embodiment of the present application obtain the original image sequence output by the image sensor, and determine the first exposure time corresponding to the original image sequence, where the first exposure time is related to the current light source environment; based on the first exposure time and the first exposure information of the image sensor, determine the second exposure information, where the first exposure information represents the exposure degree corresponding to each pixel of the image sensor, and the second exposure information includes the second exposure time corresponding to each pixel of the image sensor; based on the second exposure information, perform exposure on the original image sequence pixel by pixel to obtain a first image; perform format conversion on the first image to obtain a second image, where the second image is an image with a browsable format; here, for the original image sequence, the exposure time of each pixel is predicted in combination with the current light source environment, and exposure is performed pixel by pixel to achieve variable pixel exposure, obtain an image with normal tone relationship and balanced brightness, and the levels and details of the exposed image are richer and more natural, presenting a more realistic tone relationship, and eliminating the complex Tone Mapping algorithm. Description of the Drawings

[0020] Figure 1 is an image acquisition device provided by an embodiment of the present application

[0021] Figure 2 is an optional process schematic diagram of an image acquisition method provided by an embodiment of the present application;

[0022] Figure 3 is an optional process schematic diagram of an image acquisition method provided by an embodiment of the present application;

[0023] Figure 4 is an optional process schematic diagram of an image acquisition method provided by an embodiment of the present application;

[0024] Figure 5 is an optional process schematic diagram of an image acquisition method provided by an embodiment of the present application;

[0025] Figure 6 is an optional process schematic diagram of an image acquisition method provided by an embodiment of the present application;

[0026] Figure 7 is an optional process schematic diagram of an image acquisition method provided by an embodiment of the present application;

[0027] Figure 8 is an optional process schematic diagram of an image acquisition method provided by an embodiment of the present application;

[0028] Figure 9 is an optional process schematic diagram of an image acquisition method provided by an embodiment of the present application;

[0029] Figure 10 is a schematic diagram of different tone mapping relationships provided by an embodiment of the present application;

[0030] Figure 11 is an optional process schematic diagram of an image acquisition method provided by an embodiment of the present application

[0031] Figure 12 is a schematic diagram of gray-scale conversion provided by an embodiment of the present application;

[0032] Figure 13 is a schematic diagram of gray-scale conversion provided by an embodiment of the present application;

[0033] Figure 14 is a schematic diagram of interpolation provided by an embodiment of the present application;

[0034] Figure 15 is a schematic diagram of image fusion provided by an embodiment of the present application;

[0035] Figure 16 is a schematic diagram of a pixel array provided by an embodiment of the present application;

[0036] Figure 17 It is a deployment schematic diagram of an image sensor provided by an embodiment of the present application;

[0037] Figure 18 It is a deployment schematic diagram of an image sensor provided by an embodiment of the present application;

[0038] Figure 19 It is a schematic diagram for comparing the image tone adjustment effects provided by an embodiment of the present application;

[0039] Figure 20 It is a schematic diagram for comparing the image tone adjustment effects provided by an embodiment of the present application;

[0040] Figure 21 It is an optional schematic diagram of an image acquisition device provided by an embodiment of the present application;

[0041] Figure 22 It is an optional schematic structural diagram of an image acquisition device provided by an embodiment of the present application;

[0042] Figure 23 It is an optional schematic structural diagram of an image acquisition device provided by an embodiment of the present application;

[0043] Figure 24 It is an optional schematic structural diagram of an image acquisition device provided by an embodiment of the present application;

[0044] Figure 25 It is an optional schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0045] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0046] Embodiments of the present application can provide an image acquisition method, device, equipment and storage medium. In practical applications, the image acquisition method can be implemented by an image acquisition device, and each functional entity in the image acquisition device can be jointly implemented by hardware resources such as a processor and an image sensor.

[0047] Of course, the embodiments of the present application are not limited to being provided as a method and hardware, and there are also various implementation manners, for example, being provided as a storage medium (storing instructions for executing the image acquisition method provided by the embodiments of the present application).

[0048] Next, the embodiments of the image acquisition method, device, equipment and storage medium provided by the embodiments of the present application will be described.

[0049] The image acquisition method provided by the embodiment of the present application is applied to an image acquisition device. For example, Figure 1 as shown, the image acquisition device 100 may include:

[0050] An image sensor 101 and a processor 102. Among them, the image sensor 101 can be called a photosensitive element, which is used to convert an optical image into a digital signal. It uses the photoelectric conversion function of optoelectronic devices to convert the optical image on the photosensitive surface into a digital signal that is in a corresponding proportional relationship with the optical image. Here, the original data output by the image sensor 101, that is, the data that has not been processed or compressed, is a raw image.

[0051] The processor 102 can be used to process the raw image output by the image sensor and output the processed image.

[0052] Here, the image output by the processor 102 can be called the output image of the image sensor 101.

[0053] In the case where the image acquisition device includes an image sensor and a processor, the image acquisition device can be understood as a programmable image sensor (abbreviated as programmable Sensor).

[0054] Each pixel of the programmable Sensor in the embodiment of the present application uses a different exposure time, so as to present rich levels through exposure.

[0055] In the embodiment of the present application, the image acquisition device can be applied as a camera module for acquiring images. The image acquisition device can be applied independently or integrated in an electronic device, and the electronic device can be a device with an image acquisition function such as a mobile phone, a tablet, or a wearable device.

[0056] The image acquisition method provided by the embodiment of the present application, which is applied to an image acquisition device, may be as Figure 2 shown, and includes:

[0057] S201. Obtain the original image sequence output by the image sensor and determine the first exposure time corresponding to the original image sequence. The first exposure time is related to the current light source environment.

[0058] When the image sensor in the image acquisition device acquires an image of a target object, the collected optical signal is converted into image information that is in a proportional relationship with the optical image, so as to obtain an original image. The format of the original image is a raw image. Among them, the image information may include one or more of the following information: the gray value of each pixel information, the light intensity, and the color.

[0059] Here, the original image output by the image sensor for a period of time is the original image sequence. Among them, the original image sequence includes one or more original images, that is, one frame or multiple frames of original images.

[0060] In some embodiments, the multiple original images included in the original image sequence are directed to the same target object.

[0061] The image acquisition device determines the exposure time of the current original image sequence based on the acquired image information. Here, the determined exposure time is referred to as the first exposure time. The first exposure time is for the original image and is related to the light source environment where the image sensor is currently located. The light in the light source environment of the image sensor is related to the target object and the ambient light of the target object.

[0062] S202. Determine second exposure information based on the first exposure time and the first exposure information of the image sensor. The first exposure information represents the exposure degree corresponding to each pixel of the image sensor, and the second exposure information includes the second exposure time corresponding to each pixel of the image sensor.

[0063] After the image acquisition device determines the first exposure time, it obtains the first exposure information. The first exposure information can be understood as an exposure array or an exposure matrix. The elements of the exposure array or the exposure matrix are used to determine the second exposure time of the pixel corresponding to the element. An element in the exposure array corresponds to one or more pixels in the image sensor.

[0064] In a possible implementation manner, the size of the exposure array is the same as the size H*W of the resolution of the image sensor. At this time, an element in the exposure array corresponds to one pixel. At this time, for a pixel, the second exposure time of the pixel is determined based on the element corresponding to the pixel and the first exposure time.

[0065] In one case, the first exposure information may include the exposure coefficient corresponding to each pixel among the pixels included in the image sensor. For a pixel, the second exposure time corresponding to the pixel is the product of the first exposure time and the exposure coefficient, that is, the second exposure time is the exposure coefficient times of the first exposure time.

[0066] In one case, the first exposure information may include the exposure time corresponding to each pixel among the pixels included in the image sensor. For a pixel, the exposure coefficient corresponding to the pixel is determined based on the exposure time corresponding to the pixel and the exposure time of one frame of the image, and the second exposure time corresponding to the pixel is the product of the first exposure time and the exposure coefficient.

[0067] In a possible implementation, the resolution of the image sensor is H*W, and the size of the exposure array is (H / m)*(W / m). At this time, one element in the exposure array corresponds to one pixel region, where one pixel region includes m*m pixels. In this case, for a pixel region, the second exposure time of the elements in the pixel region is determined based on the element corresponding to the pixel region and the first exposure time.

[0068] The pixel region can be understood as the arrangement unit adopted by the pixel arrangement. For example, when the image sensor adopts the Bayer array arrangement, one pixel region includes four pixels, that is, one pixel region includes 2*2 pixels.

[0069] In the embodiments of the present application, the first exposure information may include the exposure coefficient or exposure time corresponding to each pixel region in the image sensor. Then, for a pixel region, the second exposure time corresponding to the pixel region can be determined based on the first exposure time, and the second exposure time of each pixel in a pixel region is the second exposure time corresponding to the pixel region.

[0070] In the embodiments of the present application, the second exposure time is for pixels, and the second exposure times corresponding to different pixels or different pixel regions are independent.

[0071] In the embodiments of the present application, the pixel arrangement of the image sensor may be a single (Mono) arrangement, Bayer arrangement, Quad-Bayer, RYYB arrangement, etc. In the embodiments of the present application, no limitation is imposed on the pixel arrangement of the image sensor.

[0072] In the embodiments of the present application, when the pixel arrangement is such that the information of different pixels is not relevant, it can be considered that one pixel corresponds to one element in the exposure array. When the pixel arrangement is such that the information of multiple pixels is relevant, it can be considered that the multiple pixels are one pixel region, and one pixel region corresponds to one element in the exposure array.

[0073] S203. Based on the second exposure information, expose the original image sequence at the pixel level to obtain a first image;

[0074] The second exposure information includes the second exposure time corresponding to each pixel in the image sensor. Among them, the value of the second exposure time can be an integer multiple or a floating-point multiple of the exposure time of one frame of the image. For each pixel, the target pixel value of the pixel is determined based on the pixel values within the second exposure time in the original image sequence, thereby determining the first image.

[0075] S204. Perform format conversion on the first image to obtain a second image, and the second image is an image with a browsable format.

[0076] In the embodiments of the present application, the first image is a raw image obtained after exposure and is not an image in a browsable format. The image acquisition device performs format conversion on the first image to obtain a second image in a browsable format.

[0077] In the embodiments of the present application, the format of the second image may include one of the following: Bitmap (BMP), Portable Document Format (PDF), Joint Photographic Experts Group (JPEG), Graphics Interchange Format (GIF), Tagged Image File Format (TIFF). In the embodiments of the present application, no limitation is imposed on the format of the second image.

[0078] The image acquisition method, device, equipment, and storage medium provided in the embodiments of the present application obtain a raw image sequence output by an image sensor and determine a first exposure time corresponding to the raw image sequence, where the first exposure time is related to the current light source environment; based on the first exposure time and the first exposure information of the image sensor, determine second exposure information, where the first exposure information represents the exposure degree corresponding to each pixel of the image sensor, and the second exposure information includes a second exposure time corresponding to each pixel of the image sensor; based on the second exposure information, perform exposure on the raw image sequence in pixel granularity to obtain a first image; perform format conversion on the first image to obtain a second image, where the second image is an image in a browsable format; here, for the raw image sequence, the exposure time of each pixel is predicted in combination with the current light source environment, and exposure is performed pixel by pixel to achieve variable pixel exposure, obtaining an image with normal tone relationship and balanced brightness, and the levels and details of the exposed image are richer and more natural, presenting a more real tone relationship, and eliminating the complex Tone Mapping algorithm.

[0079] In some embodiments, determining the first exposure time corresponding to the raw image sequence in S201 includes:

[0080] S2011. Determine the first exposure time according to the brightness in the current light source environment measured by the image sensor and the image information collected by the image sensor.

[0081] In the embodiments of the present application, the exposure time may be iteratively adjusted according to the brightness in the current light source environment measured by the image sensor until the exposure time converges, so that the obtained first exposure time is related to the current light source environment. At this time, the obtained exposure time may be referred to as the first exposure time.

[0082] In the embodiments of the present application, the brightness of different regions in the image collected by the image sensor can be measured and statistically analyzed to determine the brightness distribution of the entire image. After completing the brightness statistics, the brightness of the current image is evaluated to determine whether the current brightness meets the requirements. If the current brightness does not meet the requirements, the exposure time of the image sensor is adjusted according to the evaluation result to change the brightness of the image. And the brightness statistics and evaluation are performed again according to the adjusted exposure time until the current brightness reaches the target brightness threshold and the algorithm converges. At this time, the obtained exposure time is the first exposure time.

[0083] Here, an automatic exposure control (AEC) algorithm can be used to determine the first exposure time according to the brightness measured by the image sensor in the current light source environment.

[0084] In some embodiments, S203 exposes the original image sequence in pixel granularity based on the second exposure information to obtain a first image, including:

[0085] For each pixel point on each pixel, the following processing shown is respectively performed Figure 3 as follows:

[0086] S301. Determine the second exposure time corresponding to the pixel in the second exposure information;

[0087] S302. Integrate the pixel points of the pixel in the corresponding second exposure time in the original image sequence to obtain the target pixel value corresponding to the pixel; the pixel value of each pixel in the first image is the target pixel value.

[0088] For each pixel in the image sensor, the image acquisition device determines the second exposure time of the pixel in the first exposure information according to the position of the pixel; and determines the pixel value of the pixel in the second exposure time in the original image sequence as the target pixel value.

[0089] Here, there is one pixel point corresponding to one pixel in the original image sequence. Determining the pixel value of the pixel in the second exposure time in the original image sequence as the target pixel value can be understood as, for a pixel point in an original image sequence, determining the pixel value of the pixel in the second exposure time as the target pixel value of the pixel point. Different pixel points are the positions corresponding to different pixels.

[0090] When the second exposure time is an integer multiple of the exposure time of each frame of the image, the target pixel value E at the p position p can be expressed as Equation (1):

[0091]

[0092] Among them, E p,i represents the pixel value at position p on the i-th original image in the original image sequence.

[0093] When the second exposure time is a floating-point multiple of the exposure time of each frame of image, the target pixel value E at position p p can be expressed as Equation (2):

[0094]

[0095] The first image obtained by the image acquisition device based on the second exposure information for exposure can be understood as a new raw image, which can be expressed as I(p, Δt), where p represents different pixel positions and Δt represents the exposure time of the pixel at position p. Among them, the size of I(p, Δt) is the same as the resolution of the image sensor, which is H*W.

[0096] In an example, the original image sequence includes T original images (the sequence is M0, M1, M2,..., MT), the resolution of the original image is 3x3, and the second exposure information is a 3x3 matrix [[1, 2, 3], [4, 5, 6], [7, 8, 9]]. Then, after exposure, the obtained I(p, Δt) is also a 3x3 matrix, and the exposure value corresponding to 1 in I(p, Δt) is the value of the first pixel in M0, the exposure value corresponding to 2 in the matrix is the value of the second pixel in M0 plus the value of the second pixel in M1, and correspondingly, the value corresponding to 3 is the sum of the corresponding pixel values of M0, M1, and M2.

[0097] In the image acquisition method provided by the embodiments of the present application, the original image can be exposed in a discrete or continuous manner, so as not to limit the applicable scenarios of the image acquisition method.

[0098] In a possible implementation manner, as Figure 4 shown, S204 performs format conversion on the first image to obtain a second image, including:

[0099] S401. Separating channels of the first image based on a first number of first channels to obtain a first image sequence, where the first image sequence includes images on each of the first channels in the first number of first channels, and the exposure times corresponding to different first channels are different;

[0100] S402. Interpolating the images on each of the first channels in the first image sequence to obtain a second image sequence;

[0101] S403. Fusing the second image sequence based on a tone mapping model to obtain the second image.

[0102] In the embodiments of the present application, the first quantity may be a predefined fixed quantity or a quantity determined based on the second exposure information. In one example, the first quantity is the maximum exposure time in the second exposure information.

[0103] In one example, the second exposure time is a floating-point multiple of the exposure time of each frame image, and the first quantity is a predefined fixed quantity.

[0104] In one example, the second exposure time is an integer multiple of the exposure time of each frame image, and the first quantity is the maximum exposure time in the second exposure information.

[0105] The image acquisition device separates the exposure of each pixel in the first image based on the first quantity of first channels, and obtains a first image sequence. The first channel can be understood as an exposure channel, and different first channels are channels at different exposure times.

[0106] In one example, taking the first quantity as C, each pixel point is separated into C channels. For a pixel point, the actual exposure time may be C0, C0 <= C. Then, the channels from 0 to C0 are set to the corresponding pixel values on mono, and for the channels from C0 + 1 to C, the pixel values are 0.

[0107] When the first image is marked as I(p, Δt), the first image sequence can be identified as I s (p, Δt), where the size of I s (p, Δt) is H * W * C.

[0108] The image acquisition device interpolates the images on each first channel in the first image sequence to obtain a second image sequence. The first image sequence can be understood as a sparse image sequence, and the second image sequence can be understood as a dense image sequence. Here, the pigeon image in the first image sequence can be understood as a sparse image. That is, for a first channel, some pixel points have values and some pixel points do not have values. In the embodiments of the present application, through interpolation, a dense image, that is, the image in the second image sequence, can be obtained, so as to obtain the second image sequence with more image information.

[0109] The first image sequence can be identified as I s (p, Δt), and the interpolated second image sequence can be identified as I d (p, Δt).

[0110] In the embodiments of the present application, the interpolation algorithms used during image interpolation may include but are not limited to: nearest neighbor boundary value, bilinear interpolation, bicubic interpolation, adaptive interpolation, interpolation based on deep learning, etc.

[0111] After the image acquisition device determines the second image sequence, it fuses the multi-channel images included in the second image sequence to obtain the first image.

[0112] In a possible implementation, S403 fuses the second image sequence based on a tone mapping model to obtain the second image, including: inputting the second image into a predefined tone mapping model to obtain a first image output by the tone mapping model. The input of the tone mapping model here includes the multi-channel second image sequence.

[0113] In a possible implementation, S403 fuses the second image sequence based on a tone mapping model to obtain the second image, including:

[0114] S501. For the images in the second image sequence, perform global darkening and global brightening respectively to obtain a third image sequence and a fourth image sequence;

[0115] S502. Input the second image sequence, the third image sequence, and the fourth image sequence into the tone mapping model to obtain the second image output by the tone mapping model.

[0116] As Figure 6 shown, perform overall darkening on the images in the image sequence 601 to obtain the image sequence 602, perform overall brightening on the images in the image sequence 601 to obtain the image sequence 603, input the image sequence 601, the image sequence 602, and the image sequence 603 into the tone mapping model 604 to obtain the image 605 output by the tone mapping model 604, where the image sequence 601 can be understood as the second image sequence, and the image 605 can be understood as the second image.

[0117] The tone mapping model here is a neural network model that can fuse multiple image sequences into one image, and can also be called a tone fusion network.

[0118] In the embodiments of the present application, the structure of the tone mapping model is not limited in any way.

[0119] In a possible implementation, S204 performs format conversion on the first image to obtain a second image, including:

[0120] Performing channel separation on the first image based on a third number of second channels to obtain a first image sequence, where the first image sequence includes images on each of the second channels in the third number of second channels, and different second channels correspond to different colors;

[0121] S402. Perform interpolation on the images on each of the second channels in the first image sequence to obtain a second image.

[0122] The second channel can be understood as a color channel. In one example, the second number is 3, and the 3 second channels include an R channel, a G channel, and a B channel.

[0123] Here, by separating the second channels of the second quantity, the complete RGB value of each pixel is restored.

[0124] In the embodiments of the present application, the solution of separating and fusing the first image through the second channels can be understood as performing demosaicing on the first image to obtain a second image.

[0125] In some embodiments, the first exposure information is received by the image acquisition device or obtained through training of the image acquisition device.

[0126] In some embodiments, when the first exposure information is obtained through training of the image acquisition device, the image acquisition method provided by the embodiments of the present application further includes:

[0127] Initializing third exposure information, where the third exposure information includes the third exposure time corresponding to each pixel of the image sensor, and the third exposure time is the initialized exposure time;

[0128] Determining a target image from the first video;

[0129] Based on the third exposure information, exposing the original images corresponding to the respective image frames in the first video at the pixel level to obtain a third image, and performing format conversion on the third image to obtain a fourth image;

[0130] When the fourth image does not meet the stop condition, updating the third exposure information based on the fourth image and the target image to obtain updated third exposure information, and continuing to expose the respective image frames in the first video at the pixel level based on the updated third exposure information to obtain a new fourth image until the obtained fourth image meets the stop condition;

[0131] Wherein, the third exposure information when the fourth image meets the stop condition is the first exposure information.

[0132] The manner in which the image acquisition device initializes the third exposure information may include: random initialization, or initialization according to an exposure array, where the exposure array may include: a uniform random array, a Poisson random array, a Nonad array, a Quad array, etc.

[0133] The image acquisition device uses the first video as a training sample to train the third exposure information to obtain the first exposure information. Wherein, the original image corresponding to the first video is used as the input, and the target frame in the first video is used as the label to train the third exposure information.

[0134] The image acquisition device can select a video frame from the first video as the target image. In one example, the target image is a randomly selected frame image in the first video. In one example, the target image is the frame image with the best image quality in the first video.

[0135] The training method of the third exposure information can be as Figure 7 shown, including:

[0136] S701. Based on the third exposure information, expose the original images corresponding to the respective image frames in the first video in terms of pixels to obtain a third image;

[0137] S702. Perform format conversion on the third image to obtain a fourth image.

[0138] S703. Based on whether the fourth image meets the stop condition.

[0139] In the case of not meeting the stop condition, execute S704. In the case of meeting the stop condition, determine that the training is over, and the third exposure information at this time can be considered as the first exposure information.

[0140] S704. Update the third exposure information based on the fourth image and the target image to obtain the updated third exposure information.

[0141] Here, continue to execute S701 based on the updated third exposure information.

[0142] In the embodiments of the present application, the loss can be determined based on the fourth image and the target image, and the third exposure information can be updated based on the loss. Among them, the calculation method of the loss in the embodiments of the present application is not limited in any way.

[0143] The exposure method in S701 can refer to the exposure method described in S203, which will not be elaborated here.

[0144] The description of the format conversion method in S702 can refer to the format conversion method described in S204, which will not be elaborated here.

[0145] In some embodiments, before determining the loss between the fourth image and the target image, the fourth image can be tone-mapped through a tone mapping algorithm so that both the fourth image and the target image undergo tone processing, improving the exposure effect of the first exposure information.

[0146] In the embodiments of the present application, before training the first exposure information, the original image corresponding to the first video can be converted into a mono image, and the target image can be converted into a mono image, thereby simplifying the first exposure information.

[0147] Converting the original image into a mono image can be understood as converting a RAW image into a grayscale image. In one example, if the pixel arrangement of the image sensor is the RGGB Bayer arrangement, the method of converting the original image corresponding to the first video into a mono image may include: weighting the values of each pixel based on the weights of each pixel in RGGB to obtain the values of each pixel in the pixel region. The calculation process can be shown as in Equation (3):

[0148]

[0149] Among them, one pixel region contains two green components. Therefore, in the average of RGB, the weights of R and B are 1 / 3 respectively, and the weights of the two Gs are 1 / 6 respectively.

[0150] In the embodiments of the present application, the mono image is a grayscale image, and the image has no color but only a gray component. Suppose the resolution of a raw image is 4096x4096. After converting the image with the rggb arrangement into a mono image, the resolution of the generated mono image is 2048x2048, because four pixel points (rggb) are combined into one point (gray).

[0151] Converting the target image into a mono image can be understood as converting an RGB image into a grayscale image. Among them, the target image is intercepted from the first video, and the image frames in the first video are RGB images obtained after processing the RAW images.

[0152] In one example, if the pixel arrangement of the image sensor is the RGGB Bayer arrangement, the method of converting the target image into a mono image may include: averaging the values of the three RGB channels of each pixel. In order to make the positions of the pixel points match precisely, a four-in-one operation is performed on the corresponding pixels on the target image. The calculation process can be shown as in Equation (4):

[0153]

[0154] Here, for example, if the resolution of the target image is 4096x4096, it also needs to convert the three components of R, G, and B into

[0155] Then, when converting the three components into gray, only direct averaging is required. Therefore, each component is multiplied by 1 / 3 and then added together.

[0156] In some embodiments, the format conversion of the third image to obtain the fourth image includes:

[0157] Performing channel separation on the third image based on the first channels of the second quantity to obtain a third image sequence, where the third image sequence includes the images on each of the first channels in the second quantity of first channels, and the exposure times corresponding to different first channels are different;

[0158] Interpolating the images on each of the first channels in the third image sequence to obtain a fourth image sequence;

[0159] Fusing the fourth image sequence based on the initial tone mapping model to obtain the fourth image;

[0160] Wherein, in the case that the fourth image does not meet the stop condition, the fourth image and the target image are further used to update the initial color bar mapping model to obtain an updated initial tone mapping model, and the initial tone mapping model when the fourth image meets the stop condition is the tone mapping model, and the tone mapping model is used to perform format conversion on the first image.

[0161] In the process of performing format conversion in S203, a tone mapping model is used. During the process of the image acquisition device training the third exposure information, the initial tone mapping model is trained together to obtain the tone mapping model. Among them, the structure of the tone mapping model can be set according to actual needs.

[0162] Here, in the format conversion process shown in S702, based on the second quantity of first channels, channel separation is performed on the third image to obtain a third image sequence, interpolation is performed on the images on each of the first channels in the third image sequence to obtain a high-dimensional image, and the high-dimensional image is fused based on the initial tone mapping model to obtain the fourth image.

[0163] Based on Figure 7 , as Figure 8 shown, includes:

[0164] S801. Based on the third exposure information, exposing the original images corresponding to each image frame in the first video at the pixel level to obtain a third image;

[0165] S802. Performing channel separation on the third image based on the first channels of the second quantity to obtain a third image sequence;

[0166] S803. Interpolating the images on each of the first channels in the third image sequence to obtain a fourth image sequence;

[0167] S804. Fusing the fourth image sequence based on the initial tone mapping model to obtain the fourth image.

[0168] S805. Based on whether the fourth image meets the stop condition.

[0169] When the stop condition is not met, S806 is executed. When the stop condition is met, it is determined that the training is completed, and the third exposure information at this time can be regarded as the first exposure information.

[0170] S806. Update the third exposure information and the initial tone mapping model based on the fourth image and the target image to obtain the updated third exposure information and the initial tone mapping model.

[0171] Here, S801 is continued to be executed based on the updated third exposure information and the initial tone mapping model.

[0172] The image acquisition method adopted in the embodiments of the present application can be as Figure 9 shown, including: a training stage 901 and an acquisition stage 902.

[0173] In the training stage 901, the first exposure information 9012 is obtained by training based on the first video 9011.

[0174] In the acquisition stage 902, the first exposure time 9022 is determined based on the original image sequence 9021, and the second exposure information 9023 is obtained based on the first exposure time 9022 and the first exposure information 9012. The original image sequence 9021 is exposed based on the second exposure information to obtain the first image 9024, and the first image 9024 is format-converted to obtain the second image 9024.

[0175] Among them, in the training stage 901, a tone mapping network can also be trained. When format-converting the first image 9024, the first image 9024 can be separated into multiple channels to obtain a first image sequence. The images in the first image sequence are subjected to brightness enrichment processing to obtain a second image sequence and a third image sequence. The first image sequence, the second image sequence, and the third image sequence are input into the tone mapping model to obtain the second image output by the tone mapping model.

[0176] Next, the image acquisition method provided in the embodiments of the present application will be further described.

[0177] Tone is one of the important expressive forces of a photo or video, referring to the relationships such as the light and dark levels and the virtual and real contrasts in the picture. Through these relationships, the viewer can feel the flow and change of light. What it emphasizes is the light and dark of the picture rather than the color. As Figure 10 shown, Image 1001 is an image with a dull tone, Image 1002 is an image with a gray and bright tone, and Image 1003 is an image with a normal tone.

[0178] In the related art, various Tone Mapping algorithms are used in mobile phones or cameras to improve the tone effect, some are local and some are global. However, the use of Tone Mapping algorithms is limited to the raw image exposed by the image sensor. If the raw image is not dynamic enough or overexposed, it is difficult for the Tone Mapping algorithm to calculate an image with excellent tone.

[0179] Existing sensors often have a small amount of light input, and in order to prevent overexposure, the Raw image obtained by exposure often compresses a large amount of information into a very small range. This causes a lot of information in natural light to be lost, and it is extremely difficult to recover.

[0180] In addition, due to the size of the aperture, the image obtained by the image sensor exposure is often extremely unevenly distributed. Most pixels are concentrated in the dark area, and the pixels corresponding to the light source are concentrated in the bright area. This causes the contrast of the dark area to be severely compressed. The image stretched by the back-end Tone Mapping algorithm is very likely to have no layers, or the layers are extremely unnatural and not delicate. The uneven distribution of the output image has brought a great burden to the Tone Mapping in the Image Signal Processing (ISP). The information of the dark area is difficult to restore through the Tone Mapping algorithm. When processing local areas, Tone Mapping often has problems such as tone inversion, too light local contrast, or too heavy local contrast. It makes the user look very different from the image scene in the human eye. Even if the Tone Mapping algorithm is improved, it is difficult for the algorithm to restore the true tone relationship. After the Tone Mapping algorithm is processed, the calm and gentle waves may show strong contrast, like soapy water spilled on the ground, or the contrast may not be clear, like mud mixed in the sea water.

[0181] In the embodiment of the present application, a programmable sensor is proposed, which is used to change the traditional sensor exposure so that the Raw image obtained by the sensor exposure has excellent layer distribution. Specifically, each pixel of the programmable sensor uses a different exposure time. For the mid-tone area, the sensor can present rich layers through exposure. In this way, the ISP only needs to adapt the brightness to the pixel light source on the screen, without using an algorithm to adjust the tone.

[0182] In the embodiments of the present application, a method for constructing a programmable Sensor is provided. An AI model is used to learn the exposure time of the Sensor's response under different light sources, so as to obtain a Raw image with excellent brightness distribution. Compared with using the ToneMapping algorithm, the Sensor exposure results in a more rich and natural hierarchy and details, and there will be no problems such as layering in the imaged image, and it can show a more realistic tone relationship.

[0183] In the embodiments of the present application, the training of the exposure time and the fusion network is as Figure 11 shown. The raw image sequence 1101 directly output by the Sensor (image sensor) is used as the input. Each pixel point is exposed by simulating different exposure times to obtain a new Raw image 1102. Among them, the exposure times of each pixel in the sensor constitute the exposure information. The new Raw image 1102 is separated to obtain an image sequence 1103, and the image sequence 1103 is interpolated to obtain a high-dimensional image sequence 1104. The high-dimensional image sequence 1104 is input into a fusion network 1105, and the output is the final image 1106. The final image 1106 is an image with a natural tone hierarchy, and the exposure information and the fusion network are iteratively updated through backpropagation between the final image 1106 and the target image 1107.

[0184] Among them, in the video corresponding to the raw image sequence 1101, frames are selected, and the clearest and best-toned picture is converted into a grayscale image as the target image 1107.

[0185] In the embodiments of the present application, in the first round of training, the input is the raw image sequence 1101 directly output by the Sensor (image sensor). In the nth round of training, the input is the input of the (n - 1)th round and the exposure time updated in the nth round.

[0186] In the embodiments of the present application, the exposure information can be understood as Figure 11 the grid included in the spatially variable exposure image shown. It includes a continuous plurality of grids. One grid represents a fixed exposure time t. A small square on the grid represents a pixel. The non-shadowed square indicates that this pixel has been exposed for the current time t, and the shadowed square indicates that the current pixel has not been exposed at the current time t. Taking a 3x3 grid as an example, the first 3x3 grid represents 3x3 pixels. In the time period from 0 to t, all 3x3 pixels have been exposed for a time t. The second grid represents that in the time period from t to t + t, among the 3x3 pixels, the first pixel has a shadow, indicating no exposure, and the second pixel has no shadow, indicating exposure. Continuous n grids mean that the maximum exposure time is n*t, and the grid sequence represents dividing the maximum exposure time into n parts.

[0187] For Figure 11The exposure time and the training process of the fusion network shown include:

[0188] 1. The training dataset consists of video datasets. When shooting a video, the Raw obtained by the exposure of each video's corresponding Sensor is taken out and converted into Mono form as the input. Frames are selected from the video, and the clearest and best-toned picture is taken out and converted into a grayscale picture as the target image.

[0189] 2. Randomly initialize the exposure time of each pixel of the Sensor, and generate new Raw according to the exposure time of each pixel.

[0190] 3. Separate the channels of the new Raw according to the exposure time and interpolate it into a high-dimensional image sequence.

[0191] 4. Perform multi-scale decomposition on the input high-dimensional image sequence, input it into a multi-branch fusion network, calculate the loss (Loss) between the fused image and the target image, and backpropagate to optimize the exposure time of the Sensor and the parameters of the fusion network.

[0192] Next, a specific description of the training process of the exposure time and the fusion network provided by the embodiments of the present application is given, including:

[0193] Step 1. Generate input and output sequences for the video. When shooting the video, the camera is stationary. The total number of frames of a video is T, and new Raw images obtained by the exposure of each frame's Sensor are generated. Among them, the pixels are RGGB. As Figure 12 shown, the grayscale is obtained by averaging four pixels and combined into an input arranged in Mono to obtain T Mono images. Select a clear and excellent-toned picture from the video. As Figure 13 shown, the pixels of the selected picture are weighted according to the pixel positions of the input Raw picture to obtain and then the RGB is averaged to obtain a grayscale picture as the target image (Target).

[0194] It should be noted that an inverse Gamma operation may be additionally performed on the Target, and the inverse Gamma is calibrated from the screen of the video shooting device.

[0195] Step 2. Randomly initialize the exposure time of each pixel on the Sensor. The formula adopted by the exposure mathematical model can be as shown in the following formula (5):

[0196]

[0197] The pixel value obtained by exposure is the integral of the illuminance V over time, where p represents different pixel positions. In the embodiments of the present application, a video is used to simulate the light source in the real world, and the time is discretized into S = {1, 2, …, T}. For the currently used cameras, the exposure time used globally is the same, so the exposure time of each pixel on the raw is also the same. Assuming that the exposure time of one frame of the image is 1, given the exposure time Δt ∈ S, the pixel value after exposure in Equation (1) is:

[0198]

[0199] where E p,i represents the pixel value at the p position on the i-th frame of the raw image. In this way, the new raw obtained by the programmable Sensor based on the exposure time can be expressed as I(p, Δt), where p represents different pixel positions and Δt represents the exposure time of the pixel at the p position. Among them, I(p, Δt) is an image with the same resolution as the Sensor, that is, I(p, Δt) ∈ R H×W .

[0200] In the above step 2, it is assumed that the exposure time of one frame of the image is 1 ms. In practical applications, the exposure time of one frame of the image can be any value, such as 5 ms, 10 ms, etc., and the exposure time Δt of each pixel position is an integer multiple of this value. Among them, the exposure time Δt and the exposure time of one frame of the image can be understood as the exposure degree.

[0201] In step 2, an exposure time matrix can be set. The size of this matrix is the same as the resolution of the mono. Each element in the exposure matrix is initialized to a value Δt, and then at each element position, exposure is performed according to the corresponding Δt. The exposure process is to integrate the pixel values on the mono according to Δt.

[0202] In an example, the resolution of the mono image is 3x3, and a 3x3 matrix [[1, 2, 3], [4, 5, 6], [7, 8, 9]] is initialized. The corresponding resolution of T monos is also 3x3 (the sequence is M0, M1, M2, …, MT). Then, after exposure, the obtained I(p, Δt) is also a 3x3 matrix, and the exposure value corresponding to 1 in I(p, Δt) is the value of the first pixel in M0. The exposure value corresponding to 2 in the matrix is the value of the second pixel in M0 plus the value of the second pixel in M1. Correspondingly, the value corresponding to 3 is the sum of the corresponding pixel values in M0, M1, and M2.

[0203] Step 3: Separate the value of each pixel point on I(p, Δt) into multiple channels using the exposure time, that is, R H×W →R H ×W×C , where C is a certain value in the set S.

[0204] In one example, the exposure time of a pixel is 3, and the pixel values of three consecutive frames of Mono at this position are respectively placed on three channels. The expanded image is marked as I s (p, Δt).

[0205] In the embodiments of the present application, each pixel is separated into C channels, and the number of channels is obtained according to the maximum exposure time. During initialization, it is the maximum value of matrix initialization. During the training of the AI model, the maximum exposure value calculated by the model is C. The number of channels can also be a predefined value.

[0206] For a pixel, the actual exposure time may be C0, C0 <= C. Then, the pixel values corresponding to 0 to C0 channels are set to the pixel values on Mono, and the pixel values of C0 + 1 to C are 0.

[0207] Step 4, I s (p, Δt) Each channel of may be a sparse image. Use KD-Tree interpolation for the sparse channels to obtain a dense, i.e., high-dimensional image sequence I d (p, Δt).

[0208] For KD-Tree interpolation, as Figure 14 shown, when interpolating the pixel point 1401, find the N pixel points closest to the pixel point 1401, and obtain the pixel value of the interpolated pixel point 1401 by weighted summation according to the distance I p , I p can be expressed as Equation (6):

[0209]

[0210] where, I i is the pixel value of the i-th pixel point among the N points, and the weight is calculated using the Euclidean distance, that is P is the position information of the pixel point I p of, p i the position information of the pixel point I i of.

[0211] Step 5, use the fused AI model to fuse the high-dimensional image sequence I d (p, Δt).

[0212] In one example, the fused AI model, i.e., the fusion network, can be as Figure 15As shown, the high-dimensional image sequence is respectively subjected to global darkening under-exposure (Shot Exp) and brightening over-exposure (Over Exp) to obtain an over-exposed image sequence and an under-exposed image sequence. The high-dimensional image sequence, the over-exposed image sequence, and the under-exposed image sequence are fed into the fusion network. Since the dimensions of the input image sequence are not fixed, 3D convolution in the fusion network is first used to fix the input image sequence to the same dimension. The encoder is used to compress the dimension, and the merger is used to splice (concat) the images with compressed dimensions. After that, the over-exposed image sequence and the under-exposed image sequence corresponding to the dimension-compressed images and the fused feature maps are fused using a fully convolutional layer. Then the decoder is used to expand the fusion result to the original dimension. In the case where each Encoder is connected to the Decoder through a skip connection, the output of the encoder can be directly input to the decoder, and then the decoder decodes after fusing the three received feature maps. A fully convolutional tone mapper is connected in series after the decoder, and then tone mapping is performed on the output of the decoder to obtain the output of the fusion network, that is, the final image.

[0213] Step 6: The loss (loss) of the training image is denoted as L p , L p can be expressed as Equation (7):

[0214] L p = L2 + λL VGG Equation (7);

[0215] Among them, L2 is the mean square error loss, and L VGG is the VGG loss.

[0216] Based on error backpropagation, the exposure time of each pixel of the Sensor is updated, and at the same time, the fusion network is updated to obtain the exposure information S1 and the fusion network. S1 includes the exposure time of each pixel.

[0217] In the embodiment of the present application, the exposure time of one shot is randomly initialized, and the updated exposure time is used after the exposure time is updated.

[0218] In the embodiment of the present application, the obtained S1 from the training result is the first exposure information, and the obtained fusion network is the tone mapping network.

[0219] The random initialization in step 2 in the embodiment of the present application can be replaced with the following Figure 16The following exposure arrays are shown: statistical random 1601, Poisson random 1602, Nonad array 1603, Quad array 1604. In Figure 16 there are three values for the exposure time: long exposure, medium exposure, and short exposure.

[0220] In step 2 above, the exposure time in formula (2) is discrete and an integer. Here, the exposure time can be extended to a floating point number. When using a floating point number, the pixel value obtained by exposure can be changed to formula (2):

[0221]

[0222] The corresponding exposure separation can also be changed to a separation with a fixed number of channels. The purpose is to remove redundant data, and the final obtained image will be more natural.

[0223] Next, the deployment of the variable sensor provided by the embodiments of the present application will be described. As Figure 17 shown, in a natural environment, AE C is used for exposure. Among them, AEC has a convergence process. After AEC converges, a globally unified exposure time t1, that is, the first exposure time, and the Sensor image are obtained. Taking t1 as the minimum exposure time, applying S1 to obtain the exposure time t2 of each pixel point, that is, the second exposure time, so as to infer the exposure time of different pixel points of the Sensor from the image output by the Sensor under the current light source. For a pixel point, t2 = t1 * Δt, where Δt is the exposure time of the corresponding pixel point in S1. It can be understood that Δt represents the multiple of the minimum exposure time under the current light source environment.

[0224] After determining the exposure time t2 of each pixel point, the corresponding pixels are exposed based on the exposure time t2 of each pixel to obtain the exposed image.

[0225] In the embodiments of the present application, exposure, separation, interpolation, and fusion networks are combined together to obtain a decoding algorithm (DecodeAlgo). Inputting the exposed image into the decoding algorithm can obtain a raw image with excellent tone.

[0226] In the above method, all the applied Sensors are in Mono arrangement. When the Sensor used in the camera is in Bayer arrangement, for the RGGB Sensor, based on Figure 17 , as Figure 18 shown, it is still the case that first, AEC converges to obtain the minimum exposure time t1. Correspondingly, taking four pixels of RG GB as a group, a group of pixel points use the same exposure time. Then, exposure and the Decode algorithm are performed on RGGB to obtain the raw image of RGGB.

[0227] In the embodiments of the present application, the generation of the multi-frame fusion network can use more complex AI network models or relatively lightweight AI models, which depends on the computing power of the device and the deployment method. High-level AI models can also be compatible with functions such as image noise reduction, Demosai c, and dead pixel correction, thus reducing the burden on the ISP. At the same time, lightweight AI models only require image fusion functions. Since the images obtained by Sensor exposure already contain all the light and shadow information, this AI model does not belong to a generative AI model, and its role is only to make a single adjustment to the raw image, without requiring too much computing power.

[0228] In the embodiments of the present application, a sensor with variable exposure per pixel is trained. This sensor predicts the corresponding exposure time for each pixel in the current light source environment, so that a raw image with normal tone relationship and balanced brightness can be obtained by exposure. In the ISP, there is no need to use the Tone Mapping algorithm to reconstruct the tone relationship, and an exposure image that conforms to the human eye can be directly captured.

[0229] In addition, in the embodiments of the present application, a video sequence is used to simulate the real light source, and frames are selected from the video as the target. After the sensor and the decoding algorithm are trained sufficiently, based on the exposure of the current AEC algorithm, it converges to the final Sensor like variable exposure. This greatly improves the light input of the Sensor, making the exposed pictures closer to the natural brightness seen by the human eye. The presented tone relationship will not have excessive contrast, nor will it be dull and foggy. A real and beautiful tone relationship can be captured from the data source, and thus the complex Tone Mapping algorithm in the ISP can be omitted.

[0230] The image acquisition method provided by the embodiments of the present application can obtain images such as Figure 19 image 1901 shown in Figure 20 and Figure 19 image 2001 shown in Figure 20 . It can be determined that the tone relationship obtained by the Sensor changing the exposure can restore the scenery seen by the human eye in the natural environment. After uneven exposure, when using the Tone Mapping algorithm to reconstruct the tone relationship, problems such as Figure 19 image 1902 shown in Figure 20 with excessive contrast, contours appearing in the local flat area, or Figure 20 image 2002 shown in

[0231] with low global contrast, the image looking gray and unnatural are very likely to occur.

[0231] An image acquisition device according to an embodiment of the present application, as Figure 21 shown, the image acquisition device 2100 includes:

[0232] An acquisition module 2101, configured to acquire the original image sequence output by the image sensor;

[0233] The first determination module 2102 is configured to determine a first exposure time corresponding to the original image sequence, where the first exposure time is related to the current light source environment;

[0234] The second determination module 2103 is configured to determine second exposure information based on the first exposure time and first exposure information of the image sensor, where the first exposure information characterizes the exposure degree corresponding to each pixel of the image sensor, and the second exposure information includes the second exposure time corresponding to each pixel of the image sensor;

[0235] The exposure module 2104 is configured to perform exposure on the original image sequence in pixel granularity based on the second exposure information to obtain a first image;

[0236] The conversion module 2105 is configured to perform format conversion on the first image to obtain a second image, where the second image is an image with a browsable format.

[0237] In some embodiments, the first determination module 2102 is further configured to determine the first exposure time according to the brightness in the current light source environment measured by the image sensor and the image information collected by the image sensor.

[0238] In some embodiments, the exposure module 2103 is further configured to:

[0239] For each pixel point on each pixel, the following processing is respectively performed:

[0240] Determine the second exposure time corresponding to the pixel in the second exposure information;

[0241] Integrate the pixel points of the pixel in the corresponding second exposure time in the original image sequence to obtain a target pixel value corresponding to the pixel; the pixel value of each pixel in the first image is the target pixel value.

[0242] In some embodiments, the conversion module 2104 is further configured to:

[0243] Perform channel separation on the first image based on a first number of first channels to obtain a first image sequence, where the first image sequence includes images on each of the first number of first channels, and the exposure times corresponding to different first channels are different;

[0244] Perform interpolation on the images on each of the first channels in the first image sequence to obtain a second image sequence;

[0245] Fuse the second image sequence based on a tone mapping model to obtain the second image.

[0246] In some embodiments, the conversion module 2104 is further configured to:

[0247] For the images in the second image sequence, perform global darkening and global brightening respectively to obtain a third image sequence and a fourth image sequence;

[0248] Input the second image sequence, the third image sequence, and the fourth image sequence into the tone mapping model to obtain the second image output by the tone mapping model.

[0249] In some embodiments, the image acquisition device 2100 further includes: a training module configured to:

[0250] Initialize third exposure information, where the third exposure information includes the third exposure time corresponding to each pixel of the image sensor, and the third exposure time is the initialized exposure time;

[0251] Determine a target image from the first video;

[0252] Based on the third exposure information, expose the original images corresponding to each image frame in the first video at the pixel level to obtain a third image, and perform format conversion on the third image to obtain a fourth image;

[0253] In the case where the fourth image does not meet the stop condition, update the third exposure information based on the fourth image and the target image to obtain updated third exposure information, and continue to expose each image frame in the first video at the pixel level based on the updated third exposure information to obtain a new fourth image until the obtained fourth image meets the stop condition;

[0254] Wherein, the third exposure information when the fourth image meets the stop condition is the first exposure information.

[0255] In some embodiments, the training module is further configured to:

[0256] Perform channel separation on the third image based on a second number of first channels to obtain a third image sequence, where the third image sequence includes the images on each of the second number of first channels, and the exposure times corresponding to different first channels are different;

[0257] Interpolate the images on each first channel in the third image sequence to obtain a fourth image sequence;

[0258] Fuse the fourth image sequence based on an initial tone mapping model to obtain the fourth image;

[0259] Wherein, in the case that the fourth image does not meet the stop condition, the fourth image and the target image are further used to update the initial color bar mapping model to obtain an updated initial tone mapping model. The initial tone mapping model when the fourth image meets the stop condition is the tone mapping model, and the tone mapping model is used to perform format conversion on the first image.

[0260] In practical applications, the above-mentioned obtaining module, first determination module, second determination module, exposure module, conversion module, training module, etc. can be implemented by a processor located on the image acquisition device of the electronic device, specifically a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor unit (MPU), an application processor (AP), a digital signal processor (DSP), and a field programmable gate array (FPGA).

[0261] In the embodiment of the present application, the image acquisition device and the image sensor can be integrated on one chip.

[0262] Those skilled in the art should understand that the relevant descriptions of the above-mentioned image acquisition device in the embodiment of the present application can be understood with reference to the relevant descriptions of the image acquisition method in the embodiment of the present application.

[0263] Figure 22 FIG. is a schematic structural diagram of an optional implementation of an electronic device provided in an embodiment of the present application. As Figure 22 shown, an embodiment of the present application provides an electronic device 2200, including an electronic chip 2201, and the electronic chip can be implemented as the image acquisition method described in one or more of the above embodiments.

[0264] An embodiment of the present application provides an image acquisition device. Figure 23 FIG. is a schematic structural diagram of another optional implementation of an image acquisition device provided in an embodiment of the present application. As Figure 23 shown, an embodiment of the present application provides an image acquisition device 2300, including:

[0265] a processor 2301 and a storage medium 2302 storing executable instructions of the processor 2301. The storage medium 2302 depends on the processor 2301 to perform operations through a communication bus 2303. When the instructions are executed by the processor 2301, the image acquisition method executed in one or more of the above embodiments is executed.

[0266] It should be noted that in actual application, each component in the terminal is coupled together through the communication bus 2303. It can be understood that the communication bus 2303 is used to implement the connection and communication between these components. In addition to the data bus, the communication bus 2303 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 23 all kinds of buses are labeled as the communication bus 2303.

[0267] An embodiment of the present application provides a computer storage medium, and the computer-readable storage medium is used to store a computer program, and the computer program causes the computer to execute the steps of the image acquisition method as described in one or more of the above embodiments.

[0268] A schematic structural diagram of an image acquisition device 2400 provided by an embodiment of the present application. The image acquisition device 2400 can be implemented as a camera module. Figure 24 The shown image acquisition device 2400 includes a processor 2410 and an image sensor. The processor 2410 is configured to:

[0269] Obtain the original image sequence output by the image sensor 2402, and determine the first exposure time corresponding to the original image sequence, where the first exposure time is related to the current light source environment;

[0270] Based on the first exposure time and the first exposure information of the image sensor, determine the second exposure information, where the first exposure information represents the exposure degree corresponding to each pixel of the image sensor, and the second exposure information includes the second exposure time corresponding to each pixel of the image sensor;

[0271] Based on the second exposure information, perform exposure on the original image sequence in pixel granularity to obtain a first image;

[0272] Perform format conversion on the first image to obtain a second image, and the second image is an image with a browsable format.

[0273] In an embodiment of the present application, the processor 2410 can call and run a computer program from the memory to implement the image acquisition method in the embodiment of the present application.

[0274] Optionally, as Figure 24 shown, the image acquisition device 2400 may further include a memory 2420. Among them, the processor 2410 can call and run a computer program from the memory 2420 to implement the image acquisition method in the embodiment of the present application.

[0275] Among them, the memory 2420 can be a separate device independent of the processor 2410 or can be integrated into the processor 2410.

[0276] An embodiment of the present application provides an electronic device, such as Figure 25 As shown, the electronic device 2500 includes an image acquisition device 2501, and the image acquisition device 2501 can be the image acquisition device described above.

[0277] Optionally, the electronic device 1400 can implement the corresponding processes implemented by the electronic device in each method of the embodiments of the present application. For the sake of brevity, it will not be described in detail here. It should be understood that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or by instructions in the form of software. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or can be executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0278] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.

[0279] It should be understood that the above-mentioned memory is by way of example but not limitation. For example, the memory in the embodiments of the present application can also be a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate SDRAM (DDR SDRAM), an enhanced SDRAM (ESDRAM), a synch link DRAM (SLDRAM), and a direct rambus random access memory (Direct Rambus RAM, DR RAM), etc. That is to say, the memory in the embodiments of the present application is intended to include but not be limited to these and any other suitable types of memory.

[0280] The embodiments of the present application also provide a computer-readable storage medium for storing a computer program.

[0281] Optionally, the computer-readable storage medium can be applied to the electronic device in the embodiments of the present application, and the computer program causes the computer to execute the corresponding processes implemented by the electronic device in the various methods of the embodiments of the present application. For the sake of brevity, details are not described herein again.

[0282] The embodiments of the present application also provide a computer program product, including computer program instructions.

[0283] Optionally, the computer program can be applied to the electronic device in the embodiments of the present application. When the computer program runs on the computer, it causes the computer to execute the corresponding processes implemented by the electronic device in the various methods of the embodiments of the present application. For the sake of brevity, details are not described herein again.

[0284] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0285] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and details are not described herein again.

[0286] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0287] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0288] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit.

[0289] If the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0290] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An image acquisition method, characterized in that: The method comprises: Obtaining an original image sequence output by an image sensor, and determining a first exposure time corresponding to the original image sequence, wherein the first exposure time is related to a current light source environment; Determining second exposure information based on the first exposure time and first exposure information of the image sensor, wherein the first exposure information represents an exposure level corresponding to each pixel of the image sensor, and the second exposure information includes a second exposure time corresponding to each pixel of the image sensor; Based on the second exposure information, the original image sequence is exposed with pixels as granularity to obtain a first image; The first image is format converted to obtain a second image, where the second image is an image in a browsable format.

2. The method according to claim 1, characterized in that: The determining a first exposure time corresponding to the original image sequence includes: The first exposure time is determined according to the brightness under the current light source environment measured by the image sensor and the image information collected by the image sensor.

3. The method according to claim 1, characterized in that The step of exposing the original image sequence based on the second exposure information with pixels as granularity to obtain a first image includes: For each pixel, the following processing is performed: Determine a second exposure time corresponding to the pixel in the second exposure information; Integrate the pixel points of the pixel in the original image sequence within the corresponding second exposure time to obtain the target pixel value corresponding to the pixel; the pixel value of each pixel in the first image is the target pixel value.

4. The method according to claim 1, characterized in that The step of converting the format of the first image to obtain a second image includes: Based on the first number of first channels, the first image is channel-separated to obtain a first image sequence, where the first image sequence includes images on each first channel of the first number of first channels, and different first channels correspond to different exposure times; interpolating the images on each first channel in the first image sequence to obtain a second image sequence; The second image sequence is fused based on a tone mapping model to obtain the second image.

5. The method according to claim 4, characterized in that The fusing the second image sequence based on the tone mapping model to obtain the second image includes: For the images in the second image sequence, globally darken and globally brighten them respectively to obtain a third image sequence and a fourth image sequence; The second image sequence, the third image sequence and the fourth image sequence are input into the tone mapping model to obtain the second image output by the tone mapping model.

6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Initializing third exposure information, wherein the third exposure information includes a third exposure time corresponding to each pixel of the image sensor, and the third exposure time is an initialized exposure time; determining a target image from a first video; Based on the third exposure information, exposing the original image corresponding to each image frame in the first video with pixels as granularity to obtain a third image, and performing format conversion on the third image to obtain a fourth image; In a case where the fourth image does not satisfy the stop condition, the third exposure information is updated based on the fourth image and the target image to obtain updated third exposure information, and each image frame in the first video is continuously exposed based on the updated third exposure information with pixels as granularity to obtain a new fourth image, until the obtained fourth image satisfies the stop condition; Among them, the third exposure information of the fourth image that meets the stop condition is the first exposure information.

7. The method according to claim 6, characterized in that The method of obtaining a fourth image based on format conversion of the third image includes: Based on the second number of first channels, the third image is channel-separated to obtain a third image sequence, wherein the third image sequence includes images on each first channel of the second number of first channels, and different first channels correspond to different exposure times; interpolating the images on each first channel in the third image sequence to obtain a fourth image sequence; fusing the fourth image sequence based on an initial tone mapping model to obtain the fourth image; Among them, when the fourth image does not meet the stop condition, the fourth image and the target image are also used to update the initial color stripe mapping model to obtain an updated initial tone mapping model, and the initial tone mapping model when the fourth image meets the stop condition is the tone mapping model, and the tone mapping model is used to perform format conversion on the first image.

8. An image acquisition device, characterized in that: The image acquisition device comprises: a processor and an image sensor, wherein the processor is configured to: Obtaining an original image sequence output by the image sensor, and determining a first exposure time corresponding to the original image sequence, wherein the first exposure time is related to a current light source environment; Determining second exposure information based on the first exposure time and first exposure information of the image sensor, wherein the first exposure information represents an exposure level corresponding to each pixel of the image sensor, and the second exposure information includes a second exposure time corresponding to each pixel of the image sensor; Based on the second exposure information, the original image sequence is exposed with pixels as granularity to obtain a first image; The first image is format converted to obtain a second image, where the second image is an image in a browsable format.

9. An electronic device comprising the image acquisition device according to claim 8.

10. A storage medium storing an executable program, characterized in that: When the executable program is executed by a processor, the image acquisition method described in any one of claims 1 to 7 is implemented.

11. A chip comprising an image sensor and a processor, characterized in that: The processor is configured to execute the image acquisition method according to any one of claims 1 to 7.