Image processing methods and apparatus, electronic devices, storage media
By combining image processing methods with color sensor arrays and multi-scale decomposition techniques, the limitations and abrupt changes in existing image processing technologies are resolved, enabling fine control over the target region and improving image quality and processing flexibility.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2022-08-04
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, the combination of global and local tone mapping in image processing cannot perform fine processing on certain parts, resulting in poor image quality and limitations and abrupt changes.
By acquiring the image to be processed and dividing it into multiple image blocks using a color sensor array, multi-scale decomposition is performed to obtain pixel information and difference information of each image layer. The target image is generated by combining the mapped pixel information and difference information of each image layer, thereby achieving local control and multi-scale processing of the target region.
It improves the accuracy and independence of image processing, eliminates processing differences between different image patches, enhances tone mapping effects, and improves image quality and comprehensiveness.
Smart Images

Figure CN115187488B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of imaging technology, and more specifically, to an image processing method and apparatus, an electronic device, and a computer-readable storage medium. Background Technology
[0002] In image processing, tone mapping is a common method to improve image quality.
[0003] In related technologies, tone mapping is generally performed by combining global tone mapping and local tone mapping. This method cannot perform fine processing on certain parts of the content and may have certain abrupt changes, which has certain limitations and results in poor image quality.
[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this disclosure is to provide an image processing method and apparatus, electronic device, and computer-readable storage medium, thereby overcoming, at least to some extent, the problem of poor image quality caused by the limitations and defects of related technologies.
[0006] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part by practice of this disclosure.
[0007] According to a first aspect of this disclosure, an image processing method is provided, comprising: acquiring an image to be processed, and dividing the image to be processed into multiple image blocks using a color sensor array; performing multi-scale decomposition on the image to be processed to obtain multiple image layers, and acquiring pixel information and difference information of each image block in each image layer; performing tone mapping on the pixel information of each image layer to obtain mapped pixel information of a target region; and combining the mapped pixel information and difference information of each image layer to obtain target pixel information of each image layer to generate a target image.
[0008] According to a second aspect of this disclosure, an image processing apparatus is provided, comprising: an image acquisition module for acquiring an image to be processed and dividing the image to be processed into multiple image blocks in conjunction with a color sensor array; an image decomposition module for performing multi-scale decomposition on the image to be processed to obtain multiple image layers, and acquiring pixel information and difference information of each image block in each image layer; a mapping module for performing tone mapping on the pixel information of each image layer to acquire mapped pixel information of a target region; and an image generation module for combining the mapped pixel information and difference information of each image layer to acquire target pixel information of each image layer to generate a target image.
[0009] According to a third aspect of this disclosure, an electronic device is provided, comprising: an imaging module including a color sensor array; a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the image processing method described in the first aspect by executing the executable instructions.
[0010] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the image processing method described in the first aspect above.
[0011] In the technical solution provided in this disclosure, on the one hand, the gain information of each image block is determined, and the mapping pixel information of the target region is further determined, so as to realize the individual control of the local image blocks represented by the target region. This avoids the limitation of related technologies that cannot perform targeted processing on the required regions, improves the accuracy and independence of image processing, and also improves the targeting of image processing, thereby improving the effect of tone mapping. On the other hand, it can combine multiple image layers to realize the individual control of the target region in multiple scale dimensions. Compared with a single scale, it can eliminate the abrupt effect caused by the processing differences between different image blocks in the spatial region, improve the comprehensiveness, and improve the image quality obtained by tone mapping.
[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0014] Figure 1 A schematic diagram illustrates an application scenario where the image processing method of the present disclosure embodiments can be applied.
[0015] Figure 2 The diagram illustrates an image processing method according to an embodiment of the present disclosure.
[0016] Figure 3 A schematic diagram of an image block is shown in an embodiment of this disclosure.
[0017] Figure 4 The diagram illustrates a multi-scale decomposition in an embodiment of this disclosure.
[0018] Figure 5This illustration shows a flowchart of the process for obtaining pixel information in an embodiment of the present disclosure.
[0019] Figure 6 A schematic diagram illustrating the target area of an embodiment of this disclosure is shown.
[0020] Figure 7 This schematic diagram illustrates the smooth transition of gain information in different regions in an embodiment of this disclosure.
[0021] Figure 8 A schematic diagram illustrating a smooth transition method in an embodiment of this disclosure is shown.
[0022] Figure 9 The schematic diagram illustrates the structure of the image signal processor in an embodiment of this disclosure.
[0023] Figure 10 A block diagram of an image processing apparatus according to an embodiment of the present disclosure is shown schematically.
[0024] Figure 11 A block diagram of an electronic device according to an embodiment of the present disclosure is shown schematically. Detailed Implementation
[0025] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0026] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0027] In related technologies, tone mapping is generally divided into global tone mapping and local tone mapping. Global tone mapping is fast, but it cannot differentiate based on information content. Local tone mapping can adjust the mapping function based on local spatial information. Combining global and local tone mapping can effectively handle tone mapping effects. However, with the expansion of media terminal applications, more user demands have been placed on the performance of this subsystem module. Beyond simple spatial region division, there is a lack of further refined processing of information content. For example, in portrait selfies, group photos, and other situations with more interesting content, more refined processing of the important content in these target areas is needed. In some scenarios, it is necessary to segment and extract these important information regions. The segmented regions should be combined with color information from multi-window color sensors for richer color management to achieve more targeted brightness and color representation of the target areas.
[0028] To address the technical problems in related technologies, this disclosure provides an image processing method that can be applied to scenarios involving tone mapping of images during the photography process. Figure 1 A schematic diagram of a system architecture for an image processing method and apparatus applicable to embodiments of the present disclosure is shown.
[0029] like Figure 1 As shown, terminal 101 can be a smart device with image processing capabilities, such as a smartphone, computer, tablet, smart speaker, smartwatch, in-vehicle device, wearable device, monitoring device, etc. The terminal may include a camera; the type of camera can be any type, as long as it can perform image processing. The number of cameras can be at least one, for example, one, four, etc., as long as they can take pictures. The image to be processed can be a captured image or each frame of a captured video.
[0030] In this embodiment, terminal 101 may include memory 102 and processor 103. The memory stores images, and the processor processes the images, such as performing white balance processing. Memory 102 may store an image 104 to be processed. Terminal 101 retrieves the image 104 to be processed from memory 102 and sends it to processor 103. Processor 103 performs block processing on the image to be processed to obtain multiple image blocks. The image to be processed is decomposed into multiple image layers at multiple scales, and pixel information and difference information of each image block in each image layer are obtained. Next, tone mapping is performed on the pixel information of each image layer to obtain the mapped pixel information of the target region. The mapped pixel information and difference information of each image layer are combined to obtain the target pixel information of each image layer, thereby generating a tone-processed target image 105.
[0031] It should be noted that the image processing method provided in this embodiment can be executed by terminal 101. Alternatively, the image processing method can be configured within the terminal.
[0032] Next, refer to Figure 2 The image processing methods in the embodiments of this disclosure will be described in detail.
[0033] In step S210, the image to be processed is acquired and divided into multiple image blocks using a color sensor array.
[0034] In this embodiment, the image to be processed can be an image captured by the camera module of the terminal, or it can be each frame of a captured video. The image to be processed can also be an image or each frame of a video directly obtained from a photo album or other storage location. The terminal can be any of the following: a smartphone, digital camera, smartwatch, wearable device, in-vehicle device, or surveillance camera, as long as it can capture images of the object and perform image processing. A smartphone is used as an example here. The camera module can include at least one camera, such as a main camera, telephoto camera, wide-angle camera, macro camera, or a combination thereof. The image to be processed can be various types of images, such as moving images or still images, etc.
[0035] The image to be processed can be an RGB image, i.e., an RGB three-channel image. Each pixel in an RGB image is composed of three colors: RGB. When the terminal is in camera mode, the image captured by the camera module can be a RAW image. A RAW image is the raw image data information acquired by the camera module. In this embodiment, the captured image obtained by the terminal can be converted to obtain the image to be processed. For example, a general conversion algorithm can be used to convert the RAW format captured image to RGB format to obtain the image to be processed, thereby improving the convenience of subsequent processing.
[0036] After acquiring the image to be processed, it can be divided into multiple image blocks. Each image block can be a portion of the image to be processed, and the blocks do not overlap. Each image block can be the same size, and the number of image blocks can be determined based on the number of grid cells. For example, grid information can be provided and applied to the image to be processed to divide it into multiple image blocks according to the grid information, with each block corresponding one-to-one with the grid cells. (Reference) Figure 3As shown, it can contain multiple image blocks 301. For example, grid 00 corresponds to image block 00, grid 01 corresponds to image block 01, and so on. The size of the grid information can be set according to actual needs and hardware structure, that is, set according to actual needs within the limits of the hardware structure. For example, it can be row × column. Based on this, the color sensor can be associated with image blocks and grids, and the image to be processed can be divided into row × column image blocks, with each image block corresponding to each grid. Under the same field of view, the multi-window spatial range represented by the grid information is consistent with the image to be processed, so that the grid information can cover the image to be processed. Each grid can correspond to one window, so it can be called multi-window.
[0037] In this embodiment, an additional color sensor array may be provided, which may include multiple color sensors 302 arranged in an array. The number of color sensors can be determined based on the number of image blocks. Furthermore, the color sensor array can be combined with multi-window information, so each color sensor can be a multi-window color sensor. Each grid represents a color sensor, and each color sensor corresponds to each image block. Since the image to be processed is divided into multiple grids, a row×col color sensor array is formed in space, which constitutes a multi-window color sensor.
[0038] The multi-window color sensor array is an independent sensor that can be positioned to one side of any camera in the camera module, close to the camera module. The camera module can be a rear-facing camera module, and it can include at least one camera, such as a main camera, telephoto camera, wide-angle camera, macro camera, or a combination thereof. The specific placement and arrangement order of the multiple cameras can be determined according to actual needs and are not specifically limited here. For example, the color sensor array can be positioned to the left of the telephoto camera, to the right of the main camera, or below the last camera in the at least one camera setup, etc. The camera module and the color sensor array can be placed adjacent to each other or separated by a certain distance. The specific position of the color sensor array can be determined based on the calibration results during the actual application process or according to actual needs; it is not specifically limited here.
[0039] After introducing a color sensor array, each color sensor in the array can detect associated image patches to obtain the local color information of each patch. The color sensors can be used to detect scene color, color temperature, spectrum, and other related information. They can be used to detect the color information of the object corresponding to each image patch and the color temperature information of the current scene. The object can be any type of object contained in each image patch, such as a person, etc. Here, the spectrum refers to different wavelengths, and the color temperature is the response curve for different wavelengths. Wavelength corresponds to color, so the spectrum, color temperature, and color are interrelated. Therefore, this explanation uses the local color information of the object corresponding to each image patch as an example.
[0040] Next, in step S220, the image to be processed is decomposed into multiple image layers at multiple scales, and the pixel information and difference information of each image block in each image layer are obtained.
[0041] In this embodiment, the mesh information acquired by the multi-window color sensor and the mask information obtained by the segmentation module are transmitted to the backend tone mapping module. The tone mapping module calls the mesh information, and the local tone mapping uses local block mesh information to match the block mesh size and region in the mesh information of the multi-window color sensor. That is, the mesh information of the tone mapping module matches the mesh information of the color sensor. The tone mapping here can be local tone mapping.
[0042] After determining the grid information of the tone mapping module, multi-scale decomposition processing is performed on the grid information. Since the grid information corresponds to the image to be processed, the image to be processed can be decomposed into multiple image layers based on the grid information, and each image layer can include at least one image patch. Specifically, the decomposition can be performed using a pyramid decomposition method, such as a Gaussian pyramid or a Laplacian pyramid, without specific limitations here.
[0043] In some embodiments, adjacent image blocks of the next layer adjacent to the current layer in the multi-layer image system can be aggregated to obtain the image block of the current layer, until the current layer no longer meets the decomposition conditions, thereby decomposing the image to be processed into multiple image layers. The current layer can be any layer, such as the first layer or the Nth layer, and the current layer can change dynamically according to the decomposition process. The next layer refers to the adjacent layer located at the bottom of the current layer.
[0044] Figure 4 A schematic diagram illustrating multi-scale decomposition is shown below. (Refer to...) Figure 4As shown, the first layer is the original image layer representing the image to be processed, where each image block represents a minimum unit of information data. The second layer aggregates adjacent image blocks from the first layer, outputting the resulting image blocks at their corresponding spatial locations in the second layer. That is, adjacent image blocks from the first layer are aggregated to obtain image blocks for the second layer. This aggregation can be a weighted summation. After a weighted summation of all image blocks from the first layer, all corresponding image blocks for the second layer are obtained. Dividing the number of rows and columns of the image blocks in the first layer by 2 gives the number of rows and columns of the image blocks in the second layer. This process is repeated to aggregate image blocks in the second layer until the current layer no longer meets the decomposition criteria. The decomposition criteria are not met when the current layer contains only one image block (1 row, 1 column), meaning it contains only one image block and cannot be further decomposed.
[0045] refer to Figure 4 As shown, the image blocks of the first layer can be aggregated to obtain the image blocks of the second layer, and the image blocks of the second layer can be aggregated to obtain the third layer. This aggregation is repeated to obtain the Nth image layer, the N+1th image layer, the N+2th image layer, and so on, until the N+3th layer containing only one image block is obtained, so as to decompose the image to be processed into multiple image layers based on grid information.
[0046] Specifically, the current layer can be any layer, and the current layer can change from the first layer to the last layer.
[0047] We can start from the first layer and downsample each information unit to obtain the information unit of the next layer (the second layer) connected to it.
[0048] Furthermore, the difference between the upsampled pixel information of the first layer and the original pixel information of the first layer can be calculated to obtain the difference information between adjacent layers. For example, the data array of the first layer with r rows and c columns is first downsampled to obtain a data array with r / 2 rows and c / 2 columns. This data array with r / 2 rows and c / 2 columns is the data array of the second layer. The data array of the second layer is then upsampled to obtain a data array with r rows and c columns. This data array has the same number of rows and columns as the original array of the first layer. Next, the difference between the two data arrays of the same size corresponding to the first layer is calculated pixel by pixel to obtain an array as the difference information. This difference information can be stored in the corresponding layer position, such as the corresponding position in the first layer. The above steps are continued until the current layer does not meet the decomposition conditions. Not meeting the decomposition conditions means reaching the configured number of layers or the current layer containing only a single pixel. In this way, the pixel information of each layer and the difference information between each layer and the adjacent next layer can be obtained.
[0049] After obtaining the difference information, the upsampled pixel information of each layer and the difference information between each layer and the next adjacent layer can be fused to obtain the target pixel information of each layer.
[0050] exist Figure 4 Based on the multi-layered image structure, the weighted sum of the data information of the smallest information data unit represented by adjacent image blocks in each layer can be stored in the corresponding image block of the next layer. For example, the weighted sum of adjacent image blocks in the first layer can be stored in the image block of the second layer. Adjacent image blocks can be four adjacent image blocks, two adjacent image blocks, nine adjacent image blocks, etc. Here, we will use four adjacent image blocks as an example for explanation.
[0051] The data information for each image block in each image layer includes, but is not limited to: original pixel information, grid information acquired by a multi-window color sensor, mapping curves, and the distribution information of the mask information required by the segmentation module in each image block of each layer. The grid information includes color temperature, color, and other related information within each image block. Furthermore, the data information in each image block may differ slightly. Since the image to be processed is divided into multiple image layers, the mask information can also be divided into multiple image layers. The mask information scales with each image layer, and the mask information for each image layer can be determined based on the original mask information and the intersection or overlap of the multi-image layers. The distribution information of the mask information refers to the boundary information of the mask information in each image block of each image layer. Corresponding operations can be performed using the data information contained in each image block.
[0052] While dividing the image to be processed into multiple image layers, the original pixel information of the image to be processed can be decomposed into the pixel information of each image block of the multiple image layers. Figure 5 The flowchart for obtaining pixel information is illustrated in the diagram. Figure 5 As shown, the main steps include:
[0053] In step S510, the original pixel information of each image block in the next layer adjacent to the current layer is downsampled to obtain the downsampled pixel information of the current layer as the pixel information of each image block in the current layer;
[0054] In step S520, the downsampled pixel information of the current layer is upsampled to obtain the upsampled pixel information of each image block in the next layer adjacent to the current layer;
[0055] In step S530, the difference information between the upsampled pixel information and the original pixel information of the next adjacent layer of the current layer is obtained until the decomposition condition is no longer met, so as to obtain the difference information of each image layer.
[0056] In this embodiment, the decomposition process begins from the bottom layer represented by the first layer. Downsampling the original pixel information of the image blocks in the next layer adjacent to the current layer means performing a weighted average of the original pixel information of the adjacent image blocks in the next layer to obtain the pixel information of each image block in the current layer. When calculating the difference information, the difference between the upsampled pixel information and the original pixel information of each image block can be calculated. The current layer can be the (N+1)th layer. The pixel information of the Nth layer in rows r and columns c is first downsampled to obtain a data array of rows r and columns c, which is the downsampled pixel information of the (N+1)th layer. The downsampled pixel information of the (N+1)th layer is then upsampled to obtain upsampled pixel information in rows r and columns c, and the size of the upsampled pixel information is the same as the size of the original pixel information of the Nth layer. The difference between these two pixel information of the same size is calculated pixel by pixel, and the difference array is stored. When performing downsampling, the weights of the sampling filter can be Gaussian distributed weights or other low-pass filtering weight allocation methods. Furthermore, the weights of the sampling filter during upsampling can be the same as the weights of the filter during downsampling at the same information unit, which will not be elaborated further here.
[0057] For example, if the current layer is the second layer, then the next layer adjacent to the current layer is the first layer. The original pixel information of adjacent image blocks in the first layer can be downsampled to obtain the pixel information of image blocks in the second layer. The pixel information of the image blocks in the second layer is then upsampled to obtain the upsampled pixel information of each image block in the first layer. The difference between the upsampled pixel information and the original pixel information of each image block in the first layer is calculated as the difference information, and this difference information is stored in the corresponding location in the first layer.
[0058] Repeat steps S510 to S530 until the decomposition condition is no longer met, i.e., until the pixel information and difference information of the last layer are obtained, or the decomposition is carried out to the single pixel of the top layer, or the decomposition is carried out to the number of layers configured by the system, so as to obtain the pixel information of each image layer and the difference information between the layer and the next layer.
[0059] Continue to refer to Figure 2 As shown, in step S230, tone mapping is performed on the pixel information of each image layer to obtain the mapped pixel information of the target region.
[0060] In this embodiment of the disclosure, tone mapping refers to compressing the dynamic range to below the dynamic range of the output device, enabling high dynamic range (HDR) images to adapt to low dynamic range (LDR) displays. Tone mapping can be implemented using a mapping function, which can be described by a mapping curve, specifically representing the mapping relationship between input pixel information and output pixel information. The radian of the mapping curve for each image block may be different, and the radian of the mapping curve for each image block can be configured and set according to actual needs, without specific limitations here.
[0061] Furthermore, based on the mapping curve, the local gain information of each image block in each image layer can be determined. Specifically, the local gain information of each image block in each image layer can be determined according to the ratio of the intermediate pixel information output after mapping to the original pixel information.
[0062] For example, for the top level (e.g.) Figure 4 The pixel value 1 (pix1) of the (N+3)th layer is tone-mapped. This tone mapping is calculated using the data of this layer, and the mapped pixel value 2 (pix2) is obtained by mapping the data of this layer. The local gain distribution of this top layer is obtained by dividing pixel2 by pixel1. For other image layers, since there are the same or different mapping curves in each image block, the mapped intermediate pixel information corresponding to the original pixel information of each image block can be obtained according to the mapping curve. Furthermore, the local gain information of each image block in each image layer is determined according to the ratio of the intermediate pixel information to the original pixel information.
[0063] To perform localized processing on specific regions of an image, a mask can be provided to select a target region for targeted processing. (Refer to...) Figure 6 As shown, a target region 602 can be selected from the image to be processed 600 using mask information 601, and other regions besides the target region can be designated as reference regions 603. The target region can be a region that requires special attention, such as a face region or a region with many details. Here, the face region is used as an example for explanation. Since the image to be processed is layered according to the pyramid principle, the target region is also layered at the same time, and as the image to be processed is scaled from the bottom layer to the top layer, its target region is also scaled by the same proportion.
[0064] After obtaining the local gain information for each image patch in each image layer, for each image layer, the second gain information lev_gain_2 of the target region can be determined based on the local gain information of the image patches contained within the target region. Since the target region can contain complete image patches as well as partial image patches, the curvature of the mapping curve for each image patch is different, specifically determined according to the proportion of the image patch. For example, the curvature of the mapping curve can be positively correlated with the proportion of image patches contained in the target region; that is, the curvature of the mapping curve for complete image patches is the largest, and the curvature decreases as the proportion of the image patch decreases. Here, the curvature of the mapping curve can be used to represent the mapping intensity.
[0065] In addition, regions in the image to be processed other than the target region can be defined as reference regions, and the gain information of the reference regions can be defined as the first gain information lev_gain_1. The first gain information can be the local gain information of each image patch in the reference region. The first gain information can be the result of the local tone mapping and the multi-window color sensor being correlated. That is, the gain information is obtained by interpolating the gain information of the local tone mapping with the color information obtained by the multi-window color sensor according to the weight parameters.
[0066] It should be noted that, to distinguish between the target region and the reference region, the first gain information can be greater than or less than the second gain information, as long as they differ; no specific limitation is made here. Furthermore, the gain information of the target region can be adjusted according to the adjustment parameters to output the second gain information, while the first gain information can remain unchanged. The adjustment parameters are determined based on actual needs, such as user requirements or system settings. The adjustment parameters can include the image block to be adjusted and the adjustment range. During adjustment, all or part of the image block can be adjusted according to the required adjustment range, thereby adjusting the gain information of the target region. This allows for precise and independent control of the gain information of the target region, improving the flexibility of tone mapping processing.
[0067] To avoid differences in gain information between different regions, the second gain information within the target region and the first gain information in the reference region outside the target region can be smoothly transitioned. This smooth transition can be radial or other methods; radial transition will be used as an example here.
[0068] The radial transition is defined by its center. During the radial transition, the center, shape, and size of the gradient can also be specified. The shape can be circular or elliptical, and the size of the gradient can represent the furthest corner. In this embodiment, the radial origin is taken as the center region of the first gain information of the reference region outside the target region, and a radial transition is performed from the second gain information of the target region to the first gain information of the reference region. Figure 7 As shown in the figure. The transition between the two can be a series of weight transition methods such as a smooth curve weight radial distribution or a linear distribution, as referenced. Figure 8 As shown, no specific limitations are made here. A smooth curve weight radial distribution means the weights of the transition method can be distributed along a smooth curve, while a linear distribution means the weights of the transition method are distributed linearly. When it is a linear distribution, the first gain information can be used when the radial distance is less than the first threshold Th1; when the radial distance is greater than the first threshold Th1 but less than the first threshold Th2, the first gain information and the second gain information are weighted and fused; and when the radial distance is greater than the second threshold Th2, the second gain information is used.
[0069] After determining the second gain information of the target region and the first gain information of the reference region, tone mapping processing can be performed on the image to be processed based on the gain information of each location region. Specifically, the original pixel information of the associated image blocks can be mapped according to the second gain information of the target region to obtain the mapped pixel information of the target region. Specifically, the second gain information of the target region can be multiplied with the original pixel information of each image block in each image layer of the target region, specifically by multiplying the second gain information with the value of the color sub-pixel of the pixel; the first gain information of the reference region can be multiplied with the original pixel information of each image block in each image layer of the reference region, specifically by multiplying the first gain information with the value of the color sub-pixel of each pixel, to obtain the mapped pixel information of the pixels of each image block in the image to be processed.
[0070] Continue to refer to Figure 2 As shown, in step S240, the target pixel information of each image layer is obtained by combining the mapped pixel information and the difference information of each image layer to generate the target image.
[0071] In this embodiment of the disclosure, after obtaining the mapped pixel information of each image layer, reconstruction can be performed starting from the top layer to obtain the complete pixel information of each image layer, i.e., the target pixel information. Specifically, this may include the following steps: fusing the mapped pixel information and difference information of the previous layer adjacent to the current layer to obtain the target pixel information of the current layer; upsampling the target pixel information of the current layer to obtain the target pixel information of the next layer adjacent to the current layer, until the target pixel information of all image layers is determined, thereby determining the target image corresponding to the image to be processed. For example, the mapped pixel information of the previous layer adjacent to the current layer can be upsampled to obtain the upsampled pixel information of the current layer, and the upsampled pixel information and the difference information between the pixel information of the current layer and the previous layer can be fused to obtain the target pixel information of the current layer. Further, the target pixel information of the current layer can be used as input, the mapped pixel information of the next layer can be determined according to the mapping curve of the next layer, and the mapped pixel information can be upsampled and added to the difference information between the current layer and the next layer to obtain the target pixel information of the next layer. This process is repeated until the target pixel information of all image layers is obtained, that is, until the target pixel information of the bottom layer is obtained. The target pixel information of the original image represented by the first layer can be used as the final pixel information of the entire image to be processed, thereby generating the target image based on the target pixel information after tone mapping of the first layer.
[0072] For example, starting from the top level, the top level (e.g.) Figure 4 The mapped pixel information obtained from the (N+3)th layer is upsampled and interpolated to obtain the upsampled pixel information of the next layer (N+2) at that size. Then, the difference information between the next layer (N+2) and the top layer (N+3) is added to obtain the complete pixel information of the current layer. This process is repeated, upsampling and fusing the current layer (N+2) to obtain the target pixel information of the (N+1)th layer, until the target pixel information of the original data layer (layer 1) is obtained. The result of the tone mapping processing after multi-scale processing is used as the target image. The target image here can be an image obtained by specifically processing the image content of the target region.
[0073] In this embodiment of the disclosure, the target region is obtained through mask information, and then local tone mapping can be performed on the image content of the local region represented by the target region in the image to be processed. This avoids the limitation of only being able to process the whole in related technologies, improves the flexibility and targeting of image processing, and also improves image quality.
[0074] In this embodiment, by introducing a color sensor array, local color information of each image block can be obtained based on the multi-window information of the color sensor array. Furthermore, the gain in the global white balance can be further adjusted based on the local color information of each image block to obtain the target white balance gain value for each image block. This enables local adaptive adjustment, thereby obtaining the target white balance gain for local areas and improving accuracy. The segmentation module can perform more targeted color gain correction on the image content. Combining the information from the color sensor array allows for more accurate and optimized color processing of the image to be processed across multiple dimensions, including global, local, and image content dimensions.
[0075] Figure 9 A schematic diagram of the image signal processor is shown for reference. Figure 9 As shown, the image signal processor may include an image white balance algorithm module 900, and may also be divided into front-end processing 901 and back-end processing 902. The front-end processing includes a white balance module 903, and the back-end processing includes modules such as tone mapping 904 and color 905. The image signal processor may also include a segmentation module 906. In addition, it may also include a color sensor array 907 and a sensor 908.
[0076] refer to Figure 9 As shown, the mesh information acquired by the color sensor array is associated with the mesh acquired in the tone mapping module. The sensor is associated with the segmentation module, sending the acquired information to the segmentation module for processing to obtain the required mask information. The color sensor array is associated with the mask information output by the segmentation module to output the color information within the target area.
[0077] Based on the aforementioned hardware structure, the color sensor array can acquire local color information of each image block. It then obtains gain information for each image block through weighted fusion of the local color information and local tone mapping. The color sensor array acquires mask information via the segmentation module, thereby obtaining color information of the target region corresponding to the mask information. The image to be processed is then decomposed using a pyramid decomposition method, and the second gain information of the target region in each image layer is calculated. Finally, tone mapping is performed on the image to be processed by the backend tone mapping module. In this embodiment, the introduction of a color sensor array enhances the capabilities of the tone mapping module in the backend processing.
[0078] In this embodiment, by introducing a color sensor array, local color information of each image block can be obtained based on the multi-window information of the color sensor array. The target region is determined based on the mask information obtained by the segmentation module, and the target pixel information within the target region is determined, thereby determining the target pixel information of each image layer to obtain the target image. This enables local adaptive tone mapping of the target region, improving the accuracy of local tone mapping. Furthermore, it allows for the decomposition of the image to be processed to obtain the target pixel information of the target region in each image layer, thus determining the target image, increasing the frequency dimension of the image tone mapping process, and improving accuracy and comprehensiveness.
[0079] It's worth noting that during image processing, for real-time processing methods such as video or preview, temporal-space smoothing can be performed. At each time step t, an image can be fed into the image signal processing system for processing to obtain the images corresponding to Frame t-2, Frame t-1, Frame t, Frame t+1, and Frame t+2. During this process, temporal smoothing filtering can be performed on each image in the image sequence direction. Temporal smoothing filtering of each image can be understood as smoothing filtering according to the image's temporal sequence direction. Specifically, smoothing filtering can employ IIR filtering in the temporal domain, where the output result can be calculated as I = A*w + B*(1-w). Here, I represents the current frame, i.e., the output result of Frame t; A is the data or parameters of the current frame; w is the weight of the current frame; B is the data or parameters of the previous Frame-1 adjacent to the current frame on the time axis; and 1-w is the weight of Frame-1. This method allows for temporal (image sequence) smoothing to reduce the differences between different times and avoid abrupt changes. In this embodiment, various temporal smoothing techniques can be employed, such as smoothing the processed single-scale output or smoothing each layer of data across multiple scales; no specific limitations are specified here.
[0080] In summary, the technical solution in this disclosure, by determining the local color information of each image block based on the color sensor array, thereby determining the gain information of each image block, and further achieving individual control of local image blocks through the target region, avoids the limitation of only being able to process as a whole in related technologies, improving the accuracy and independence of image processing. Furthermore, it can combine multiple image layers to achieve individual control of the target region at multiple scales, eliminating abrupt changes in effect caused by processing differences between different image blocks within a spatial region compared to a single scale, thus improving image quality. In addition, by smoothly transitioning the gain information between different regions, abrupt regions are avoided in the entire image, improving smoothness and image quality. After introducing the segmentation module, the target region can be obtained in the image to be processed through mask information, thereby enabling local tone mapping processing of the image content of the local region represented by the target region in the image to be processed, avoiding the limitation of only being able to process as a whole in related technologies, improving the flexibility and targeting of image processing, and also improving image quality. It can effectively improve the accuracy of color reproduction in the image and improve the image quality of the target image.
[0081] This disclosure provides an image processing apparatus, with reference to... Figure 10 As shown, the image processing apparatus 1000 may include:
[0082] The image acquisition module 1001 is used to acquire the image to be processed and, in conjunction with the color sensor array, divide the image to be processed into multiple image blocks;
[0083] The image decomposition module 1002 is used to perform multi-scale decomposition on the image to be processed to obtain multiple image layers, and to obtain pixel information and difference information of each image block in each image layer.
[0084] The mapping module 1003 is used to perform tone mapping on the pixel information of each image layer to obtain the mapped pixel information of the target region;
[0085] The image generation module 1004 is used to combine the mapped pixel information and difference information of each image layer to obtain the target pixel information of each image layer in order to generate the target image.
[0086] In one exemplary embodiment of this disclosure, the image decomposition module includes an aggregation module, configured to aggregate adjacent image blocks of the next layer adjacent to the current layer in the multi-layer image layers to obtain the image blocks of the current layer, until the current layer no longer meets the decomposition conditions, so as to decompose the image to be processed into multi-layer image layers.
[0087] In one exemplary embodiment of this disclosure, the aggregation module includes a weighted summation module, configured to perform a weighted summation on all adjacent image blocks in the next layer adjacent to the current layer to obtain the image blocks of the current layer; the number of image blocks in the current layer is less than the number of image blocks in the adjacent next layer.
[0088] In an exemplary embodiment of this disclosure, the image decomposition module includes: a downsampling module, configured to downsample the original pixel information of image blocks in the next layer adjacent to the current layer, to obtain downsampled pixel information of the current layer as pixel information of each image block; an upsampling module, configured to upsample the downsampled pixel information of the current layer, to obtain upsampled pixel information of each image block in the next layer adjacent to the current layer; and a difference information acquisition module, configured to acquire difference information between the upsampled pixel information and the original pixel information of the next layer adjacent to the current layer, until the decomposition condition is no longer met, to acquire difference information of each image layer.
[0089] In one exemplary embodiment of this disclosure, the difference information acquisition module includes a difference calculation module, which is used to perform pixel-by-pixel difference calculation between the upsampled pixel information and the original pixel information to obtain difference information.
[0090] In one exemplary embodiment of this disclosure, the mapping module includes: a gain determination module, configured to acquire first gain information of a reference region other than the target region, and acquire second gain information of the target region corresponding to the mask information of each image layer; and a local mapping module, configured to map the original pixel information of the target region in each image layer according to the second gain information to obtain the mapped pixel information.
[0091] In one exemplary embodiment of this disclosure, the gain determination module includes: a local gain determination module, configured to perform tone mapping on the original pixel information of each image layer, and determine the local gain information of each image block in each image layer based on the ratio of the intermediate pixel information obtained by tone mapping to the original pixel information; and a target region gain determination module, configured to determine the second gain information of the target region based on the local gain information of the image blocks contained in the target region.
[0092] In one exemplary embodiment of this disclosure, the image generation module includes: a fusion module, configured to fuse the mapped pixel information and difference information of the previous layer adjacent to the current layer to obtain the target pixel information of the current layer; and an upsampling module, configured to upsample the target pixel information of the current layer to obtain the target pixel information of the next layer adjacent to the current layer, until the target pixel information of multiple image layers is determined.
[0093] In one exemplary embodiment of this disclosure, the fusion module includes: an upsampling module, configured to upsample the mapped pixel information of the previous layer adjacent to the current layer to obtain the upsampled pixel information of the current layer; and a difference fusion module, configured to fuse the upsampled pixel information and the difference information between the pixel information of the current layer and the previous layer to obtain the target pixel information of the current layer.
[0094] It should be noted that the specific details of each part of the above-mentioned image processing apparatus have been described in detail in the implementation of the image processing method section. For any undisclosed details, please refer to the implementation of the method section, and therefore will not be repeated here.
[0095] An exemplary embodiment of this disclosure also provides an electronic device. This electronic device may be the terminal 101 described above. Generally, the electronic device may include a processor and a memory, the memory being used to store executable instructions of the processor, the processor being configured to perform the image processing method described above by executing the executable instructions.
[0096] The following is based on Figure 11 Taking the mobile terminal 1100 as an example, the construction of this electronic device will be described by way of example. Those skilled in the art will understand that, apart from components specifically designed for mobile purposes, Figure 11 The structure can also be applied to fixed types of equipment.
[0097] like Figure 11 As shown, the mobile terminal 1100 may specifically include: a processor 1101, a memory 1102, a bus 1103, a mobile communication module 1104, an antenna 1, a wireless communication module 1105, an antenna 2, a display screen 1106, a camera module 1107, an audio module 1108, a power module 1109, and a sensor module 1110.
[0098] Processor 1101 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, an encoder, a decoder, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). The image denoising method in this exemplary embodiment can be executed by an AP, GPU, or DSP. When the method involves neural network-related processing, it can be executed by an NPU. For example, the NPU can load neural network parameters and execute neural network-related algorithm instructions.
[0099] An encoder encodes (compresses) images or videos to reduce data size for easier storage or transmission. A decoder decodes (decompresses) the encoded data to restore the original image or video data. The mobile terminal 1100 can support one or more encoders and decoders, such as image formats like JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), and BMP (Bitmap), and video formats like MPEG (Moving Picture Experts Group) 1, MPEG10, H.1063, H.1064, and HEVC (High Efficiency Video Coding).
[0100] The processor 1101 can be connected to the memory 1102 or other components via the bus 1103.
[0101] The memory 1102 can be used to store computer executable program code, which includes instructions. The processor 1101 executes various functional applications and data processing of the mobile terminal 1100 by running the instructions stored in the memory 1102. The memory 1102 can also store application data, such as images, videos, and other files.
[0102] The communication function of mobile terminal 1100 can be implemented through mobile communication module 1104, antenna 1, wireless communication module 1105, antenna 2, modem processor, and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 1104 can provide 3G, 4G, 5G and other mobile communication solutions for mobile terminal 1100. Wireless communication module 1105 can provide wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication for mobile terminal 1100.
[0103] The display screen 1106 is used to implement display functions, such as displaying user interfaces, images, and videos. The camera module 1107 is used to implement shooting functions, such as capturing images and videos, and may include a color sensor array. The audio module 1108 is used to implement audio functions, such as playing audio and capturing voice. The power module 1109 is used to implement power management functions, such as charging the battery, supplying power to the device, and monitoring battery status. The sensor module 1110 may include one or more sensors to implement corresponding sensing and detection functions. For example, the sensor module 1110 may include an inertial sensor, which is used to detect the motion posture of the mobile terminal 1100 and output inertial sensing data.
[0104] It should be noted that the present disclosure also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist alone and not assembled into the electronic device.
[0105] Computer-readable storage media can be, for example—but not limited to—electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0106] A computer-readable storage medium can be sent, propagated, or transmitted for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.
[0107] A computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to perform the methods described in the following embodiments.
[0108] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0109] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0110] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0111] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims. It should be understood that this disclosure is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An image processing method, characterized in that, include: The image to be processed is acquired and divided into multiple image blocks using a color sensor array; The image to be processed is decomposed into multiple image layers at multiple scales, and the pixel information and difference information of each image block in each image layer are obtained. Obtain the second gain information of the target region corresponding to the mask information of each image layer, and map the original pixel information of the target region in each image layer according to the second gain information to obtain the mapped pixel information of the target region; The target pixel information of the current layer is obtained by fusing the mapped pixel information and difference information of the previous layer adjacent to the current layer. The target pixel information of the current layer is used as input to determine the mapped pixel information of the next layer according to the mapping curve of the next layer. The mapped pixel information is upsampled and the difference information between the current layer and the next layer is added. This process is repeated to generate the target image. The step of obtaining the second gain information of the target region corresponding to the mask information of each image layer includes: Tone mapping is performed on the original pixel information of each image layer according to the mapping curve, and the local gain information of each image block in each image layer is determined according to the ratio of the intermediate pixel information obtained by tone mapping to the original pixel information. Based on the local gain information of the image blocks contained in the target region, the second gain information of the target region is determined.
2. The image processing method according to claim 1, characterized in that, The step of performing multi-scale decomposition on the image to be processed to obtain multiple image layers includes: The adjacent image blocks of the next layer adjacent to the current layer in the multi-layer image layers are aggregated to obtain the image block of the current layer, until the current layer no longer meets the decomposition conditions, so as to decompose the image to be processed into multi-layer image layers.
3. The image processing method according to claim 2, characterized in that, The step of aggregating adjacent image blocks from the next layer adjacent to the current layer to obtain the image block of the current layer includes: The image blocks of the current layer are obtained by weighted summation of all adjacent image blocks in the next layer; the number of image blocks in the current layer is less than the number of image blocks in the adjacent next layer.
4. The image processing method according to claim 1, characterized in that, The acquisition of pixel information and difference information of each image block in each image layer includes: The original pixel information of the image block in the next layer adjacent to the current layer is downsampled to obtain the downsampled pixel information of the current layer as the pixel information of each image block; Upsample the downsampled pixel information of the current layer to obtain the upsampled pixel information of each image block in the next layer adjacent to the current layer; The difference information between the upsampled pixel information and the original pixel information of the next adjacent layer of the current layer is obtained until the decomposition condition is no longer met, so as to obtain the difference information of each image layer.
5. The image processing method according to claim 4, characterized in that, The step of obtaining the difference information between the upsampled pixel information and the original pixel information of the next adjacent layer of the current layer includes: The difference information is obtained by calculating the difference between the upsampled pixel information and the original pixel information pixel by pixel.
6. The image processing method according to claim 1, characterized in that, The step of performing tone mapping on the pixel information of each image layer to obtain the mapped pixel information of the target region includes: Obtain the first gain information for the reference region other than the target region.
7. The image processing method according to claim 1, characterized in that, The step of fusing the mapped pixel information and difference information of the previous layer adjacent to the current layer to obtain the target pixel information of the current layer includes: The mapped pixel information of the previous layer adjacent to the current layer is upsampled to obtain the upsampled pixel information of the current layer; The upsampled pixel information and the difference information between the pixel information of the current layer and the previous layer are fused to obtain the target pixel information of the current layer.
8. An image processing apparatus, characterized in that, include: The image acquisition module is used to acquire the image to be processed and, in conjunction with the color sensor array, divide the image to be processed into multiple image blocks; The image decomposition module is used to decompose the image to be processed into multiple image layers at multiple scales, and to obtain the pixel information and difference information of each image block in each image layer. The mapping module is used to obtain the second gain information of the target region corresponding to the mask information of each image layer, and to map the original pixel information of the target region in each image layer according to the second gain information to obtain the mapped pixel information of the target region. The image generation module is used to fuse the mapped pixel information and difference information of the previous layer adjacent to the current layer to obtain the target pixel information of the current layer. The target pixel information of the current layer is used as input to determine the mapped pixel information of the next layer according to the mapping curve of the next layer. The mapped pixel information is upsampled and the difference information between the current layer and the next layer is added. The process is repeated to generate the target image. The step of obtaining the second gain information of the target region corresponding to the mask information of each image layer includes: Tone mapping is performed on the original pixel information of each image layer according to the mapping curve, and the local gain information of each image block in each image layer is determined according to the ratio of the intermediate pixel information obtained by tone mapping to the original pixel information. Based on the local gain information of the image blocks contained in the target region, the second gain information of the target region is determined.
9. An electronic device, characterized in that, include: Imaging module, including color sensor array; processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the image processing method according to any one of claims 1-7 by executing the executable instructions.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the image processing method according to any one of claims 1-7.