Image processing method, electronic equipment and computer readable storage medium
By acquiring and enhancing global and local features of an image, and combining semantic segmentation, color extraction, and frequency processing, the color and lighting of local areas of the image are optimized, solving the problems of low local smoothness and color cast, and achieving an improvement in the overall and local enhancement effects of the image.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-20
- Publication Date
- 2026-03-27
AI Technical Summary
During image enhancement, low smoothness or color cast in local areas can lead to poor image enhancement results.
By acquiring the global features and target local features of the first image, enhancement processing is performed to optimize the color and lighting of the overall and local regions of the image. Semantic segmentation, color extraction, frequency extraction, and feature extraction modules are used, combined with color compensation and enhancement modules, to improve the image quality of local regions.
It enhances the overall image and local areas, making the overall image colors more vivid, the local details richer, the edges smoother, and the visual effect stronger.
Smart Images

Figure CN121746226A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and particularly relates to an image processing method, an electronic device and a computer readable storage medium. BACKGROUND
[0002] When an electronic device takes a photo, an enhanced image can be obtained by using an image enhancement function. The enhanced image has more distinct colors, higher contrast and higher definition, and thus has a stronger visual effect. However, in some scenarios, the smoothness of a local region of the image is low or color cast occurs, resulting in poor image enhancement effect. SUMMARY
[0003] In view of this, the embodiments of the present application provide an image processing method, an electronic device and a computer readable storage medium, aiming to solve the problem of how to improve the image enhancement effect.
[0004] The first aspect of the embodiments of the present application provides an image processing method, which comprises: obtaining a first image, the first image containing a first region; and performing enhancement processing on the first image to obtain a second image, the second image including a second region. The position of the first region in the first image corresponds to the position of the second region in the second image, and the image quality of the second region is better than that of the first region.
[0005] In the embodiments, the second image contains global features and target local features of the first image. The global features contain color information and light and shadow information of the first image as a whole, and the target local features contain color information and frequency band information of the first region. The global features are enhanced to optimize the color and light and shadow of the first image as a whole, and the target local features are enhanced to optimize the image quality of the first region, thereby improving the enhancement effect of the image as a whole and the local region.
[0006] In one embodiment, the image quality of the second region being better than that of the first region comprises: the contrast of the first region being less than that of the second region, and / or the definition of the first region being less than that of the second region, and / or the saturation of the first region being less than that of the second region, and / or the smoothness of the first region being less than that of the second region, and / or the brightness of the first region being less than that of the second region.
[0007] In the embodiments, the image quality of the local region comprises at least one of contrast, definition, saturation, smoothness and brightness. Correspondingly, the enhancement processing comprises at least one of enhancing local contrast, enhancing local definition, enhancing local saturation, enhancing local smoothness and enhancing local brightness.
[0008] In another embodiment, the enhancing the first image to obtain the second image comprises: performing semantic segmentation on the first image to obtain m segmentation mask images, m being a positive integer. Main color data of the first image is extracted, and a main color mask image is obtained, the main color data containing color data of the first region. A target frequency band mask image is obtained according to the first image. A target contrast mask image is obtained according to the target frequency band mask image and the main color mask image. m feature images are obtained according to the first image, the m segmentation mask images, the main color data and the target contrast mask image. The second image is obtained according to the first image and the m feature images.
[0009] In the embodiment, the second region contains color information and frequency band information of the first region, the direction of color enhancement is guided by the main color, and the color enhancement region is determined, and the color enhancement region is subjected to frequency band constraint by the target frequency band. The image quality of the first region is optimized by enhancing the color of the first region under the frequency band constraint, thereby improving the enhancement effect of the local region of the image. The color enhancement of the local region of the image makes the color of the whole image more vivid, thereby improving the enhancement effect of the whole image.
[0010] In another embodiment, the second image is obtained according to the first image and the m feature images, comprising: performing color space conversion on the m feature images to obtain m feature compensation images. m target feature images are obtained according to the main color data and the m feature compensation images. The second image is obtained according to the first image and the m target feature images.
[0011] In the embodiment, the feature compensation image contains more rich color information, and therefore the target feature image contains more comprehensive color enhancement region.
[0012] In another embodiment, the second image is obtained according to the first image and the m target feature images, comprising: performing enhancement processing on the m target feature images to obtain m local enhancement images. The first image and the m local enhancement images are merged, or the first image, the main color matrix and the m local enhancement images are merged to obtain the second image. The main color matrix is generated from the main color data.
[0013] In another embodiment, the m target feature images are obtained according to the main color data and the m feature compensation images, comprising: generating a main color matrix according to the main color data. The i-th target feature image is obtained by merging the i-th feature compensation image and the main color matrix, until the iteration is completed. i is a positive integer, and i≤m.
[0014] In another embodiment, the color space conversion of the m feature images to obtain m feature compensation images comprises: converting the m feature images from RGB color space to HSV color space or LAB color space respectively to obtain the m feature compensation images.
[0015] In another embodiment, the obtaining of the second image according to the first image and the m feature images comprises: performing enhancement processing on the m feature images to obtain m local enhancement images. The first image and the m local enhancement images are merged, or the first image, the main color matrix and the m local enhancement images are merged to obtain the second image. The main color matrix is generated according to the main color data.
[0016] In another embodiment, the obtaining of the m feature images according to the first image, the m segmentation mask images, the main color data and the target contrast mask image comprises: generating a main color matrix according to the main color data. The i-th segmentation mask image, the first image, the main color matrix and the target contrast mask image are merged to obtain the i-th feature mask image, until the traversal is completed. i is a positive integer, and i≤m.
[0017] In another embodiment, the obtaining of the target contrast mask image according to the target frequency band mask image and the main color mask image comprises: multiplying the target frequency band mask image and the main color mask image to obtain the target contrast mask image.
[0018] In another embodiment, the extracting of the main color data of the first image and the obtaining of the main color mask image comprises: extracting the main color data of the first image. A target color range is set according to the main color data. The main color mask image is generated according to the target color range.
[0019] In another embodiment, the obtaining of the target frequency band mask image according to the first image comprises: performing frequency domain transformation on the first image to obtain a frequency domain image. A target frequency band mask template is set. The target frequency band mask template is applied to the frequency domain image to obtain a frequency domain mask image. The target frequency band mask image is obtained by performing inverse frequency domain transformation on the frequency domain mask image.
[0020] In another embodiment, the target frequency band mask image comprises a high frequency mask image and / or a low frequency mask image, the high frequency mask image containing high frequency information of the first image, and the low frequency mask image containing low frequency information of the first image.
[0021] In the embodiment, the color enhancement region is made to contain high frequency information by applying a high frequency constraint to the color enhancement region, so that the edges and textures of the objects in the color enhancement region are enhanced, and the definition of the color enhancement region is improved. The color enhancement region is made to contain low frequency information by applying a low frequency constraint to the color enhancement region, so that the content inside the edges of the objects in the color enhancement region is enhanced, and the contrast, saturation, smoothness and brightness of the color enhancement region are improved.
[0022] The second aspect of the embodiment of the present application provides an electronic device, which comprises a memory and a processor, and the processor implements the image processing method provided in the first aspect when executing computer instructions stored in the memory.
[0023] The third aspect of the embodiment of the present application provides a computer readable storage medium, which stores computer instructions, and the processor implements the image processing method provided in the first aspect when executing the computer instructions.
[0024] The fourth aspect of the embodiment of the present application provides a computer program product, which comprises computer instructions, and the processor implements the image processing method provided in the first aspect when executing the computer instructions.
[0025] It can be understood that the electronic device provided in the second aspect, the computer readable storage medium provided in the third aspect and the computer program product provided in the fourth aspect have approximately the same beneficial effects as the image processing method provided in the first aspect, and details are not repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 FIG. 1 is a schematic diagram of a software structure of an electronic device provided in an example.
[0027] Figure 2 FIG. 2 is a schematic diagram of turning on an image enhancement function in a photographing scene provided in an example.
[0028] Figure 3 FIG. 3 is a schematic diagram of turning on an image enhancement function in a photographing scene provided in another example.
[0029] Figure 4 FIG. 4 is a schematic diagram of turning on an image enhancement function in a proactive triggering scene provided in an example.
[0030] Figure 5 FIG. 5 is a schematic diagram of turning on an image enhancement function in a proactive triggering scene provided in another example.
[0031] Figure 6 FIG. 6 is a schematic diagram of a logic structure of an image enhancement algorithm provided in an example.
[0032] Figure 7 FIG. 7 is a schematic diagram of a user interface of image enhancement in a photographing scene provided in an example.
[0033] Figure 8 is a user interface diagram of image enhancement in an active trigger scenario provided by an example.
[0034] Figure 9 is a timing diagram of an implementation process of an image enhancement function in a photographing scenario provided by an example.
[0035] Figure 10 is a timing diagram of an implementation process of an image enhancement function in an active trigger scenario provided by an example.
[0036] Figure 11 is a flowchart of an implementation of an image enhancement algorithm provided by an example.
[0037] Figure 12 is a hardware structure diagram of an electronic device provided by an example. DETAILED DESCRIPTION
[0038] It should be noted that “at least one” in the embodiments of the present application means one or more, and “multiple” means two or more than two. The terms “first”, “second”, “third”, “fourth” and the like in the specification and claims of the present application and the drawings are used to distinguish similar objects, and are not used to describe a specific order or sequence. The method disclosed in the embodiments of the present application or the method shown in the flowchart includes one or more steps for implementing the method, and the execution order of the multiple steps can be interchanged with each other without departing from the scope of the claims, and some steps can also be deleted.
[0039] In this application embodiment, the electronic device has a shooting function, including, but not limited to, smartphones, tablets, handheld computers, laptops, cameras, camcorders, intelligent robots, drones, mobile internet devices (MID), virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, Session Initiation Protocol (SIP) phones, Wireless Local Loop (WLL) stations, and personal digital assistants (PDAs). Assistant (PDA), handheld device with wireless communication capabilities, computing device or other processing device connected to a wireless modem, in-vehicle device, wearable device, terminal device in a 5G network or terminal device in a Public Land Mobile Network (PLMN).
[0040] The software system of the electronic device is described in detail below.
[0041] The software system of electronic devices can adopt a layered architecture, which divides the software into several layers, each with a clear role and division of labor, and the layers communicate with each other through software interfaces. For example... Figure 1 As shown, the software system of an electronic device is divided into four layers, from top to bottom: Application (APP) layer, Framework (FWK) layer, Hardware Abstraction Layer (HAL) and Hardware layer.
[0042] The application layer comprises a series of application packages. These packages include a camera app and a gallery app. The camera app and / or gallery app feature image enhancement capabilities. These enhancements improve the color and contrast of images, making colors more vibrant and details more prominent, thereby improving the overall visual appeal of the image.
[0043] Users can enable image enhancement features through the camera app or gallery app. For example, in a photo-taking scenario, such as... Figure 2 As shown, when the user taps the camera app, the electronic device's screen displays the shooting interface. On the shooting interface, when the user taps the shooting control 11 to take a picture, the camera app activates the image enhancement function. For example, in a photo-taking scenario, such as... Figure 3 As shown, on the shooting interface, when the user presses volume button 12 on the electronic device to take a picture (volume button 12 includes volume up and volume down buttons), the camera application activates the image enhancement function. For example, in actively triggered scenarios, such as... Figure 4 As shown, on the shooting interface, when the user clicks the gallery control 13, image A and a toolbar are displayed on the screen. The toolbar includes an image enhancement control 14. When the user clicks the image enhancement control 14, the gallery application enables the image enhancement function. For example, in an actively triggered scenario, such as... Figure 5 As shown, when a user clicks the Gallery app, the album interface is displayed on the screen. On the album interface, when the user clicks image B, image B and a toolbar are displayed on the screen, including image enhancement control 14. When the user clicks image enhancement control 14, the Gallery app activates the image enhancement function.
[0044] The framework layer provides application programming interfaces (APIs) and programming frameworks for various apps in the application layer. The framework layer includes the camera framework (CameraFWK), which provides the camera API (CameraAPI) for camera applications.
[0045] The hardware abstraction layer is used to execute various application functions and data processing in apps. The hardware abstraction layer includes the camera hardware abstraction layer (CameraHAL), which comprises the processor driver, camera driver, and camera algorithm library.
[0046] A processor driver is a processor-oriented control node used to control the processor to implement image processing algorithms.
[0047] A camera driver is a control node for a camera, used to control the camera to capture images.
[0048] The camera algorithm library stores a series of image processing algorithms, such as image enhancement, image blurring, automatic exposure (AE), automatic focus (AF), electronic image stabilization (EIS), and image quality (PQ).
[0049] The hardware layer provides hardware support for the various applications and data processing of the hardware abstraction layer. The hardware layer includes processors and cameras. Processors include a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a Neural Processing Unit (NPU). The CPU handles control logic and serial computation tasks. The GPU handles image algorithm logic and parallel computation tasks. The NPU handles neural network computation tasks. Cameras can be wide-angle, ultra-wide-angle, or telephoto cameras. Each camera has a corresponding field of view (FOV) and focal length.
[0050] exist Figure 1 In the camera algorithm library shown, the image enhancement algorithm extracts the main color data of the original image and uses the main color to guide the artificial intelligence (AI) network to learn the global features and target local features of the original image. The global features reflect the overall color and lighting of the image, while the target local features reflect the color information and frequency band information of the local area of the image. The AI network can learn to enhance both the overall image and the local image, thereby improving the enhancement effect of the overall image and the local area.
[0051] In this embodiment, the AI network includes, but is not limited to, Convolutional Neural Network (CNN), Generative Adversarial Network (GAN), Fully Convolutional Network (FCN), Autoencoder, and Zero Reference Deep Curve Estimation (Zero-DCE) network.
[0052] like Figure 6 As shown, the image enhancement algorithm includes a semantic segmentation module, a color extraction module, a frequency extraction module, a feature extraction module, a color compensation module, and an enhancement module.
[0053] The semantic segmentation module performs semantic segmentation on the original image, resulting in m segmentation mask images, where m is a positive integer. Semantic segmentation classifies each pixel in the image, dividing it into multiple regions and assigning a category label to each region. The results of semantic segmentation are presented as mask images, with the number of mask images equal to the number of segmented regions. The categories for semantic segmentation can be set as needed.
[0054] In this embodiment, the size of the mask image is the same as the size of the original image. The mask image includes a foreground region and a background region, where the foreground region is the Region of Interest (ROI). The foreground and background regions are represented using different colors; for example, the foreground region is represented by white, and the background region is represented by black. The mask image is enhanced to different degrees for the foreground and background regions.
[0055] The semantic segmentation module may employ semantic segmentation networks, including, but not limited to, Pyramid Scene Parsing Network (PSPNet), Mask Region-based Convolutional Neural Network (Mask R-CNN), FCN, U-Net, and DeepLab.
[0056] The color extraction module extracts the primary color data from the original image, sets the target color range based on the primary color data, and generates a primary color mask image based on the target color range. The primary color refers to the most frequently occurring color in the image or the color that best represents the overall tone of the image. The primary color guides the direction of color enhancement and determines the color enhancement area. The primary color is presented as an array, which is the primary color data.
[0057] In this embodiment, clustering algorithms, such as the K-means algorithm, can be used to extract the main color data of the original image.
[0058] The target color range can be set as needed.
[0059] For example, in OpenCV, the relevant code is as follows:
[0060] colors=kmeans.cluster_centers_.astype(int)
[0061] main_color = colors[x]
[0062] tolerance = y
[0063] lower_bound = np.clip(main_color - tolerance, 0, 255)
[0064] upper_bound = np.clip(main_color + tolerance, 0, 255)
[0065] mask = cv2.inRange(image, lower_bound, upper_bound)
[0066] Among them, colors is the main color cluster extracted by the K-means algorithm. main_color is the main color data, which represents the x-th color data in the main color cluster, where x is a natural number and can be set as needed. tolerance is the color tolerance, y is a positive integer, and 0 < y < 255. lower_bound is the lower limit of the target color range, and upper_bound is the upper limit of the target color range. image is the original image, and mask is the main color mask image.
[0067] It can be understood that the main color data main_color can be a color data in the main color cluster colors, or a union of multiple color data in the main color cluster colors.
[0068] The frequency extraction module is connected to the color extraction module. The frequency extraction module is used to obtain the target frequency band mask image according to the original image, and obtain the target contrast mask image according to the target frequency band mask image and the main color mask image. The frequency extraction module施加s frequency band constraints on the color enhancement area through the target frequency band.
[0069] In this embodiment, obtaining the target frequency band mask image according to the original image includes: performing a frequency domain transformation on the original image to obtain a frequency domain image. Setting a target frequency band mask template. Applying the target frequency band mask template to the frequency domain image to obtain a frequency domain mask image. Performing an inverse frequency domain transformation on the frequency domain mask image to obtain the target frequency band mask image.
[0070] Among them, the frequency domain transformation is to convert the image from the spatial domain to the frequency domain, and the inverse frequency domain transformation is to convert the image from the frequency domain to the spatial domain. The frequency domain transformation may include, but is not limited to, Fourier Transform (FT), Fast Fourier Transform (FFT), Discrete Cosine Transform (DCT), and Wavelet Transform (WT).
[0071] It should be noted that there is an unclear expression "施加s" in the original text. I have translated it as "施加" as it seems to be a misspelling. If this is not what you intended, please correct the original text for a more accurate translation.The target frequency band mask template includes a high-frequency mask template and / or a low-frequency mask template, and the target frequency band mask image includes a high-frequency mask image and / or a low-frequency mask image. The foreground region of the high-frequency mask image is the area in the original image where grayscale values change drastically; it contains high-frequency information from the original image, such as object edges and textures. The foreground region of the low-frequency mask image is the area in the original image where grayscale values change slowly; it contains low-frequency information from the original image, such as content within object edges.
[0072] The target contrast mask image includes a high-contrast mask image and / or a low-contrast mask image. Multiplying the high-frequency mask image by the primary color mask image yields the high-contrast mask image. Multiplying the low-frequency mask image by the primary color mask image yields the low-contrast mask image. The foreground region of the high-contrast mask image is a color enhancement region containing high-frequency information. The foreground region of the low-contrast mask image is a color enhancement region containing low-frequency information. By applying high-frequency constraints to the color enhancement regions, making them contain high-frequency information, the edges and textures of objects within the color enhancement regions can be enhanced, thereby improving the sharpness of the color enhancement regions. By applying low-frequency constraints to the color enhancement regions, making them contain low-frequency information, the content within the edges of objects within the color enhancement regions can be enhanced, thereby improving the contrast, saturation, smoothness, and brightness of the color enhancement regions.
[0073] In this embodiment, image multiplication refers to the operation of multiplying the pixel values at corresponding positions in multiple images element by element.
[0074] The feature extraction module is connected to the semantic segmentation module, color extraction module, and frequency extraction module. The feature extraction module is used to obtain m feature images based on the original image, m segmentation mask images, main color data, and target contrast mask image. The m feature images contain global features of the original image and local features of the target. The global features contain the overall color information and lighting information of the image, while the local features of the target contain the color information and frequency band information of the color enhancement region.
[0075] In this embodiment, obtaining m feature maps based on the original image, m segmentation mask images, main color data, and target contrast mask image includes: generating a main color matrix based on the main color data; traversing the m segmentation mask images, merging the i-th segmentation mask image, the original image, the main color matrix, and the target contrast mask image to obtain the i-th feature image, until traversal is complete. i is a positive integer, and i ≤ m. That is, for each segmentation mask image, it is merged with the original image, the main color matrix, and the target contrast mask image to obtain its corresponding feature image.
[0076] Generating a primary color matrix from primary color data can involve using a CNN to convolve the primary color data to generate the primary color matrix. The size of the primary color matrix is the same as the size of the original image.
[0077] Merging the original image, segmentation mask image, main color matrix, and target contrast mask image may include: superimposing the segmentation mask image, main color matrix, and target contrast mask image onto the original image.
[0078] The color compensation module is connected to the color extraction module and the feature extraction module. The color compensation module is used to perform color space conversion on m feature images to obtain m feature-compensated images, and to obtain m target feature images based on the main color data and the m feature-compensated images.
[0079] In this embodiment, obtaining m target feature images based on the main color data and m feature compensation images includes: generating a main color matrix based on the main color data; traversing the m feature compensation images, merging the i-th feature compensation image with the main color matrix to obtain the i-th target feature image, until the traversal is complete. i is a positive integer, and i ≤ m. That is, for each feature compensation image, it is merged with the main color matrix to obtain its corresponding target feature image.
[0080] Merging the feature-compensated image and the primary color matrix may include overlaying the primary color matrix onto the feature-compensated image.
[0081] Color space conversion of a feature image may include converting the feature image from the RGB (Red, Green, Blue) color space to the HSV (Hue, Saturation, Value) color space or the LAB (Lightness, Green-Red Axis, Blue-Yellow Axis) color space.
[0082] For example, in OpenCV, the relevant code is as follows:
[0083] hsv_image=cv2.cvtColor(image,cv2.COLOR_BGR2HSV)
[0084] lab_image=cv2.cvtColor(image,cv2.COLOR_BGR2Lab)
[0085] Wherein, image is the feature image, hsv_image is the image obtained by converting the feature image from RGB color space to HSV color space, and lab_image is the image obtained by converting the feature image from RGB color space to LAB color space.
[0086] It's understandable that the RGB color space uses a linear combination of red, green, and blue color components to represent color. However, when image colors change, the magnitude of the change in each component is uneven, making it impossible to accurately estimate the change in each component. Therefore, the RGB color space is not suitable for direct application to image processing tasks. Images acquired in natural environments are easily affected by natural lighting, occlusion, and shadows. To eliminate these effects, separating luminance or lightness information is crucial, but the RGB color space does not have separate luminance or lightness components. Unlike the RGB color space, the HSV color space includes three channels: hue, saturation, and lightness. Separating hue and lightness information allows for targeted processing of low-light images, and the separated hue channel can guide AI networks to generate more realistic colors. The LAB color space has a wider color gamut, fully preserving the color information of an image, and its separated lightness channel makes it suitable for low-light image enhancement scenarios.
[0087] In this embodiment, since the feature compensation image contains richer color information, the target feature image contains a more comprehensive color enhancement area.
[0088] The enhancement module is connected to the color extraction module and the color compensation module. The enhancement module is used to enhance m target feature images to obtain m local enhanced images, and to obtain a global enhanced image based on the original image, main color data and m local enhanced images.
[0089] In this embodiment, the enhancement module can use an AI network to enhance m target feature images, generate a main color matrix based on the main color data, and then merge the original image, the main color matrix, and m local enhanced images to obtain a global enhanced image.
[0090] Merging the original image, the main color matrix, and m locally enhanced images can include superimposing the main color matrix and m locally enhanced images onto the original image.
[0091] Taking the Zero-DCE network as an example, the Zero-DCE network predicts the depth value corresponding to each pixel in each target feature image, estimates a depth map, generates a set of depth curves based on the depth map, and uses the depth curves to adjust the brightness and contrast of the target feature image. The depth curve of each target feature image is enhanced n times through n iterations, resulting in n sets of depth curves, where n is a positive integer. A total of n*m sets of depth curves are generated from m target feature images.
[0092] The loss function of the Zero-DCE network can include spatial consistency loss, exposure control loss, color preservation loss, and illumination smoothing loss. Spatial consistency loss maintains the spatial consistency between the input and output images. Exposure control loss measures the distance between the average intensity value of a local region and a good exposure level, which can be represented by an exposure reference image. Color preservation loss corrects potential color deviations and maintains the relationship between the R, G, and B color channels. Illumination smoothing loss maintains the monotonicity relationship between adjacent pixels.
[0093] For example, in OpenCV, the relevant code is as follows:
[0094] spatial_consistency_loss=torch.mean(abs_left_diff)+torch.mean(abs_right_diff)+
[0095] torch.mean(abs_top_diff)+torch.mean(abs_bottom_diff)
[0096] exposure_loss=torch.mean((enhanced-torch.mean(enhanced))**2)
[0097] color_constancy_loss=torch.mean((enhanced-original)**2)
[0098] brightness_smoothness_loss=torch.mean(torch.abs(enhanced[:,:,:,:-1]-enhanced[:,:,:,1:]))
[0099] Here, `spatial_consistency_loss` is the spatial consistency loss, calculated using the `torch.mean` function. `abs_left_diff`, `abs_right_diff`, `abs_top_diff`, and `abs_bottom_diff` are the absolute values of the deviations between adjacent pixels in the left, right, top, and bottom directions, respectively. `exposure_loss` is the exposure control loss, and `enhanced` enhances the image. `color_constancy_loss` is the color preservation loss, and `original` represents the original image. `brightness_smoothness_loss` is the illumination smoothing loss, calculated using the `torch.abs` function.
[0100] In other embodiments, the AI network can use a 3D LUT (Look Up Table) to enhance m target feature images. The 3D LUT uses a three-dimensional color space mapping table to transform the color values of the image, thereby achieving precise control and adjustment of the color.
[0101] The logical structure of the image enhancement algorithm has been explained in detail above. The following section will combine... Figure 1 The software system of the electronic device shown is described in detail, outlining the implementation process of the image enhancement function.
[0102] In one embodiment, such as Figure 7 As shown, the process of implementing image enhancement includes the following steps:
[0103] S101. When the shooting control is triggered, the camera application sends a shooting command to the CameraAPI.
[0104] Among them, the shooting controls are controls for taking pictures, for example... Figure 2 The shooting control 11 shown, or Figure 3 Volume button 12 is shown. The shooting command is used to instruct someone to take a picture.
[0105] S102, CameraAPI sends shooting commands to the camera driver.
[0106] S103, The camera driver responds to the shooting command and controls the camera to capture images.
[0107] S104: The processor driver acquires images from the camera driver, calls image enhancement algorithms from the camera algorithm library, and controls the CPU, GPU, and GPU to perform image enhancement processing to obtain an enhanced image.
[0108] S105, the processor driver sends enhanced images to the Camera API.
[0109] S106, Camera API sends enhanced images to camera applications.
[0110] S107, The camera app will enhance the image and store it in the gallery app.
[0111] In this embodiment, the camera application is launched to take a picture, and the image enhancement function is enabled by default. For example, in a photo-taking scenario, such as... Figure 8 As shown, when the user clicks the shooting control 11 to take a picture, the camera application activates the image enhancement function and saves the enhanced image to the gallery application. On the shooting interface, when the user clicks the gallery control 13, the screen displays the enhanced image. The enhanced image includes a person area 21, a ground area 22, a sky area 23, and a building area 24, each with a different primary color. Compared to the original image, the enhanced image has more vivid overall colors, richer details in local areas, smoother edges, and a stronger visual effect.
[0112] In another embodiment, such as Figure 9 As shown, the process of implementing image enhancement includes the following steps:
[0113] S201. When the image enhancement control is triggered, the gallery application sends an enhancement command to the Camera API.
[0114] Among them, the image enhancement control is a control that controls image enhancement functions, for example... Figure 4 The image enhancement control 14 shown, or Figure 5 The image enhancement control 14 is shown. Enhancement commands are used to instruct the target image to undergo enhancement processing.
[0115] S202, CameraAPI sends enhancement instructions to the processor driver.
[0116] S203: The processor driver responds to the enhancement instruction, calls the image enhancement algorithm from the camera algorithm library, and controls the CPU, GPU, and GPU to perform enhancement processing on the target image to obtain the enhanced image.
[0117] S204, The processor driver sends enhanced images to the Camera API.
[0118] S205, Camera API sends enhanced images to gallery applications.
[0119] S206, Gallery application stores enhanced images.
[0120] In this embodiment, the gallery application is launched, a target image is selected from the album, and the image enhancement function is actively activated. The target image can be any image stored in the gallery application, including images that have already been enhanced by the camera application. For example, in an actively triggered scenario, such as... Figure 10 As shown, when a user clicks on the target image, the screen displays the target image and a toolbar, which includes an image enhancement control 14. When the user clicks on the image enhancement control 14, the gallery application activates the image enhancement function, and the enhanced image is displayed on the screen. The enhanced image includes a people area 21, a ground area 22, a sky area 23, and a building area 24, each with a different primary color. Compared to the original image, the enhanced image has more vibrant overall colors, richer details in local areas, smoother edges, and a stronger visual effect.
[0121] The above provides a detailed explanation of the implementation process of image enhancement. The implementation of image enhancement relies on image enhancement algorithms; the following section describes the implementation process of these algorithms.
[0122] like Figure 11 As shown, the image enhancement algorithm includes the following steps:
[0123] S301. Perform semantic segmentation on the original image to obtain m segmentation mask images.
[0124] Semantic segmentation classifies each pixel in an image, dividing it into multiple regions and assigning a category label to each region. The number of categories in semantic segmentation can be set as needed. The number of segmentation mask images is equal to the number of segmented regions. m is a positive integer. The size of the segmentation mask images is the same as the size of the original image.
[0125] In this embodiment, a semantic segmentation network may be used, including, but not limited to, PSPNet, MaskR-CNN, FCN, U-Net, and DeepLab.
[0126] S302. Extract the main color data of the original image, set the target color range based on the main color data, and generate a main color mask image based on the target color range.
[0127] In this context, the primary color refers to the most frequently occurring color in the image or the color that best represents the overall tone of the image. The primary color guides the direction of color enhancement and identifies the areas to be enhanced. The primary colors are presented as an array, which is the primary color data.
[0128] In this embodiment, a clustering algorithm can be used to extract the main color data of the original image. The main color data can be a single color or a union of multiple color data. The target color range can be set as needed.
[0129] S303. Obtain the target frequency band mask image based on the original image, and obtain the target contrast mask image based on the target frequency band mask image and the main color mask image.
[0130] In this embodiment, obtaining the target frequency band mask image from the original image includes: performing a frequency domain transformation on the original image to obtain a frequency domain image; setting a target frequency band mask template; applying the target frequency band mask template to the frequency domain image to obtain a frequency domain mask image; and performing an inverse frequency domain transformation on the frequency domain mask image to obtain the target frequency band mask image. Here, the frequency domain transformation converts the image from the spatial domain to the frequency domain, and the inverse frequency domain transformation converts the image from the frequency domain back to the spatial domain.
[0131] The target frequency band mask template includes a high-frequency mask template and / or a low-frequency mask template, and the target frequency band mask image includes a high-frequency mask image and / or a low-frequency mask image. The high-frequency mask image contains the high-frequency information of the original image, and the low-frequency mask image contains the low-frequency information of the original image.
[0132] The target contrast mask image includes a high-contrast mask image and / or a low-contrast mask image. The high-contrast mask image is obtained by multiplying the high-contrast mask image with the primary color mask image. The low-contrast mask image is obtained by multiplying the low-contrast mask image with the primary color mask image. The size of the target contrast mask image is the same as the size of the original image.
[0133] S304. Obtain m feature images based on the original image, m segmentation mask images, main color data, and target contrast mask image.
[0134] In this embodiment, for each segmented mask image, it is merged with the original image, the main color matrix, and the target contrast mask image to obtain its corresponding feature image. The main color matrix is generated from the main color data. A CNN is used to convolve the main color data to generate the main color matrix. The size of the main color matrix is the same as the size of the original image.
[0135] Merging the original image, segmentation mask image, main color matrix, and target contrast mask image may include: superimposing the segmentation mask image, main color matrix, and target contrast mask image onto the original image.
[0136] S305. Perform color space conversion on m feature images to obtain m feature-compensated images, and obtain m target feature images based on the main color data and m feature-compensated images.
[0137] In this embodiment, color space conversion of the feature image may include converting the feature image from the RGB color space to the HSV color space or the LAB color space. For each feature compensation image, it is merged with the primary color matrix to obtain the target feature image. Since the feature compensation image contains richer color information, the target feature image contains a more comprehensive color enhancement area.
[0138] Merging the feature-compensated image and the primary color matrix may include overlaying the primary color matrix onto the feature-compensated image.
[0139] S306. Enhance m target feature images to obtain m local enhanced images, and obtain a global enhanced image based on the original image, main color data and m local enhanced images.
[0140] In this embodiment, an AI network can be used to enhance m target feature images. A primary color matrix is generated based on the primary color data. Then, the original image, the primary color matrix, and the m locally enhanced images are merged to obtain a global enhanced image. The global enhanced image contains global features from the original image and target local features. Global features include overall color and lighting information, while target local features include color and frequency band information for the color-enhanced regions. Global features reflect the overall color and lighting of the image, while target local features reflect the contrast and smoothness of the color-enhanced regions. By enhancing both global and target local features, the AI network learns to enhance both the overall color and lighting of the image and the contrast and smoothness of the color-enhanced regions, thereby improving the overall visual effect of the image.
[0141] Merging the original image, the main color matrix, and m locally enhanced images can include superimposing the main color matrix and m locally enhanced images onto the original image.
[0142] AI networks include, but are not limited to, CNN, GAN, FCN, Autoencoder, and Zero-DCE networks.
[0143] In some embodiments, step S305 may be omitted. That is, the m feature images are enhanced to obtain m locally enhanced images, and a global enhanced image is obtained based on the original image, the main color data, and the m locally enhanced images.
[0144] In some other embodiments, in step S306, a global enhanced image can be obtained based on the original image and m local enhanced images, that is, the original image and m local enhanced images are merged to obtain a global enhanced image.
[0145] Based on the above image enhancement algorithm implementation process, in this embodiment, an original image is acquired, which includes a first region. The original image is then enhanced to obtain a globally enhanced image, which includes a second region. The position of the first region in the original image corresponds to the position of the second region in the globally enhanced image, and the image quality of the second region is superior to that of the first region. The globally enhanced image contains global features and target local features of the original image. The global features include the overall color and lighting information of the original image, while the target local features include the color and frequency band information of the first region. By enhancing the global features, the overall color and lighting of the original image are optimized; by enhancing the target local features, the image quality of the first region is optimized, thereby improving the enhancement effect of both the overall image and local regions.
[0146] The image quality of the second region being superior to that of the first region includes: the contrast of the first region being lower than that of the second region, and / or the sharpness of the first region being lower than that of the second region, and / or the saturation of the first region being lower than that of the second region, and / or the smoothness of the first region being lower than that of the second region, and / or the brightness of the first region being lower than that of the second region. The image quality of a local region includes at least one of contrast, sharpness, saturation, smoothness, and brightness. Accordingly, the enhancement processing includes at least one of enhancing local contrast, enhancing local sharpness, enhancing local saturation, enhancing local smoothness, and enhancing local brightness.
[0147] The hardware structure of the electronic device is described in detail below.
[0148] like Figure 12 As shown, the electronic device includes a processor 110, an external memory interface 120, an internal memory 121, a sensor module 130, a display screen 140, and a camera module 150. The sensor module 130 includes a touch sensor 131 and an image sensor 132. The camera module 150 includes at least one camera 151.
[0149] The processor 110 is used to execute the various functions or steps performed by the electronic device in the above embodiments. The processor 110 includes a CPU, a GPU, and a GPU. The processor 110 may also include an application processor (AP), an image signal processor (ISP), a digital signal processor (DSP), a modem processor, and a video codec, etc.
[0150] The external memory interface 120 is used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 110 through the external memory interface 120 to perform data storage functions.
[0151] Internal memory 121 stores executable program code, including instructions. Processor 110 executes the various functions or steps performed by the electronic device in the above embodiments by running the instructions stored in internal memory 121. Internal memory 121 includes a program storage area and a data storage area. The program storage area may store the operating system, at least one application required for a function, etc. The data storage area may store data created during the use of the electronic device. Internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, and Universal Flash Storage (UFS), etc.
[0152] The display screen 140 is used to display an interface. The display screen 140 includes a display panel. The display panel may be a liquid crystal display (LCD), a light-emitting diode (LED), or an organic light-emitting diode (OLED), etc.
[0153] If the display screen 140 integrates the touch sensor 131, then the display screen 140 can be referred to as a touch screen. The touch sensor 131 can also be referred to as a "touch panel". That is, the display screen 140 may include a display panel and a touch panel. The touch sensor 131 is used to detect touch operations applied to or near it. After the touch sensor 131 detects a touch operation, it can trigger the driver of the hardware abstraction layer of the electronic device to periodically scan the touch parameters generated by the touch operation. Then, the driver of the hardware abstraction layer sends the touch parameters to the relevant modules in the upper layer so that the relevant modules can determine the touch event corresponding to the touch parameters.
[0154] The electronic device can realize display functions through GPU, display screen 140 and AP, and can realize shooting functions through GPU, ISP, camera module 150, image sensor 132, video codec, display screen 140 and AP.
[0155] Electronic devices can capture light through any one or more cameras 151 in the camera module 150 and transmit the light signal to the image sensor 132. The light signal is converted into a raw image (or RAW image) by the image sensor 132. Then, the RAW image is converted into a YUV image by the ISP, and then the YUV image is converted into an RGB image, thus completing the image acquisition. Here, RAW is a raw image data format that has not been compressed or modified. YUV is a color encoding format, where Y represents luminance, and U and V represent chrominance. An RGB image is an image composed of three color channels: red, green, and blue. The RGB image can be directly displayed on the display screen 140 or subjected to subsequent editing and processing.
[0156] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements.
[0157] The functions or steps performed by the electronic device in the above embodiments can also be applied to chips, computer-readable storage media, or computer program products.
[0158] The chip includes a processor and an interface circuit, with the processor and interface circuit electrically connected. The interface circuit can read computer instructions stored in the memory and send the computer instructions to the processor. When the processor executes the computer instructions, it implements the various functions or steps performed by the electronic device in the above embodiments.
[0159] The computer-readable storage medium stores computer instructions, which, when executed by a processor, implement the various functions or steps performed by the electronic device in the above embodiments.
[0160] Computer-readable storage media include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer-readable storage media include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer.
[0161] The computer program product includes computer instructions, which, when executed by a processor, implement the various functions or steps performed by the electronic device in the above embodiments.
[0162] The embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of this application.
Claims
1. An image processing method, characterized in that, The method includes: Acquire a first image, the first image containing a first region; The first image is enhanced to obtain a second image, which includes a second region; The position of the first region in the first image corresponds to the position of the second region in the second image, and the image quality of the second region is better than that of the first region.
2. The image processing method as described in claim 1, characterized in that, The image quality of the second region is superior to that of the first region in the following ways: The contrast of the first region is less than that of the second region, and / or the sharpness of the first region is less than that of the second region, and / or the saturation of the first region is less than that of the second region, and / or the smoothness of the first region is less than that of the second region, and / or the brightness of the first region is less than that of the second region.
3. The image processing method as described in claim 1 or 2, characterized in that, The enhancement process of the first image to obtain the second image includes: Semantic segmentation is performed on the first image to obtain m segmentation mask images, where m is a positive integer; Extract the main color data of the first image and obtain the main color mask image, wherein the main color data includes the color data of the first region; Obtain the target frequency band mask image based on the first image; Obtain a target contrast mask image based on the target frequency band mask image and the primary color mask image; Based on the first image, the m segmentation mask images, the main color data, and the target contrast mask image, obtain m feature images; The second image is obtained based on the first image and the m feature images.
4. The image processing method as described in claim 3, characterized in that, The step of obtaining the second image based on the first image and the m feature images includes: The m feature images are converted to a different color space to obtain m feature-compensated images; Based on the main color data and the m feature compensation images, obtain m target feature images; The second image is obtained based on the first image and the m target feature images.
5. The image processing method as described in claim 4, characterized in that, The step of obtaining the second image based on the first image and the m target feature images includes: The m target feature images are enhanced to obtain m locally enhanced images; The first image and the m locally enhanced images are merged, or the first image, the main color matrix, and the m locally enhanced images are merged to obtain the second image; the main color matrix is generated from the main color data.
6. The image processing method as described in claim 4 or 5, characterized in that, The step of obtaining m target feature images based on the main color data and the m feature compensation images includes: Generate a primary color matrix based on the primary color data; traverse the m feature compensation images, merge the i-th feature compensation image with the primary color matrix to obtain the i-th target feature image, until the traversal is complete; i is a positive integer and i≤m.
7. The image processing method according to any one of claims 4-6, characterized in that, The step of performing color space conversion on the m feature images to obtain m feature-compensated images includes: The m feature images are converted from RGB color space to HSV color space or LAB color space respectively to obtain the m feature compensation images.
8. The image processing method as described in claim 3, characterized in that, The step of obtaining the second image based on the first image and the m feature images includes: The m feature images are enhanced to obtain m locally enhanced images; The first image and the m locally enhanced images are merged, or the first image, the main color matrix, and the m locally enhanced images are merged to obtain the second image; the main color matrix is generated from the main color data.
9. The image processing method according to any one of claims 3-8, characterized in that, The step of obtaining m feature images based on the first image, the m segmentation mask images, the main color data, and the target contrast mask image includes: Generate a main color matrix based on the main color data; traverse the m segmentation mask images, and merge the i-th segmentation mask image, the first image, the main color matrix, and the target contrast mask image to obtain the i-th feature image, until the traversal is complete; i is a positive integer, and i≤m.
10. The image processing method according to any one of claims 3-9, characterized in that, The step of obtaining the target contrast mask image based on the target frequency band mask image and the primary color mask image includes: The target frequency band mask image and the primary color mask image are multiplied together to obtain the target contrast mask image.
11. The image processing method according to any one of claims 3-10, characterized in that, The step of extracting the main color data of the first image and obtaining the main color mask image includes: Extract the main color data from the first image; Set the target color range based on the primary color data; The primary color mask image is generated based on the target color range.
12. The image processing method according to any one of claims 3-11, characterized in that, The step of obtaining the target frequency band mask image based on the first image includes: Perform a frequency domain transformation on the first image to obtain a frequency domain image; Set the target frequency band mask template; The target frequency band mask template is applied to the frequency domain image to obtain a frequency domain mask image; The target frequency band mask image is obtained by performing an inverse frequency domain transformation on the frequency domain mask image.
13. The image processing method according to any one of claims 3-12, characterized in that, The target frequency band mask image includes a high-frequency mask image and / or a low-frequency mask image. The high-frequency mask image contains high-frequency information of the first image, and the low-frequency mask image contains low-frequency information of the first image.
14. An electronic device, characterized in that, It includes a memory and a processor, which implements the image processing method as described in any one of claims 1-13 when the processor executes computer instructions stored in the memory.
15. A computer-readable storage medium, characterized in that, It stores computer instructions, which, when executed by the processor, implement the image processing method as described in any one of claims 1-13.