Image processing methods, apparatuses, electronic devices, storage media and chips
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2026-08-14
AI Technical Summary
然而,相关技术中的图像处理方法依然存在成像效果差的问题
[0046] The technical solutions provided by the embodiments of this disclosure can include the following beneficial effects: After obtaining the image to be processed, the attribute information of the image to be processed can be extracted first, and then the image processing strategy corresponding to the image to be processed can be determined according to the attribute information. Finally, the image to be processed can be processed based on the image processing strategy to obtain the target image. Since the image processing strategy can be selected specifically according to the characteristics of the image to be processed, the imaging effect of the obtained target image can be improved.
Smart Images

Figure CN117197260B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image technology, and in particular to an image processing method, apparatus, electronic device, storage medium, and chip. Background Technology
[0002] With the gradual development of hardware and software in imaging devices (such as mobile phone cameras and digital cameras), people's requirements for imaging quality are also increasing. However, image processing methods in related technologies still suffer from poor imaging results. Summary of the Invention
[0003] To overcome the problems existing in related technologies, this disclosure provides an image processing method, apparatus, electronic device, storage medium, and chip.
[0004] According to a first aspect of the present disclosure, an image processing method is provided, comprising:
[0005] Obtain the image to be processed;
[0006] Extract the attribute information of the image to be processed;
[0007] Based on the attribute information, determine the image processing strategy corresponding to the image to be processed;
[0008] The image to be processed is processed based on the image processing strategy to obtain the target image.
[0009] In some implementations, the attribute information of the image to be processed includes the global semantic attributes of the image to be processed, and determining the image processing strategy corresponding to the image to be processed based on the attribute information includes:
[0010] At least based on the global semantic attributes of the image to be processed, the first scene context corresponding to the image to be processed is determined;
[0011] Obtain the first color adjustment strategy corresponding to the first scene context;
[0012] The step of processing the image to be processed based on the image processing strategy to obtain the target image includes:
[0013] Based on the first color adjustment strategy, the image to be processed is color adjusted to obtain the target image.
[0014] In some implementations, determining the first scene context corresponding to the image to be processed, at least based on the global semantic attributes of the image to be processed, includes:
[0015] Based on the global semantic attributes of the image to be processed, determine the first scene context corresponding to the image to be processed; or
[0016] The first scene context corresponding to the image to be processed is determined based on the global semantic attributes of the image to be processed and the global semantic attributes of images in a preset number of adjacent frames.
[0017] In some implementations, the attribute information of the image to be processed includes global semantic attributes and temporal semantic attributes of the image to be processed. The temporal semantic attributes of the image to be processed are determined based on images adjacent to the image to be processed for a preset number of frames. Determining the image processing strategy corresponding to the image to be processed based on the attribute information includes:
[0018] Based on the global semantic attributes and temporal semantic attributes of the image to be processed, the second scene context corresponding to the image to be processed is determined;
[0019] Obtain the second color adjustment strategy corresponding to the second scene context;
[0020] The step of processing the image to be processed based on the image processing strategy to obtain the target image includes:
[0021] Based on the second color adjustment strategy, the image to be processed is color adjusted to obtain the target image.
[0022] In some implementations, the attribute information of the image to be processed includes local semantic attributes of each image region in the image to be processed, and determining the image processing strategy corresponding to the image to be processed based on the attribute information includes:
[0023] Based on the local semantic attributes of each image region, the target image enhancement strategy corresponding to each image region is determined;
[0024] The step of processing the image to be processed based on the image processing strategy to obtain the target image includes:
[0025] By utilizing the target image enhancement strategies corresponding to each image region, image enhancement is performed on each image region to obtain the enhanced image corresponding to each image region.
[0026] The enhanced images corresponding to each of the image regions are stitched together to obtain the target image.
[0027] In some implementations, the type of the local semantic attribute includes at least one of the following attribute types: noise intensity, detail richness, brightness, and edge sharpness.
[0028] In some implementations, the attribute information of the image to be processed includes the alignment difficulty of each image region to be aligned and each image region to be aligned. Determining the image processing strategy corresponding to the image to be processed based on the attribute information includes:
[0029] Based on the alignment difficulty of each image region to be aligned, determine the target image alignment strategy corresponding to each image region to be aligned;
[0030] The step of processing the image to be processed based on the image processing strategy to obtain the target image includes:
[0031] Based on the target image alignment strategy corresponding to each image region to be aligned, the corresponding image regions to be aligned in the image to be processed and the candidate image are aligned to obtain the pixel correspondence between the corresponding image regions to be aligned. The candidate image is a frame image adjacent to the image to be processed.
[0032] Based on the preset fusion strategy and the pixel correspondence between the corresponding image regions to be aligned, the corresponding image regions to be aligned are fused to obtain the fused image of the image to be processed.
[0033] Based on the preset fusion strategy, the corresponding aligned image regions in the image to be processed and the candidate image are fused to obtain the fused image of each aligned image region in the image to be processed;
[0034] The target image is obtained by stitching together the image after fusing the image regions to be aligned in the image to be processed and the image after fusing the image regions to be aligned in the image to be processed.
[0035] In some implementations, the image to be processed is a raw image acquired by an image sensor, or a target image obtained after processing the raw image using at least one of the following image processing strategies: color adjustment strategy, image enhancement strategy, and image alignment strategy.
[0036] According to a second aspect of the present disclosure, an image processing apparatus is provided, comprising:
[0037] The acquisition module is configured to acquire the image to be processed.
[0038] The extraction module is configured to extract attribute information from the image to be processed;
[0039] The determination module is configured to determine the image processing strategy corresponding to the image to be processed based on the attribute information;
[0040] The processing module is configured to process the image to be processed based on the image processing strategy to obtain the target image.
[0041] According to a third aspect of the present disclosure, an electronic device is provided, comprising:
[0042] A memory on which computer programs are stored;
[0043] A processor is configured to execute the computer program in the memory to implement the steps of the image processing method provided in the first aspect of this disclosure.
[0044] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that stores computer program instructions thereon, which, when executed by a processor, implement the steps of the image processing method provided in the first aspect of the present disclosure.
[0045] According to a fifth aspect of the present disclosure, a chip is provided, including a processor and an interface; the processor is configured to read instructions to execute the steps of the image processing method provided in the first aspect of the present disclosure.
[0046] The technical solutions provided by the embodiments of this disclosure can include the following beneficial effects: After obtaining the image to be processed, the attribute information of the image to be processed can be extracted first, and then the image processing strategy corresponding to the image to be processed can be determined according to the attribute information. Finally, the image to be processed can be processed based on the image processing strategy to obtain the target image. Since the image processing strategy can be selected specifically according to the characteristics of the image to be processed, the imaging effect of the obtained target image can be improved.
[0047] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0048] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0049] Figure 1 This is a flowchart illustrating an image processing method according to an exemplary embodiment.
[0050] Figure 2 This is a schematic flowchart illustrating a color adjustment stage according to an exemplary embodiment.
[0051] Figure 3 This is a schematic diagram illustrating a process of an image enhancement stage according to an exemplary embodiment.
[0052] Figure 4 This is a schematic diagram of an alignment and fusion stage according to an exemplary embodiment.
[0053] Figure 5 This is a block diagram illustrating an image processing apparatus according to an exemplary embodiment.
[0054] Figure 6 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment. Detailed Implementation
[0055] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0056] It should be noted that all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located, and with the authorization granted by the owner of the relevant device.
[0057] As mentioned in the background section, people have increasingly higher requirements for imaging quality. Among these, image color, noise, sharpness, brightness, and inter-frame stability are all important factors affecting imaging results. Currently, imaging devices typically employ ISP (Image Signal Processing) systems to improve image quality.
[0058] However, during image color restoration, the AWB (Auto White Balance) module in the ISP system adjusts the colors by using the brightness statistics of the input frame image. The color processing result has a significant impact on the final visual experience, and different scenes have different color processing requirements, such as warm and cool tones, different hue saturation levels, etc.
[0059] Furthermore, in ISP systems, the processing of noise, detail, and edge sharpness involves considering that each image is essentially composed of various brightness levels, noise patterns, and frequencies. For example, in a city night scene image, building areas are characterized by high brightness, low noise, and high frequency, while the sky is characterized by low brightness, high noise, and low frequency. However, the corresponding processing algorithms commonly used in ISP systems process the entire image in a one-size-fits-all manner, lacking specificity. This results in all regions being processed in an intermediate state or biased towards one side, rather than all being in an ideal state. For example, high-noise areas may not be processed completely, while details in low-noise areas may be erased.
[0060] Furthermore, for video frame stabilization algorithms, the strategy in the ISP system is to align first and then fuse. To ensure imaging speed, the alignment algorithm generally uses simple and fast traditional alignment algorithms, such as the Homograpy algorithm for global alignment and the optical flow method for pixel-by-pixel alignment. However, for some areas that are difficult to align, such as foreground motion areas, local large motion areas, and dark motion areas, traditional alignment algorithms cannot effectively extract feature points, resulting in poor alignment effects. Subsequent fusion requires high accuracy in alignment; if the alignment algorithm is not well-processed, ghosting and other phenomena will appear during fusion, severely affecting the image quality.
[0061] It is evident that the image processing methods in related technologies still suffer from poor imaging results.
[0062] In view of this, embodiments of the present disclosure provide an image processing method, apparatus, electronic device, storage medium, and chip, which can select image processing strategies in a targeted manner according to the characteristics of the image to be processed, thereby improving the imaging effect of the obtained target image.
[0063] Figure 1 This is a flowchart illustrating an image processing method according to an exemplary embodiment, such as... Figure 1 As shown, the image processing method can be used in electronic devices, such as mobile phones, cameras, laptops, tablets, and smart wearable devices with camera functions. The image processing method includes:
[0064] S110, Obtain the image to be processed.
[0065] In some implementations, the electronic device can acquire the image to be processed using its built-in image sensor. In this case, the image to be processed can be the raw image acquired by the image sensor, that is, the raw RAW (RAW Image Format) image output by the image sensor can be used as the image to be processed.
[0066] Furthermore, considering that in practice, after obtaining the raw RAW image output by the sensor, various intermediate processing steps can be performed on the raw RAW image to obtain a satisfactory RAW image for output, in some embodiments, the image to be processed can also be an intermediate RAW image after processing the raw RAW image using at least one image processing strategy.
[0067] In some implementations, the image to be processed can be the target image obtained after processing the original image using at least one of the image processing strategies, namely color adjustment strategy, image enhancement strategy, and image alignment strategy.
[0068] For example, in a single-frame image imaging scenario, such as a photo-taking scenario, after obtaining the raw RAW image output by the sensor, color adjustment strategies and image enhancement strategies can be applied to process the raw RAW image to obtain a satisfactory RAW image for output. In this case, the raw RAW image or the intermediate RAW image after applying color adjustment strategies can be used as the image to be processed.
[0069] For example, in a multi-frame image imaging scenario, such as a video shooting scenario, after obtaining the multiple frames of raw RAW images output by the sensor, color adjustment strategies, image enhancement strategies, and image alignment strategies can be applied to process the raw RAW images to obtain multi-frame RAW images that meet the requirements for video output. In this case, the original RAW image, or an intermediate RAW image after processing the original RAW image using color adjustment strategies, or an intermediate RAW image after sequentially processing the original RAW image using color adjustment strategies and image enhancement strategies, can be used as the image to be processed.
[0070] It should be noted that the order of the image processing strategies applied in the above examples can be changed as needed. For example, image enhancement strategies can be applied to the original RAW image first, and then the resulting intermediate RAW image can be used as the image to be processed, and color adjustment strategies can be applied, etc.
[0071] It should be noted that, depending on the actual needs, one or more image processing strategies can be applied to the original RAW image to obtain the target image that meets the requirements for output. That is, after obtaining the target image that meets the requirements, it is not necessary to process the target image as the image to be processed.
[0072] It should be noted that the original RAW image or intermediate RAW image can come from the electronic device that acquires the image to be processed, or from other electronic devices that are communicatively connected to the electronic device that acquires the image to be processed.
[0073] S120, Extract the attribute information of the image to be processed.
[0074] Among them, attribute information can be understood as information that reflects the global, local and temporal characteristics of the image to be processed.
[0075] In some implementations, attribute information of the image to be processed can be extracted based on a deep learning extraction model.
[0076] It is understandable that in some implementations, the type of attribute information may differ, and the deep learning-based extraction model may be different.
[0077] S130, Based on the attribute information, determine the image processing strategy corresponding to the image to be processed.
[0078] S140: The image to be processed is processed based on the image processing strategy to obtain the target image.
[0079] In this embodiment of the disclosure, after determining the attribute information of the image to be processed, it is equivalent to obtaining the characteristics of the image to be processed. Therefore, an image processing strategy can be selected in a targeted manner according to the characteristics of the image to be processed to process the image, thereby improving the imaging effect of the obtained target image.
[0080] Using the above method, after obtaining the image to be processed, the attribute information of the image to be processed can be extracted first. Then, the image processing strategy corresponding to the image to be processed can be determined based on the attribute information. Finally, the image to be processed can be processed based on the image processing strategy to obtain the target image. Since the image processing strategy can be selected specifically according to the characteristics of the image to be processed, compared with the related technologies that use a uniform image processing strategy for different images, the imaging effect of the obtained target image can be improved.
[0081] It should be noted that the above image processing method can be applied to each stage of processing the original RAW image into a RAW image that meets the requirements. In this case, the extracted attribute information of the image to be processed can be the attribute information corresponding to each stage. Therefore, based on the attribute information, the image processing strategy determined for the image to be processed can be the image processing strategy specific to that stage. For example, in the color adjustment stage, the attribute information corresponding to the color adjustment stage can be obtained, and the corresponding color adjustment strategy can be determined based on the attribute information of the color adjustment stage. Similarly, in the image enhancement stage, the attribute information corresponding to the image enhancement stage can be obtained, and the corresponding image enhancement strategy can be determined based on the attribute information of the image enhancement stage.
[0082] Furthermore, the aforementioned image processing method can also be applied to the overall process of processing the original RAW image into a RAW image that meets the requirements. In this case, the extracted attribute information of the image to be processed can be a set of attribute information corresponding to each stage. Therefore, based on the attribute information, the image processing strategy determined for the image to be processed can be a set of image processing strategies corresponding to each attribute information in the set. For example, the attribute information corresponding to the color adjustment stage and the attribute information corresponding to the image enhancement stage can be obtained at once. Then, based on the attribute information of the color adjustment stage and the image enhancement stage, the corresponding color adjustment strategy and image enhancement strategy can be determined. Finally, the image to be processed can be processed according to the color adjustment strategy and the image enhancement strategy respectively to obtain the target image that meets the requirements.
[0083] As can be seen from the foregoing, different attribute information can exist at different stages of processing the original RAW image. Optionally, the attribute information of the image to be processed may include the global semantic attributes of the image to be processed, the temporal semantic attributes of the image to be processed, the local semantic attributes of each image region in the image to be processed, the alignment difficulty of each image region to be aligned in the image to be processed, and the image region to be aligned, etc.
[0084] Optionally, corresponding to the color adjustment stage, the global semantic attributes of the image to be processed can be selected. Therefore, in some embodiments, the attribute information of the image to be processed includes the global semantic attributes of the image to be processed. In this case, determining the image processing strategy corresponding to the image to be processed based on the attribute information includes the following steps:
[0085] Based at least on the global semantic attributes of the image to be processed, determine the first scene context corresponding to the image to be processed; obtain the first color adjustment strategy corresponding to the first scene context.
[0086] In this case, the image to be processed is processed based on an image processing strategy to obtain the target image, including the following steps:
[0087] Based on the first color adjustment strategy, the image to be processed is color adjusted to obtain the target image.
[0088] The first color adjustment strategy can be understood as the color adjustment strategy corresponding to the first scene context.
[0089] Global semantic attributes can be understood as information about the actual linguistic meaning expressed by a frame of an image as a whole. For example, global semantic attributes include rainy day, city streets, only a few streetlights, and no one.
[0090] Scene context can be understood as further expressing deeper information about the shooting scene. For example, by using the overall semantic attributes of a rainy day, city streets, only a few streetlights, and no people, a somber, cold, dark, and lonely scene context can be depicted.
[0091] Among these methods, there are multiple ways to determine the first scene context corresponding to the image to be processed, based at least on the global semantic attributes of the image to be processed.
[0092] In some implementations, determining the first scene context corresponding to the image to be processed, at least based on the global semantic attributes of the image, may include the following steps:
[0093] The first scene context corresponding to the image to be processed is determined based on the global semantic attributes of the image to be processed; or the first scene context corresponding to the image to be processed is determined based on the global semantic attributes of the image to be processed and the global semantic attributes of images in a preset number of adjacent frames.
[0094] In this embodiment of the disclosure, for single-frame or multi-frame scenes, the first scene context corresponding to the image to be processed can be determined separately based on the global semantic attributes of the image to be processed.
[0095] In addition to determining the first scene context corresponding to the image to be processed based solely on the global semantic attributes of the image to be processed, for multi-frame scenes, the first scene context corresponding to the image to be processed can also be determined based on the global semantic attributes of the image to be processed and the global semantic attributes of images in a preset number of adjacent frames.
[0096] Among them, the image adjacent to the image to be processed at a preset number of frames can be an image in the video frame sequence that is before the image to be processed, or an image in the video frame sequence that is after the image to be processed, or an image in the video frame sequence that is both before and after the image to be processed.
[0097] For example, if the image to be processed is the 5th frame in a video frame sequence, the global semantic attributes of frames 1 through 9 in the video frame sequence can be obtained. By integrating and refining the global semantic attributes of these 9 frames, the common first scene context corresponding to these 9 frames can be obtained. Alternatively, the first scene context of the 5th frame can be determined based on the global semantic attributes of frames 1 through 9, and the first scene context of the 6th frame can be determined based on the global semantic attributes of frames 2 through 10, and the first scene context of the 6th frame can be determined based on the global semantic attributes of frames 2 through 10, and the first scene context of the 6th frame can be determined by integrating and refining the global semantic attributes of these 9 frames.
[0098] In this embodiment of the disclosure, the first scene context corresponding to the image to be processed is determined based on the global semantic attributes of the image to be processed and the global semantic attributes of the images of a preset number of adjacent frames. This can eliminate random errors between single-frame images and improve the accuracy of the determined first scene context corresponding to the image to be processed.
[0099] In some implementations, the first color adjustment strategy may be a color correction model trained using deep learning.
[0100] In this case, the relationship between the first scene context and the key parameters of the color correction model can be established in advance. After determining the corresponding first scene context, the parameters of the corresponding color correction model can be obtained based on the relationship. Then, the color of the image to be processed can be adjusted based on the color correction model under these parameters to obtain the target image.
[0101] In addition, by pre-establishing the association between the first scene context and different color correction models, after determining the corresponding first scene context, the corresponding color correction model can be obtained based on the association, and the color of the image to be processed can be adjusted based on the color correction model to obtain the target image.
[0102] Optionally, corresponding to the color adjustment stage, global semantic attributes and temporal semantic attributes of the image to be processed can also be selected. Therefore, in some embodiments, the attribute information of the image to be processed includes global semantic attributes and temporal semantic attributes. In this case, determining the image processing strategy corresponding to the image to be processed based on the attribute information includes the following steps:
[0103] Based on the global semantic attributes and temporal semantic attributes of the image to be processed, the second scene context corresponding to the image to be processed is determined; and the second color adjustment strategy corresponding to the second scene context is obtained.
[0104] In this case, the image to be processed is processed based on an image processing strategy to obtain the target image, including the following steps:
[0105] Based on the second color adjustment strategy, the image to be processed is color-adjusted to obtain the target image.
[0106] The second color adjustment strategy can be understood as the color adjustment strategy corresponding to the second scene context.
[0107] The temporal semantic attributes of the image to be processed are determined based on a preset number of images in the video frame sequence that are adjacent to the image to be processed.
[0108] Temporal semantic attributes can be understood as information expressing the actual linguistic meaning based on the data stream temporal sequence of a given image frame. Here, data stream temporal sequence refers to the sequential order of a given image frame with a predetermined number of adjacent frames. For example, if an image frame contains a basketball, without considering the data stream temporal sequence, only the semantics of a basketball can be obtained. However, if the data stream temporal sequence between this frame and other adjacent frames is considered, the semantics of a shooting motion can be obtained. The semantics of shooting is thus a temporal semantic attribute.
[0109] Therefore, in this embodiment of the present disclosure, when determining the second scene context corresponding to the image to be processed, the global semantic attributes and temporal semantic attributes of the image to be processed can be considered simultaneously, so as to make the determined second scene context richer and more accurate.
[0110] In this embodiment, the process of obtaining the second color adjustment strategy corresponding to the second scene context is similar to the process of obtaining the first color adjustment strategy corresponding to the first scene context in the aforementioned embodiments. The process of adjusting the color of the image to be processed based on the second color adjustment strategy to obtain the target image is similar to the process of adjusting the color of the image to be processed based on the first color adjustment strategy to obtain the target image in the aforementioned embodiments. Similarities can be found in the aforementioned embodiments, and will not be repeated here.
[0111] Optionally, corresponding to the image enhancement stage, local semantic attributes of each image region in the image to be processed can be selected. Therefore, in some embodiments, the attribute information of the image to be processed includes the local semantic attributes of each image region in the image to be processed. In this case, determining the image processing strategy corresponding to the image to be processed based on the attribute information includes the following steps:
[0112] Based on the local semantic attributes of each image region, the target image enhancement strategy corresponding to each image region is determined.
[0113] In this case, the image to be processed is processed based on an image processing strategy to obtain the target image, including the following steps:
[0114] By utilizing the target image enhancement strategies corresponding to each image region, image enhancement is performed on each image region to obtain the enhanced image corresponding to each image region; the enhanced images corresponding to each image region are then stitched together to obtain the target image.
[0115] Local semantic attributes can be understood as the semantic attributes of local regions in an image. For example, the types of local semantic attributes include at least one of the following: noise intensity, detail richness, brightness, and edge sharpness. Furthermore, each attribute type can have a specific attribute classification. For example, high noise intensity, low noise intensity; high brightness, low brightness, etc.
[0116] In this embodiment of the disclosure, considering that different image regions in the same frame may correspond to different specific attribute classifications under the same attribute type, targeted image enhancement strategies can be adopted for different attribute classifications under the same attribute type to improve the imaging effect.
[0117] In other words, the target image enhancement strategies corresponding to each image region can be used to enhance each image region and obtain the enhanced image corresponding to each image region. Since the obtained image is the enhanced image corresponding to each image region, the enhanced images corresponding to each image region can be further stitched together to obtain the target image.
[0118] For example, in some implementations, local semantic attributes may include three attribute types: brightness, noise intensity, and edge sharpness. Regarding the noise intensity attribute type, region 1 in the image to be processed may have high noise intensity, thus requiring a strong denoising strategy; however, region 2 may have relatively low noise intensity. If the same strong denoising strategy is used, the details in this region will be erased as noise, so a weaker denoising strategy that retains more detail is needed.
[0119] Similarly, for the brightness attribute, region 3 in the image to be processed may have low brightness, so a strong brightness adjustment strategy is needed; but for region 4, its brightness may be relatively high, so a weaker brightness adjustment strategy can be adopted.
[0120] Similarly, for the edge sharpness attribute, region 5 in the image to be processed may have high edge sharpness, so a strong sharpening strategy is needed; but for region 6, its edge sharpness may be relatively low, so a weaker sharpening strategy can be adopted.
[0121] It should be noted that regions 1, 3, and 5 mentioned above can be the same region in the image to be processed, or they can be different regions in the image to be processed. Similarly, regions 2, 4, and 6 mentioned above can be the same region in the image to be processed, or they can be different regions in the image to be processed. In addition, in practice, more regions can be divided.
[0122] As can be seen from the foregoing, the region division in the image to be processed may be inconsistent under different attribute types. Therefore, in some implementations, when the local semantic attributes include multiple attribute types, the above method can be performed on each attribute type in turn.
[0123] For example, for the noise intensity attribute type, we can first obtain the local semantic attributes of each image region in the image to be processed under the noise intensity attribute type, and determine the denoising strategy corresponding to each image region. Then, we can use the denoising strategy corresponding to each image region to denoise each image region to obtain the denoised image corresponding to each image region. Then, we can stitch the denoised images corresponding to each image region together to obtain the denoised image.
[0124] Next, for the brightness attribute type, we can first obtain the local semantic attributes of each image region in the denoised intermediate image under the brightness attribute type, and determine the brightness adjustment strategy corresponding to each image region. Then, we can use the brightness adjustment strategy corresponding to each image region to adjust the brightness of each image region, and obtain the brightness-adjusted image corresponding to each image region. Finally, we can stitch the brightness-adjusted images corresponding to each image region together to obtain the brightness-adjusted image.
[0125] Next, for the edge sharpness attribute type, we can first obtain the local semantic attributes of each image region in the denoised intermediate image under the edge sharpness attribute type, and determine the sharpening strategy corresponding to each image region. Then, we can use the sharpening strategy corresponding to each image region to sharpen each image region, and obtain the sharpened image corresponding to each image region. Finally, we can stitch together the sharpened images corresponding to each image region to obtain the sharpened image.
[0126] It should be noted that the order of noise reduction, brightness adjustment, and sharpening processes can be changed as needed.
[0127] In some implementations, a semantic region segmentation model trained based on deep learning can be used to obtain the local semantic attributes of each image region in the image to be processed. Optionally, the semantic region segmentation model may include a semantic region segmentation model for noise intensity, a semantic region segmentation model for brightness, and a semantic region segmentation model for edge sharpness, etc.
[0128] In some implementations, the target image enhancement strategy can also be based on an image enhancement model trained using deep learning.
[0129] In some implementations, the image enhancement model may be, for example, a denoising model, a brightness adjustment model, or a sharpening model.
[0130] In some implementations, for a local semantic attribute of a certain attribute type, the association between the local semantic attribute and the key parameters of the image enhancement model corresponding to the attribute type can be established in advance. After determining the local semantic attributes of each image region, the parameters of the corresponding image enhancement model can be obtained according to the association. Then, the corresponding image region can be enhanced based on the image enhancement model under the parameters to obtain the enhanced image of each image region.
[0131] In addition, by pre-establishing the association between local semantic attributes of a certain attribute type and different image enhancement models under the corresponding attribute type, after determining the local semantic attributes of each corresponding image region, the corresponding image enhancement model can be obtained based on the association. Then, the corresponding image region can be enhanced based on the image enhancement model to obtain the enhanced image corresponding to each image region.
[0132] Understandably, for some image imaging processes, such as video recording, image alignment and fusion processes are needed to improve inter-frame stability. Currently, algorithms with good alignment results are based on deep learning alignment strategies, but their speed is relatively slow and sometimes cannot meet real-time processing requirements. Therefore, to improve alignment results while meeting real-time processing requirements, in some implementations, corresponding to the image alignment stage, the attribute information of the image to be processed can be selected, including the alignment difficulty of each image region to be aligned and each aligned image region. Thus, in some implementations, the attribute information of the image to be processed can include the alignment difficulty of each image region to be aligned and each aligned image region. In this case, based on the attribute information, the image processing strategy corresponding to the image to be processed is determined, including the following steps:
[0133] Based on the alignment difficulty of each image region to be aligned, determine the target image alignment strategy corresponding to each image region to be aligned.
[0134] In this case, the image to be processed is processed based on an image processing strategy to obtain the target image, including the following steps:
[0135] Based on the target image alignment strategy corresponding to each image region to be aligned, the corresponding image regions to be aligned in the image to be processed and the candidate image are aligned to obtain the pixel correspondence between the corresponding image regions to be aligned.
[0136] Based on the preset fusion strategy and the pixel correspondence between the corresponding image regions to be aligned, the corresponding image regions to be aligned are fused to obtain the fused image of the image to be aligned in the image to be processed.
[0137] Based on a preset fusion strategy, the corresponding aligned image regions in the image to be processed and the candidate images are fused to obtain the fused image of each aligned image region in the image to be processed.
[0138] The target image is obtained by stitching together the image after fusing the individual image regions to be aligned in the image to be processed, and the image after fusing the individual image regions aligned in the image to be processed.
[0139] In two adjacent frames, the moving region is the image region that needs to be aligned, which can be called the image region to be aligned, while the stationary region is the image region that does not need to be aligned, which can be called the aligned image region, that is, the image region that has already been aligned.
[0140] In some implementations, a local motion detection model trained based on deep learning can be used to detect moving and stationary regions in the image to be processed and an adjacent frame. For example, the image to be processed and the previous frame adjacent to the image to be processed can be input into the local motion detection model to obtain the corresponding moving and stationary regions in the two frames.
[0141] In some implementations, after obtaining the motion region and using it as the image region to be aligned, the alignment difficulty of each image region to be aligned can be further classified by an alignment difficulty classification model trained based on deep learning. For example, it can be classified into two categories: difficult alignment and easy alignment.
[0142] Therefore, through the above process, the alignment difficulty of each image region to be aligned in the image to be processed and each image region to be aligned can be obtained.
[0143] In some implementations, the correlation between alignment difficulty and different alignment strategies can be established in advance. After determining the alignment difficulty of each region to be aligned, the target image alignment strategy corresponding to each region to be aligned can be obtained based on the alignment difficulty of each region to be aligned.
[0144] In some implementations, different image alignment strategies result in different alignment speeds. For example, image alignment strategies may include deep learning-based alignment strategies and traditional alignment strategies, where traditional alignment strategies may be, for example, the Homograpy algorithm for global alignment or optical flow methods for pixel-by-pixel alignment.
[0145] For example, suppose the image to be processed has an image region 1 to be aligned and an image region 2 to be aligned, wherein the alignment of the image region 1 is more difficult and the alignment of the image region 2 is easier. In this case, the target image alignment strategy corresponding to the determined image region 1 to be aligned can be a deep learning-based alignment strategy, while the target image alignment strategy corresponding to the determined image region 2 to be aligned can be a traditional alignment strategy.
[0146] In this embodiment of the disclosure, after determining the target alignment strategy corresponding to each region to be aligned, the corresponding regions to be aligned in the image to be processed and the candidate images can be aligned based on the target image alignment strategy corresponding to each region to be aligned, thereby obtaining the pixel correspondence between the corresponding regions to be aligned. The candidate image can be an image that shares the same input to the local motion detection model as the image to be processed, such as the previous frame image adjacent to the image to be processed in a video frame, or the next frame image adjacent to the image to be processed in a video frame.
[0147] In some implementations, fusion processing can be performed by summing and averaging corresponding pixels. For example, the weighted average of two corresponding pixels in two adjacent frames can be taken as the actual value of that pixel in the image to be processed, thereby obtaining the fused image of each image region to be aligned in the image to be processed and the fused image of each image region aligned in the image to be processed.
[0148] In some implementations, the weight ratio can be 1:1. In other implementations, considering that the image to be processed is being imaged, more reference can be made to the pixels of the image to be processed. Therefore, the weight of the image to be processed can be set relatively large, while the weight of the candidate image can be set relatively small. For example, the weight between the image to be processed and the candidate image can be set to 1.5:1.
[0149] Finally, the image obtained by merging the various image regions to be aligned in the image to be processed, and the image obtained by merging the various image regions aligned in the image to be processed, can be stitched together to obtain the target image.
[0150] For example, suppose there are N adjacent image frames in a video frame, namely the first frame, the second frame, ... the Nth frame. When the first frame is received, since there is only one frame, the first frame can be cached.
[0151] When the second frame image is received, it can be used as the image to be processed, and the first frame image can be used as a candidate image. At this time, the first frame image and the second frame image can be input into the local motion detection model to detect the moving and stationary regions in the first frame image and the second frame image. Suppose that the first frame image and the second frame image both contain the image regions of a dog running, a person walking, and a stationary building. At this time, the image regions of the dog running and the person walking can be identified as the image regions to be aligned, and the image regions of the building can be identified as the image regions to be aligned.
[0152] Furthermore, the image regions of the dog running and the person walking in the second frame are input into the alignment difficulty classification model respectively. Assuming that the alignment difficulty of the image region of the dog running is high, while the alignment difficulty of the image region of the person walking is low, a deep learning-based alignment strategy can be used to align the image regions of the dog running in the first frame and the dog running in the second frame to obtain the pixel correspondence between the two image regions of the dog running. Similarly, a traditional alignment strategy can be used to align the image regions of the person walking in the first frame and the person walking in the second frame to obtain the pixel correspondence between the two image regions of the person walking.
[0153] Next, the images can be further fused according to the preset weight relationship and the pixel correspondence of the image regions of the dog running in the two frames to obtain the fused image of the dog running, and according to the preset weight relationship and the pixel correspondence of the image regions of the person walking in the two frames to obtain the fused image of the person walking, and according to the preset weight relationship and the pixel correspondence of the image regions of the building in the two frames to obtain the fused image of the building.
[0154] In other words, the image regions of the dog running in the first frame and the dog running in the second frame can be understood as corresponding image regions to be aligned; the image regions of the person walking in the first frame and the person walking in the second frame can be understood as another corresponding image region to be aligned; and the image regions of the building in the first frame and the building in the second frame can be understood as corresponding aligned image regions.
[0155] When the third frame is received, it can be used as the image to be processed, and the second frame as a candidate image. Alternatively, when the Nth frame is received, it can be used as the image to be processed, and the (N-1)th frame as a candidate image. After determining the image to be processed and the candidate images, the process of determining the target image alignment strategy corresponding to each region to be aligned in the image to be processed, and processing based on the target image alignment strategy and the preset fusion strategy, can be referred to the above content, and will not be repeated here.
[0156] In this embodiment, the alignment strategy can be dynamically adjusted according to the different alignment difficulties of different image regions of the image to be processed, thereby balancing the alignment effect and speed, so as to avoid redundant operations while meeting the alignment requirements. The overall solution is efficient and fast.
[0157] Furthermore, considering that there are a certain number of deep learning-based models in the embodiments of this disclosure, in some implementations, model architecture search techniques can be used to simplify the structure of neural network models, and quantization, distillation, pruning, and other methods can be used to further optimize the processing speed. The optimization process described above can be found in related technologies and will not be elaborated here.
[0158] Furthermore, considering that the ISP system in the related technology can already process images well, but the image processing effect is poor only in some special scenarios, such as night scenes and rainy scenes, in some implementations, the image processing method of the present disclosure embodiment can be activated in night scenes or rainy scenes to improve the imaging effect.
[0159] Therefore, in some implementations, in scenarios such as nighttime or rainy scenes, the electronic device can respond to the user's trigger operation, enable the image processing method of this disclosure embodiment, and execute steps S110-S140.
[0160] In other embodiments, the electronic device can detect that the environment is in a special scene such as a night scene or a rainy scene, and then automatically activate the image processing method of the present disclosure embodiment to execute steps S110-S140.
[0161] The image processing method of this disclosure embodiment will be illustrated below using a complete nighttime video shooting scenario as an example.
[0162] When recording video at night, the electronic device can automatically activate the image processing method of this disclosure embodiment. The image sensor can continuously acquire video frame sequences, i.e., multi-frame RAW images. After acquiring the first frame image, the first frame image can be input as the image to be processed into the image processing device. In the image processing device, such as... Figure 2As shown, the global semantic attributes of the image to be processed can be extracted first using a global semantic attribute extractor trained by deep learning. Then, the color adjustment model corresponding to the image to be processed can be determined based on the global semantic attributes of the image to be processed. The determined color adjustment model can then be used to adjust the color of the image to be processed, resulting in the color-adjusted target image.
[0163] Next, after obtaining the color-adjusted target image, as follows: Figure 3 As shown, the image processing device can use the color-adjusted target image as the image to be processed, and use a deep learning-based local noise intensity extractor to extract the noise intensity (e.g., ...) of each image region in the image to be processed. Figure 3 As shown in the image, region 1 corresponds to noise intensity 1, region 2 corresponds to noise intensity 2, and region M corresponds to noise intensity M. Then, based on the noise intensity of each image region, the denoising model corresponding to each image region is determined (e.g., Figure 3 As shown in the figure, region 1 corresponds to denoising model 1, region 2 corresponds to denoising model 2, and region M corresponds to denoising model M. The determined denoising model is used to denoise each region of the image to be processed, and the intermediate image after denoising is obtained for each region.
[0164] Next, a local brightness extractor based on deep learning is used to extract the brightness of each image region in the image to be processed. Then, based on the brightness of each image region, the brightness adjustment model corresponding to each image region is determined. The determined brightness adjustment model is then used to perform brightness adjustment processing on the image to be processed, resulting in an intermediate image after brightness adjustment.
[0165] Next, a deep learning-based local edge sharpness extractor is used to extract the sharpness of each image region in the image to be processed. Then, based on the sharpness of each image region, the sharpening model corresponding to each image region is determined, and the determined sharpening model is used to sharpen the image to be processed, so as to obtain the sharpened target image, which is the first target image after image enhancement.
[0166] At this point, the first frame image can be cached, waiting for the second frame image acquired by the image sensor to complete the above-mentioned color adjustment and image enhancement process. After the second frame image completes the above-mentioned color adjustment and image enhancement process, the enhanced second target image can be obtained.
[0167] After obtaining the first target image and the second target image, as follows Figure 4 As shown, the first target image can be used as a candidate image, and the second target image can be used as the object to be processed. Both images are input into a deep learning-based local motion detection model to detect motion regions (such as...) in the object to be processed and in the candidate images. Figure 4The regions to be aligned are shown as 1, 2, and P, and the stationary region is shown as alignment region 1.
[0168] Next, after obtaining the motion region and using it as the image region to be aligned, the process continues as follows: Figure 4 As shown, the alignment difficulty classification model trained based on deep learning can be used to classify the alignment difficulty of each image region to be aligned in the image to be processed, thereby obtaining the alignment difficulty of each image region to be aligned in the image to be processed.
[0169] Then, continue as follows Figure 4 As shown, based on the alignment difficulty of each region to be aligned, the target image alignment model corresponding to each region to be aligned is obtained. Based on the target image alignment model corresponding to each region to be aligned, the corresponding regions to be aligned in the image to be processed and the candidate image are aligned to obtain the pixel correspondence between the corresponding regions to be aligned.
[0170] Then, continue as follows Figure 4 As shown, based on the preset fusion strategy and the pixel correspondence between the corresponding image regions to be aligned, the corresponding image regions to be aligned are further fused to obtain the fused image of the corresponding image regions in the image to be processed. Also, based on the preset fusion strategy, the corresponding aligned image regions in the image to be processed and the candidate image are fused to obtain the fused image of the corresponding aligned image regions in the image to be processed.
[0171] Finally, continue as follows Figure 4 As shown, the image obtained by merging the image regions to be aligned in the image to be processed and the image obtained by merging the image regions aligned in the image to be processed can be stitched together to obtain the target image for output.
[0172] This completes the image processing for the first and second frames in the video frame sequence. Subsequently, for the third frame, color adjustment and image enhancement can be performed using the aforementioned method. After image enhancement, the target image obtained from enhancing the third frame can be used as the image to be processed, and the target image obtained from enhancing the second frame can be used as a candidate image. Inter-frame stabilization processing is then performed to obtain the target image corresponding to the third frame, which is then output.
[0173] It is understandable that the processing methods for the fourth, fifth, and Nth frames can refer to the above process, and will not be repeated here.
[0174] Therefore, after processing the Nth frame image to obtain the target image, the entire video frame sequence after image processing can be obtained.
[0175] Figure 5 This is a block diagram of an image processing apparatus 500 according to an exemplary embodiment. (Refer to...) Figure 5 The device 500 includes an acquisition module 510, an extraction module 520, a determination module 530, and a processing module 540.
[0176] The acquisition module 510 is configured to acquire the image to be processed.
[0177] The extraction module 520 is configured to extract attribute information of the image to be processed.
[0178] The determining module 530 is configured to determine the image processing strategy corresponding to the image to be processed based on the attribute information.
[0179] The processing module 540 is configured to process the image to be processed based on the image processing strategy to obtain the target image.
[0180] In some embodiments, the attribute information of the image to be processed includes the global semantic attributes of the image to be processed, and the determining module 530 includes:
[0181] The first context determination submodule is configured to determine the first scene context corresponding to the image to be processed, at least based on the global semantic attributes of the image to be processed.
[0182] The first acquisition submodule is configured to acquire the first color adjustment strategy corresponding to the first scene context.
[0183] Correspondingly, the processing module 540 includes:
[0184] The first color adjustment submodule is configured to adjust the color of the image to be processed based on the first color adjustment strategy to obtain the target image.
[0185] In some implementations, the first context determining submodule includes:
[0186] The first determining unit is configured to determine the first scene context corresponding to the image to be processed based on the global semantic attributes of the image to be processed;
[0187] The second determining unit is configured to determine the first scene context corresponding to the image to be processed based on the global semantic attributes of the image to be processed and the global semantic attributes of images of a preset number of frames adjacent to the image to be processed.
[0188] In some embodiments, the attribute information of the image to be processed includes global semantic attributes and temporal semantic attributes of the image to be processed. The temporal semantic attributes of the image to be processed are determined based on images adjacent to the image to be processed for a preset number of frames. The determining module 530 includes:
[0189] The second scene context determination submodule is configured to determine the second scene context corresponding to the image to be processed based on the global semantic attributes and temporal semantic attributes of the image to be processed.
[0190] The second acquisition submodule is configured to acquire the second color adjustment strategy corresponding to the second scene context.
[0191] Correspondingly, the processing module 540 includes:
[0192] The second color adjustment submodule is configured to adjust the color of the image to be processed based on the second color adjustment strategy to obtain the target image.
[0193] In some embodiments, the attribute information of the image to be processed includes local semantic attributes of each image region in the image to be processed, and the determining module 530 includes:
[0194] The target image enhancement strategy determination submodule is configured to determine the target image enhancement strategy corresponding to each image region based on the local semantic attributes of each image region.
[0195] Correspondingly, the processing module 540 includes:
[0196] The image enhancement submodule is configured to use the target image enhancement strategy corresponding to each image region to enhance each image region and obtain the enhanced image corresponding to each image region.
[0197] The first stitching submodule is configured to stitch together the enhanced images corresponding to each image region to obtain the target image.
[0198] In some implementations, the type of the local semantic attribute includes at least one of the following attribute types: noise intensity, detail richness, brightness, and edge sharpness.
[0199] In some embodiments, the attribute information of the image to be processed includes the alignment difficulty of each image region to be aligned in the image to be processed and each image region to be aligned. The determining module 530 includes:
[0200] The target image alignment strategy determination submodule is configured to determine the target image alignment strategy corresponding to each image region to be aligned based on the alignment difficulty of each image region to be aligned.
[0201] Correspondingly, the processing module 540 includes:
[0202] The alignment submodule is configured to align the corresponding image regions in the image to be processed and the candidate image based on the target image alignment strategy corresponding to each image region to be aligned, thereby obtaining the pixel correspondence between the corresponding image regions to be aligned. The candidate image is a frame image adjacent to the image to be processed.
[0203] The first fusion submodule is configured to perform fusion processing on the corresponding image regions to be aligned based on a preset fusion strategy and the pixel correspondence between the corresponding image regions to be aligned, so as to obtain the fused image of the image to be processed.
[0204] The second fusion submodule is configured to perform fusion processing on the image to be processed and the corresponding aligned image regions in the candidate image based on the preset fusion strategy, so as to obtain the fused image of each aligned image region in the image to be processed.
[0205] The second stitching submodule is configured to stitch together the image after merging the image regions to be aligned in the image to be processed and the image after merging the image regions aligned in the image to be processed to obtain the target image.
[0206] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0207] This disclosure also provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the steps of the image processing method provided in this disclosure.
[0208] Figure 6 This is a block diagram illustrating an electronic device according to an exemplary embodiment. For example, the electronic device 600 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0209] Reference Figure 6 The electronic device 600 may include one or more of the following components: processing component 602, memory 604, power supply component 606, multimedia component 608, audio component 610, input / output interface 612, sensor component 614, and communication component 616.
[0210] Processing component 602 typically controls the overall operation of electronic device 600, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 602 may include one or more processors 620 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 602 may include one or more modules to facilitate interaction between processing component 602 and other components. For example, processing component 602 may include a multimedia module to facilitate interaction between multimedia component 608 and processing component 602.
[0211] Memory 604 is configured to store various types of data to support the operation of electronic device 600. Examples of this data include instructions for any application or method operating on electronic device 600, contact data, phonebook data, messages, pictures, videos, etc. Memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0212] Power supply component 606 provides power to various components of electronic device 600. Power supply component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 600.
[0213] Multimedia component 608 includes a screen that provides an output interface between the electronic device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 608 includes a front-facing camera and / or a rear-facing camera. When the electronic device 600 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0214] Audio component 610 is configured to output and / or input audio signals. For example, audio component 610 includes a microphone (MIC) configured to receive external audio signals when electronic device 600 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 604 or transmitted via communication component 616. In some embodiments, audio component 610 also includes a speaker for outputting audio signals.
[0215] Input / output interface 612 provides an interface between processing component 602 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, start buttons, and lock buttons.
[0216] Sensor assembly 614 includes one or more sensors for providing state assessments of various aspects of electronic device 600. For example, sensor assembly 614 can detect the on / off state of electronic device 600, the relative positioning of components such as the display and keypad of electronic device 600, changes in position of electronic device 600 or a component of electronic device 600, the presence or absence of user contact with electronic device 600, orientation or acceleration / deceleration of electronic device 600, and temperature changes of electronic device 600. Sensor assembly 614 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 614 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 614 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0217] Communication component 616 is configured to facilitate wired or wireless communication between electronic device 600 and other devices. Electronic device 600 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 616 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 616 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0218] In an exemplary embodiment, the electronic device 600 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0219] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, which can be executed by a processor 620 of an electronic device 600 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0220] The aforementioned device can be a standalone electronic device or a part of a standalone electronic device. For example, in one embodiment, the device can be an integrated circuit (IC) or a chip, wherein the integrated circuit can be a single IC or a collection of multiple ICs. The chip can include, but is not limited to, the following types: GPU (Graphics Processing Unit), CPU (Central Processing Unit), FPGA (Field Programmable Gate Array), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), and SoC (System on Chip). The aforementioned integrated circuit or chip can be used to execute executable instructions (or code) to implement the aforementioned image processing method. The executable instructions can be stored in the integrated circuit or chip or obtained from other devices or equipment. For example, the integrated circuit or chip includes a processor, memory, and an interface for communicating with other devices. The executable instructions can be stored in the memory, and when the executable instructions are executed by the processor, the above-described image processing method can be implemented; alternatively, the integrated circuit or chip can receive the executable instructions through the interface and transmit them to the processor for execution to implement the above-described image processing method.
[0221] In another exemplary embodiment, a computer program product is also provided, which includes a computer program executable by a programmable device, the computer program having a code portion for performing the image processing method described above when executed by the programmable device.
[0222] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of this disclosure. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0223] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An image processing method, characterized in that, include: Obtain the image to be processed; Extract the attribute information of the image to be processed; Based on the attribute information, determine the image processing strategy corresponding to the image to be processed; The image to be processed is processed based on the image processing strategy to obtain the target image; The attribute information of the image to be processed includes the alignment difficulty of each image region to be aligned and each image region to be aligned. Determining the image processing strategy corresponding to the image to be processed based on the attribute information includes: Based on the alignment difficulty of each image region to be aligned, determine the target image alignment strategy corresponding to each image region to be aligned; The step of processing the image to be processed based on the image processing strategy to obtain the target image includes: Based on the target image alignment strategy corresponding to each image region to be aligned, the corresponding image regions to be aligned in the image to be processed and the candidate image are aligned to obtain the pixel correspondence between the corresponding image regions to be aligned. The candidate image is a frame image adjacent to the image to be processed. Based on the preset fusion strategy and the pixel correspondence between the corresponding image regions to be aligned, the corresponding image regions to be aligned are fused to obtain the fused image of the image to be processed. Based on the preset fusion strategy, the corresponding aligned image regions in the image to be processed and the candidate image are fused to obtain the fused image of each aligned image region in the image to be processed; The target image is obtained by stitching together the image after fusing the image regions to be aligned in the image to be processed and the image after fusing the image regions to be aligned in the image to be processed.
2. The method according to claim 1, characterized in that, The attribute information of the image to be processed includes the global semantic attributes of the image to be processed. Determining the image processing strategy corresponding to the image to be processed based on the attribute information includes: At least based on the global semantic attributes of the image to be processed, the first scene context corresponding to the image to be processed is determined; Obtain the first color adjustment strategy corresponding to the first scene context; The step of processing the image to be processed based on the image processing strategy to obtain the target image includes: Based on the first color adjustment strategy, the image to be processed is color adjusted to obtain the target image.
3. The method according to claim 2, characterized in that, Determining the first scene context corresponding to the image to be processed, at least based on the global semantic attributes of the image to be processed, includes: Based on the global semantic attributes of the image to be processed, determine the first scene context corresponding to the image to be processed; or The first scene context corresponding to the image to be processed is determined based on the global semantic attributes of the image to be processed and the global semantic attributes of images in a preset number of adjacent frames.
4. The method according to claim 1, characterized in that, The attribute information of the image to be processed includes global semantic attributes and temporal semantic attributes of the image to be processed. The temporal semantic attributes of the image to be processed are determined based on images adjacent to the image to be processed for a preset number of frames. Determining the image processing strategy corresponding to the image to be processed based on the attribute information includes: Based on the global semantic attributes and temporal semantic attributes of the image to be processed, the second scene context corresponding to the image to be processed is determined; Obtain the second color adjustment strategy corresponding to the second scene context; The step of processing the image to be processed based on the image processing strategy to obtain the target image includes: Based on the second color adjustment strategy, the image to be processed is color adjusted to obtain the target image.
5. The method according to claim 1, characterized in that, The attribute information of the image to be processed includes the local semantic attributes of each image region in the image to be processed. Determining the image processing strategy corresponding to the image to be processed based on the attribute information includes: Based on the local semantic attributes of each image region, the target image enhancement strategy corresponding to each image region is determined; The step of processing the image to be processed based on the image processing strategy to obtain the target image includes: By utilizing the target image enhancement strategies corresponding to each image region, image enhancement is performed on each image region to obtain the enhanced image corresponding to each image region. The enhanced images corresponding to each of the image regions are stitched together to obtain the target image.
6. The method according to claim 5, characterized in that, The types of local semantic attributes include at least one of the following attribute types: noise intensity, detail richness, brightness, and edge sharpness.
7. The method according to any one of claims 1-6, characterized in that, The image to be processed is either the original image acquired by an image sensor, or the target image obtained after processing the original image using at least one of the image processing strategies, namely color adjustment strategy, image enhancement strategy, and image alignment strategy.
8. An image processing apparatus, characterized in that, include: The acquisition module is configured to acquire the image to be processed. The extraction module is configured to extract attribute information from the image to be processed; The determination module is configured to determine the image processing strategy corresponding to the image to be processed based on the attribute information; The processing module is configured to process the image to be processed based on the image processing strategy to obtain the target image; The attribute information of the image to be processed includes the alignment difficulty of each image region to be aligned and each image region to be aligned. The determining module includes: The target image alignment strategy determination submodule is configured to determine the target image alignment strategy corresponding to each image region to be aligned based on the alignment difficulty of each image region to be aligned. The processing module includes: The alignment submodule is configured to align the corresponding image regions in the image to be processed and the candidate image based on the target image alignment strategy corresponding to each image region to be aligned, thereby obtaining the pixel correspondence between the corresponding image regions to be aligned. The candidate image is a frame image adjacent to the image to be processed. The first fusion submodule is configured to perform fusion processing on the corresponding image regions to be aligned based on a preset fusion strategy and the pixel correspondence between the corresponding image regions to be aligned, so as to obtain the fused image of the image to be processed. The second fusion submodule is configured to perform fusion processing on the image to be processed and the corresponding aligned image regions in the candidate image based on the preset fusion strategy, so as to obtain the fused image of each aligned image region in the image to be processed. The second stitching submodule is configured to stitch together the image after merging the image regions to be aligned in the image to be processed and the image after merging the image regions aligned in the image to be processed to obtain the target image.
9. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When executed by a processor, the program instructions implement the steps of the method described in any one of claims 1 to 7.
11. A chip, characterized in that, It includes a processor and an interface; the processor is used to read instructions to execute the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Image signal processing method, device and equipment
CN109688351A